Research cutoff · 18 July 2026

Where AI helps. Where people remain accountable.

A task-first decision map of illustrative work cases across healthcare, technology, and a cross-industry baseline. Select a task to inspect all six score records, sources, exact model uses, oversight, scenarios, and recommendations.

64illustrative task cases32 occupation contexts · 1/3/5-year scenarios
Task evidence64 illustrative task casesprimary unit of analysis
Occupation context32two cases per mapped occupation
Model identity6stable product identities; benchmark settings separate
Weighting boundaryNot a task mixequal illustrative case weights

Task decision map

Capability meets accountable judgment

Each point is an illustrative task, not an occupation. Horizontal position is its AI-capability assessment; vertical position combines its human-judgment and accountability assessments only for visual placement.

Human-led, high stakesAI–human opportunityLimited AI fitAI-led potential
Healthcare Technology Cross-industry baseline

Text alternative

Decision map data

All values are task-level 0–100 ordinal assessments with per-score evidence metadata. They are not probabilities or task shares.

TaskOccupation contextTypeSectorAI capabilityHuman judgmentAccountability & safetyClassificationOpen evidence
Educate a patient about a care planRegistered Nursedocumentation communicationhealthcare758294ai human
Administer medication at the bedsideRegistered Nursephysical interventionhealthcare2094100human essential
Form a differential diagnosisFamily Medicine Physiciananalysis decisionhealthcare6598100human led
Conduct a shared treatment decisionFamily Medicine Physicianrelationship coordinationhealthcare42100100human essential
Draft a clinical encounter notePhysician Assistantdocumentation communicationhealthcare707892ai human
Perform a physical examinationPhysician Assistantphysical interventionhealthcare259498human essential
Summarize laboratory trendsNurse Practitionerdocumentation communicationhealthcare658896ai human
Prescribe a treatment planNurse Practitioneranalysis decisionhealthcare3598100human essential
Screen a medication list for interactionsPharmacistanalysis decisionhealthcare709199ai human
Counsel a patient about medication usePharmacistrelationship coordinationhealthcare559699human led
Prioritize an imaging worklistRadiologistanalysis decisionhealthcare729199ai human
Finalize an imaging reportRadiologistmonitoring governancehealthcare6098100human led
Analyze laboratory quality-control resultsClinical Laboratory Technologistanalysis decisionhealthcare609399human led
Process a patient specimenClinical Laboratory Technologistphysical interventionhealthcare228898human essential
Adjust ventilator settingsRespiratory Therapistphysical interventionhealthcare2598100human essential
Teach and verify inhaler techniqueRespiratory Therapistanalysis decisionhealthcare559296human led
Design an individualized rehabilitation planPhysical Therapistanalysis decisionhealthcare259894human led
Physically assist gait trainingPhysical Therapistphysical interventionhealthcare129998human essential
Summarize a behavioral-health intakeMental Health Counselordocumentation communicationhealthcare659396ai human
Respond to an acute behavioral-health crisisMental Health Counseloranalysis decisionhealthcare18100100human essential
Assign diagnosis codes from a completed recordMedical Records Specialistdocumentation communicationhealthcare607285ai led
Validate a medical record for completenessMedical Records Specialistmonitoring governancehealthcare608291ai human
Analyze a clinical staffing planMedical and Health Services Managermonitoring governancehealthcare709395ai human
Resolve a care-delivery escalationMedical and Health Services Managerrelationship coordinationhealthcare3899100human led
Implement a routine API endpointSoftware Developersoftware tool worktechnology857682ai human
Design a service architectureSoftware Developeranalysis decisiontechnology609892human led
Generate regression tests for a defined changeSoftware Quality Assurance Analystsoftware tool worktechnology757078ai led
Conduct exploratory software testingSoftware Quality Assurance Analystanalysis decisiontechnology609086ai human
Triage security alertsInformation Security Analystmonitoring governancetechnology609199ai human
Lead a cybersecurity incident responseInformation Security Analystmonitoring governancetechnology4599100human led
Clean a structured analysis datasetData Scientistanalysis decisiontechnology707375ai led
Interpret a causal model for a decisionData Scientistanalysis decisiontechnology659793human led
Draft system requirements from stakeholder notesComputer Systems Analystdocumentation communicationtechnology659488ai human
Recommend system tradeoffs to stakeholdersComputer Systems Analystanalysis decisiontechnology629992human led
Generate a database migrationDatabase Administratorsoftware tool worktechnology728291ai human
Recover a production database after failureDatabase Administratormonitoring governancetechnology4299100human led
Generate a network configuration templateComputer Network Architectsoftware tool worktechnology658394ai human
Design an enterprise network architectureComputer Network Architectanalysis decisiontechnology599996human led
Prototype a user interface from a briefWeb and Digital Interface Designersoftware tool worktechnology858372ai human
Synthesize usability researchWeb and Digital Interface Designeranalysis decisiontechnology609578ai human
Answer a routine technical-support ticketComputer Support Specialistanalysis decisiontechnology857172ai led
Troubleshoot a novel system failureComputer Support Specialistanalysis decisiontechnology709588human led
Translate a complete specification into codeComputer Programmerdocumentation communicationtechnology857884ai led
Debug a complex software defectComputer Programmersoftware tool worktechnology809592ai human
Draft API documentation from code and testsTechnical Writersoftware tool worktechnology708074ai led
Validate technical documentation against behaviorTechnical Writersoftware tool worktechnology659587ai human
Prioritize a technology portfolioComputer and Information Systems Managermonitoring governancetechnology549994human led
Approve a high-risk technology changeComputer and Information Systems Managermonitoring governancetechnology25100100human essential
Stock retail shelvesStocking Associatephysical interventionbaseline185845human essential
Investigate an inventory discrepancyStocking Associateanalysis decisionbaseline657968ai human
Diagnose an electrical fault on siteElectriciananalysis decisionbaseline2599100human led
Install electrical wiringElectricianphysical interventionbaseline1098100human essential
Inspect a plumbing system on sitePlumberanalysis decisionbaseline259899human led
Repair a plumbing fixturePlumberphysical interventionbaseline899100human essential
Draft a differentiated lesson planElementary School Teacherdocumentation communicationbaseline709290ai human
Manage a live elementary classroomElementary School Teacheranalysis decisionbaseline1210099human essential
Forecast restaurant demand and purchasingFood Service Manageranalysis decisionbaseline658278ai human
Supervise food-safety practices during serviceFood Service Manageranalysis decisionbaseline2099100human essential
Answer a routine customer questionCustomer Service Representativeanalysis decisionbaseline856566ai led
Resolve a vulnerable customer's escalationCustomer Service Representativerelationship coordinationbaseline489894human led
Reconcile financial accountsAccountant and Auditordocumentation communicationbaseline728291ai led
Exercise audit judgment on material riskAccountant and Auditoranalysis decisionbaseline609998human led
Draft personalized sales outreachServices Sales Representativedocumentation communicationbaseline707768ai led
Negotiate a client agreementServices Sales Representativerelationship coordinationbaseline489986human led

Analyst overlay · July 2026

My real assessment of current AI capability

These ranges estimate frontier AI's practical technical performance on bounded tasks. They are reasoned estimates—not universal benchmark percentages, employment forecasts or permission to deploy autonomously.

Reasoning capability

Can the AI determine what should be done on a bounded task?

Execution capability

Can the AI complete the workflow using available tools and data?

Reliability

How often is the result correct without an undetected material error?

Deployment authority

Is the AI legally and organizationally permitted to act?

Work categoryTechnical capabilityBest execution modelDecision authority
Writing, summarization & documentation90–98%AI-led draft and transformationHuman review when consequential
Research & evidence synthesis85–95%AI-led search, comparison and synthesisHuman validates sources and conclusions
Coding & software maintenance80–95%AI-led implementation with testsHuman owns architecture, security and release
Data analysis & reporting85–95%AI-led analysis on governed dataHuman verifies definitions and decisions
Scheduling & administrative processing90–98%Highly automatableHuman handles exceptions and appeals
Clinical documentation & chart review85–95%AI drafts, extracts and flagsLicensed clinician verifies the record
Prior authorization & utilization review80–95%AI assembles and compares evidenceHuman controls denials and appeals
Differential-diagnosis support75–90%AI proposes and ranks possibilitiesClinician examines, decides and signs
Complex independent diagnosis60–80%AI support; autonomy remains unreliableLicensed human decision
Treatment-plan generation70–90%AI prepares options and checksClinician individualizes and authorizes
Bedside assessment & physical examination20–45%Limited without sensors and roboticsQualified human leads
Procedures & physical care10–40%Specialized robotics onlyQualified human leads

Adoption scenarios

How fast capability becomes normal work practice

Technical capability can arrive years before organizations permit it. These scenarios separate what AI can do from what workplaces will allow.

1–5 years

Conservative

Institutions move slowly. Humans remain workflow owners while AI drafts, extracts, monitors and recommends.

  1. 1 year: broad supervised copilots
  2. 3 years: selective workflow automation
  3. 5 years: human-led structures remain common
1–5 years

Expected

AI agents perform most bounded digital work. Humans handle exceptions, relationships and consequential authorization.

  1. 1 year: AI-first digital production
  2. 3 years: agents execute multi-step workflows
  3. 5 years: smaller human teams supervise systems
1–5 years

Accelerated

Capability and adoption compound quickly. AI executes nearly all routine digital tasks and escalates uncertainty.

  1. 1 year: rapid agent adoption
  2. 3 years: most routine digital execution automated
  3. 5 years: humans concentrate on authority, trust and edge cases

Stable identities · benchmark settings separate · cutoff 18 July 2026

Model benchmark comparison — bounded evidence

Every model/category cell records a compatible result, bounded qualitative evidence, or a cited researched-unavailable finding. Only identical named evaluation groups are compared directly.

OpenAI

GPT-5.6 Sol

GPT-5.6 SolStable product identity · released 2026-07-09

Anthropic

Claude Fable 5

Claude Fable 5Stable product identity · released 2026-06-09

Google DeepMind

Gemini 3.1 Pro Preview

Gemini 3.1 Pro PreviewStable product identity · released 2026-02-19

SpaceXAI

Grok 4.5

Grok 4.5Stable product identity · released 2026-07-16

DeepSeek

DeepSeek V4 Pro Preview

DeepSeek V4 Pro PreviewStable product identity · released 2026-04-24

Meta

Muse Spark 1.1

Muse Spark 1.1Stable product identity · released 2026-07-09

Compatible comparison · Artificial Analysis Intelligence Index v4.1

  • Gemini 3.1 Pro Preview: Artificial Analysis' stable v4.1 report records a 46 Intelligence Index score and 1.6 minutes per task. Configuration: Gemini 3.1 Pro Preview, Thinking (High)
  • DeepSeek V4 Pro Preview: Artificial Analysis' stable v4.1 report records 44 on its Intelligence Index for V4 Pro at max effort. Configuration: DeepSeek V4 Pro, Max Effort (Artificial Analysis harness)

Other evidence remains separate because harnesses, settings, or measures differ. Provenance is retained as vendor-reported or independent.

Evidence categoryGPT-5.6 SolClaude Fable 5Gemini 3.1 Pro PreviewGrok 4.5DeepSeek V4 Pro PreviewMuse Spark 1.1
Reasoning
Researched — no comparable result

The reviewed provider source does not publish a compatible reasoning result for this configuration.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible reasoning result for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18)
Reported result

Google reports 44.4% on Humanity's Last Exam without tools and 51.4% with search and code.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Developer-run academic benchmarks can be sensitive to prompt, tools, effort, contamination controls, and scoring.Source: Google DeepMind (accessed 2026-07-18)
Reported result

Artificial Analysis' stable v4.1 report records a 46 Intelligence Index score and 1.6 minutes per task.

External evaluationConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-06-15Limit: The composite is methodology-specific and does not directly measure clinical or workplace reliability.Source: Artificial Analysis (accessed 2026-07-18)
Reported result

Artificial Analysis reports 54 on its Intelligence Index at high effort in its pre-release evaluation.

External evaluationConfiguration: Grok 4.5, high effort in Grok Build · Evaluated 2026-07-08Benchmark settings: effort highLimit: The July 8 evaluation predates the July 16 public launch and does not establish reliability on unseen organizational workflows.Source: Artificial Analysis (accessed 2026-07-18)
Reported result

Artificial Analysis' stable v4.1 report records 44 on its Intelligence Index for V4 Pro at max effort.

External evaluationConfiguration: DeepSeek V4 Pro, Max Effort (Artificial Analysis harness) · Evaluated 2026-06-15Benchmark settings: effort max · harness Artificial AnalysisLimit: The result follows a v4.1 methodology change and is not a workplace success or replacement rate.Source: Artificial Analysis (accessed 2026-07-18)
Reported result

Artificial Analysis reports 51 on its Intelligence Index for Muse Spark 1.1 at xhigh effort.

External evaluationConfiguration: Muse Spark 1.1, xhigh effort / Thinking mode · Evaluated 2026-07-10Benchmark settings: effort xhighLimit: The evaluator had pre-release access and the composite is sensitive to harness, task mix, token budget, and graders.Source: Artificial Analysis (accessed 2026-07-18)
Coding
Reported result

Artificial Analysis reports 80 on its Coding Agent Index for GPT-5.6 Sol in Codex.

External evaluationConfiguration: GPT-5.6 Sol, max reasoning · Evaluated 2026-07-09Limit: The result is specific to max reasoning, the Codex harness, and the evaluator's current task and grading mix.Source: Artificial Analysis (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible coding result for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible coding result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18)
Reported result

The July 16 launch report lists 83.3% on Terminal-Bench 2.1 and 64.7% resolution on SWE-Bench Pro.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: The provider page combines its own runs and cited leaderboard figures across differing harnesses and configurations.Source: SpaceXAI (accessed 2026-07-18)
Bounded qualitative evidence

DeepSeek reports open-model-leading agentic coding performance for V4 Pro.

Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: The accessible launch text gives a relative vendor claim without the full numeric chart methodology.Source: DeepSeek (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible coding result for this configuration.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18)
Writing
Bounded qualitative evidence

OpenAI reports improved document generation and structured drafting for GPT-5.6 Sol.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: The provider uses selected document examples and customer evaluations rather than a portable workplace-writing score.Source: OpenAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible writing result for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible writing result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible writing result for this configuration.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible writing result for this configuration.

Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible writing result for this configuration.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18)
Data analysis
Researched — no comparable result

The reviewed provider source does not publish a compatible data analysis result for this configuration.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18)
Reported result

Artificial Analysis reports a leading 1932 GDPval-AA score with fallback on 2% of tasks.

External evaluationConfiguration: Claude Fable 5, adaptive reasoning max, Opus 4.8 fallback · Evaluated 2026-06-09Limit: This is one evaluator's agentic knowledge-work benchmark at max effort and is not evidence of whole-job replacement.Source: Artificial Analysis (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible data analysis result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible data analysis result for this configuration.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible data analysis result for this configuration.

Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible data analysis result for this configuration.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18)
Healthcare support
Reported result

OpenAI reports 60.5% on HealthBench Professional for GPT-5.6 Sol.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: HealthBench Professional is a rubric-scored evaluation and does not establish clinical outcomes or autonomous-care safety.Source: OpenAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible healthcare support result for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible healthcare support result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible healthcare support result for this configuration.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible healthcare support result for this configuration.

Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible healthcare support result for this configuration.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18)
Tool use
Researched — no comparable result

The reviewed provider source does not publish a compatible tool use result for this configuration.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible tool use result for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible tool use result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible tool use result for this configuration.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible tool use result for this configuration.

Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18)
Bounded qualitative evidence

Meta reports gains in computer use, tool orchestration, coding, and one-million-token context management.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Several results are internal or vendor-run and the API is a public preview whose behavior may change.Source: Meta (accessed 2026-07-18)
Agentic execution
Bounded qualitative evidence

OpenAI's page contains conflicting Agents' Last Exam values: the narrative states 53.6 while the results table lists 52.7 for GPT-5.6 Sol.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: The source disagreement is unresolved, and neither value is a production success rate across occupations.Source disagreement: The same OpenAI release reports 53.6 in narrative text and 52.7 in its Agents' Last Exam results table; neither value is selected as authoritative here.Source: OpenAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible agentic execution result for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible agentic execution result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible agentic execution result for this configuration.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible agentic execution result for this configuration.

Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible agentic execution result for this configuration.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18)
Reliability
Researched — no comparable result

The reviewed provider source does not publish a compatible reliability result for this configuration.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible reliability result for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible reliability result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible reliability result for this configuration.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18)
Reported result

Artificial Analysis reports -10 on AA-Omniscience and a 94% hallucination rate when the model did not know an answer.

External evaluationConfiguration: DeepSeek V4 Pro, Max Effort (Artificial Analysis harness) · Evaluated 2026-04-24Benchmark settings: effort max · harness Artificial AnalysisLimit: These are evaluator-specific knowledge and non-hallucination measures, not a universal factual-error rate.Source: Artificial Analysis (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible reliability result for this configuration.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18)
Oversight
Researched — no comparable result

The reviewed provider source does not publish a compatible oversight result for this configuration.

Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: OpenAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a comparable human-review, intervention, or escalation requirement for this configuration.

Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Fallback routing is not human oversight; the release reports model-to-model safeguard routing, so no oversight requirement is inferred.Source: Anthropic (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible oversight result for this configuration.

Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: Google DeepMind (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible oversight result for this configuration.

Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: SpaceXAI (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible oversight result for this configuration.

Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: DeepSeek (accessed 2026-07-18)
Researched — no comparable result

The reviewed provider source does not publish a compatible oversight result for this configuration.

Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: Meta (accessed 2026-07-18)

Guided research

Turn benchmark evidence into work decisions

Use this sequence to distinguish a bounded model result from responsible deployment, workforce planning, or personal career action.

2. A task is not a job

Jobs combine many tasks with context, handoffs, relationship work, physical presence, and formal authority. This map contains two equally weighted illustrative cases per occupation—not a measured task mix—and does not publish occupation ratings or predict that an occupation disappears.

Sources: National Center for O*NET Development, U.S. Department of Labor: Registered Nurses, O*NET OnLine occupation summary; National Center for O*NET Development, U.S. Department of Labor: Software Developers, O*NET OnLine occupation summary

3. Adoption has constraints

Demonstrated capability is only one condition. Adoption also depends on workflow integration, data quality, procurement, cost, latency, privacy, security, training, regulation, labor agreements, acceptance, and a workable route for exceptions and accountability.

4. Healthcare and technology cases

Healthcare evidence in this map includes simulations and bounded workflows, not patient outcomes. It does not support autonomous diagnosis, prescribing, crisis care, or clinical sign-off. In technology, code and tool benchmarks still leave requirements, security, integration, review, and ownership with people.

Sources: OpenAI: GPT-5.6: Frontier intelligence that scales with your ambition; Artificial Analysis: GPT-5.6 benchmarks across Intelligence, Speed and Cost; METR: Task-Completion Time Horizons of Frontier AI Models

5. Workforce action

Start with a bounded task, define a quality baseline, name who reviews output, log exceptions, and measure whether the workflow improves. Expand only when evidence supports the next use case; preserve accountable human authority for consequential work.

6. Career strategy

Build fluency in supervised AI use alongside domain judgment, communication, verification, exception handling, and collaboration. Treat the one-, three-, and five-year views as planning scenarios—not guarantees about hiring, pay, or displacement.

Research recommendations

Supported conclusions: current evidence supports bounded task assistance and task-specific human-control decisions; it does not support occupation replacement rates.

Disagreements: the GPT-5.6 provider page reports different Agents' Last Exam values in narrative text and its result table, so this map records the conflict without choosing a number.

Evidence gaps: most model/category pairs lack compatible measures, occupation task importance weights are not compiled, and many task scores transfer from related evaluations rather than direct workplace studies.

Uncertainty and refresh: model configurations, benchmark methods, deployment evidence, regulation, and adoption can change quickly. Recheck model evidence by 18 January 2027 and before any consequential decision.

Task-first decision explorer

Plan from a task and its evidence

Filter occupation contexts and task types, then select a task for scores, sources, exact model uses, oversight, scenarios, and separate career, workforce, and research recommendations.

Occupation, task-type, classification, horizon, and selected-task state are stored in the URL for sharing and Back/Forward restoration.

32 matching occupation contexts

Occupation context · healthcare

Registered Nurse

Illustrative cases

Paired illustrative cases — not a measured task mix. Equal case weights are not importance, frequency, or time-use weights and must not be read as an occupation task mix.

Choose a task

Skills to strengthen

  • Task-grounded practice: accountable judgment and exception handling for “Educate a patient about a care plan”.
  • Task-grounded practice: accountable judgment and exception handling for “Administer medication at the bedside”.

Potentially routinized activities

No routinization claim is supported by these two cases.

AI tool guidance

  • Evaluate bounded assistance for “Educate a patient about a care plan”; this is not a deployment approval. The model record supplies category-level evidence only; task-specific quality, workflow fit, and safety still require local evaluation.

Transitions — not assessed

Transitions are not assessed because transferable skills, education, credentials, licensure, training time, work setting, and labor evidence were not compiled for Registered Nurse.

These prompts derive from two O*NET-linked illustrative tasks; they are not measured rising or declining skill trends. Assessment date: 2026-07-18. Sources: onet-registered-nurse.

Task evidence · Documentation & communication

Task evidence: Educate a patient about a care plan

AI–human partnership

Scores are ordinal task assessments, not automation percentages or employment forecasts. Each record retains its rationale, confidence, assessment date, and source trail.

AI capability
75/100 · Confidence: medium

Direct task evidence: HealthBench: Evaluating Large Language Models Towards Improved Human Health evaluates the same workflow family as “Educate a patient about a care plan.” A score of 75 (60–79 rubric band) reflects that strong drafting and plain-language transformation when the clinical facts are supplied. The cited result is bounded and is not a claim of autonomous workplace execution.

Assessment date: 2026-07-18 · Sources: openai-healthbench-2025

Automation exposure
60/100 · Confidence: low

Analyst automation inference: the reviewed direct capability evidence and O*NET's Registered Nurse task context place “Educate a patient about a care plan” at 60 (60–79 rubric band) because a substantial drafting step can be assisted, while tailoring and delivery remain clinician work. This is not an adoption, job-loss, or replacement percentage.

Assessment date: 2026-07-18 · Sources: openai-healthbench-2025, onet-registered-nurse

Human judgment
82/100 · Confidence: medium

Occupation-specific inference: O*NET's Registered Nurse task context places human judgment for “Educate a patient about a care plan” at 82 (80–89 rubric band) because comprehension, culture, preferences, and changing symptoms require clinician judgment. The score is an analyst rubric placement, not a measured human-performance rate.

Assessment date: 2026-07-18 · Sources: onet-registered-nurse

Accountability & safety
94/100 · Confidence: medium

Occupation-specific safety inference: O*NET's Registered Nurse duties place accountability and consequence severity for “Educate a patient about a care plan” at 94 (90–100 rubric band) because incorrect or misunderstood instructions can cause direct patient harm. This measures task stakes, not model safety performance.

Assessment date: 2026-07-18 · Sources: onet-registered-nurse, openai-healthbench-2025

Collaboration advantage
82/100 · Confidence: medium

Analyst workflow inference: combining the reviewed direct capability evidence with O*NET's Registered Nurse task context places AI–human collaboration on “Educate a patient about a care plan” at 82 (80–89 rubric band) because AI drafting plus nurse verification can improve consistency without transferring accountability. No cited study measures this exact deployment design.

Assessment date: 2026-07-18 · Sources: openai-healthbench-2025, onet-registered-nurse

Career durability
90/100 · Confidence: low

Analyst durability inference: O*NET's Registered Nurse task requirements place the persistent human contribution to “Educate a patient about a care plan” at 90 (90–100 rubric band) because trust, bedside assessment, and licensed responsibility preserve a large human contribution. This forward-looking ordinal value is a scenario input, not an employment forecast.

Assessment date: 2026-07-18 · Sources: onet-registered-nurse

Exact model configurations
  • GPT-5.6 SolEvaluate bounded assistance for “Educate a patient about a care plan”; this is not a deployment approval.
    GPT-5.6 Sol (vendor mode and effort not stated for these claims)
    Evidence fit: transfer · Oversight: human-review. The model record supplies category-level evidence only; task-specific quality, workflow fit, and safety still require local evaluation.
1-year scenario

Supervised AI assistance for educate a patient about a care plan may expand where inputs and review responsibilities are clear.

Career recommendation
  • Strengthen accountable judgment, context, and communication within “Educate a patient about a care plan” while testing only bounded assistance.

This recommendation is scoped to one illustrative task and is not an occupation-wide labor-market forecast.

Workforce recommendation

augment · human review

Use AI for bounded assistance while a qualified person reviews, contextualizes, and owns the result.

Research recommendation

Supported: ai human is the supported task-level classification under the cited evidence and stated rubric.

Disagreements: None recorded for this task; source-level conflicts remain visible in the model evidence section.

Gaps: No production deployment or outcome study is cited for this exact workplace workflow.

Uncertainty: Scores are analyst ordinal assessments; adoption, costs, regulation, integration, and real-world outcomes remain uncertain.

Refresh by: 2027-01-18

Task sources