Research cutoff · 18 July 2026
Where AI helps. Where people remain accountable.
A task-first decision map of illustrative work cases across healthcare, technology, and a cross-industry baseline. Select a task to inspect all six score records, sources, exact model uses, oversight, scenarios, and recommendations.
Task decision map
Capability meets accountable judgment
Each point is an illustrative task, not an occupation. Horizontal position is its AI-capability assessment; vertical position combines its human-judgment and accountability assessments only for visual placement.
Text alternative
Decision map data
All values are task-level 0–100 ordinal assessments with per-score evidence metadata. They are not probabilities or task shares.
| Task | Occupation context | Type | Sector | AI capability | Human judgment | Accountability & safety | Classification | Open evidence |
|---|---|---|---|---|---|---|---|---|
| Educate a patient about a care plan | Registered Nurse | documentation communication | healthcare | 75 | 82 | 94 | ai human | |
| Administer medication at the bedside | Registered Nurse | physical intervention | healthcare | 20 | 94 | 100 | human essential | |
| Form a differential diagnosis | Family Medicine Physician | analysis decision | healthcare | 65 | 98 | 100 | human led | |
| Conduct a shared treatment decision | Family Medicine Physician | relationship coordination | healthcare | 42 | 100 | 100 | human essential | |
| Draft a clinical encounter note | Physician Assistant | documentation communication | healthcare | 70 | 78 | 92 | ai human | |
| Perform a physical examination | Physician Assistant | physical intervention | healthcare | 25 | 94 | 98 | human essential | |
| Summarize laboratory trends | Nurse Practitioner | documentation communication | healthcare | 65 | 88 | 96 | ai human | |
| Prescribe a treatment plan | Nurse Practitioner | analysis decision | healthcare | 35 | 98 | 100 | human essential | |
| Screen a medication list for interactions | Pharmacist | analysis decision | healthcare | 70 | 91 | 99 | ai human | |
| Counsel a patient about medication use | Pharmacist | relationship coordination | healthcare | 55 | 96 | 99 | human led | |
| Prioritize an imaging worklist | Radiologist | analysis decision | healthcare | 72 | 91 | 99 | ai human | |
| Finalize an imaging report | Radiologist | monitoring governance | healthcare | 60 | 98 | 100 | human led | |
| Analyze laboratory quality-control results | Clinical Laboratory Technologist | analysis decision | healthcare | 60 | 93 | 99 | human led | |
| Process a patient specimen | Clinical Laboratory Technologist | physical intervention | healthcare | 22 | 88 | 98 | human essential | |
| Adjust ventilator settings | Respiratory Therapist | physical intervention | healthcare | 25 | 98 | 100 | human essential | |
| Teach and verify inhaler technique | Respiratory Therapist | analysis decision | healthcare | 55 | 92 | 96 | human led | |
| Design an individualized rehabilitation plan | Physical Therapist | analysis decision | healthcare | 25 | 98 | 94 | human led | |
| Physically assist gait training | Physical Therapist | physical intervention | healthcare | 12 | 99 | 98 | human essential | |
| Summarize a behavioral-health intake | Mental Health Counselor | documentation communication | healthcare | 65 | 93 | 96 | ai human | |
| Respond to an acute behavioral-health crisis | Mental Health Counselor | analysis decision | healthcare | 18 | 100 | 100 | human essential | |
| Assign diagnosis codes from a completed record | Medical Records Specialist | documentation communication | healthcare | 60 | 72 | 85 | ai led | |
| Validate a medical record for completeness | Medical Records Specialist | monitoring governance | healthcare | 60 | 82 | 91 | ai human | |
| Analyze a clinical staffing plan | Medical and Health Services Manager | monitoring governance | healthcare | 70 | 93 | 95 | ai human | |
| Resolve a care-delivery escalation | Medical and Health Services Manager | relationship coordination | healthcare | 38 | 99 | 100 | human led | |
| Implement a routine API endpoint | Software Developer | software tool work | technology | 85 | 76 | 82 | ai human | |
| Design a service architecture | Software Developer | analysis decision | technology | 60 | 98 | 92 | human led | |
| Generate regression tests for a defined change | Software Quality Assurance Analyst | software tool work | technology | 75 | 70 | 78 | ai led | |
| Conduct exploratory software testing | Software Quality Assurance Analyst | analysis decision | technology | 60 | 90 | 86 | ai human | |
| Triage security alerts | Information Security Analyst | monitoring governance | technology | 60 | 91 | 99 | ai human | |
| Lead a cybersecurity incident response | Information Security Analyst | monitoring governance | technology | 45 | 99 | 100 | human led | |
| Clean a structured analysis dataset | Data Scientist | analysis decision | technology | 70 | 73 | 75 | ai led | |
| Interpret a causal model for a decision | Data Scientist | analysis decision | technology | 65 | 97 | 93 | human led | |
| Draft system requirements from stakeholder notes | Computer Systems Analyst | documentation communication | technology | 65 | 94 | 88 | ai human | |
| Recommend system tradeoffs to stakeholders | Computer Systems Analyst | analysis decision | technology | 62 | 99 | 92 | human led | |
| Generate a database migration | Database Administrator | software tool work | technology | 72 | 82 | 91 | ai human | |
| Recover a production database after failure | Database Administrator | monitoring governance | technology | 42 | 99 | 100 | human led | |
| Generate a network configuration template | Computer Network Architect | software tool work | technology | 65 | 83 | 94 | ai human | |
| Design an enterprise network architecture | Computer Network Architect | analysis decision | technology | 59 | 99 | 96 | human led | |
| Prototype a user interface from a brief | Web and Digital Interface Designer | software tool work | technology | 85 | 83 | 72 | ai human | |
| Synthesize usability research | Web and Digital Interface Designer | analysis decision | technology | 60 | 95 | 78 | ai human | |
| Answer a routine technical-support ticket | Computer Support Specialist | analysis decision | technology | 85 | 71 | 72 | ai led | |
| Troubleshoot a novel system failure | Computer Support Specialist | analysis decision | technology | 70 | 95 | 88 | human led | |
| Translate a complete specification into code | Computer Programmer | documentation communication | technology | 85 | 78 | 84 | ai led | |
| Debug a complex software defect | Computer Programmer | software tool work | technology | 80 | 95 | 92 | ai human | |
| Draft API documentation from code and tests | Technical Writer | software tool work | technology | 70 | 80 | 74 | ai led | |
| Validate technical documentation against behavior | Technical Writer | software tool work | technology | 65 | 95 | 87 | ai human | |
| Prioritize a technology portfolio | Computer and Information Systems Manager | monitoring governance | technology | 54 | 99 | 94 | human led | |
| Approve a high-risk technology change | Computer and Information Systems Manager | monitoring governance | technology | 25 | 100 | 100 | human essential | |
| Stock retail shelves | Stocking Associate | physical intervention | baseline | 18 | 58 | 45 | human essential | |
| Investigate an inventory discrepancy | Stocking Associate | analysis decision | baseline | 65 | 79 | 68 | ai human | |
| Diagnose an electrical fault on site | Electrician | analysis decision | baseline | 25 | 99 | 100 | human led | |
| Install electrical wiring | Electrician | physical intervention | baseline | 10 | 98 | 100 | human essential | |
| Inspect a plumbing system on site | Plumber | analysis decision | baseline | 25 | 98 | 99 | human led | |
| Repair a plumbing fixture | Plumber | physical intervention | baseline | 8 | 99 | 100 | human essential | |
| Draft a differentiated lesson plan | Elementary School Teacher | documentation communication | baseline | 70 | 92 | 90 | ai human | |
| Manage a live elementary classroom | Elementary School Teacher | analysis decision | baseline | 12 | 100 | 99 | human essential | |
| Forecast restaurant demand and purchasing | Food Service Manager | analysis decision | baseline | 65 | 82 | 78 | ai human | |
| Supervise food-safety practices during service | Food Service Manager | analysis decision | baseline | 20 | 99 | 100 | human essential | |
| Answer a routine customer question | Customer Service Representative | analysis decision | baseline | 85 | 65 | 66 | ai led | |
| Resolve a vulnerable customer's escalation | Customer Service Representative | relationship coordination | baseline | 48 | 98 | 94 | human led | |
| Reconcile financial accounts | Accountant and Auditor | documentation communication | baseline | 72 | 82 | 91 | ai led | |
| Exercise audit judgment on material risk | Accountant and Auditor | analysis decision | baseline | 60 | 99 | 98 | human led | |
| Draft personalized sales outreach | Services Sales Representative | documentation communication | baseline | 70 | 77 | 68 | ai led | |
| Negotiate a client agreement | Services Sales Representative | relationship coordination | baseline | 48 | 99 | 86 | human led |
Analyst overlay · July 2026
My real assessment of current AI capability
These ranges estimate frontier AI's practical technical performance on bounded tasks. They are reasoned estimates—not universal benchmark percentages, employment forecasts or permission to deploy autonomously.
Reasoning capability
Can the AI determine what should be done on a bounded task?
Execution capability
Can the AI complete the workflow using available tools and data?
Reliability
How often is the result correct without an undetected material error?
Deployment authority
Is the AI legally and organizationally permitted to act?
| Work category | Technical capability | Best execution model | Decision authority |
|---|---|---|---|
| Writing, summarization & documentation | 90–98% | AI-led draft and transformation | Human review when consequential |
| Research & evidence synthesis | 85–95% | AI-led search, comparison and synthesis | Human validates sources and conclusions |
| Coding & software maintenance | 80–95% | AI-led implementation with tests | Human owns architecture, security and release |
| Data analysis & reporting | 85–95% | AI-led analysis on governed data | Human verifies definitions and decisions |
| Scheduling & administrative processing | 90–98% | Highly automatable | Human handles exceptions and appeals |
| Clinical documentation & chart review | 85–95% | AI drafts, extracts and flags | Licensed clinician verifies the record |
| Prior authorization & utilization review | 80–95% | AI assembles and compares evidence | Human controls denials and appeals |
| Differential-diagnosis support | 75–90% | AI proposes and ranks possibilities | Clinician examines, decides and signs |
| Complex independent diagnosis | 60–80% | AI support; autonomy remains unreliable | Licensed human decision |
| Treatment-plan generation | 70–90% | AI prepares options and checks | Clinician individualizes and authorizes |
| Bedside assessment & physical examination | 20–45% | Limited without sensors and robotics | Qualified human leads |
| Procedures & physical care | 10–40% | Specialized robotics only | Qualified human leads |
Adoption scenarios
How fast capability becomes normal work practice
Technical capability can arrive years before organizations permit it. These scenarios separate what AI can do from what workplaces will allow.
Conservative
Institutions move slowly. Humans remain workflow owners while AI drafts, extracts, monitors and recommends.
- 1 year: broad supervised copilots
- 3 years: selective workflow automation
- 5 years: human-led structures remain common
Expected
AI agents perform most bounded digital work. Humans handle exceptions, relationships and consequential authorization.
- 1 year: AI-first digital production
- 3 years: agents execute multi-step workflows
- 5 years: smaller human teams supervise systems
Accelerated
Capability and adoption compound quickly. AI executes nearly all routine digital tasks and escalates uncertainty.
- 1 year: rapid agent adoption
- 3 years: most routine digital execution automated
- 5 years: humans concentrate on authority, trust and edge cases
Stable identities · benchmark settings separate · cutoff 18 July 2026
Model benchmark comparison — bounded evidence
Every model/category cell records a compatible result, bounded qualitative evidence, or a cited researched-unavailable finding. Only identical named evaluation groups are compared directly.
OpenAI
GPT-5.6 Sol
GPT-5.6 SolStable product identity · released 2026-07-09Anthropic
Claude Fable 5
Claude Fable 5Stable product identity · released 2026-06-09Google DeepMind
Gemini 3.1 Pro Preview
Gemini 3.1 Pro PreviewStable product identity · released 2026-02-19SpaceXAI
Grok 4.5
Grok 4.5Stable product identity · released 2026-07-16DeepSeek
DeepSeek V4 Pro Preview
DeepSeek V4 Pro PreviewStable product identity · released 2026-04-24Meta
Muse Spark 1.1
Muse Spark 1.1Stable product identity · released 2026-07-09Compatible comparison · Artificial Analysis Intelligence Index v4.1
- Gemini 3.1 Pro Preview: Artificial Analysis' stable v4.1 report records a 46 Intelligence Index score and 1.6 minutes per task. Configuration: Gemini 3.1 Pro Preview, Thinking (High)
- DeepSeek V4 Pro Preview: Artificial Analysis' stable v4.1 report records 44 on its Intelligence Index for V4 Pro at max effort. Configuration: DeepSeek V4 Pro, Max Effort (Artificial Analysis harness)
Other evidence remains separate because harnesses, settings, or measures differ. Provenance is retained as vendor-reported or independent.
| Evidence category | GPT-5.6 Sol | Claude Fable 5 | Gemini 3.1 Pro Preview | Grok 4.5 | DeepSeek V4 Pro Preview | Muse Spark 1.1 |
|---|---|---|---|---|---|---|
| Reasoning | Researched — no comparable result The reviewed provider source does not publish a compatible reasoning result for this configuration. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible reasoning result for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18) | Reported result Google reports 44.4% on Humanity's Last Exam without tools and 51.4% with search and code. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Developer-run academic benchmarks can be sensitive to prompt, tools, effort, contamination controls, and scoring.Source: Google DeepMind (accessed 2026-07-18)Reported result Artificial Analysis' stable v4.1 report records a 46 Intelligence Index score and 1.6 minutes per task. External evaluationConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-06-15Limit: The composite is methodology-specific and does not directly measure clinical or workplace reliability.Source: Artificial Analysis (accessed 2026-07-18) | Reported result Artificial Analysis reports 54 on its Intelligence Index at high effort in its pre-release evaluation. External evaluationConfiguration: Grok 4.5, high effort in Grok Build · Evaluated 2026-07-08Benchmark settings: effort highLimit: The July 8 evaluation predates the July 16 public launch and does not establish reliability on unseen organizational workflows.Source: Artificial Analysis (accessed 2026-07-18) | Reported result Artificial Analysis' stable v4.1 report records 44 on its Intelligence Index for V4 Pro at max effort. External evaluationConfiguration: DeepSeek V4 Pro, Max Effort (Artificial Analysis harness) · Evaluated 2026-06-15Benchmark settings: effort max · harness Artificial AnalysisLimit: The result follows a v4.1 methodology change and is not a workplace success or replacement rate.Source: Artificial Analysis (accessed 2026-07-18) | Reported result Artificial Analysis reports 51 on its Intelligence Index for Muse Spark 1.1 at xhigh effort. External evaluationConfiguration: Muse Spark 1.1, xhigh effort / Thinking mode · Evaluated 2026-07-10Benchmark settings: effort xhighLimit: The evaluator had pre-release access and the composite is sensitive to harness, task mix, token budget, and graders.Source: Artificial Analysis (accessed 2026-07-18) |
| Coding | Reported result Artificial Analysis reports 80 on its Coding Agent Index for GPT-5.6 Sol in Codex. External evaluationConfiguration: GPT-5.6 Sol, max reasoning · Evaluated 2026-07-09Limit: The result is specific to max reasoning, the Codex harness, and the evaluator's current task and grading mix.Source: Artificial Analysis (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible coding result for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible coding result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18) | Reported result The July 16 launch report lists 83.3% on Terminal-Bench 2.1 and 64.7% resolution on SWE-Bench Pro. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: The provider page combines its own runs and cited leaderboard figures across differing harnesses and configurations.Source: SpaceXAI (accessed 2026-07-18) | Bounded qualitative evidence DeepSeek reports open-model-leading agentic coding performance for V4 Pro. Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: The accessible launch text gives a relative vendor claim without the full numeric chart methodology.Source: DeepSeek (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible coding result for this configuration. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18) |
| Writing | Bounded qualitative evidence OpenAI reports improved document generation and structured drafting for GPT-5.6 Sol. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: The provider uses selected document examples and customer evaluations rather than a portable workplace-writing score.Source: OpenAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible writing result for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible writing result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible writing result for this configuration. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible writing result for this configuration. Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible writing result for this configuration. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18) |
| Data analysis | Researched — no comparable result The reviewed provider source does not publish a compatible data analysis result for this configuration. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18) | Reported result Artificial Analysis reports a leading 1932 GDPval-AA score with fallback on 2% of tasks. External evaluationConfiguration: Claude Fable 5, adaptive reasoning max, Opus 4.8 fallback · Evaluated 2026-06-09Limit: This is one evaluator's agentic knowledge-work benchmark at max effort and is not evidence of whole-job replacement.Source: Artificial Analysis (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible data analysis result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible data analysis result for this configuration. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible data analysis result for this configuration. Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible data analysis result for this configuration. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18) |
| Healthcare support | Reported result OpenAI reports 60.5% on HealthBench Professional for GPT-5.6 Sol. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: HealthBench Professional is a rubric-scored evaluation and does not establish clinical outcomes or autonomous-care safety.Source: OpenAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible healthcare support result for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible healthcare support result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible healthcare support result for this configuration. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible healthcare support result for this configuration. Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible healthcare support result for this configuration. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18) |
| Tool use | Researched — no comparable result The reviewed provider source does not publish a compatible tool use result for this configuration. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible tool use result for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible tool use result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible tool use result for this configuration. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible tool use result for this configuration. Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18) | Bounded qualitative evidence Meta reports gains in computer use, tool orchestration, coding, and one-million-token context management. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Several results are internal or vendor-run and the API is a public preview whose behavior may change.Source: Meta (accessed 2026-07-18) |
| Agentic execution | Bounded qualitative evidence OpenAI's page contains conflicting Agents' Last Exam values: the narrative states 53.6 while the results table lists 52.7 for GPT-5.6 Sol. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: The source disagreement is unresolved, and neither value is a production success rate across occupations.Source disagreement: The same OpenAI release reports 53.6 in narrative text and 52.7 in its Agents' Last Exam results table; neither value is selected as authoritative here.Source: OpenAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible agentic execution result for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible agentic execution result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible agentic execution result for this configuration. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible agentic execution result for this configuration. Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: DeepSeek (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible agentic execution result for this configuration. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18) |
| Reliability | Researched — no comparable result The reviewed provider source does not publish a compatible reliability result for this configuration. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: OpenAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible reliability result for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Anthropic (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible reliability result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Google DeepMind (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible reliability result for this configuration. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: SpaceXAI (accessed 2026-07-18) | Reported result Artificial Analysis reports -10 on AA-Omniscience and a 94% hallucination rate when the model did not know an answer. External evaluationConfiguration: DeepSeek V4 Pro, Max Effort (Artificial Analysis harness) · Evaluated 2026-04-24Benchmark settings: effort max · harness Artificial AnalysisLimit: These are evaluator-specific knowledge and non-hallucination measures, not a universal factual-error rate.Source: Artificial Analysis (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible reliability result for this configuration. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: Absence in the reviewed provider source is an evidence gap, not a zero score or evidence of incapability.Source: Meta (accessed 2026-07-18) |
| Oversight | Researched — no comparable result The reviewed provider source does not publish a compatible oversight result for this configuration. Provider evidenceConfiguration: GPT-5.6 Sol (vendor mode and effort not stated for these claims) · Evaluated 2026-07-09Limit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: OpenAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a comparable human-review, intervention, or escalation requirement for this configuration. Provider evidenceConfiguration: Claude Fable 5 with adaptive reasoning and Opus 4.8 fallback · Evaluated 2026-06-09Benchmark settings: mode adaptive reasoning · routing Opus 4.8 fallbackLimit: Fallback routing is not human oversight; the release reports model-to-model safeguard routing, so no oversight requirement is inferred.Source: Anthropic (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible oversight result for this configuration. Provider evidenceConfiguration: Gemini 3.1 Pro Preview, Thinking (High) · Evaluated 2026-02-19Benchmark settings: mode Thinking · effort HighLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: Google DeepMind (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible oversight result for this configuration. Provider evidenceConfiguration: Grok 4.5 in Grok Build (vendor effort not stated) · Evaluated 2026-07-16Benchmark settings: harness Grok BuildLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: SpaceXAI (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible oversight result for this configuration. Provider evidenceConfiguration: DeepSeek V4 Pro Preview, Thinking mode · Evaluated 2026-04-24Benchmark settings: mode ThinkingLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: DeepSeek (accessed 2026-07-18) | Researched — no comparable result The reviewed provider source does not publish a compatible oversight result for this configuration. Provider evidenceConfiguration: Muse Spark 1.1, Thinking mode · Evaluated 2026-07-09Benchmark settings: mode ThinkingLimit: No comparable human-review, intervention, or escalation requirement is published; model routing is not treated as human oversight.Source: Meta (accessed 2026-07-18) |
Guided research
Turn benchmark evidence into work decisions
Use this sequence to distinguish a bounded model result from responsible deployment, workforce planning, or personal career action.
1. What a benchmark means
A benchmark measures performance on its own defined tasks, harness, tool access, effort setting, and grading method. It can show useful capability, but it is not a production reliability, cost, safety, or job-replacement measure.
Sources: Artificial Analysis: GPT-5.6 benchmarks across Intelligence, Speed and Cost; METR: Task-Completion Time Horizons of Frontier AI Models
2. A task is not a job
Jobs combine many tasks with context, handoffs, relationship work, physical presence, and formal authority. This map contains two equally weighted illustrative cases per occupation—not a measured task mix—and does not publish occupation ratings or predict that an occupation disappears.
Sources: National Center for O*NET Development, U.S. Department of Labor: Registered Nurses, O*NET OnLine occupation summary; National Center for O*NET Development, U.S. Department of Labor: Software Developers, O*NET OnLine occupation summary
3. Adoption has constraints
Demonstrated capability is only one condition. Adoption also depends on workflow integration, data quality, procurement, cost, latency, privacy, security, training, regulation, labor agreements, acceptance, and a workable route for exceptions and accountability.
4. Healthcare and technology cases
Healthcare evidence in this map includes simulations and bounded workflows, not patient outcomes. It does not support autonomous diagnosis, prescribing, crisis care, or clinical sign-off. In technology, code and tool benchmarks still leave requirements, security, integration, review, and ownership with people.
Sources: OpenAI: GPT-5.6: Frontier intelligence that scales with your ambition; Artificial Analysis: GPT-5.6 benchmarks across Intelligence, Speed and Cost; METR: Task-Completion Time Horizons of Frontier AI Models
5. Workforce action
Start with a bounded task, define a quality baseline, name who reviews output, log exceptions, and measure whether the workflow improves. Expand only when evidence supports the next use case; preserve accountable human authority for consequential work.
6. Career strategy
Build fluency in supervised AI use alongside domain judgment, communication, verification, exception handling, and collaboration. Treat the one-, three-, and five-year views as planning scenarios—not guarantees about hiring, pay, or displacement.
Research recommendations
Supported conclusions: current evidence supports bounded task assistance and task-specific human-control decisions; it does not support occupation replacement rates.
Disagreements: the GPT-5.6 provider page reports different Agents' Last Exam values in narrative text and its result table, so this map records the conflict without choosing a number.
Evidence gaps: most model/category pairs lack compatible measures, occupation task importance weights are not compiled, and many task scores transfer from related evaluations rather than direct workplace studies.
Uncertainty and refresh: model configurations, benchmark methods, deployment evidence, regulation, and adoption can change quickly. Recheck model evidence by 18 January 2027 and before any consequential decision.
Task-first decision explorer
Plan from a task and its evidence
Filter occupation contexts and task types, then select a task for scores, sources, exact model uses, oversight, scenarios, and separate career, workforce, and research recommendations.
Occupation, task-type, classification, horizon, and selected-task state are stored in the URL for sharing and Back/Forward restoration.
32 matching occupation contexts
Occupation context · healthcare
Registered Nurse
Paired illustrative cases — not a measured task mix. Equal case weights are not importance, frequency, or time-use weights and must not be read as an occupation task mix.
Choose a task
Skills to strengthen
- Task-grounded practice: accountable judgment and exception handling for “Educate a patient about a care plan”.
- Task-grounded practice: accountable judgment and exception handling for “Administer medication at the bedside”.
Potentially routinized activities
No routinization claim is supported by these two cases.
AI tool guidance
- Evaluate bounded assistance for “Educate a patient about a care plan”; this is not a deployment approval. The model record supplies category-level evidence only; task-specific quality, workflow fit, and safety still require local evaluation.
Transitions — not assessed
Transitions are not assessed because transferable skills, education, credentials, licensure, training time, work setting, and labor evidence were not compiled for Registered Nurse.
These prompts derive from two O*NET-linked illustrative tasks; they are not measured rising or declining skill trends. Assessment date: 2026-07-18. Sources: onet-registered-nurse.
Task evidence · Documentation & communication
Task evidence: Educate a patient about a care plan
Scores are ordinal task assessments, not automation percentages or employment forecasts. Each record retains its rationale, confidence, assessment date, and source trail.
AI capability
75/100 · Confidence: mediumDirect task evidence: HealthBench: Evaluating Large Language Models Towards Improved Human Health evaluates the same workflow family as “Educate a patient about a care plan.” A score of 75 (60–79 rubric band) reflects that strong drafting and plain-language transformation when the clinical facts are supplied. The cited result is bounded and is not a claim of autonomous workplace execution.
Assessment date: 2026-07-18 · Sources: openai-healthbench-2025
Automation exposure
60/100 · Confidence: lowAnalyst automation inference: the reviewed direct capability evidence and O*NET's Registered Nurse task context place “Educate a patient about a care plan” at 60 (60–79 rubric band) because a substantial drafting step can be assisted, while tailoring and delivery remain clinician work. This is not an adoption, job-loss, or replacement percentage.
Assessment date: 2026-07-18 · Sources: openai-healthbench-2025, onet-registered-nurse
Human judgment
82/100 · Confidence: mediumOccupation-specific inference: O*NET's Registered Nurse task context places human judgment for “Educate a patient about a care plan” at 82 (80–89 rubric band) because comprehension, culture, preferences, and changing symptoms require clinician judgment. The score is an analyst rubric placement, not a measured human-performance rate.
Assessment date: 2026-07-18 · Sources: onet-registered-nurse
Accountability & safety
94/100 · Confidence: mediumOccupation-specific safety inference: O*NET's Registered Nurse duties place accountability and consequence severity for “Educate a patient about a care plan” at 94 (90–100 rubric band) because incorrect or misunderstood instructions can cause direct patient harm. This measures task stakes, not model safety performance.
Assessment date: 2026-07-18 · Sources: onet-registered-nurse, openai-healthbench-2025
Collaboration advantage
82/100 · Confidence: mediumAnalyst workflow inference: combining the reviewed direct capability evidence with O*NET's Registered Nurse task context places AI–human collaboration on “Educate a patient about a care plan” at 82 (80–89 rubric band) because AI drafting plus nurse verification can improve consistency without transferring accountability. No cited study measures this exact deployment design.
Assessment date: 2026-07-18 · Sources: openai-healthbench-2025, onet-registered-nurse
Career durability
90/100 · Confidence: lowAnalyst durability inference: O*NET's Registered Nurse task requirements place the persistent human contribution to “Educate a patient about a care plan” at 90 (90–100 rubric band) because trust, bedside assessment, and licensed responsibility preserve a large human contribution. This forward-looking ordinal value is a scenario input, not an employment forecast.
Assessment date: 2026-07-18 · Sources: onet-registered-nurse
Exact model configurations
- GPT-5.6 Sol — Evaluate bounded assistance for “Educate a patient about a care plan”; this is not a deployment approval.
GPT-5.6 Sol (vendor mode and effort not stated for these claims)
Evidence fit: transfer · Oversight: human-review. The model record supplies category-level evidence only; task-specific quality, workflow fit, and safety still require local evaluation.
1-year scenario
Supervised AI assistance for educate a patient about a care plan may expand where inputs and review responsibilities are clear.
Career recommendation
- Strengthen accountable judgment, context, and communication within “Educate a patient about a care plan” while testing only bounded assistance.
This recommendation is scoped to one illustrative task and is not an occupation-wide labor-market forecast.
Workforce recommendation
augment · human review
Use AI for bounded assistance while a qualified person reviews, contextualizes, and owns the result.
Research recommendation
Supported: ai human is the supported task-level classification under the cited evidence and stated rubric.
Disagreements: None recorded for this task; source-level conflicts remain visible in the model evidence section.
Gaps: No production deployment or outcome study is cited for this exact workplace workflow.
Uncertainty: Scores are analyst ordinal assessments; adoption, costs, regulation, integration, and real-world outcomes remain uncertain.
Refresh by: 2027-01-18
Task sources
- OpenAI: HealthBench: Evaluating Large Language Models Towards Improved Human Healthpublished 2025-05-12 · evaluated 2025-05-12 · accessed 2026-07-18
- National Center for O*NET Development, U.S. Department of Labor: Registered Nurses, O*NET OnLine occupation summaryaccessed 2026-07-13
- OpenAI: GPT-5.6: Frontier intelligence that scales with your ambitionpublished 2026-07-09 · evaluated 2026-07-09 · accessed 2026-07-18