Opens in a new tab
needsahuman

How we score jobs: methodology

We score every job by asking three plain questions about its tasks, then combine the answers into one number: Still needs a human, from 0 to 100. Higher is safer. Status of this method. Score version , last changed 2 October 2026. Every change is listed under Method changes below. What the method cannot do…

We score every job by asking three plain questions about its tasks, then combine the answers into one number: Still needs a human, from 0 to 100. Higher is safer.

Status of this method. Score version 1.1.0, last changed 2 October 2026. Every change is listed under Method changes below. What the method cannot do yet is under What this cannot tell you, and how our scores compare with official data is under Checking ourselves.

Why we score tasks, not jobs

A job is a bundle of tasks. An accountant reconciles accounts, explains results to clients, checks for fraud and signs off returns. AI may be good at some of those and poor at others.

So we start from the task list for each job in O*NET, the US Department of Labor’s occupation database. O*NET lists the tasks for each occupation and rates how important and how frequent each one is. We score each task, then weight it by importance × frequency, so the work that fills most of the day counts most.

This is also why our answer is rarely “yes” or “no”. Most jobs are a mix of tasks AI does now, tasks AI helps with, and tasks that still need a person.

The three questions

1. Can AI do it?

Capability Coverage, 0–100: the share of the job’s weighted task time that AI can handle today. It blends three kinds of evidence for each task:

  • Observed (40%): how much people already use AI assistants for the task at work, from the Anthropic Economic Index and Microsoft Research’s Working with AI data.
  • Theoretical (35%): what current AI could do, rated task by task by an AI model (Claude Sonnet 5.5) against a published four-level rubric.
  • Measured (25%): how AI performs on benchmarks built from real work. No benchmark is matched to tasks yet, so for now the other two share its weight.

Hands-on tasks are capped by what robots can do today. The cap can only hold a task’s AI score down, never push it up. Read the coverage method.

2. Is it better than a person?

Quality Parity, 0–100, where 50 means “as good as a typical qualified professional”. Every quality score carries an evidence grade from A to D. We have not yet entered any direct test of AI against people, so every job is currently grade D: the page shows “Not yet measured”, and the calculation assumes parity. Read the quality method.

3. When could it be replaced?

Estimated Replacement Year: a median year with an 80% range, from 10,000 simulated futures. “Replaced” has a strict meaning: AI doing at least 90% of the job’s task time, at least as well as a person, with at least half of employers using it. Ranges that run past 2060 read “Not foreseeable before 2060”. Read the timeline method.

One number, one word

Still needs a human is the headline score. It combines how much of the job AI can do with how well AI does it. It is always shown with the three answers beside it, never alone. The figures on each job page fill with amber to the same number. Read how the headline score works.

Each score band has a verdict word, so the answer to “Will AI replace this job?” fits in one word. The word is read from the score as shown, rounded to a whole number.

Still needs a humanWill AI replace it?What it means
80–100Nah.AI handles little of this job, or does not do it well enough to replace a person.
60–79A little.AI handles a real share of the work, but the job still rests on people.
40–59Partly.A large part of the work is within AI’s reach. Expect the job to change.
20–39Mostly.Most of the work is within AI’s reach, at or near human quality.
0–19Largely.Nearly all of the work is within AI’s reach, at or above human quality.

The bands are fixed. We do not grade on a curve: if the evidence puts no job in a band, that band stays empty.

Where jobs fall today. In the current release (2026-Q4, 923 jobs), scores run from 51.4 to 88.6, with a median of 75.4. 325 jobs mostly need a person (our top band, Nah.); 562 are jobs where AI could do a little of the work (A little.); 36 are jobs AI could partly do (Partly.); 0 are jobs AI could mostly do (Mostly.); and 0 are jobs AI could largely do (Largely.). In the October 2026 release no job scored below about 50. AI is still mostly used for parts of jobs, and we hold no study yet that shows AI doing a job’s work at a person’s quality. A job would need both before it could reach the bottom two bands (Mostly. and Largely.). Why no job scores below about 50.

Evidence grades

GradeWhat the evidence is
AA blind, expert-graded, head-to-head test of AI against people on this job’s own work.
BA head-to-head test on closely related tasks or the job’s family, with a human baseline.
CBenchmarks without a human baseline, or productivity trials. Shown as low confidence.
DNothing usable yet. Shown as “Not yet measured”.

What else is on a job page

Alongside the scores, each job page breaks the answer down:

  • What AI can and cannot do: the job’s tasks split into “AI does this now”, “AI helps” and “Still needs a human”.
  • What’s stopping it: the main blockers for this job, such as licensing, regulation, liability, client preference for a person, physical work and gaps in the evidence. Each has a strength and one line on why.
  • What it costs: the cost of AI doing the AI-capable share of the work, set against the human wage for the same hours, with the assumptions shown and always as a range.
  • Robots and humanoids: how much of the job is physical, what kind of robot would be needed (fixed automation, mobile robots or dexterous humanoids) and where that technology stands.
  • Which AI skills matter: the job’s tasks mapped to AI capability areas (writing, analysis, coding, vision, speech, planning and agents, physical manipulation, care and persuasion), with today’s level for each.
  • Job market and pay: US employment, the Bureau of Labor Statistics 2025–35 projection, US pay and UK pay, each with its source and year.

Sources

SourceWhat we use it forLicence
O*NET 31.0, US Department of LaborOccupations, tasks, importance and frequency, work context, alternate titlesCC BY 4.0
Anthropic Economic IndexObserved AI use by task, automation vs augmentation, physical task flags and robot tiersCC BY
Microsoft Research, Working with AIObserved AI applicability by work activityCC BY 4.0
Our own task ratings (rubric r1, Claude Sonnet 5.5, October 2026)Theoretical capability by taskCC BY 4.0, published with the dataset
Benchmarks such as GDPval and the Remote Labor Index, and domain studiesMeasured capability and quality evidence, once entered. None is used in the current scores.Cited; results quoted
METR task time horizonsPace of capability growthCited
US Census Bureau Business Trends and Outlook SurveyBusiness AI adoptionPublic domain
BLS Occupational Employment and Wage Statistics; Employment ProjectionsUS jobs, pay and outlookPublic domain
ONS Annual Survey of Hours and EarningsUK payOpen Government Licence

Full credits and versions for each release are on the open data page.

What this cannot tell you

  • Your job is not the average job. We score occupations as the US Department of Labor defines them. Your employer, your sector and your own tasks may differ.
  • Usage data shows AI assistants, not all automation. The observed data comes from people using AI chat assistants, and it carries 40% of each task’s score. Automation that runs through other systems, such as document scanning, machine translation, imaging software or self-checkout, is not in it, so some jobs may look safer than they are.
  • The score measures software AI. Robots only set a ceiling. Robot data caps how much AI is credited with on hands-on tasks, but never adds to it. So factory machinery and other automation of physical work never lower a job’s score.
  • Coding is under-counted in usage data, because some coding tools are outside the observed sources.
  • The task ratings come from one AI model. Claude Sonnet 5.5 rated every task under our published rubric. An earlier, separate rating of 573 tasks by the same model family agreed exactly on 81% of them and within one level on 99.7%. That shows the rubric is applied consistently, not that the ratings are right. No person has rated a sample yet.
  • No job has a direct quality test yet. Every job is grade D and takes a neutral “parity” value. That lowers every score by the same amount, and it is the main reason no job scores below about 50.
  • No benchmark results are in the scores yet, and the timeline does not yet respond to each job’s blockers (see the timeline method).
  • The future is uncertain. Our timelines are model outputs, not promises. That is why we always show the range, and why the range is often wide.
  • We score capability and quality, not demand. A job can be highly automatable and still grow, if demand for the work grows faster.

Checking ourselves

We compare our scores with independent data and publish the result here, whatever it shows. The first check uses the October 2026 release and the Bureau of Labor Statistics (BLS) 2025–35 projections, which also place each occupation in an AI exposure group from Low to Very high.

  • BLS AI exposure groups: strong agreement. The more exposed BLS says a job is, the lower our score (rank correlation −0.90 across 892 jobs). The median score is 85.1 for BLS Low, 79.7 for Moderate, 72.8 for High and 65.5 for Very high. Every job AI could partly do (Partly.) is in BLS’s Very high group, and no Very high job is in our top band (Nah.). Part of this agreement comes from shared inputs: BLS built its groups partly from the same usage studies we use.
  • BLS projected job growth: weak, and the wrong way overall. Across the 922 jobs with a projection, a safer score goes very slightly with slower growth (rank correlation −0.13). BLS’s own exposure groups show the same pattern against its own projections. AI-exposed jobs are mostly skilled desk jobs, and those are growing. The steepest projected declines are in production work, such as foundry and sewing jobs, lost to machinery and trade rather than AI, which our score does not measure. Among desk jobs (under 10% physical) the link runs the expected way (+0.19), and among mixed jobs (10% to 50% physical) more so (+0.27).

We have not yet tested the scores against job postings trends from Indeed Hiring Lab or early-career employment research from the Stanford Digital Economy Lab, and we have not yet tested the replacement years against past outcomes. When we do, we will set the pass marks before we look and publish the result here.

Method changes

Every change to the method raises the score version. Changes to the data alone do not; those are in the data changelog.

VersionDateChange
1.1.02 October 2026Adoption moved out of the headline score and into the timeline, and the score now spans the full 0–100 range. Jobs with no quality evidence take a neutral parity value. Every task is rated against a published rubric. Hands-on tasks are capped by robot exposure tier. The timeline adds a slow-progress scenario and starts adoption from how much AI the job already uses.
1.0.02 October 2026First version of the method.

Independence and corrections

No commercial relationship ever changes a score. If you think a score is wrong, tell us and include a source if you can. Fixes are logged on the corrections page.