Opens in a new tab
needsahuman

Is it better than a person? How we measure quality parity

AI being able to attempt a task is not the same as AI doing it well. Quality Parity asks how AI's work compares with the work of a typical qualified professional in the same job. It is a score from 0 to 100: 50 means parity: AI's work is as good as a typical professional's.…

AI being able to attempt a task is not the same as AI doing it well. Quality Parity asks how AI’s work compares with the work of a typical qualified professional in the same job.

It is a score from 0 to 100:

  • 50 means parity: AI’s work is as good as a typical professional’s.
  • Above 50 means AI’s work is judged better more often than not.
  • Below 50 means people still do better.

Every quality score carries an evidence grade, because the strength of the evidence matters as much as the result.

Evidence grades

GradeWhat countsWhat the page shows
AA blind test where experts in the job compare AI’s work with professionals’ work on this job’s own tasks, without knowing which is which.The score.
BA head-to-head test on closely related tasks, or on the job’s wider family, with a human baseline.The score.
CBenchmarks with no human baseline, or productivity trials.The score, marked low confidence.
DNothing usable yet.“Not yet measured”, and what evidence would settle it.

A job’s grade is the best grade among its usable evidence.

Today every job is grade D. Direct, expert-graded tests of AI against people exist for only a few dozen occupations, and we have not yet entered any of them into our evidence register. We would rather say “Not yet measured” than guess.

How evidence is weighted

Each study or benchmark result is converted to the 0–100 scale (50 = parity) and weighted three ways:

  • By recency. AI changes fast, so a result loses half its weight every 10.5 months.
  • By independence. Independent tests count most, then academic studies, then tests run by the company that makes the AI. Current weights: independent 1.0, academic 0.85, vendor 0.6.
  • By grade. Grade A counts fully, grade B at 0.8 and grade C at 0.5. Grade D evidence is not used.

The job’s score is the weighted average. For grades A and B, we also estimate a trend: how fast the score has been rising, in points per year. The timeline model uses it.

Evidence we will use

These are the kinds of evidence the register will hold. None is in the current scores yet.

  • GDPval, from OpenAI: experts in each occupation blind-compare AI deliverables with professionals’ deliverables on real work tasks. Used at grade A for the occupations it covers. It is a vendor study, so it is weighted as one.
  • The Remote Labor Index: AI agents attempt real freelance projects, judged against the human-made work. Used at grade B for related job families.
  • Domain studies in law, medicine and software, where AI and professionals were tested on the same cases.

When there is no evidence

A grade D job shows “Not yet measured”. It does not show a number.

The timeline and the headline score still need a value to calculate with. For those calculations only, a grade D job uses:

  1. the average for its job family, once any job in that family has usable evidence, or
  2. until then, a neutral default of 50: parity, meaning no evidence either way.

Either way the value is marked as imputed in the dataset and never shown on the page as the job’s own result.

In the current release every job uses the neutral default. It takes the same 11.4 points off every job’s Still needs a human score, so it moves where the verdict bands fall without changing the order of jobs. It is also the main reason no job scores below about 50. The bottom band (Largely.) needs evidence that AI’s work beats people’s, and the band above it (Mostly.) needs either that or far more coverage than any job has today.

How to read it

  • Grade A or B, score above 50: strong evidence that AI’s work matches or beats professionals’ on this kind of task.
  • Grade A or B, score below 50: strong evidence that people still do better, for now.
  • Grade C: a hint, not a verdict.
  • Grade D: we have not measured it yet.

Quality is judged on the work product. It does not measure trust, accountability, bedside manner or the legal right to sign off, which are often why a person stays in the loop. Those show up in What’s stopping it.

Known limits

  • No evidence is entered yet. Every job’s quality is the neutral default, so quality does not yet separate one job from another.
  • Benchmarks are narrow. A test of 30 tasks says little about a job with 300.
  • Vendor tests may favor the vendor. That is why they carry less weight.
  • Results age quickly. A test from 18 months ago may understate today’s AI.
  • Experts disagree. Even blind grading has noise.

Sources

The studies used for each job are listed, with dates, grades and links, in the “Is AI better than a person?” section of that job’s page and in the evidence file on the open data page.