LIVE
The AI Parity Index0.1%Tasks measured1 of 18,796Unemployment4.1% · Jul 20260.1 pt vs Jun 2026Payrolls−23K m/m · Jul 2026158.9M employedJob openings7.4M · Jun 20260.2M vs May 2026Wages$37.62/h avg · Jul 2026+3.2% y/yProductivity+1.4% q/q ann. · Q2 2026Inflation+3.4% y/y · Jul 2026Next upSoftware Developers · 16 more tasksNext releaseQ4 2026
The AI Parity Index0.1%Tasks measured1 of 18,796Unemployment4.1% · Jul 20260.1 pt vs Jun 2026Payrolls−23K m/m · Jul 2026158.9M employedJob openings7.4M · Jun 20260.2M vs May 2026Wages$37.62/h avg · Jul 2026+3.2% y/yProductivity+1.4% q/q ann. · Q2 2026Inflation+3.4% y/y · Jul 2026Next upSoftware Developers · 16 more tasksNext releaseQ4 2026

Method

Method version 1. This page states what the number is, how a run qualifies, which figures are estimates, and where the method stops.

01Definition

The AI Parity Index is the share of US labor-hours in tasks where AI has been measured at or above the human bar. The Index moves only when a verified run lands. Past readings are never revised.

Current reading0.1% · release 2026-08
Numeratorannual labor-hours in tasks with a verified at-parity run, weighted by task importance within each occupation
Denominator303.2 billion annual US labor-hours at release 2026-08. The denominator tracks the federal data: it updates when BLS employment and the O*NET task catalog update, and each release records the denominator it used. Changes are mechanical from the federal sources, never discretionary.
Coverage1 of 18,796 O*NET tasks measured
Cadencequarterly, aligned with the federal data releases; next release Q4 2026
Direction of errorthe Index is a lower bound: unmeasured work counts as zero, so the reading can only be too low, never too high

02Unit of account

The unit is the O*NET task: one of 18,796 task statements that federal data uses to describe US jobs. Each task joins to employment, wages, and output through its occupation. This is the same decomposition the economics literature uses.

TasksO*NET 30.1, US Department of Labor (CC-BY). Importance weights set each task’s share of its job.
Jobs and wagesBLS Occupational Employment and Wage Statistics, May 2025. The occupation-level survey is annual; May 2026 data publishes in spring 2027.
Monthly seriesBLS Current Employment Statistics and Current Population Survey, fetched from the BLS public API and refreshed twice a day on the ticker.
OutputBEA GDP by Industry, 2025. Occupation output = wages × value added per dollar of pay in its industries, scaled so occupations sum to GDP. Housing rent is excluded. This figure is an estimate and is labeled as one.

03What a verified run requires

A run moves the Index only when every condition below holds. A run that misses any of them is published as context and counts zero.

Task setpublic materials only. Anyone can re-run the test. Private data never enters the Index.
Human baselinemeasured from real work records or published studies. Inferred figures, such as time and cost estimates, are labeled est. and shown gray.
Pass criterionstated before the run and executed by machine where possible; never judged after the fact.
Repetitionsthree or more independent runs per instance. Same input, same answer is reported alongside success.
Quality bara run below the human quality bar contributes nothing. Fast, bad work is not output.
Receiptsthe task as written, every patch or output, checksums, and per-run results are published with the score.

04Evidence tiers

Four kinds of evidence exist about AI and work. The Index cites the first three and is built from the fourth.

Predictedexperts rate which work AI could do. OECD, MIT AI Labor Index, academic exposure scores. Context only.
Observedusage logs show what people already ask AI to do. Anthropic Economic Index (CC-BY), OpenAI. Context, and validation for the task rubric.
Benchmarkedmodel scores on realistic tasks without a human baseline. APEX, public leaderboards. Cited where they map to our tasks.
ProvenAI did the task and was scored against a measured human baseline, with receipts. Only this tier moves the Index.

05Measured and estimated

Two colors carry the distinction everywhere on the site. Green is measured: the Index reading and verified run results. Gray is estimated, and every estimate is labeled at the point of use.

AI Parity columnthe share of a job’s work testable against a clear pass-fail bar. Rubric v1.0, model-applied, not yet human-reviewed. Verified runs replace it with measurement.
GDP outputderived from BLS wages and BEA value added as stated in section 02.
Human time and costwhere a baseline team’s time and cost are not in the record, they are inferred and labeled est.

06Verification queue

Tasks are tested in order of the output sitting in work that already has a clear pass-fail bar. The ordering uses the rubric estimate; it never touches published numbers.

Cashiers$143B
Software Developers$139B
Retail Salespersons$135B
Sales Representatives, Wholesale and Manufacturing, Except Technical and Scientific Products$121B
Bookkeeping, Accounting, and Auditing Clerks$120B
Accountants and Auditors$111B

07Limitations

Task authorshipAn O*NET statement is a job duty, not a test. Turning one into a gradeable instance takes judgment, and two reasonable authors can produce different instances. Every instance is published and versioned so that judgment is inspectable. Disputed instances are re-authored without breaking the series.
Jobs are not task sumsAn occupation at full task parity is not an automatable job. Coordination, context, and knowing what to do next live between the tasks, and much production is multiplicative, so the weakest step dominates. As runs land, three statistics are reported: task parity, the bottleneck (the lowest-parity task, named), and job parity, measured by running the whole chain as one scenario. The difference between task parity and job parity is the coordination gap.
The method grows by versionMethod v1 verifies machine-checkable work, about 15% of US labor-hours. That bound belongs to the method version, not to the Index: expert grading, scenario testing, and operator field data each extend the verifiable share in later versions, and physical work already publishes field baselines (autonomous driving crash rates against human benchmarks) that a later version can admit under the same receipts rule. The headline stays denominated on all US work-hours at every version.

08Rules

Below the bar is zeroa run that misses the human quality bar contributes nothing.
Floor, not forecastanything untested counts as unchanged.
Versionedmethod changes carry a version number; old numbers stay reproducible.
Receiptsevery number links to the test, the baseline, and the runs that produced it.

09Prior work and sources

Task decompositionAutor, Levy & Murnane (2003); Brynjolfsson, Mitchell & Rock (2018); Eloundou et al. (2023).
Coordination and processMalone, the MIT Process Handbook. Coordination cost and hand-off structure inform the multi-agent benches.
Agent team orchestrationMassa, SCORE-AI: a human-compatible framework for AI agent team orchestration. SPARK AI Consortium Working Papers, Vol. 1 Nos. 7–8 (2026). Coordination patterns and role structure inform the multi-agent benches.
Bottleneck statisticGoldratt’s theory of constraints (1984); Kremer’s O-ring model (1993). The coordination gap operationalizes Autor’s Polanyi’s-paradox argument (2015). Task creation follows Acemoglu & Restrepo.
Productivity interpretationShort, Micro Gains, Macro Patience: interpreting early evidence on agentic AI productivity. SPARK working draft (2026). Frames how micro task gains relate to the aggregate statistics on the ticker.
Usage validationAnthropic Economic Index (CC-BY).
Grading conventionsAPEX (Mercor, 2025) for expert-graded work. Full-time-equivalent convention follows the McKinsey Global Institute.
DataO*NET 30.1 (CC-BY) · BLS OEWS, CES, CPS · BEA GDP by Industry 2025 · full method, versioned, at github.com/recursivlabs/spark-lab

Working prototype by Recursiv Labs · not an official SPARK AI or UC San Diego publication · Data · © 2026