Method version 1. This page states what the number is, how a run qualifies, which figures are estimates, and where the method stops.
The AI Parity Index is the share of US labor-hours in tasks where AI has been measured at or above the human bar. The Index moves only when a verified run lands. Past readings are never revised.
| Current reading | 0.1% · release 2026-08 |
| Numerator | annual labor-hours in tasks with a verified at-parity run, weighted by task importance within each occupation |
| Denominator | 303.2 billion annual US labor-hours at release 2026-08. The denominator tracks the federal data: it updates when BLS employment and the O*NET task catalog update, and each release records the denominator it used. Changes are mechanical from the federal sources, never discretionary. |
| Coverage | 1 of 18,796 O*NET tasks measured |
| Cadence | quarterly, aligned with the federal data releases; next release Q4 2026 |
| Direction of error | the Index is a lower bound: unmeasured work counts as zero, so the reading can only be too low, never too high |
The unit is the O*NET task: one of 18,796 task statements that federal data uses to describe US jobs. Each task joins to employment, wages, and output through its occupation. This is the same decomposition the economics literature uses.
| Tasks | O*NET 30.1, US Department of Labor (CC-BY). Importance weights set each task’s share of its job. |
| Jobs and wages | BLS Occupational Employment and Wage Statistics, May 2025. The occupation-level survey is annual; May 2026 data publishes in spring 2027. |
| Monthly series | BLS Current Employment Statistics and Current Population Survey, fetched from the BLS public API and refreshed twice a day on the ticker. |
| Output | BEA GDP by Industry, 2025. Occupation output = wages × value added per dollar of pay in its industries, scaled so occupations sum to GDP. Housing rent is excluded. This figure is an estimate and is labeled as one. |
A run moves the Index only when every condition below holds. A run that misses any of them is published as context and counts zero.
| Task set | public materials only. Anyone can re-run the test. Private data never enters the Index. |
| Human baseline | measured from real work records or published studies. Inferred figures, such as time and cost estimates, are labeled est. and shown gray. |
| Pass criterion | stated before the run and executed by machine where possible; never judged after the fact. |
| Repetitions | three or more independent runs per instance. Same input, same answer is reported alongside success. |
| Quality bar | a run below the human quality bar contributes nothing. Fast, bad work is not output. |
| Receipts | the task as written, every patch or output, checksums, and per-run results are published with the score. |
Four kinds of evidence exist about AI and work. The Index cites the first three and is built from the fourth.
| Predicted | experts rate which work AI could do. OECD, MIT AI Labor Index, academic exposure scores. Context only. |
| Observed | usage logs show what people already ask AI to do. Anthropic Economic Index (CC-BY), OpenAI. Context, and validation for the task rubric. |
| Benchmarked | model scores on realistic tasks without a human baseline. APEX, public leaderboards. Cited where they map to our tasks. |
| Proven | AI did the task and was scored against a measured human baseline, with receipts. Only this tier moves the Index. |
Two colors carry the distinction everywhere on the site. Green is measured: the Index reading and verified run results. Gray is estimated, and every estimate is labeled at the point of use.
| AI Parity column | the share of a job’s work testable against a clear pass-fail bar. Rubric v1.0, model-applied, not yet human-reviewed. Verified runs replace it with measurement. |
| GDP output | derived from BLS wages and BEA value added as stated in section 02. |
| Human time and cost | where a baseline team’s time and cost are not in the record, they are inferred and labeled est. |
Tasks are tested in order of the output sitting in work that already has a clear pass-fail bar. The ordering uses the rubric estimate; it never touches published numbers.
| Task authorship | An O*NET statement is a job duty, not a test. Turning one into a gradeable instance takes judgment, and two reasonable authors can produce different instances. Every instance is published and versioned so that judgment is inspectable. Disputed instances are re-authored without breaking the series. |
| Jobs are not task sums | An occupation at full task parity is not an automatable job. Coordination, context, and knowing what to do next live between the tasks, and much production is multiplicative, so the weakest step dominates. As runs land, three statistics are reported: task parity, the bottleneck (the lowest-parity task, named), and job parity, measured by running the whole chain as one scenario. The difference between task parity and job parity is the coordination gap. |
| The method grows by version | Method v1 verifies machine-checkable work, about 15% of US labor-hours. That bound belongs to the method version, not to the Index: expert grading, scenario testing, and operator field data each extend the verifiable share in later versions, and physical work already publishes field baselines (autonomous driving crash rates against human benchmarks) that a later version can admit under the same receipts rule. The headline stays denominated on all US work-hours at every version. |
| Below the bar is zero | a run that misses the human quality bar contributes nothing. |
| Floor, not forecast | anything untested counts as unchanged. |
| Versioned | method changes carry a version number; old numbers stay reproducible. |
| Receipts | every number links to the test, the baseline, and the runs that produced it. |
| Task decomposition | Autor, Levy & Murnane (2003); Brynjolfsson, Mitchell & Rock (2018); Eloundou et al. (2023). |
| Coordination and process | Malone, the MIT Process Handbook. Coordination cost and hand-off structure inform the multi-agent benches. |
| Agent team orchestration | Massa, SCORE-AI: a human-compatible framework for AI agent team orchestration. SPARK AI Consortium Working Papers, Vol. 1 Nos. 7–8 (2026). Coordination patterns and role structure inform the multi-agent benches. |
| Bottleneck statistic | Goldratt’s theory of constraints (1984); Kremer’s O-ring model (1993). The coordination gap operationalizes Autor’s Polanyi’s-paradox argument (2015). Task creation follows Acemoglu & Restrepo. |
| Productivity interpretation | Short, Micro Gains, Macro Patience: interpreting early evidence on agentic AI productivity. SPARK working draft (2026). Frames how micro task gains relate to the aggregate statistics on the ticker. |
| Usage validation | Anthropic Economic Index (CC-BY). |
| Grading conventions | APEX (Mercor, 2025) for expert-graded work. Full-time-equivalent convention follows the McKinsey Global Institute. |
| Data | O*NET 30.1 (CC-BY) · BLS OEWS, CES, CPS · BEA GDP by Industry 2025 · full method, versioned, at github.com/recursivlabs/spark-lab |
Working prototype by Recursiv Labs · not an official SPARK AI or UC San Diego publication · Data · © 2026