LIVE
The AI Parity Index0.1%Tasks measured1 of 18,796Unemployment4.1% · Jul 20260.1 pt vs Jun 2026Payrolls−23K m/m · Jul 2026158.9M employedJob openings7.4M · Jun 20260.2M vs May 2026Wages$37.62/h avg · Jul 2026+3.2% y/yProductivity+1.4% q/q ann. · Q2 2026Inflation+3.4% y/y · Jul 2026Next upSoftware Developers · 16 more tasksNext releaseQ4 2026
The AI Parity Index0.1%Tasks measured1 of 18,796Unemployment4.1% · Jul 20260.1 pt vs Jun 2026Payrolls−23K m/m · Jul 2026158.9M employedJob openings7.4M · Jun 20260.2M vs May 2026Wages$37.62/h avg · Jul 2026+3.2% y/yProductivity+1.4% q/q ann. · Q2 2026Inflation+3.4% y/y · Jul 2026Next upSoftware Developers · 16 more tasksNext releaseQ4 2026
Computer and MathematicalSoftware DevelopersTask 21670

Modify existing software to correct errors, adapt it to new hardware, or upgrade interfaces and improve performance.

1.7M jobs · $71/h avg · $46B of GDP a year · 6.2% of the job · verification method: machine-checked

AllModelsHarnessesMulti-agent teams
win zone: at least as reliable, cheaper0%10%20%30%40%50%60%70%80%90%100%$0.01$0.03$0.1$0.3$1$3$10$30$100$300Cost per fix (log scale)Success rateHuman baseline · ~$225 per fixGemini 3.1 Pro · 0% · $0.01Gemini 3.1 ProGPT-5.5 · 22% · $0.03GPT-5.5Claude Opus · 11% · $0.04Claude Opus

Leaderboard

AllModelsHarnessesMulti-agent teams

Each frontier model, handed the ticket cold: one pass, no tools, no retries. This is what the model alone can do.

#Modelranked by successSuccess rateshare of runs at the barConsistentruns passedCost per tasklist-price est.Time per taskinferencevs human× output · − pts
Human baseline100%~$225 est.~25 min est.
1GPT-5.5 22%6/27$0.0330s-78 pts
2Claude Opus 11%3/27$0.046s-89 pts
3Gemini 3.1 Pro 0%0/27$0.0197s-100 pts
·Claude Sonnetinvited · not yet run
·Grok 4invited · not yet run
·DeepSeekinvited · not yet run
Verification methodMachine-checked · rubric v1.0, model-applied, not yet human-reviewed
IndustriesProfessional and Technical Services 41% · Information 21% · Finance and Insurance 10% · Manufacturing 8%
SourceO*NET task 21670 · Core · importance 3.74 of 5 · jobs and wages BLS OEWS May 2025

Working prototype by Recursiv Labs · not an official SPARK AI or UC San Diego publication · Data · © 2026