1.7M jobs · $71/h avg · $46B of GDP a year · 6.2% of the job · verification method: machine-checked
The coding agents people use every day, on the same ticket. The measured loop is our minimal reference harness: test, patch, retry.
| # | Harnessranked by success | Success rateshare of runs at the bar | Consistentruns passed | Cost per tasklist-price est. | Time per taskinference | vs human× output · − pts |
|---|---|---|---|---|---|---|
| Human baseline | 100% | — | ~$225 est. | ~25 min est. | — | |
| 1 | 93% | 25/27 | $0.29 | 19s | -7 pts | |
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| Verification method | Machine-checked · rubric v1.0, model-applied, not yet human-reviewed |
| Industries | Professional and Technical Services 41% · Information 21% · Finance and Insurance 10% · Manufacturing 8% |
| Source | O*NET task 21670 · Core · importance 3.74 of 5 · jobs and wages BLS OEWS May 2025 |
Working prototype by Recursiv Labs · not an official SPARK AI or UC San Diego publication · Data · © 2026