1.7M jobs · $71/h avg · $46B of GDP a year · 6.2% of the job · verification method: machine-checked
| # | Entryranked by success | Success rateshare of runs at the bar | Consistentruns passed | Cost per tasklist-price est. | Time per taskinference | vs human× output · − pts |
|---|---|---|---|---|---|---|
| Human baseline | 100% | — | ~$225 est. | ~25 min est. | — | |
| 1 | 100% | 27/27 | $0.29 | 19s | ×79 | |
| 2 | 93% | 25/27 | $0.29 | 19s | -7 pts | |
| 3 | 22% | 6/27 | $0.03 | 30s | -78 pts | |
| 4 | 11% | 3/27 | $0.04 | 6s | -89 pts | |
| 5 | 0% | 0/27 | $0.01 | 97s | -100 pts | |
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | invited · not yet run | |||||
| · | Plan → fix → review chain · mixed models | invited · not yet run | ||||
| · | Cross-model pair · Opus + GPT-5.5 | invited · not yet run | ||||
| Verification method | Machine-checked · rubric v1.0, model-applied, not yet human-reviewed |
| Industries | Professional and Technical Services 41% · Information 21% · Finance and Insurance 10% · Manufacturing 8% |
| Source | O*NET task 21670 · Core · importance 3.74 of 5 · jobs and wages BLS OEWS May 2025 |
Working prototype by Recursiv Labs · not an official SPARK AI or UC San Diego publication · Data · © 2026