AI progress varies by task, with gains changing over time
A practical look at measuring AI progress using human task times and where the method may mislead or miss differences.
- Publication
- arXiv
- Stage
- Preprint
- What we read
- Summary of the paper
- Authors
- Drew T. Nguyen, William Fithian
- Universities and research institutions
- Not yet supplied in verified metadata; the Brief does not guess.
Institution metadata: OpenAlex record ↗
What they did and found
Experts reviewed data from 228 software tasks and 26 AI systems. Measuring AI difficulty by how long a human works is not straightforward: AI ability shifts differently over time, with a flat window of 2–30 minutes and faster changes outside it.
Why it matters
Don’t rely on one horizon number. Use simple charts to see if time changes are similar across AI types and avoid drawing conclusions from a single figure.
The study shows a non‑linear link between human work time and AI difficulty. The 2–30 minute flat region means small time changes can reflect different capability levels.
Managers should show horizon data with checks that show how predictable and comparable a jump is across AI models, tasks, and settings. Are these jumps reliably comparable across models?
What remains uncertain
Findings depend on the task set and modeling choices; future tasks or AI designs may yield different results.
Original sources · 1
- On the estimation and validity of AI time horizons---a statistical look at the METR plot ↗arXiv · 2026-10-08
Check the original paper for its authors, methods, version and access terms.
What is your take?
Ask a question, add useful context or share a different perspective. Keep the conversation respectful and grounded.