Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Drew T. Nguyen, William Fithian
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

Institution metadata: OpenAlex record ↗

What they did and found

Experts reviewed data from 228 software tasks and 26 AI systems. Measuring AI difficulty by how long a human works is not straightforward: AI ability shifts differently over time, with a flat window of 2–30 minutes and faster changes outside it.

Why it matters

Don’t rely on one horizon number. Use simple charts to see if time changes are similar across AI types and avoid drawing conclusions from a single figure.

The study shows a non‑linear link between human work time and AI difficulty. The 2–30 minute flat region means small time changes can reflect different capability levels.

Managers should show horizon data with checks that show how predictable and comparable a jump is across AI models, tasks, and settings. Are these jumps reliably comparable across models?

What remains uncertain

Findings depend on the task set and modeling choices; future tasks or AI designs may yield different results.

Read the paper PDF ↗

Original sources · 1
  1. On the estimation and validity of AI time horizons---a statistical look at the METR plot ↗arXiv · 2026-10-08

Check the original paper for its authors, methods, version and access terms.