
METR’s AI time horizon metric says frontier models now complete, with 50% reliability, tasks that take human professionals about 110 minutes.
And that this capability has doubled roughly every seven months since 2019.
If that trend holds, models will handle month-long tasks with 50% reliability sometime between late 2028 and early 2031. That is the single most quoted number in AI forecasting right now.
And I think most people quoting it have not read the caveats attached to it.
This post walks through what the metric actually measures, where the forecast is solid, where it is not.
And what you should measure yourself before you let an agent run unsupervised for more than an hour.
What the 50% time horizon actually measures
METR defines the 50%-task-completion time horizon as the amount of time human professionals typically need to complete tasks that an AI model completes with a 50% success rate (arXiv:2503.14499). Two words in that definition do most of the work, and both get skipped in the hot takes.
The first is “autonomously.” METR’s measurement concerns tasks completed without routine human assistance, not tasks where a developer nudges the model every ten minutes.
That is a harder and more honest bar than most demos clear.
The second is the task pool itself. METR’s published work focuses on software and research tasks, prototyped on datasets totaling 170 tasks. On those tasks, frontier models including OpenAI’s o3 hit a 50% time horizon of around 110 minutes. That is a real, useful number. It is also a number about software work specifically. And METR itself reports that time horizons can differ by a large factor depending on the task domain and the reference human population. Software results are not a universal capability measure.
And treating them as one is the first place these forecasts go wrong.
The trend is steeper than the error bars
Here is why this metric caught fire.
METR found the 50% time horizon grew exponentially from 2019 through 2025, with a doubling time of approximately seven months. A separate presentation of the estimate puts the doubling time at 212 days, with a 95% confidence interval of 171 to 249 days. And METR notes the trend may have accelerated since 2024. METR attributes the growth primarily to better reliability, adaptation to mistakes, logical reasoning.
And tool use, which matches what anyone shipping agents has watched happen in their own logs.
METR’s boldest claim is about the measurement error itself.
Their argument: even if the absolute measurements are off by a factor of 10, the trend still predicts agents that independently complete a large fraction of software tasks currently taking humans days or weeks, in under a decade (arXiv:2503.14499, Section 5). That is a clever defense, and mostly fair. An exponential this steep absorbs a lot of sloppiness. But notice what it does not defend: the assumption that the exponential continues.
Nothing in a confidence interval covers that.
Where the forecast gets thin
Naive extrapolation of the trend puts a one-month horizon, defined as 167 work hours, somewhere between mid-2028 and mid-2031. The UK AI Security Institute’s version of the same math lands on 2030. Redwood Research’s August 2025 forecast assumed doubling times of around 170 days on METR’s task suite over the following two years, which implies a two-week 50%-reliability horizon around the start of 2028 (Redwood Research).
Those dates all assume the trend survives contact with reality.
One follow-up analysis tries to model what happens if compute growth slows. And its projections are explicitly conditional on no software-only singularity (arXiv:2511.19492). Read that twice. The careful version of this forecast contains a clause that says “assuming nothing weird happens.” Every exponential AI curve of the last decade has eventually met the thing that was supposed to be impossible.
My honest read: the direction is real, the historical doubling is well supported. And the mid-2028-to-early-2031 window is a reasonable central estimate. It is not a schedule. I have watched my own automations go from “handles a 10-minute task” to “handles a 45-minute task” in about a year of swapping models, which rhymes with the curve. But I as well have a graveyard of pipelines that worked at demo scale and fell apart the week I pointed them at real inputs. A trend line measures what was tested. It does not measure your task, your data, or your Tuesday.
The forecasters are doing better than you’d expect
The contrarian point almost nobody makes: recent capability forecasts have actually been landing. Epoch AI’s review of 2025 forecasts found the median was basically correct on RE-Bench and FrontierMath. And fairly close on OSWorld and SWE-Bench Verified (Epoch AI). An AI forecasting update reported a model reaching an 80% time horizon of 3 hours and 6 minutes, already within the range of median expert and superforecaster predictions with more than seven months remaining in 2026 (Forecasting Research Institute).
That matters more than the trend debate. The crowd of people making structured predictions about AI capability has a track record that beats the vibe-based discourse by a wide margin. When a forecaster and a headline disagree, I have started defaulting to the forecaster. And so far that has not burned me.
What you should measure instead
The UK AI Security Institute points at benchmarks that track the failure modes that actually kill long autonomous runs: long serial-reasoning suites like RE-Bench and HCAST. And hallucination measures like HalluEval, HalluLens, and HHEM (AISI). None of those tell you whether your agent is ready. They tell you whether the model class you’re renting is improving on the failure modes that matter.
For your own stack, the operator move is boring and it works. Pick one real task from your business. Time how long it takes you. Run it through your agent twenty times and count successes. That number, your success rate at your task length, is the only time horizon that affects your bill. METR’s seven-month doubling tells you when to re-run that test. It does not replace the test.
If your measured success rate is under 90% on a task that takes you an hour, keep a human in the loop and spend your effort on the reliability drivers METR actually credits: recovery from mistakes and tool use. Those improve with scaffolding, retries, and checkpoints you control, not with a model swap.
The takeaway
AI time horizon research is the best capability metric we have, the historical trend held through 2025. And forecasters have been right more often than the skeptics.
It is too a software-task metric with an extrapolation clause that quietly assumes nothing disruptive happens.
Use the seven-month doubling as a calendar reminder to re-test your own automations, not as permission to let an agent run unsupervised on work you have not personally stress-tested. Run your twenty-task test this week, log the success rate. And re-run it every time a new frontier model ships. That log, not the trend line, is the forecast you can bank on.
