Introduction
"AI's abilities are doubling every seven months" has become one of the most repeated claims about AI progress. It comes from a real study, and the study says something narrower and more careful than the slogan.
This example reads the study's abstract closely and turns its central claim into one sentence that an accurate newspaper could print.
The Problem
The study is "Measuring AI Ability to Complete Long Software Tasks" (Kwa et al. 2025). An abstract is the one-paragraph summary at the top of a research paper. This one is free to read on the paper's arXiv page, listed in the References.
Here is a typical headline version of its claim:
AI can now do hour-long tasks, and its abilities are doubling every seven months.
The problem is to work out what the abstract supports and to rewrite the claim so that it says that and no more. The quoted phrases below are from the abstract.
Working It Through
Step 1: State exactly what was measured
Three things need pinning down: the tasks, the bar for success, and the comparison.
The abstract names its measure the "50%-task-completion time horizon." It defines this as "the time humans typically take to complete tasks that AI models can complete with 50% success rate."
- The tasks. The title says "software tasks." The abstract lists two existing collections of tasks and "66 novel shorter tasks." Nothing in it concerns writing, teaching, law, or customer service.
- The bar. Success half the time. A model with a one-hour time horizon fails about as often as it succeeds on tasks of that length.
- The comparison. The clock is a human one. The researchers "timed humans with relevant domain expertise" on the same tasks. "An hour-long task" means a task that takes a skilled person an hour. It says nothing about how long the AI system takes.
Step 2: Identify who ran the study and what they'd gain or lose
The authors work at METR, a nonprofit organization that tests AI systems. METR doesn't sell an AI model, so it has no product whose score it would want to raise. That counts in the study's favor.
It isn't a disinterested party in every respect. METR's purpose is to assess whether AI systems are becoming capable enough to be dangerous, and the abstract mentions "the implications of increased autonomy for dangerous capabilities." An organization built around that concern gains attention and support when its findings show capability rising fast.
Neither point settles whether the result is right. Together they indicate how to read it: the measurement comes from a group with no model to sell and with a stake in the question being important.
Step 3: Separate the measurement from the extrapolation
The abstract contains two different kinds of claim.
The first is a measurement of the past: the time horizon of leading models "has been doubling approximately every seven months since 2019." This describes what the researchers found when they tested models released over about six years.
The second is an extrapolation. The abstract's final sentence begins "If these results generalize to real-world software tasks" and continues that "extrapolation of this trend predicts that within 5 years" AI systems will be able to automate many software tasks that now take people a month.
That sentence has two conditions built in. The results must carry over from test tasks to real ones, and the trend must continue. The first sentence reports what happened, and the second is a forecast of what would follow if two things hold. The headline's "are doubling" merges them, presenting a past trend as a standing fact about the future.
Step 4: List the limits the authors themselves give
The abstract states several limits in its own words.
- The trend is approximate. The doubling is "approximately" every seven months, and the abstract adds that "the trend may have accelerated in 2024." The rate isn't fixed.
- The results may not transfer. The authors say they "discuss the limitations of our results" and name "their degree of external validity." External validity is the degree to which a result holds outside the conditions of the study.
- The forecast is conditional. It opens with "If."
- The domain is software. The word appears in the title and again in the forecast.
One further limit follows from the definition in Step 1. A 50 percent success rate is far from dependable.
Step 5: Rewrite the claim in one careful sentence
A careful version needs the items from Steps 1 to 4: who measured, on what tasks, at what bar, compared with whom, over what period, and with the forecast marked as a forecast. Here is one.
Researchers at METR, a nonprofit that tests AI systems, found that on a set of software tasks the length of task leading AI models could complete half the time, measured by how long the task takes a skilled person, doubled about every seven months from 2019 to early 2025; the authors caution that the result may not carry over to other kinds of work, and their projection of the trend is a forecast.
It's long. Most of the added length is the limits.
Key Considerations
The common mistake is to read "50 percent success on tasks that take people an hour" as "can do an hour of anyone's job." Three separate errors are packed into that reading.
- "An hour" is the human's time on a defined, self-contained task. A job is a stream of tasks that depend on one another and on other people.
- "50 percent" means failure half the time. Few employers would accept that from a person.
- "Anyone's job" swaps software tasks for all work. The study didn't test other fields.
A second mistake is to treat the doubling as a law. A trend measured over six years describes those six years. Whether it holds for the next six depends on things the study didn't measure, and the authors don't claim otherwise.
Rejecting the study because the headline overstated it would also be a mistake. The measurement is careful, and the authors are open about its limits. The overstatement was added afterward by others.
Summary
The headline and the careful sentence describe the same study. Set side by side, they differ in five places.
| Headline version | Careful version |
|---|
| 1. Who measured | Not stated | METR, a nonprofit that tests AI systems |
| 2. Which tasks | "Tasks" | A set of software tasks |
| 3. Success bar | "Can do" | Completes half the time |
| 4. Whose time | "Hour-long" | Time a skilled person would take |
| 5. Past or future | "Are doubling" | Doubled from 2019 to early 2025; projection marked as a forecast |
- Rows 1 to 4 restore what was measured. Each comes from the abstract's definition of its measure.
- Row 5 separates the finding from the forecast. The abstract does this itself by starting its projection with "If."
- Check: each statement about the measurement in the careful sentence can be matched to a phrase in the abstract, and the sentence contains no claim about fields other than software, about reliability above 50 percent, or about what will happen next.
References
- Kwa, Thomas, Ben West, Joel Becker, and 23 others. 2025. "Measuring AI Ability to Complete Long Software Tasks." arXiv:2503.14499. NeurIPS 2025.