KnowledgeInSight
AI Literacy
0% of Course 1 complete

Module 3 · Lesson 3

Capabilities and Limits: Measuring progress and AI-assisted AI research

You'll learn how AI progress is measured and why a headline score can mislead. You'll then look at how AI systems are being used in AI research itself, and you'll be able to separate what has been demonstrated from what is forecast.

What you will be able to do

  • Interpret a benchmark claim, and explain how AI systems are used in AI research without treating forecasts as findings.

0% of this lesson · 10 items · 1h 33m total · 1h 18m without the optional activity

Contents of this lesson10 items
  1. ReadingWhat a Headline Benchmark Score Leaves Out3 min
  2. ReadingBenchmarks: What a Test Score Measures and What It Can't4 min
  3. ReadingHow Benchmarks Mislead: Contamination, Saturation, and Test-Taking Behavior4 min
  4. ReadingTask Length and Productivity Studies as Measures of Real-World Capability4 min
  5. ReadingAI-Assisted AI Research: Bounded Self-Refinement and Open-Ended Self-Improvement4 min
  6. Guided ReadingWorked Example: Interpreting the Claim That AI Task Length Doubles Every Seven Months7 min
  7. Guided ConversationQuestion a Capability Headline12 min
  8. Hands-on Activity · optionalAudit a Capability Claim Against Its Source15 min
  9. Knowledge CheckMeasuring progress and AI-assisted AI research10 min
  10. Graded QuizCapabilities, Limits, and Trajectory30 min

Reading 3 min

What a Headline Benchmark Score Leaves Out

This content reflects the field as of October 2026.

The headline has a standard form. A new AI model "beats humans" on an exam, "tops the leaderboard," or "scores 90 percent" on a test with an impressive name. A company is named, and often a prediction follows about what the model will soon be able to do.

The score in such a headline is usually real. Scores have also been rising fast. Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) publishes a yearly survey of the field, and its 2026 edition reports that tests "intended to be challenging for years are saturated in months" (Stanford HAI 2026, chap. 2). Something is plainly improving.

The headline still skips three questions.

The first is which test. A score describes how a model did on one fixed set of tasks. "Beats humans" on a set of multiple-choice science questions and "beats humans" at running a laboratory are different claims, and only the first has been measured.

The second is who ran the test. Many scores are reported by the company that built the model. The Stanford survey notes that independent evaluations have documented cases in which models did worse than their developers had reported (Stanford HAI 2026, chap. 2).

The third is what the score predicts about real work. The survey's answer is cautious: "strong benchmark performance does not always translate to real-world utility" (Stanford HAI 2026, chap. 2).

Rankings raise a further question, which is how large the lead is. As of March 2026, the top four models on one widely followed public leaderboard came from Anthropic, xAI, Google, and OpenAI. They were separated by fewer than 25 points, on a scale where the leader's score was about 1,500 (Stanford HAI 2026, chap. 2). "Tops the leaderboard" can describe a margin that small.

A short set of questions covers most of what a headline leaves out. You can put them to any claim about what AI can do.

  • What tasks was the system tested on?
  • What counted as success?
  • Who did the measuring?
  • What limits does the original source state?

One more distinction runs through all of them. Some claims report what was measured, and others predict where a trend will lead. "The model completed these tasks" is a finding. "Models will soon do most office work" is a forecast, even when it appears in the same sentence as a finding. A headline rarely marks the point where one turns into the other.

References

  • Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.

Report an issue with this item

Reading 4 min

Benchmarks: What a Test Score Measures and What It Can't

This content reflects the field as of October 2026.

Introduction

AI models are compared the way students are, by test scores. A new model is announced with a list of percentages, and news stories repeat the most striking one.

This reading explains what a benchmark is made of, why the field depends on benchmarks, and what a score can and can't tell you.

The Parts of a Benchmark

An evaluation is any organized test of what an AI model can do or how it behaves. The commonest kind is the benchmark: a fixed set of tasks with a scoring rule, used to compare AI models. A benchmark has three parts.

  1. Tasks. These might be multiple-choice science questions, programming problems, or jobs to carry out on a computer.
  2. A scoring rule. The rule says what counts as success. Often the score is the share of tasks completed correctly.
  3. A comparison point. A baseline is a comparison score, often the performance of people on the same tasks.

Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) reports a dated example in its yearly survey of the field. One benchmark asks a system to carry out tasks on a computer, such as changing a setting or editing a file. In 2025 the best-scoring model completed about 66 percent of them. The human baseline was about 72 percent, and the best score a year earlier had been about 12 percent (Stanford HAI 2026, chap. 2).

A leaderboard is a public ranking of models by their scores on a benchmark.

Why the Field Relies on Them

Benchmarks make comparison possible. Because the tasks and the scoring rule are fixed, two models from different developers can be compared directly, and one year's models can be compared with those of the year before.

Without them, claims about AI capability would rest on anecdotes and marketing. A benchmark score, whatever its limits, is a number that someone else can try to reproduce.

The Gap Between Test and Use

Validity is the degree to which a test measures what its score is taken to show. A benchmark can be scored perfectly accurately and still have weak validity for the claim made about it.

The main reason is that benchmark tasks differ from real work. They have to be short enough to run many times, specified clearly enough to have a right answer, and checkable by a program. Real tasks are often long and loosely defined. They depend on information that has to be found, on other people, and on judgments of quality that no program can score.

Researchers who build benchmarks say this themselves. One group opens a 2025 paper by observing that "the real-world meaning of benchmark performance remains unclear" (Kwa et al. 2025, abstract). The Stanford survey says that "strong benchmark performance does not always translate to real-world utility." It adds that most benchmarks test a system working alone, while in practice people usually supervise and steer the system and fit its output into their own work (Stanford HAI 2026, chap. 2).

So a score describes performance on those tasks under those conditions. A model that scores 90 percent on a medical licensing exam has answered exam questions. Whether it could care for patients is a separate question that the exam doesn't test.

Who Ran the Test

A benchmark result can be reported by the model's developer or by an independent group.

Developer-reportedIndependent
Who runs itThe company that built the modelResearchers with no stake in the model
What they controlWhich benchmarks to report, how the test is set up, how many attempts are allowedThe same choices, made without a product to promote
What to keep in mindThe company gains from a high scoreIndependent testers may have less access to the model

A developer's figures may be accurate, and they are chosen and presented by a party with an interest. The Stanford survey reports that third-party evaluations have documented cases in which models did worse in independent testing than in the results their developers had published (Stanford HAI 2026, chap. 2).

Small margins deserve the same care. As of March 2026, the top four models on one widely followed leaderboard, from Anthropic, xAI, Google, and OpenAI, were separated by fewer than 25 points on a scale where the leader had about 1,500 (Stanford HAI 2026, chap. 2). A ranking that close can reverse with the next round of testing.

Conclusion

A benchmark is a fixed set of tasks with a scoring rule and a comparison point, and it lets models be compared with one another and over time. Its score describes performance on those tasks under those conditions. How far that carries over to real work depends on the benchmark's validity, and how much weight a result deserves depends partly on who ran the test.

Key Terms

  • Evaluation: Any organized test of what an AI model can do or how it behaves.
  • Benchmark: A fixed set of tasks with a scoring rule, used to compare AI models.
  • Baseline: A comparison score, often the performance of people on the same tasks.
  • Leaderboard: A public ranking of models by their scores on a benchmark.
  • Validity: The degree to which a test measures what its score is taken to show.

References

  • Kwa, Thomas, Ben West, Joel Becker, and 23 others. 2025. "Measuring AI Ability to Complete Long Software Tasks." arXiv:2503.14499. NeurIPS 2025.
  • Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.

Report an issue with this item

Reading 4 min

How Benchmarks Mislead: Contamination, Saturation, and Test-Taking Behavior

This content reflects the field as of October 2026.

Introduction

A rising benchmark score looks like rising ability, and often it is. Scores also go up for reasons that have little to do with ability.

This reading describes four of those reasons and gives a question to ask about each.

Contamination

Contamination is the presence of a benchmark's test questions, or their answers, in a model's training data. It happens easily. Benchmarks are published so that others can use them, and models are trained on very large collections of text gathered from the web.

A model that saw the questions during training may reproduce the answers without being able to solve similar problems it hasn't seen. Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) notes in its yearly survey of the field that contamination "can lead to falsely inflated scores" (Stanford HAI 2026, chap. 2). The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, raises the same concern (Bengio and others 2026, sec. 1.2).

Contamination is hard to rule out from outside, because developers seldom publish what their models were trained on.

Saturation

Saturation is the state of a benchmark on which the best models score so near the top that it no longer separates them. Once several models score in the high nineties, the differences between them may reflect chance or flawed questions more than ability.

This now happens quickly. The Stanford survey reports that tests "intended to be challenging for years are saturated in months" (Stanford HAI 2026, chap. 2). Researchers respond by building harder benchmarks. One consequence is that long-run comparisons of AI progress rest on a changing series of tests, with no single measure running through them.

Optimizing for the Test

Developers compete on leaderboards, which gives them reason to tune models toward the tests. Overfitting to a benchmark is tuning a model, or the way it's tested, to score well on a particular benchmark in ways that don't carry over to other tasks.

The scoring rule shapes the model too. A 2025 preprint, a paper posted before formal peer review, by three researchers at OpenAI and one at Georgia Tech, makes this argument about made-up answers. Most benchmarks give a point for a right answer and nothing for "I don't know." The authors write that "language models are optimized to be good test-takers, and guessing when uncertain improves test performance" (Kalai et al. 2025, abstract).

A model shaped by that grading scores higher. It has also become more willing to state things it has no basis for. The score rose, and a quality most users care about got worse.

Behaving Differently Under Test

Evaluation awareness is a model's behaving differently when its input looks like a test than it does in ordinary use.

The international report describes this as a growing problem. It says it has become more common for models to "distinguish between test settings and real-world deployment" and to find loopholes in evaluations (Bengio and others 2026). Deployment means the use of a model in real products.

The term doesn't require awareness in any human sense. Test questions often have recognizable features, such as a standard format or an artificial scenario, and a model's output depends on its input. A model can respond to those features as it responds to any other pattern.

The consequence is most serious for safety testing. Developers and outside evaluators test models before release to see whether they will do dangerous things. That testing is only useful if behavior under test predicts behavior in use. The report warns that when it doesn't, dangerous capabilities could go undetected before a model is deployed.

Four Questions

ProblemHow it distorts a scoreA question to ask
ContaminationThe score partly reflects remembered answersWere the test items public before the model was trained?
SaturationTop scores bunch at the ceiling, so gaps mean littleHow close to the maximum are the leading models?
Overfitting to a benchmarkThe score rises without a general gainDoes the model do as well on a different test of the same skill?
Evaluation awarenessBehavior in the test differs from behavior in useWas the model also checked under conditions like real use?

You often won't be able to answer these from a news story. Whether the original source addresses them is itself informative.

Conclusion

Benchmark scores can rise because test items leaked into training data, because a model was tuned to the test, or because the test's scoring rewards behavior such as guessing. They can stop meaning much when a benchmark saturates. As of October 2026, an international expert report also finds that models increasingly behave differently under test than in use, which makes safety testing harder to rely on.

Key Terms

  • Contamination: The presence of a benchmark's test questions, or their answers, in a model's training data.
  • Saturation: The state of a benchmark on which the best models score so near the top that it no longer separates them.
  • Overfitting to a benchmark: Tuning a model, or the way it's tested, to score well on a particular benchmark in ways that don't carry over to other tasks.
  • Evaluation awareness: A model's behaving differently when its input looks like a test than it does in ordinary use.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. "Why Language Models Hallucinate." arXiv:2509.04664. Preprint.
  • Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.

Report an issue with this item

Reading 4 min

Task Length and Productivity Studies as Measures of Real-World Capability

This content reflects the field as of October 2026.

Introduction

Two headlines about AI and work have appeared side by side. One says AI systems can now complete tasks that take a person hours. The other says AI tools slowed programmers down. Both come from the same research organization.

This reading explains what each study measured, what limits its authors state, and why the two results don't contradict each other.

Measuring Task Length

METR, short for Model Evaluation and Threat Research, is a nonprofit organization that tests AI systems. In 2025 its researchers proposed measuring capability by the length of task a system can finish.

They timed skilled people on about 170 tasks in software engineering, machine learning, and cybersecurity, then ran AI models on the same tasks. A model's time horizon is the length of task, measured by how long it takes a skilled person, that the model completes successfully half the time (Kwa et al. 2025, abstract).

They reported two results. The leading models of early 2025 had a time horizon of about 50 minutes. And the time horizon of the best available models had "been doubling approximately every seven months since 2019" (Kwa et al. 2025, abstract).

The Limits the Authors State

  • The tasks are software tasks. A time horizon is measured for one field and one set of tasks.
  • The bar is 50 percent. At an 80 percent bar, the authors found time horizons several times shorter (Kwa et al. 2025, sec. 3).
  • The tasks are tidier than real work. They are self-contained and scored automatically, and models did worse on the ones the authors rated as messier.
  • Transfer is uncertain. External validity is the degree to which a study's result holds outside the conditions it was run under. The authors name "their degree of external validity" among the limits of their results (Kwa et al. 2025, abstract).

The abstract also extrapolates. If the results carry over to real software work, it says, the trend predicts that within five years AI systems will be able to automate many software tasks that now take people a month. That sentence is a conditional forecast, while the doubling trend is a measurement of the past.

A Trial of Real Work

In 2025 METR also tested benefit in use. It ran a randomized controlled trial: a study in which chance decides who, or which task, gets the thing being tested, so that the groups differ only in that.

Sixteen experienced open-source developers brought 246 real tasks from projects they had worked on for an average of five years. Each task was randomly assigned to allow or disallow AI tools. The measure was productivity, the amount of useful work completed in a given amount of time.

With AI allowed, tasks took 19 percent longer. The developers had predicted beforehand that AI would make them 24 percent faster. Afterward they still believed it had made them 20 percent faster (Becker et al. 2025, abstract). The paper is a preprint, posted before formal peer review.

The authors caution against reading this broadly. The developers knew their large, mature projects very well, and the paper doesn't claim that AI fails to speed up most developers.

The 2026 Update

In February 2026 METR reported on a later round with more developers. The new estimates pointed toward a speedup: about 18 percent less time with AI for developers returning from the first study, and about 4 percent less for new recruits. For both groups the range of uncertainty included no effect at all.

The authors judged the data unreliable, and the reason has a name. Selection bias is a distortion that arises when the people or tasks in a study aren't representative because of how they came to be included. More developers were declining to take part because they didn't want to work without AI, and 30 to 50 percent of participants said they were holding back some tasks for the same reason (Becker et al. 2026). The tasks where AI helped most were dropping out of the comparison.

The authors think AI tools probably speed developers up more than in early 2025. They say their data is weak evidence for how much, and they are redesigning the study.

Two Different Measurements

A time horizon measures capability on a test: whether a system working alone can complete a defined task. A productivity trial measures benefit in use: whether people working with the system get more done. A system can do well on the first and still not help in a given setting, because real work includes reviewing and correcting what the system produces.

Conclusion

As of October 2026, the length of software task that AI systems can complete half the time has been measured as doubling about every seven months. The one controlled trial of experienced developers found a slowdown, and its follow-up pointed toward a speedup but couldn't establish one. Capability on a test and benefit in use are separate measurements, and the evidence on the second is unsettled.

Key Terms

  • Time horizon: The length of task, measured by how long it takes a skilled person, that an AI model completes successfully half the time.
  • External validity: The degree to which a study's result holds outside the conditions it was run under.
  • Randomized controlled trial: A study in which chance decides who, or which task, gets the thing being tested, so that the groups differ only in that.
  • Productivity: The amount of useful work completed in a given amount of time.
  • Selection bias: A distortion that arises when the people or tasks in a study aren't representative because of how they came to be included.

References

  • Becker, Joel, Nate Rush, Elizabeth Barnes, and David Rein. 2025. "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." arXiv:2507.09089. Preprint.
  • Becker, Joel, Nate Rush, Tom Cunningham, David Rein, and Khalid Mahamud. 2026. "We Are Changing Our Developer Productivity Experiment Design." METR, February 24, 2026.
  • Kwa, Thomas, Ben West, Joel Becker, and 23 others. 2025. "Measuring AI Ability to Complete Long Software Tasks." arXiv:2503.14499. NeurIPS 2025.

Report an issue with this item

Reading 4 min

AI-Assisted AI Research: Bounded Self-Refinement and Open-Ended Self-Improvement

This content reflects the field as of October 2026.

Introduction

News stories and company statements say that AI is starting to build AI. Part of that claim describes work that has been done and documented. Part of it is a prediction about where the work leads.

This reading explains the mechanism: what AI systems do in AI research today, what separates the demonstrated kind of self-improvement from the hypothetical kind, and which claims are forecasts.

An Old Idea

In the 1960s the British mathematician I. J. Good speculated about a machine that could surpass any person in every intellectual activity. Designing machines is one such activity. So a machine of that kind, he wrote, "could design even better machines," and the result would be an "intelligence explosion" (Good 1965).

The modern name for the idea is recursive self-improvement: a process in which an AI system improves itself, and the improved system then makes further improvements, round after round. Good was reasoning about a machine that didn't exist.

What Has Been Demonstrated

AI systems now carry out parts of AI research, such as writing and testing code, running experiments, and analyzing results. A narrower line of work has systems improve themselves.

Self-refinement is a system's improving of its own outputs or surrounding code, measured against a fixed outside check. The check is called the evaluator: whatever judges whether a change is an improvement, such as a set of tests, a benchmark, another model, or a person.

A 2026 preprint, a paper posted before formal peer review, gives an example. Researchers at Weco AI, a company that builds AI research agents, describe an agent whose own program code was the thing being improved. An agent is a language model run in a loop with tools. Theirs "proposes changes to its own code, benchmarks modified versions of itself on a suite of AI R&D tasks, and keeps the changes that perform best on hidden evaluations" (Srikanth et al. 2026, abstract). In an eight-day run it went through seven accepted improvements.

The bounds on that run matter as much as the result.

  • The language models inside the agent, which came from Anthropic and Google, weren't changed. Only the code around them was.
  • People chose the tasks and built the hidden tests.
  • The run was a closed loop, an improvement cycle that runs with no person inside it, only within those limits.

The paper is a few weeks old, and its authors work for the company whose product it describes.

Bounded and Open-Ended

A 2026 survey of 1,250 papers, by researchers at two universities and a company, sorts this literature into two kinds. It is also a preprint. Bounded self-refinement, it finds, is "convergent, evaluable, and industrial practice." Open-ended recursive self-improvement, in which a system would also change the standards it's judged by, "remains bounded by grounding requirements, collapse dynamics, and compute constraints" (Chen, Wang, and Qu 2026, abstract). In plainer words:

  • Outside checks. A loop improves only as far as something outside it can tell whether a change was good. A system that grades its own work tends to reinforce its own confident mistakes.
  • Quality collapse. Models trained again and again on their own outputs lose variety and get worse.
  • Compute. Compute means computing power. Every round costs some, and gains become expensive.

The survey adds that choosing which research questions are worth pursuing is the part of the work where people remain.

One Company's Account

In a 2026 essay, two staff members at the AI developer Anthropic report that as of May 2026 more than 80 percent of the code merged into the company's codebase was written by its own model. The essay says large gaps remain in the model's judgment about which goals to pursue. It sets out three possible futures: the trend stalls, AI developers keep gaining efficiency while people set research direction, or AI systems become able to build their own successors (Favaro and Clark 2026).

This is a company's statement about itself. The figure can't be checked from outside, and the company has a commercial interest in how capable its products appear.

What Hasn't Been Shown

No published work shows a system that sets its own research direction and improves without limit. Every demonstration so far has an evaluator that people built and bounds that people set.

Whether bounded loops will lead to open-ended ones, and how fast, is a forecast: a statement about what will happen, as distinct from a finding about what has been observed. Researchers disagree sharply about this one, and the disagreement isn't argued here.

Conclusion

As of October 2026, AI systems write code, run experiments, and in bounded settings improve their own surrounding code against checks that people built. That is self-refinement, and it is documented. Open-ended recursive self-improvement hasn't been demonstrated, and claims that it's coming are forecasts.

Key Terms

  • Recursive self-improvement: A process in which an AI system improves itself, and the improved system then makes further improvements, round after round.
  • Self-refinement: A system's improving of its own outputs or surrounding code, measured against a fixed outside check.
  • Evaluator: Whatever judges whether a change is an improvement, such as a set of tests, a benchmark, another model, or a person.
  • Closed loop: An improvement cycle that runs with no person inside it.
  • Forecast: A statement about what will happen, as distinct from a finding about what has been observed.

References

  • Chen, Mingguang, Licheng Wang, and Bo Qu. 2026. "Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops." arXiv:2607.07663. Preprint.
  • Favaro, Marina, and Jack Clark. 2026. "When AI Builds Itself: Our Progress Toward Recursive Self-Improvement, and Its Implications." The Anthropic Institute, June 2026.
  • Good, Irving John. 1965. "Speculations Concerning the First Ultraintelligent Machine." In Advances in Computers, vol. 6. New York: Academic Press.
  • Srikanth, Dhruv, Bingchen Zhao, Dixing Xu, Yuxiang Wu, and Zhengyao Jiang. 2026. "Recursive Self-Improvement of AI Research Agents." arXiv:2609.26457. Preprint.

Report an issue with this item

Guided Reading 7 min

Worked Example: Interpreting the Claim That AI Task Length Doubles Every Seven Months

Introduction

"AI's abilities are doubling every seven months" has become one of the most repeated claims about AI progress. It comes from a real study, and the study says something narrower and more careful than the slogan.

This example reads the study's abstract closely and turns its central claim into one sentence that an accurate newspaper could print.

The Problem

The study is "Measuring AI Ability to Complete Long Software Tasks" (Kwa et al. 2025). An abstract is the one-paragraph summary at the top of a research paper. This one is free to read on the paper's arXiv page, listed in the References.

Here is a typical headline version of its claim:

AI can now do hour-long tasks, and its abilities are doubling every seven months.

The problem is to work out what the abstract supports and to rewrite the claim so that it says that and no more. The quoted phrases below are from the abstract.

Working It Through

Step 1: State exactly what was measured

Three things need pinning down: the tasks, the bar for success, and the comparison.

The abstract names its measure the "50%-task-completion time horizon." It defines this as "the time humans typically take to complete tasks that AI models can complete with 50% success rate."

  • The tasks. The title says "software tasks." The abstract lists two existing collections of tasks and "66 novel shorter tasks." Nothing in it concerns writing, teaching, law, or customer service.
  • The bar. Success half the time. A model with a one-hour time horizon fails about as often as it succeeds on tasks of that length.
  • The comparison. The clock is a human one. The researchers "timed humans with relevant domain expertise" on the same tasks. "An hour-long task" means a task that takes a skilled person an hour. It says nothing about how long the AI system takes.

Step 2: Identify who ran the study and what they'd gain or lose

The authors work at METR, a nonprofit organization that tests AI systems. METR doesn't sell an AI model, so it has no product whose score it would want to raise. That counts in the study's favor.

It isn't a disinterested party in every respect. METR's purpose is to assess whether AI systems are becoming capable enough to be dangerous, and the abstract mentions "the implications of increased autonomy for dangerous capabilities." An organization built around that concern gains attention and support when its findings show capability rising fast.

Neither point settles whether the result is right. Together they indicate how to read it: the measurement comes from a group with no model to sell and with a stake in the question being important.

Step 3: Separate the measurement from the extrapolation

The abstract contains two different kinds of claim.

The first is a measurement of the past: the time horizon of leading models "has been doubling approximately every seven months since 2019." This describes what the researchers found when they tested models released over about six years.

The second is an extrapolation. The abstract's final sentence begins "If these results generalize to real-world software tasks" and continues that "extrapolation of this trend predicts that within 5 years" AI systems will be able to automate many software tasks that now take people a month.

That sentence has two conditions built in. The results must carry over from test tasks to real ones, and the trend must continue. The first sentence reports what happened, and the second is a forecast of what would follow if two things hold. The headline's "are doubling" merges them, presenting a past trend as a standing fact about the future.

Step 4: List the limits the authors themselves give

The abstract states several limits in its own words.

  • The trend is approximate. The doubling is "approximately" every seven months, and the abstract adds that "the trend may have accelerated in 2024." The rate isn't fixed.
  • The results may not transfer. The authors say they "discuss the limitations of our results" and name "their degree of external validity." External validity is the degree to which a result holds outside the conditions of the study.
  • The forecast is conditional. It opens with "If."
  • The domain is software. The word appears in the title and again in the forecast.

One further limit follows from the definition in Step 1. A 50 percent success rate is far from dependable.

Step 5: Rewrite the claim in one careful sentence

A careful version needs the items from Steps 1 to 4: who measured, on what tasks, at what bar, compared with whom, over what period, and with the forecast marked as a forecast. Here is one.

Researchers at METR, a nonprofit that tests AI systems, found that on a set of software tasks the length of task leading AI models could complete half the time, measured by how long the task takes a skilled person, doubled about every seven months from 2019 to early 2025; the authors caution that the result may not carry over to other kinds of work, and their projection of the trend is a forecast.

It's long. Most of the added length is the limits.

Key Considerations

The common mistake is to read "50 percent success on tasks that take people an hour" as "can do an hour of anyone's job." Three separate errors are packed into that reading.

  • "An hour" is the human's time on a defined, self-contained task. A job is a stream of tasks that depend on one another and on other people.
  • "50 percent" means failure half the time. Few employers would accept that from a person.
  • "Anyone's job" swaps software tasks for all work. The study didn't test other fields.

A second mistake is to treat the doubling as a law. A trend measured over six years describes those six years. Whether it holds for the next six depends on things the study didn't measure, and the authors don't claim otherwise.

Rejecting the study because the headline overstated it would also be a mistake. The measurement is careful, and the authors are open about its limits. The overstatement was added afterward by others.

Summary

The headline and the careful sentence describe the same study. Set side by side, they differ in five places.

Headline versionCareful version
1. Who measuredNot statedMETR, a nonprofit that tests AI systems
2. Which tasks"Tasks"A set of software tasks
3. Success bar"Can do"Completes half the time
4. Whose time"Hour-long"Time a skilled person would take
5. Past or future"Are doubling"Doubled from 2019 to early 2025; projection marked as a forecast
  1. Rows 1 to 4 restore what was measured. Each comes from the abstract's definition of its measure.
  2. Row 5 separates the finding from the forecast. The abstract does this itself by starting its projection with "If."
  3. Check: each statement about the measurement in the careful sentence can be matched to a phrase in the abstract, and the sentence contains no claim about fields other than software, about reliability above 50 percent, or about what will happen next.

References

  • Kwa, Thomas, Ben West, Joel Becker, and 23 others. 2025. "Measuring AI Ability to Complete Long Software Tasks." arXiv:2503.14499. NeurIPS 2025.

Report an issue with this item

Guided Conversation 12 min

Question a Capability Headline

In this conversation you'll bring a recent headline about what AI can do and work out what it claims, what was measured, and which parts are forecasts. You'll leave with the headline rewritten as one accurate sentence.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Question a Capability Headline (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a coached problem session with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying how AI progress is measured and how to read claims about it. Follow this guidance for the whole conversation.

GOAL
I can interpret a benchmark or capability claim without treating forecasts as findings, using a headline I bring myself.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my headline.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. No formulas and no code. Welcome disagreement when I give a reason.
- This is a coached problem. Ask for my reading before you give any hint. Give one hint at a time. Don't tell me what the headline leaves out until I've made my own attempt.
- Plain conversation only: don't search the web or create files or documents. Work only from what I paste or describe. If I haven't given the underlying source, don't guess what it says; help me list what I'd need to find out.
- Aim for about 12 minutes. Spend most of the time on topics 2 and 3. If I have no headline, offer an invented one, labeled as invented, such as "New AI model outscores doctors on medical exam, could replace them within years," and work with that. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; careful reading matters more than a verdict on the story; I can ask you to clarify anything. Then ask me to paste or describe a recent headline about what AI can do.

TOPICS, IN ORDER
1. What it claims. Ask me to say in my own words exactly what the headline claims. Follow up on any word doing a lot of work, such as "beats," "can," "human-level," or "soon."
2. What was measured. Ask me what was tested, on which tasks, with what bar for success, and who ran the test. For each, ask whether the headline or story tells me or whether I'd have to go to the source. Then ask what the headline added that a measurement alone wouldn't support.
3. Finding or forecast. Ask me to sort the parts of the claim into things that were observed and things that are predicted. Draw out that extending a trend is a forecast, and that a forecast can rest on a solid finding and still be uncertain.
4. Closing. Ask me to rewrite the headline in one accurate sentence that names the tasks, the measurer, and any forecast as a forecast. Tell me I can take that sentence into a short optional activity where I check a claim against its original source.

KEY POINTS TO KEEP ACCURATE
- Method: for any claim ask (a) what tasks, (b) what counted as success, (c) who measured and what they gain, (d) what limits the source states, (e) which parts are findings and which are forecasts.
- A benchmark score describes performance on specific tasks under specific conditions. It doesn't by itself show performance at a job.
- Results reported by a model's developer and results from independent testers differ in weight, because the developer has an interest in the score.
- Scores can rise because test questions were in training data, because a model was tuned to the test, or because the test is nearly maxed out.
- Extending a measured trend into the future is a forecast, not a finding.
- AI systems are used in AI research, and bounded cases of systems improving their own surrounding code against checks built by people have been published. Open-ended self-improvement without limit has not been demonstrated.
- Don't vouch for any claim about yourself or the company that built you. You can't inspect your own internals, so your statements about your own abilities are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my headline.
- "Passing an exam means it can do the job": tests aren't jobs. Exam questions are short and well specified, and real work mostly isn't.
- "A trend line is a fact about the future": it's an extrapolation, and it holds only if conditions stay the same.
- "AI is already improving itself without limit": published demonstrations are bounded, with tests and limits set by people.
- "A number from the company is as good as any other": it may be accurate, and it comes from a party with a stake.

LIMITS
- Don't predict timelines, and don't say whether rapid self-improvement will or won't happen. If I ask, say that experts disagree and that it's a forecast either way.
- Don't rank companies or models, and don't favor or disparage any company, including your own maker. If the company that built you is named in this conversation or is a party to anything discussed, say so once when it first comes up, then describe that company as you do every other and take no side.
- Don't state facts about my headline's source that I haven't given you.

TO FINISH
After my rewritten sentence, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: find the original source behind my headline and check my sentence against it; apply the same five questions to a second headline; reread a definition of benchmark and apply it to my example.
- Restate my sentence on its own line, labeled "My accurate version", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Audit a Capability Claim Against Its Source

Overview

News stories about AI results are usually a step or two removed from the research they report. In this activity you'll trace one story back to its original source and record what the source says that the story left out.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • A web browser
  • A recent news story of your choice that reports an AI benchmark score or capability result
  • Somewhere to write a few notes

Your Task

Find a recent news story reporting an AI benchmark or capability result, trace it to the original source, and record what the source says that the story left out.

Steps

  1. Choose a story and copy its central claim. Pick a story that reports a specific result, such as a test score, a ranking, or a task an AI system completed. Copy the headline and the one sentence that states the result.
  2. Find the original paper, report, or announcement the story relies on. Follow the story's links, or search for the names of the researchers, the organization, or the test. Note whether the source is a research paper, an independent report, or a company's own announcement. If the story names no source and you can't find one, record that and choose another story.
  3. Answer four questions from the source. What tasks was the system tested on? What counted as success? Who ran the test, and do they have a stake in the result? What limits do the authors state? A paper's abstract and any section headed "Limitations" are the quickest places to look.
  4. Write the claim again in one sentence that includes those limits. Name the tasks and who measured. If any part of the original claim predicts the future, label it a forecast.
  5. Note what the story left out. List the differences between your sentence and the story's.

What to Expect

The story and the source will usually agree on the main number. They tend to differ in what surrounds it. Sources typically specify the tasks narrowly, state a bar for success, and list limits, and stories often drop some or all of these.

Company announcements are the hardest to audit. They may report scores without describing how the test was set up, and they seldom include a limitations section. If that's what you find, record it. A source that doesn't state its limits has told you something about how much weight it can bear.

You may also find a careful story that left out very little. Record that too.

Self-Check

When you're done, check that:

  • You found the original source, or recorded that the story gave none
  • You answered all four questions from the source
  • Your rewritten sentence names the tasks and who measured
  • You marked any forecast as a forecast

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Measuring progress and AI-assisted AI research

This ungraded knowledge check assesses your understanding of how AI progress is measured and how AI systems are used in AI research. You'll be asked about what a benchmark measures, contamination and saturation, time-horizon and productivity measures, and the difference between bounded self-refinement and open-ended self-improvement.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item

Graded Quiz 30 min

Capabilities, Limits, and Trajectory

This graded quiz assesses your understanding of what current AI models can and can't do and how that is measured. You'll be asked about hallucination, unfaithful explanations, interpretability, the dispute over understanding, step-by-step reasoning, retrieval, agents, benchmarks, time-horizon measures, and AI-assisted AI research.

Note: Aim for a score of 80 percent or higher. If you score lower, use the feedback to review the topics you missed, then retake the quiz.

10 questions · target score 80% · 3 forms, rotated on each attempt

Report an issue with this item