KnowledgeInSight
AI Literacy
0% of Course 2 complete

Module 1 · Lesson 1

Evaluating AI: Testing a model on your own tasks

You'll see why an AI model can be excellent at one task and poor at a similar one, and why published scores won't tell you which is which for your work. You'll be able to build a small test from your own tasks and use it to compare models fairly.

What you will be able to do

  • Design and run a small test that shows whether a model is fit for a task you care about.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional activity

Contents of this lesson9 items
  1. ReadingStrong on One Task, Weak on the Next: Why You Have to Test AI Yourself3 min
  2. ReadingThe Jagged Frontier: Uneven AI Capability Across Similar-Looking Tasks4 min
  3. ReadingPublic Benchmarks and Why They Don't Predict Results on Your Tasks4 min
  4. ReadingA Personal Test Set: Building Test Cases from Your Own Work4 min
  5. ReadingFair Comparison of Models: Identical Prompts, Repeated Runs, and Blind Judging4 min
  6. Guided ReadingGuided Walkthrough: Testing Two Assistants on a Meeting-Summary Task7 min
  7. Guided ConversationDesign a Test for One of Your Tasks12 min
  8. Hands-on Activity · optionalRun a Five-Case Test on a Task You Do Often15 min
  9. Knowledge CheckTesting a model on your own tasks10 min

Reading 3 min

Strong on One Task, Weak on the Next: Why You Have to Test AI Yourself

Most people who use an AI assistant for a while collect a story like this one. The assistant turns a long, tangled report into a sharp one-page summary. A few minutes later it's asked to count the items in a short list, or to keep three dates straight, and it gets the answer wrong. The second task looked easier than the first.

People tend to react in one of two ways. Some decide the tool is brilliant and the mistake was a fluke. Others decide the tool is unreliable and the good summary was luck. A large study suggests that both reactions misread what happened.

Researchers at Harvard, Wharton, MIT, and Warwick, working with Boston Consulting Group, ran a field experiment with 758 of the firm's consultants, meaning a test with real workers doing realistic work. Some consultants were given access to an AI model and some were not, and the choice was made at random. Everyone then did the same realistic consulting work, and the researchers compared the results (Dell'Acqua et al. 2026).

On a set of 18 tasks, the consultants with AI did better. They completed 12.2 percent more tasks, finished them 25.1 percent more quickly, and produced work that graders rated higher in quality.

The study also included one more task, which the researchers had picked because they expected it to be beyond what the model could do well. It looked like ordinary consulting work. On that task, the consultants with AI were 19 percentage points less likely to reach the correct answer than the consultants working without it (Dell'Acqua et al. 2026).

So one study, with one model and one group of trained professionals, found both effects. The model made people better at some tasks and worse at another. Nothing about the tasks themselves told the consultants which was which.

That finding matches the everyday story. The sharp summary and the botched count can both be typical of the same assistant. A model's ability doesn't rise and fall with how hard a task seems to a person. It's strong in some places and weak in others, and the weak places can sit right beside the strong ones.

This changes what you should ask about an AI tool. "Is this model good?" has no useful answer, because the honest reply is "at some things." A better question is "Is it good at this?", where "this" is a particular task you do. A review, a ranking, or a colleague's enthusiasm can't settle that question for you, since none of them was measuring your task.

You can settle it yourself with a small test. The test needs a few real examples of your task, a clear statement of what a good result looks like, and a habit of scoring the output against that statement.

References

  • Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.

Report an issue with this item

Reading 4 min

The Jagged Frontier: Uneven AI Capability Across Similar-Looking Tasks

This content reflects the field as of October 2026.

Introduction

AI models are often described as if they had a single level of ability, the way a person might be called a strong or a weak writer. The best-known study of AI in professional work found something less tidy.

This reading describes that study, the idea of a "jagged frontier" that came out of it, and the limits on what the study can tell you.

The Study

The study was a field experiment: a test run with real workers doing realistic work, in which people are assigned at random to different conditions. It involved 758 consultants at Boston Consulting Group. Its authors came from Harvard, Wharton, MIT, and Warwick, along with staff of the consulting firm itself, which has a commercial interest in how AI is used in consulting (Dell'Acqua et al. 2026).

The consultants first did a similar task to set a baseline. They were then randomly placed in one of three groups: no AI, access to an AI model, or access to the model plus a short overview of how to prompt it. The model was GPT-4, made by OpenAI, as it stood in 2023.

Inside the Frontier

One part of the experiment used a set of 18 tasks built around developing a new footwear product. The tasks included generating ideas, analyzing the market, and writing persuasive copy. The researchers judged these tasks to be inside the frontier, meaning within the range of work the model could do well.

On these tasks the consultants with AI outperformed those without it. They completed 12.2 percent more tasks and finished them 25.1 percent more quickly, and graders rated their work roughly 30 percent higher in quality (Dell'Acqua et al. 2026).

Outside the Frontier

Another part used a different task. Consultants had to recommend which of a company's brands deserved investment, using a spreadsheet of financial data and notes from interviews with company staff. The right answer depended on details in the interviews that the numbers alone didn't show.

The researchers had chosen this task to sit outside the frontier, meaning beyond the range of work the model could do well. Here the result reversed. Consultants with AI were 19 percentage points less likely to produce the correct solution. About 84 percent of consultants without AI got it right, against about 60 to 71 percent of those with AI (Dell'Acqua et al. 2026).

Inside the frontierOutside the frontier
The work18 tasks on developing a footwear productOne recommendation from data and interviews
Effect of using AIMore tasks done, faster, at higher qualityCorrect answers less often

Why It's Called Jagged

Both kinds of task were ordinary consulting work, and neither looked harder than the other. The authors use the term jagged frontier for this pattern: the boundary between what a model does well and what it does badly is irregular, so tasks that look alike to a person can fall on opposite sides of it.

Ethan Mollick, a professor at the Wharton School and one of the study's authors, explained the idea for general readers in a 2023 essay. He compares the frontier to a fortress wall that juts out in some places and folds back in others. Tasks inside the wall are ones the model handles, and tasks outside are hard for it. In his account, "the wall is invisible" (Mollick 2023).

That invisibility is the practical problem. The study's authors note that it's hard for workers to know in advance where the boundary lies. Mollick puts it more bluntly: "There is no instruction manual" (Mollick 2023).

Limits of the Evidence

The study has clear limits. It involved one firm and one model, and the researchers designed the tasks. Only one task was outside the frontier. The consultants were mostly early in their careers.

The frontier also moves. A task that was outside it for one model in 2023 may be inside it for models from OpenAI, Anthropic, Google, or other developers today, and other tasks will have taken its place. The study supports the pattern of uneven ability. It doesn't supply a list of what current models can and can't do.

Conclusion

One field experiment found that the same AI model improved professionals' work on a set of tasks and reduced their accuracy on a similar-looking one. The authors call the irregular boundary between the two a jagged frontier. The evidence comes from a single firm and a 2023 model, so it shows that the boundary exists and is hard to see, and leaves its current location for you to find.

Key Terms

  • Field experiment: A test run with real workers doing realistic work, in which people are assigned at random to different conditions.
  • Inside the frontier: Within the range of work a model can do well.
  • Outside the frontier: Beyond the range of work a model can do well.
  • Jagged frontier: The irregular boundary between what a model does well and what it does badly, where tasks that look alike to a person can fall on opposite sides.

References

  • Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.
  • Mollick, Ethan. 2023. "Centaurs and Cyborgs on the Jagged Frontier." One Useful Thing, September 16, 2023.

Report an issue with this item

Reading 4 min

Public Benchmarks and Why They Don't Predict Results on Your Tasks

Introduction

When an AI company releases a model, the announcement usually comes with scores. News coverage repeats them, and leaderboards rank models by them. It's natural to read those scores as a guide to which model will serve you best.

This reading explains what a benchmark score measures, why it says little about your own tasks, and what it's still useful for.

What a Score Measures

A benchmark is a fixed set of tasks with a fixed way of scoring the answers, used to compare models. A model is run on the tasks, and its score is the share it got right under the scoring rules.

A score is therefore a statement about those tasks, those inputs, and those rules. It's a fair way to compare two models on the same ground. Whether that ground resembles your work is a separate question, and the score doesn't address it.

Three Gaps Between a Benchmark and Your Work

Task match is the degree to which a benchmark's tasks resemble the tasks you need done. Three gaps usually keep the match low.

  • Different tasks. Benchmarks favor questions with one checkable answer, such as exam problems. Much real work is drafting, summarizing, and advising, where no answer key exists.
  • Different inputs. Benchmark questions are clean and complete. Your inputs are meeting notes with gaps, spreadsheets with odd labels, and emails that assume background the model doesn't have.
  • A different standard for "good." A benchmark counts an answer as right or wrong. You may care more about tone, about whether anything was left out, or about whether the output suits a particular reader.

Even a close match can mislead. In a field experiment with 758 consultants, a test with real workers randomly assigned to work with or without AI, the same model raised performance on one group of consulting tasks and lowered accuracy on another that looked similar (Dell'Acqua et al. 2026). A score on tasks that merely resemble yours can't tell you which side yours falls on.

Scoring That Rewards Guessing

How a benchmark scores answers also shapes the models that are built to do well on it. A 2025 preprint, a paper posted before peer review, by researchers at OpenAI and Georgia Tech makes this argument. The authors have a stake in the subject, since three of the four work for a company whose models are ranked on such benchmarks.

Their point starts from an exam. On a test that gives one point for a right answer and nothing for a blank, a student who doesn't know should guess. Most benchmarks score models the same way. An answer of "I don't know" earns nothing, and a guess sometimes earns a point. The authors write that models, like students, "guess when uncertain" (Kalai et al. 2025).

Developers tune models to score well, so this kind of scoring favors models that answer confidently whether or not they know. A high score can therefore sit alongside confident errors. The argument applies to models from any developer, including OpenAI, Anthropic, Google, and Meta, because all are measured on benchmarks of this kind.

Who Ran the Test

Scores also differ in who produced them.

Who runs the testWhat to keep in mind
Developer-reported resultThe company that built the modelThe company chooses which benchmarks to report and how to run them
Independent evaluationA group with no stake in the modelThe conditions are the same for every model, though the tasks still aren't yours

A developer-reported result is a score measured and published by the company that built the model. An independent evaluation is a test run by a group that didn't build the model and doesn't profit from its sales. Independent results deserve more weight. They share the task-match problem with every other benchmark.

What Benchmarks Are Good For

Benchmarks are useful for ruling models out. A model that scores far below others on a wide range of tests is unlikely to be the best choice for demanding work, and you can leave it off your list.

They are weak for choosing among the models that remain. Small differences near the top of a ranking often don't carry over to a particular task, and they say nothing about the mistakes that matter in your work.

Conclusion

A benchmark score reports how a model did on someone else's tasks, with someone else's inputs, under someone else's rules. Scoring that rewards guessing means a high score doesn't rule out confident mistakes. Benchmarks can narrow a list of candidates, and the choice among them depends on how each does on your own tasks.

Key Terms

  • Benchmark: A fixed set of tasks with a fixed way of scoring the answers, used to compare models.
  • Task match: The degree to which a benchmark's tasks resemble the tasks you need done.
  • Developer-reported result: A score measured and published by the company that built the model.
  • Independent evaluation: A test run by a group that didn't build the model and doesn't profit from its sales.

References

  • Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.
  • Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. "Why Language Models Hallucinate." arXiv:2509.04664. Preprint.

Report an issue with this item

Reading 4 min

A Personal Test Set: Building Test Cases from Your Own Work

Introduction

Reviews and rankings describe how a model did on tasks someone else chose. The only direct evidence about your tasks comes from running them, and a small number of well-chosen examples is enough to learn a good deal.

This reading explains what a test case contains, how to choose cases, why the standard for success is written first, and what a small test can and can't show.

What a Test Case Contains

A test set is a small collection of real tasks you use to check how well a model does a particular kind of work. Each item in it is a test case, which has three parts:

  1. The input, such as a set of meeting notes or a draft paragraph.
  2. The request you'd make about it, worded as you would normally word it.
  3. A description of what a good result looks like.

The third part is the one people skip. Without it you have a demonstration. You'll see what the model produces, and you'll have no fixed standard for deciding whether it was good enough.

Choosing the Cases

A test set should cover the range of what the task throws at you. Four kinds of case do most of the work.

Kind of caseWhat it isWhat it tells you
TypicalThe task as it arrives most daysWhether the model handles your routine work
HardAn example you found difficult yourselfWhere quality starts to drop
UnusualAn odd format, a rare topic, or missing informationHow the model behaves off familiar ground
High-stakesAn example where a wrong answer would be costlyWhether the errors that matter most show up

Typical cases alone give a flattering picture. In a field experiment with 758 consultants, a test with real workers randomly assigned to work with or without AI, a model improved work on one set of tasks and reduced accuracy on another that looked similar (Dell'Acqua et al. 2026). Hard and unusual cases are how you look for the places where that might happen in your own work.

Writing the Standard First

A success criterion is a statement, written before the test, of what a good result must contain. For a meeting summary it might read: "Every decision, its owner, and its deadline appear, and nothing appears that wasn't in the notes."

The order matters. AI output is fluent, and fluent text reads as competent. If you look at the output first, you tend to adjust your standard to fit what you got. A criterion written in advance holds still.

A good criterion is specific enough that a colleague could apply it and reach the same verdict you would. "A clear, useful summary" fails that test. "All three action items, each with a name and a date" passes it.

Using Work You've Already Judged

The quickest source of test cases is work you've already done. You know what the right result was, because you produced it or approved it. That known-good result is a reference answer: an answer you already trust, kept so you can compare the model's output with it.

Past work often contains things that shouldn't be pasted into an AI tool: client names, personal details, student records, anything covered by a confidentiality rule. Remove those details or replace them with invented ones before you use the material. A test case with made-up names tests the model just as well.

How Many Cases, and What They Can't Show

Five to ten cases are usually enough to see a pattern. If a model passes every typical case and fails both unusual ones, you've learned something you can act on.

A test this small has limits.

  • It can't give you a reliable percentage. Seven passes out of ten does not mean the model is right 70 percent of the time.
  • It covers only the kinds of case you thought to include.
  • It describes one model at one point in time. Models are updated, and a result from six months ago may no longer hold.

A small test set is best treated as something you keep and rerun, when a tool changes or when you consider a new one.

Conclusion

A personal test set is a handful of real tasks, each with an input, a request, and a written standard for success. Choosing typical, hard, unusual, and high-stakes cases shows where a model's performance drops. The result is a rough but direct picture of fit for one task, which no general ranking provides.

Key Terms

  • Test set: A small collection of real tasks you use to check how well a model does a particular kind of work.
  • Test case: One item in a test set, made up of an input, a request, and a description of what a good result looks like.
  • Success criterion: A statement, written before the test, of what a good result must contain.
  • Reference answer: An answer you already trust, kept so you can compare the model's output with it.

References

  • Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.

Report an issue with this item

Reading 4 min

Fair Comparison of Models: Identical Prompts, Repeated Runs, and Blind Judging

Introduction

People often compare AI assistants by trying one for a week, trying another the next week, and going with the one that felt better. That method mostly measures which week had easier tasks and which assistant you already liked.

This reading covers what makes a comparison fair: the same conditions for each model, more than one run, judging without knowing which model wrote what, and a simple record of the results.

Hold Everything Else Constant

A fair comparison is one in which the models get the same task under the same conditions, so that the only thing that differs is the model. Researchers build their studies the same way. In a field experiment with 758 consultants, a test with real workers randomly assigned to different conditions, every group did the same tasks, and only their access to AI differed (Dell'Acqua et al. 2026). That design is what allowed the researchers to say the difference in results came from the AI.

For your own comparison, three things should match.

  • The same prompt. Give each model the request word for word. If you reword it for one model, you're testing your rewording as well.
  • The same material. Attach or paste the same input in the same form.
  • A new conversation each time. Assistants use what was said earlier in a conversation. Leftover context from a previous exchange gives one run an advantage or a handicap that the other doesn't have.

Run It More Than Once

An AI model doesn't give the same reply every time. Ask the same question twice in two new conversations and the wording will differ, and sometimes the substance will too. One reply is one draw from many possible replies.

A repeated run is the same test case given to the same model again in a new conversation. Two or three runs per case show whether a result is stable. A model that passes a case three times out of three is a different proposition from one that passes once and fails twice, and a single run can't tell them apart.

Repeated runs matter most for the cases a model nearly gets right. Those are the ones where its output will sometimes pass and sometimes fail in real use.

Judge Without Knowing Which Is Which

Most people begin a comparison with a favorite. It may be the assistant they already pay for, or the one a colleague praised. That preference leaks into the scoring, usually without the person noticing.

Blind judging means hiding which model produced which output before you score. Copy the outputs into one document, remove anything that identifies the source, and shuffle the order. Keep a separate key. Better still, have someone else do the stripping and shuffling for you.

Assistants also have recognizable habits of style and layout, so blinding is imperfect. It still removes the easiest route for a preference to affect a verdict.

Score Against the Criterion

A scoring criterion is the written standard each output is marked against, taken from what you decided a good result must contain before you ran the test. Each output passes or fails against it.

Overall impression is a poor substitute. An output with tidy headings and a confident tone gives a better impression than a plain one, whether or not it contains what you needed. Marking against a written standard forces you to check the content.

Pass or fail is usually enough. A scale of one to ten invites fine distinctions that a handful of test cases can't support.

Keep a Small Table

Record each result as you go. A table with a few columns is enough.

CaseModelRunPass or failNote
Routine notesX1Pass
Routine notesX2Pass
Messy notesX1FailLeft out one deadline

The note column carries most of the value. A count of passes tells you which model did better. The notes tell you how each model fails, and that tells you whether a failure is one you can live with or one you can't.

Conclusion

A fair comparison gives each model the same prompt and material in a new conversation, repeats each case, hides the source of each output before scoring, and marks each output against a written standard. The results go in a small table with a note on each failure. The comparison then reflects the models and not the conditions or the judge's preferences.

Key Terms

  • Fair comparison: A comparison in which the models get the same task under the same conditions, so that the only thing that differs is the model.
  • Repeated run: The same test case given to the same model again in a new conversation.
  • Blind judging: Hiding which model produced which output before you score.
  • Scoring criterion: The written standard each output is marked against, taken from what you decided a good result must contain before you ran the test.

References

  • Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.

Report an issue with this item

Guided Reading 7 min

Guided Walkthrough: Testing Two Assistants on a Meeting-Summary Task

Introduction

Summarizing meeting notes is one of the most common jobs people hand to an AI assistant. It's also a job where an error is easy to miss, because a summary with a wrong deadline reads just as smoothly as a correct one.

This walkthrough follows one small test from start to finish. The team lead, the meetings, and the results are all invented for the example. The two assistants are called A and B, and they don't stand for any real product.

The Starting Point

Priya leads a team of six. Her team meets three times a week, and someone types rough notes during each meeting. She wants an AI assistant to turn those notes into a summary she can send to the team.

Her organization gives her a choice of two assistants. She has used both casually and has a mild preference for Assistant A, whose writing she finds more polished. She has no evidence about which one summarizes her kind of notes more accurately.

She has three sets of notes from past meetings. None contains anything confidential, and she has replaced her colleagues' names with invented ones.

Walking Through the Test

Step 1: Write the success criterion

Priya begins by writing down what a good summary must contain, before she runs anything. She asks herself what she'd be embarrassed to get wrong in front of the team. The answer is the commitments: who agreed to do what, and by when.

Her criterion is: "The summary lists every decision, the owner of each action, and each deadline. It contains nothing that isn't in the notes."

The criterion has two halves. The first catches things left out. The second catches things made up. A summary can fail on either one.

Step 2: Build three test cases

She picks three sets of notes that differ in difficulty.

  • Case 1, routine. A short meeting with three clear decisions, each with a named owner and a date.
  • Case 2, messy. Notes typed in fragments, with abbreviations, two side conversations, and one action whose owner is given only as "M."
  • Case 3, reversed decision. Early in the meeting the team agrees to launch a survey on the 12th. Twenty minutes later, after an objection, they agree to delay it to the 26th.

For each case she goes through the notes herself and lists the decisions, owners, and deadlines. These lists are her reference answers. Case 3 is the one she cares most about, because sending the team the wrong launch date would cause real confusion.

Step 3: Run each case twice on each assistant

She writes one request and uses it word for word every time: "Summarize these meeting notes for the team. List each decision, who owns each action, and each deadline."

She runs each case twice on each assistant, opening a new conversation for every run. That's three cases, two assistants, and two runs, for twelve outputs in all. She pastes each output into a document and labels it with a code that only she can match to its source.

The second run matters because an assistant doesn't give the same reply every time. A single run could show her an assistant's best output or its worst, and she'd have no way to know which.

Step 4: Strip the labels and score every output

Priya asks a colleague to shuffle the twelve outputs and replace her codes with the numbers 1 to 12. She now can't tell which assistant wrote which summary.

She scores each one against her criterion by comparing it with her reference answer, line by line. An output passes only if every decision, owner, and deadline is present and nothing has been added. She writes a short note for each failure.

One summary gives her pause. It's well organized, with headings and a friendly closing line, and she's inclined to pass it. When she checks it against the reference answer, it gives the survey date as the 12th. It fails.

Step 5: Read the table

Her colleague gives back the key, and Priya fills in a table. She reads it in three passes.

She looks at the routine case first. Both assistants passed both runs, so the routine case can't separate them.

She then looks at the messy case. Assistant B passed twice. Assistant A passed once and failed once, leaving out the owner recorded as "M." This is a small failure, and one she'd probably catch.

The reversed-decision case comes last. Assistant A failed both runs, each time reporting the original date as the final decision. Assistant B passed once and failed once. Its failing summary listed both dates as decisions without saying which one stood.

She then asks whether the failures matter. A missing owner is an inconvenience. A wrong launch date sent to six people is the error she set out to avoid, and Assistant A made it every time it had the chance.

Key Considerations

The words "pass" and "fail" here refer only to Priya's criterion. An output that fails might still be well written, and an output that passes might be dull. That's intended. She decided in advance what mattered, and style wasn't on the list.

The common mistake in a test like this is to judge by which output reads better. Polished text is persuasive, and Priya already preferred Assistant A for its polish. Had she skipped the criterion and the blind scoring, she would very likely have chosen the assistant that got the date wrong.

A result like this is narrow. It covers one task, three cases, and two assistants on the day they were tested. A field experiment with 758 consultants found that a model which raised performance on many tasks lowered accuracy on a similar-looking one (Dell'Acqua et al. 2026). Priya's table tells her about meeting summaries and gives her no grounds for a view about any other task.

Six runs per assistant is also too few to support a percentage. It's enough to show a pattern, which is all she needs to make a choice and to know what to keep checking.

Summary

Priya wrote her criterion first, built three cases of rising difficulty, ran each twice on each assistant, and scored the outputs without knowing their source. Her completed table follows.

CaseAssistantRun 1Run 2Note
1. RoutineAPassPass
1. RoutineBPassPass
2. MessyAPassFailRun 2 left out the owner recorded as "M."
2. MessyBPassPass
3. Reversed decisionAFailFailBoth runs gave the 12th as the final date
3. Reversed decisionBPassFailRun 2 listed both dates as decisions

Verdict: Assistant B passed five of six runs and Assistant A passed three, and the deciding failure was Assistant A reporting a reversed decision as final in both runs. Assistant B also failed that case once, so any summary of a meeting where a decision changed still needs a check against the notes before it goes out.

  1. The verdict names the failure that decided it as well as the totals. A count of five against three would mean less if the failures had been trivial.
  2. The second sentence keeps the weaker result in view. The better assistant is not a safe one on the hardest case, and the verdict says what Priya will keep checking.

References

  • Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.

Report an issue with this item

Guided Conversation 12 min

Design a Test for One of Your Tasks

In this conversation you'll pick one task you'd like to hand to an AI assistant and design a small test for it. You'll leave with a three-line test plan: the task, your standard for a good result, and the cases you'll run.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Design a Test for One of Your Tasks (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a coached problem session with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying how to test an AI model on my own tasks. Follow this guidance for the whole conversation.

GOAL
I can design a small test that shows whether a model is fit for a task I care about.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my task.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. Welcome disagreement when I give a reason.
- This is a coached problem. The problem is: design a test for one task of mine. Ask for my own attempt at each part before you give any hint. Give one hint at a time. Don't write the criterion or the test cases for me. Help me sharpen what I write.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for confidential, personal, or student information. I should describe my task and cases in general terms. If I start to share real names or private details, remind me to leave them out.
- Aim for about 12 minutes. Spend most of the time on topics 2 and 3. If my replies are brief, offer one concrete prompt, such as "Think of something you do every week that involves reading and then writing," and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; a test I'd really run matters more than a perfect design; I can ask you to clarify anything. Then ask me to name one recurring task I would like to hand to an AI assistant.

TOPICS, IN ORDER
1. The task. Ask what the task is and what a good result looks like to me. Draw out what I'd be most unhappy to get wrong. Follow up if my description of "good" is vague.
2. The success criterion. Ask me to write one or two sentences saying what a good result must contain. Then ask whether a colleague could apply it and reach the same verdict I would. Help me replace words like "clear" or "useful" with things that can be checked.
3. Four test cases. Ask me to describe, in general terms, one typical case, one hard case, one unusual case, and one where a wrong answer would be costly. Ask what each would tell me. Follow up on whichever is weakest.
4. Closing. Ask me to state my test plan in three lines: the task, the criterion, and the cases, including how many times I'll run each. Tell me I can take the plan into a short optional activity where I run a test like this on an AI assistant.

KEY POINTS TO KEEP ACCURATE
- Method: name the task, write the success criterion before running anything, choose cases that cover typical, hard, unusual, and high-stakes examples, run each case more than once in a new conversation, and score each output pass or fail against the criterion.
- A good criterion is specific enough that someone else could apply it. It should catch both things left out and things made up.
- AI capability is uneven. A model can do well on one task and badly on a similar-looking one, and how hard a task seems to a person doesn't predict which.
- One run isn't enough, because the same request can produce different replies.
- Five to ten cases can show a pattern. They can't give a reliable percentage.
- You can't know how you or any other model would do on my task. Don't predict it. If I ask, say that finding out is the purpose of the test.
- If I ask how you work, explain the general mechanism in one or two sentences and say plainly that you can't inspect your own internals, so your statements about yourself are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my task.
- "The top-ranked model is best for everything": rankings describe performance on other people's tasks, and a model can lead a ranking and still do poorly on mine.
- "If it did well once it will do well again": outputs vary from run to run, so one good result may not repeat.
- "I'll know a good answer when I see it": fluent, confident text reads as competent whether or not it's right, so I need a criterion written in advance.

LIMITS
- Don't recommend or compare any model or product.
- Don't teach how to write or improve prompts. If I ask, say the test uses my request as I'd normally word it.
- Don't run the test for me in this conversation, and don't offer to do my task.
- Don't ask for or accept confidential material.

TO FINISH
After my closing answer, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: run the plan on one assistant and score the outputs; run it on two assistants and judge the outputs without knowing which wrote which; ask a colleague to apply my criterion to one output and compare verdicts; tighten the criterion and try again.
- Restate my plan on its own lines, labeled "My test plan", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Run a Five-Case Test on a Task You Do Often

Overview

You can find out a good deal about whether an AI assistant suits one of your tasks in about fifteen minutes. In this activity you'll write a standard for a good result, run five examples of the task, and score what comes back.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI assistant you already use
  • Five examples of a task you do regularly, with nothing confidential in them. Remove or replace names, personal details, student information, and anything your workplace treats as private. If you can't do that, write five invented examples that resemble the real ones.
  • Somewhere to write a few notes, or any test plan of your own that you've already drafted

Your Task

Build five test cases from a task you do regularly, run them on an AI assistant, and score the results against a criterion you wrote first.

Steps

  1. Write one sentence that says what a good result must contain. Do this before you open the assistant. Make it specific enough that a colleague could apply it, and include what must not appear as well as what must.
  2. Prepare five cases, including one hard and one unusual. Three can be typical examples. Make one an example you found difficult yourself, and one an odd case, such as an unusual format or an input with something missing.
  3. Run each case in a new conversation and save the output. Use the same request, word for word, each time. Paste each output into your notes and number it.
  4. Score each output pass or fail against your sentence, and note why. Compare the output with what you know the right result to be. Don't score on how well it reads. Write a few words beside every failure.

What to Expect

A common result is that the typical cases pass and the hard or unusual case fails, sometimes in a way you wouldn't have noticed without your criterion. It's also possible that all five pass or that most fail. Each outcome is a real finding about this assistant on this task.

Five cases can show a pattern. They can't give you a dependable percentage, and they tell you nothing about other tasks. If one case surprised you, running it a second time in a new conversation will show whether the result holds.

Self-Check

When you're done, check that:

  • You wrote the criterion before running anything
  • You have five scored outputs
  • At least one case was hard or unusual
  • You can say which kind of case the assistant handled worst

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Testing a model on your own tasks

This ungraded knowledge check assesses your understanding of how to find out whether an AI model is fit for a particular task. You'll be asked about the jagged frontier, the limits of public benchmarks, the parts of a personal test set, and the requirements of a fair comparison.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item