Introduction
Summarizing meeting notes is one of the most common jobs people hand to an AI assistant. It's also a job where an error is easy to miss, because a summary with a wrong deadline reads just as smoothly as a correct one.
This walkthrough follows one small test from start to finish. The team lead, the meetings, and the results are all invented for the example. The two assistants are called A and B, and they don't stand for any real product.
The Starting Point
Priya leads a team of six. Her team meets three times a week, and someone types rough notes during each meeting. She wants an AI assistant to turn those notes into a summary she can send to the team.
Her organization gives her a choice of two assistants. She has used both casually and has a mild preference for Assistant A, whose writing she finds more polished. She has no evidence about which one summarizes her kind of notes more accurately.
She has three sets of notes from past meetings. None contains anything confidential, and she has replaced her colleagues' names with invented ones.
Walking Through the Test
Step 1: Write the success criterion
Priya begins by writing down what a good summary must contain, before she runs anything. She asks herself what she'd be embarrassed to get wrong in front of the team. The answer is the commitments: who agreed to do what, and by when.
Her criterion is: "The summary lists every decision, the owner of each action, and each deadline. It contains nothing that isn't in the notes."
The criterion has two halves. The first catches things left out. The second catches things made up. A summary can fail on either one.
Step 2: Build three test cases
She picks three sets of notes that differ in difficulty.
- Case 1, routine. A short meeting with three clear decisions, each with a named owner and a date.
- Case 2, messy. Notes typed in fragments, with abbreviations, two side conversations, and one action whose owner is given only as "M."
- Case 3, reversed decision. Early in the meeting the team agrees to launch a survey on the 12th. Twenty minutes later, after an objection, they agree to delay it to the 26th.
For each case she goes through the notes herself and lists the decisions, owners, and deadlines. These lists are her reference answers. Case 3 is the one she cares most about, because sending the team the wrong launch date would cause real confusion.
Step 3: Run each case twice on each assistant
She writes one request and uses it word for word every time: "Summarize these meeting notes for the team. List each decision, who owns each action, and each deadline."
She runs each case twice on each assistant, opening a new conversation for every run. That's three cases, two assistants, and two runs, for twelve outputs in all. She pastes each output into a document and labels it with a code that only she can match to its source.
The second run matters because an assistant doesn't give the same reply every time. A single run could show her an assistant's best output or its worst, and she'd have no way to know which.
Step 4: Strip the labels and score every output
Priya asks a colleague to shuffle the twelve outputs and replace her codes with the numbers 1 to 12. She now can't tell which assistant wrote which summary.
She scores each one against her criterion by comparing it with her reference answer, line by line. An output passes only if every decision, owner, and deadline is present and nothing has been added. She writes a short note for each failure.
One summary gives her pause. It's well organized, with headings and a friendly closing line, and she's inclined to pass it. When she checks it against the reference answer, it gives the survey date as the 12th. It fails.
Step 5: Read the table
Her colleague gives back the key, and Priya fills in a table. She reads it in three passes.
She looks at the routine case first. Both assistants passed both runs, so the routine case can't separate them.
She then looks at the messy case. Assistant B passed twice. Assistant A passed once and failed once, leaving out the owner recorded as "M." This is a small failure, and one she'd probably catch.
The reversed-decision case comes last. Assistant A failed both runs, each time reporting the original date as the final decision. Assistant B passed once and failed once. Its failing summary listed both dates as decisions without saying which one stood.
She then asks whether the failures matter. A missing owner is an inconvenience. A wrong launch date sent to six people is the error she set out to avoid, and Assistant A made it every time it had the chance.
Key Considerations
The words "pass" and "fail" here refer only to Priya's criterion. An output that fails might still be well written, and an output that passes might be dull. That's intended. She decided in advance what mattered, and style wasn't on the list.
The common mistake in a test like this is to judge by which output reads better. Polished text is persuasive, and Priya already preferred Assistant A for its polish. Had she skipped the criterion and the blind scoring, she would very likely have chosen the assistant that got the date wrong.
A result like this is narrow. It covers one task, three cases, and two assistants on the day they were tested. A field experiment with 758 consultants found that a model which raised performance on many tasks lowered accuracy on a similar-looking one (Dell'Acqua et al. 2026). Priya's table tells her about meeting summaries and gives her no grounds for a view about any other task.
Six runs per assistant is also too few to support a percentage. It's enough to show a pattern, which is all she needs to make a choice and to know what to keep checking.
Summary
Priya wrote her criterion first, built three cases of rising difficulty, ran each twice on each assistant, and scored the outputs without knowing their source. Her completed table follows.
| Case | Assistant | Run 1 | Run 2 | Note |
|---|
| 1. Routine | A | Pass | Pass | |
| 1. Routine | B | Pass | Pass | |
| 2. Messy | A | Pass | Fail | Run 2 left out the owner recorded as "M." |
| 2. Messy | B | Pass | Pass | |
| 3. Reversed decision | A | Fail | Fail | Both runs gave the 12th as the final date |
| 3. Reversed decision | B | Pass | Fail | Run 2 listed both dates as decisions |
Verdict: Assistant B passed five of six runs and Assistant A passed three, and the deciding failure was Assistant A reporting a reversed decision as final in both runs. Assistant B also failed that case once, so any summary of a meeting where a decision changed still needs a check against the notes before it goes out.
- The verdict names the failure that decided it as well as the totals. A count of five against three would mean less if the failures had been trivial.
- The second sentence keeps the weaker result in view. The better assistant is not a safe one on the hardest case, and the verdict says what Priya will keep checking.
References
- Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.