Deliverable Overview
You'll write a one-page protocol of 400 to 600 words for one task you do regularly. A protocol, as the word is used here, is a short written set of rules you follow each time you use AI on that task.
The page has six parts: the task and its standard, the test and what it showed, your verification steps, your division of labor, your rules for disclosure and data, and one skill you mean to keep. A small results table sits on the page and doesn't count toward the word range.
The test comes first because AI performance is uneven across tasks that look alike. In a field experiment with 758 consultants, a test with real workers randomly assigned to work with or without AI, those using a model produced work rated roughly 30 percent higher in quality on tasks within its ability. On one task beyond its ability, which looked like ordinary consulting work, they were 19 percentage points less likely to reach the correct answer than colleagues working alone (Dell'Acqua et al. 2026). Running your own task is the direct way to learn which kind it is.
Nobody collects or grades the protocol. You check it against the success criteria below. Nothing here is legal or employment advice.
Detailed Requirements
Use non-sensitive material throughout. Nothing confidential, no personal details about other people, and no student information goes into the AI assistant or onto the page. Where a real example contains any of these, replace the details with invented ones.
The task and the standard
Name one recurring task in a sentence, such as summarizing meeting notes, drafting a newsletter to parents, or sorting survey comments into themes. Then give your success criterion: a statement, written before the test, of what a good result must contain. A colleague should be able to apply it and reach your verdict. "Every decision, its owner, and its deadline appear, and nothing appears that wasn't in the notes" can be scored. "A clear, useful summary" can't. Both the task and the criterion belong in the first three lines of the page.
The test
Report a test of at least four test cases. A test case is made up of an input, a request, and a description of what a good result looks like. Include a typical case, a hard one, and an unusual one, such as an input with something missing. Run each case twice, in a new conversation each time and with the same request word for word, because a model's replies vary from one run to the next. Score every output pass or fail against your criterion.
Put the results in a small table. Here is an invented example for a meeting-summary task.
| Case | Kind | Run 1 | Run 2 | Note |
|---|
| Weekly team notes | Typical | Pass | Pass | |
| Long notes on two projects | Hard | Pass | Fail | Run 2 left out one deadline |
| Notes with no owner named | Unusual | Fail | Fail | Supplied an owner both times |
| Notes with a budget decision | High-stakes | Pass | Pass | |
Follow the table with one sentence stating the decision the results support. For this table it might read: "I'll use the assistant for routine notes, check every owner and deadline against my notes, and write the summary myself when an owner is missing."
Verification steps
List the kinds of checkable claim your task's outputs contain. A checkable claim is a statement that could be shown true or false by looking something up. Common kinds are named sources, quotations, numbers and dates, and statements about a document you supplied. For each kind, write the minimum check you'll do every time.
Fluent output can be wrong. NIST, the US government's standards agency, lists this risk as "confabulation," meaning false content that is stated confidently (NIST 2024, "Confabulation"). Asking the assistant to confirm its own answer doesn't count as a check.
The division of labor
State your division of labor: the decision about which parts of the task you do and which parts you give to AI. The part that stays with you should include any judgment you'd be asked to defend.
Disclosure and data rules
Write one rule for disclosure, which means telling the people who rely on a piece of work that AI was used and how. Write one rule for data, saying what may and may not go into the tool.
Tie each rule to its source. The disclosure rule names the policy, contract, or syllabus it comes from. If nothing is written, say so, and say whom you asked. The data rule names the kind of account you use: a personal account, one your organization provides, or a model that runs on your own machine. It also names the terms or the approval that covers that account.
One skill to keep
End with one sentence naming a skill this task exercises that you mean to keep, and the habit that protects it. Attempting the task yourself before you ask for help is one such habit.
Success Criteria
Your protocol is complete when:
- The task and the success criterion are stated in the first three lines
- The test has at least four cases, each run twice, with results in a table
- The decision follows from the results
- Each kind of claim has a minimum check
- The protocol says what AI does and what you do
- The disclosure and data rules name the policy or terms they come from
- One skill to keep is named, with a habit that protects it
- The protocol fits on one page and contains nothing confidential
Step-by-Step Guidance
The times below add up to an hour.
Step 1: Choose one recurring task and write what a good result must contain
Pick a task small enough that one run takes a minute or two. Write the criterion before you open the assistant, and say what must not appear as well as what must. Allow about 5 minutes.
Step 2: Build and run a small test: at least four cases, each run twice, scored against your criterion
Work you've already judged is the quickest source of cases, once names and private details are replaced. Save each output and score it against the criterion, setting aside how well it reads. Write a few words beside every failure, since those notes show how the assistant fails. Allow about 20 minutes.
Step 3: Decide from the results whether and how to use AI for the task
Read the table before you write any rule. A case that passed once and failed once is unstable. Your protocol has to cover that kind of case with a check or keep it for you to do. If most cases failed, "I won't use AI for this task yet" is a sound decision. Allow about 5 minutes.
Step 4: Write your verification steps and your division of labor
Take the checks from your failures. If the assistant supplied an owner nobody named, the check is to compare every owner with your notes. Then cut any check you wouldn't do on your busiest day, and narrow the AI's part of the task to match. Allow about 10 minutes.
Step 5: Find the policy and the data terms that apply, and write your disclosure and data rules from them
Look for your employer's policy, a client contract, or a syllabus. Read it for whether AI use is allowed, for what, whether it must be disclosed, and to whom. Then confirm which kind of account you're signed in to, and find what its terms or settings say about storing conversations and using them for training. Allow about 10 minutes.
Step 6: Add the skill you mean to keep, then cut the protocol to one page
Write the skill sentence. Then cut explanation and keep rules, so that each remaining sentence states something you do or don't do. Read the page once for any name, figure, or detail that shouldn't be there. Allow about 10 minutes.
Resources
- An AI assistant you already use
- Four or more examples of your task, with nothing confidential, personal, or student-related in them
- The AI policy, contract clause, or syllabus statement that applies to you, or a note of whom you asked
- The data terms or settings page for the account you use
- Any notes, test plans, or practice work of your own on testing, checking, or disclosure
A protocol doesn't need citations. If you name a source, you're welcome to cite it by author, year, and title.
Common Pitfalls
- Writing rules before running the test. Rules written first describe the assistant you expect. The test may show failures you didn't predict, and your checks have to cover those.
- A criterion too vague to score against. If two people could disagree on whether an output passed, rewrite the criterion.
- Verification steps you won't do. Readers of the page will assume every promised check happened, so promise only the ones you'll do.
- Data rules that ignore which kind of account you're using. A personal account and one your organization provides can look the same on screen and come under different terms.
- Including confidential material in the protocol itself. Describe your test cases by kind, such as "notes from a long meeting," and leave out client names, real figures, and anything about identifiable people.
References
- Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.
- National Institute of Standards and Technology. 2024. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Gaithersburg, MD: NIST, July 2024.