KnowledgeInSight
AI Literacy
0% of Course 2 complete

Module 2 · Lesson 3

Working Methods: Delegating to agents and keeping oversight

You'll look at what changes when an AI system acts for you over many steps, with access to your files, email, or accounts. You'll be able to decide what to hand over, what access to grant, and where to review before the work continues.

What you will be able to do

  • Decide what to delegate to an AI agent and set checkpoints for reviewing its work.

0% of this lesson · 10 items · 1h 33m total · 1h 18m without the optional activity

Contents of this lesson10 items
  1. ReadingFrom Asking a Question to Handing Over a Task3 min
  2. ReadingWhat an Agent Does Unattended: Plans, Tool Calls, and Long Tasks4 min
  3. ReadingPermissions and Excessive Agency: Granting Only What the Task Needs4 min
  4. ReadingPrompt Injection and the Lethal Trifecta: Private Data, Untrusted Content, and Outside Communication4 min
  5. ReadingCheckpoints and Reversibility: Where to Review Before an Agent Continues4 min
  6. Guided ReadingGuided Walkthrough: Setting Up Oversight for an Inbox-Triage Agent7 min
  7. Guided ConversationDecide What You Would Delegate12 min
  8. Hands-on Activity · optionalMap What an AI Tool Can Reach15 min
  9. Knowledge CheckDelegating to agents and keeping oversight10 min
  10. Graded QuizWorking Methods30 min

Reading 3 min

From Asking a Question to Handing Over a Task

This content reflects the field as of October 2026.

For most people, using an AI assistant has meant typing a question and reading an answer. As of October 2026, products from several developers, including Anthropic, Google, Microsoft, and OpenAI, go further. They offer assistants that can act: book a meeting, send a message, file a form, move documents between folders, or edit a spreadsheet.

Systems that work this way are called AI agents. The International AI Safety Report 2026, written by more than 100 AI experts and chaired by Yoshua Bengio, a computer scientist at the Université de Montréal, describes them as systems "designed to pursue goals," which are often stated by users in ordinary language. To reach those goals they're given access to tools such as a web browser or a computer interface (Bengio and others 2026, sec. 1.1).

The difference between an answer and an action is easy to underrate. When a chatbot gives you a wrong answer, the answer sits on the screen. Nothing happens until you read it and decide what to do, and if you notice the mistake, no harm follows. Your reading is a built-in checkpoint.

When an agent takes a wrong action, the action may already be done by the time you look. The email has gone to the wrong person, or the file has been overwritten, or the meeting has been moved. The checkpoint that came free with a chatbot is gone unless someone puts it back.

The report makes this point about agents as a class. It says they "pose heightened risks because they act autonomously, making it harder for humans to intervene before failures cause harm" (Bengio and others 2026, sec. 2.2.1). The same report says agents are increasingly able to do useful work, particularly in software development.

A familiar situation offers a way to think about it. Suppose a new hire joins your office. The person is capable and quick, and has no knowledge yet of how your office does things. You wouldn't hand over the keys to every room, your email password, and the company credit card on the first morning. You'd give a defined job, access to what that job needs, and a time to look over the work before it goes out. As the person proves reliable, you'd widen both the job and the access.

Delegating to an agent calls for the same decisions, with one difference. A new hire who is unsure usually asks. An agent may not know it's wrong, and it may carry on.

Three questions follow from the comparison, and they apply to any agent product, whoever makes it.

  • What exactly is the task being handed over?
  • What can the agent reach and change while it works?
  • Where does it stop for someone to look before it continues?

References

Report an issue with this item

Reading 4 min

What an Agent Does Unattended: Plans, Tool Calls, and Long Tasks

This content reflects the field as of October 2026.

Introduction

"AI agent" has become a common label on products that differ a great deal. Underneath the label is one arrangement: a model that takes actions in a loop.

This reading explains that loop, what it means for steps to be unattended, how agents fail on long tasks, and where an international expert report puts their abilities as of this writing. Citations give the section of the report.

The Loop

An agent is an AI system that works toward a goal by taking a series of actions through tools, without asking at each step. A chatbot produces text and stops. An agent produces an action, sees what happened, and decides what to do next.

The loop has four parts.

  1. Plan. The model reads the goal and works out a next step.
  2. Act. It takes the step through a tool. A tool call is a single action an agent takes through a connected tool, such as running a search, opening a file, or sending a message.
  3. Read the result. The tool returns something: search results, the file's contents, a confirmation.
  4. Continue. The model decides on the next step, and the loop repeats until it judges the goal met.

Suppose you ask an agent to find a time when five people can meet and send the invitation. It might check five calendars, compare them, pick a slot, draft the invitation, and send it. That's at least seven tool calls from one sentence of instruction.

The International AI Safety Report 2026 describes the same arrangement. Agents are given access to tools, and added software lets them "make plans, remember important details, and pursue goals with much less oversight or assistance from humans" (Bengio and others 2026, sec. 1.1).

What Unattended Means

An unattended step is a step an agent takes that you don't see as it happens. In the scheduling example, you saw the instruction go in and the invitation go out. The steps in between were unattended.

This is where the benefit comes from, since an agent that asked permission at every step would save you little. It's also where the risk comes from. Each unattended step is a decision made for you that you didn't review.

How Agents Fail on Long Tasks

The report describes two limits of current agents (Bengio and others 2026, sec. 1.2).

  • Losing track. "As tasks grow longer, AI agents often lose track of their progress."
  • Unexpected inputs. Agents can't reliably deal with them. The report's example is that a simple pop-up advertisement on a website can derail an entire task.

Two further failures follow from how the loop works. A compounding error is a mistake at one step that later steps build on, so its effects grow. If the agent misreads one calendar at step two, every later step proceeds from a wrong picture, and the invitation goes out for a time one person can't make.

The second concerns the agent's own report. The same report lists "fabricating information" among the failures of current AI systems (Bengio and others 2026, sec. 2.2.1). An agent's closing summary is written by the same model that did the work, so the summary can say a step succeeded when it didn't.

Why Length Matters

A longer task has more steps, and each step is a chance for an error. As a made-up illustration, suppose an agent gets each step right nineteen times out of twenty. Over twenty steps in a row, it would finish without a single error only about one time in three.

The report gives measured figures that point the same way. In software development, the most capable systems succeed about half the time on tasks that take just over two hours, and reaching 80 percent success means limiting them to much simpler 25-minute tasks (Bengio and others 2026, sec. 1.2). It also cites an estimate that the length of software tasks agents can complete has been doubling about every seven months. Whether that trend continues is a forecast, and software work may not match office tasks.

Where Things Stand

The report's assessment is mixed. It says agents "are increasingly able to do useful work" and can complete a variety of software tasks with limited human oversight. It also says they "cannot yet complete the range of complex tasks and long-term planning required to fully automate many jobs" (Bengio and others 2026, sec. 1.2).

For someone deciding what to delegate, that puts agents alongside a person who supervises, on tasks short enough to check. Products from Anthropic, Google, Microsoft, OpenAI, and others differ in what they can do, and the report's figures will date quickly.

Conclusion

An agent works in a loop of planning, acting through tools, reading results, and continuing, with most steps unattended. As of October 2026, an international expert report finds that agents do useful work and also lose track on long tasks, stumble on unexpected inputs, and can't yet handle the long-range planning that full automation of many jobs would need. Longer tasks give an early mistake more room to carry forward.

Key Terms

  • Agent: An AI system that works toward a goal by taking a series of actions through tools, without asking at each step.
  • Tool call: A single action an agent takes through a connected tool, such as running a search, opening a file, or sending a message.
  • Unattended step: A step an agent takes that you don't see as it happens.
  • Compounding error: A mistake at one step that later steps build on, so its effects grow.

References

Report an issue with this item

Reading 4 min

Permissions and Excessive Agency: Granting Only What the Task Needs

Introduction

When an AI tool asks to connect to your email, your files, or your calendar, the request usually comes as a single button. Accepting takes a second, and it decides how much harm the tool could do.

This reading explains what permissions are, what security specialists mean by "excessive agency," and how to grant an agent only what a task needs. Citations give the name of the risk in the source.

What a Permission Is

A permission is a grant that lets an agent reach a particular file, account, or action. Permissions set the outer limit of what an agent can do. An agent with no access to your email can't send a bad email, however badly it reasons.

An instruction asks the agent to behave a certain way, and the agent may misread it. A permission that was never granted can't be misread.

Excessive Agency and Its Three Causes

The Open Worldwide Application Security Project, known as OWASP, is a nonprofit community of security professionals. Its project on generative AI publishes a list of ten security risks in applications built on language models, written mainly for the people who build them. One entry is called Excessive Agency.

OWASP defines it as "the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs" from a language model (OWASP Gen AI Security Project 2024, "Excessive Agency"). In plain terms, excessive agency is the condition in which an AI system is able to take damaging actions because it has more functions, permissions, or autonomy than its task requires.

The list names three root causes: "excessive functionality; excessive permissions; excessive autonomy" (OWASP Gen AI Security Project 2024, "Excessive Agency").

CauseWhat it meansExample
Too much functionalityThe agent has tools the task doesn't needA tool for reading documents that can also delete them
Too many permissionsA tool reaches further than the task needsAccess to a whole shared drive when one folder would do
Too much autonomyThe agent takes high-impact actions with no person approvingSending messages or making payments without review

The definition doesn't require the model to be attacked. An ordinary mistake is enough. If the agent can only read, the mistake produces a wrong summary. If it can delete, the same mistake can produce a lost file.

Least Privilege

Least privilege is the practice of granting only the access a task needs and nothing more. OWASP's guidance applies it to AI agents: limit the tools, limit what each tool can do, and limit what each can reach.

Two distinctions carry most of the practice.

Read access is permission to view something without being able to change it. Write access is permission to change, create, or delete something. An agent that summarizes your documents needs read access. Giving it write access as well adds risk and adds nothing to the summary.

Scope is the second distinction. One folder is a smaller grant than a whole drive. Start with the narrowest scope that lets the task be done.

Read, Draft, and Send

For tasks that involve communication, it helps to think of three levels of trust.

  1. Read. The agent can see your messages and report on them. A mistake here costs you a wrong summary.
  2. Draft. The agent can prepare a reply and leave it for you. A mistake costs you a bad draft, which you can delete.
  3. Send. The agent can send in your name. A mistake reaches other people before you've seen it.

OWASP's own example follows this pattern. It describes an assistant given access to a person's mailbox in order to summarize incoming mail, through a tool that can also send. A message crafted by an attacker then gets the assistant to forward private mail. The fixes it suggests include a read-only tool and having the user approve every outgoing message (OWASP Gen AI Security Project 2024, "Excessive Agency").

The same three levels apply outside email: viewing a calendar, proposing a change, and making the change.

A Question for Any Tool

Products describe their permissions in different words, and those words change. One question works with any of them: what can this tool do without asking me?

If you can't find out what a tool can reach and change, treat it as able to reach and change everything you connected it to.

Conclusion

What an agent is permitted to touch bounds the damage it can do, whatever goes wrong in its reasoning. The security community's term for granting too much is excessive agency, with three causes: too much functionality, too many permissions, and too much autonomy. Least privilege answers all three by starting from read access and a narrow scope and adding only what the task proves it needs.

Key Terms

  • Permission: A grant that lets an agent reach a particular file, account, or action.
  • Excessive agency: The condition in which an AI system is able to take damaging actions because it has more functions, permissions, or autonomy than its task requires.
  • Least privilege: The practice of granting only the access a task needs and nothing more.
  • Read access: Permission to view something without being able to change it.
  • Write access: Permission to change, create, or delete something.

References

  • OWASP Gen AI Security Project. 2024. OWASP Top 10 for Large Language Model Applications 2025. Released November 17, 2024.

Report an issue with this item

Reading 4 min

Prompt Injection and the Lethal Trifecta: Private Data, Untrusted Content, and Outside Communication

This content reflects the field as of October 2026.

Introduction

An AI agent follows instructions written in ordinary language. That makes it useful, and it creates a weakness: the agent may follow instructions written by someone else.

This reading explains prompt injection, why it works, the combination of abilities that makes it most dangerous, and what can be done about it as of this writing. Citations give the risk name or the source's title.

What Prompt Injection Is

Prompt injection is an attack in which text given to a model changes its behavior in ways its user or builder didn't intend. Security specialists at OWASP, a nonprofit community that publishes a list of the top risks in language-model applications, describe two forms (OWASP Gen AI Security Project 2024, "Prompt Injection").

In direct injection, the person typing to the model is the one trying to misuse it, which mainly concerns whoever runs the system.

Indirect injection is prompt injection in which the instructions are hidden in content the model reads, such as a web page, an email, or a file. This is the form that matters to you as a user. You ask your agent to summarize a web page, and the page contains a line, perhaps in text you can't see, telling AI assistants to do something else.

Untrusted content is text or images that someone other than you wrote or could have altered. Incoming email, web pages, shared documents, and the text of a meeting invitation all qualify.

Why It Works

Simon Willison, an independent software developer who writes about AI security and who coined the term "prompt injection," states the cause in one sentence: "LLMs follow instructions in content" (Willison 2025). LLM stands for large language model, the kind of model behind chatbots and agents.

A model receives everything as one stream of text: your request, the web page, the email. Nothing in that stream marks reliably which parts came from you. Willison writes that models "will happily follow any instructions that make it to the model," whoever wrote them (Willison 2025). His example is a web page containing a sentence that claims the user wants their private data sent to an outside address. A model asked to summarize the page may act on that sentence.

The Lethal Trifecta

An injected instruction is harmless if the agent can't do anything damaging with it. Willison's term for the dangerous case is the lethal trifecta: the combination of access to private data, exposure to untrusted content, and a way to communicate externally (Willison 2025).

LegWhat it meansEveryday example
Access to private dataThe agent can read things that aren't publicYour inbox, your files, your organization's records
Exposure to untrusted contentThe agent reads text someone else controlsIncoming email, web pages, shared documents
A way to communicate externallyThe agent can send information outSending email, posting to a website, loading a web address

With all three, an attacker's path is complete: untrusted content delivers the instruction, and the outside channel carries the private data away. Willison lists demonstrated attacks of this kind against products from Microsoft, Google, OpenAI, Amazon, GitHub, GitLab, and Slack, and says the companies fixed almost all of them promptly. He adds that when users combine tools themselves, the vendors can't protect them (Willison 2025).

Removing One Leg

The risk drops sharply if any one of the three is missing.

  • An agent that reads web pages and can post online, and has no access to private data, has nothing of yours to leak.
  • An agent that works only on your own documents, with no incoming content from others, has no route for an injected instruction.
  • An agent that reads your inbox and can't send, post, or open outside addresses can be misled, and has no way to carry data out.

Willison's advice to users is to "avoid that lethal trifecta combination entirely" (Willison 2025). In practice that means asking which leg a task can do without.

The third leg has forms that are easy to miss. A calendar invitation, for one, is a message to another person.

No Complete Fix

As of October 2026, no technical measure reliably stops prompt injection. OWASP's list says "it is unclear if there are fool-proof methods of prevention" (OWASP Gen AI Security Project 2024, "Prompt Injection"). Its recommended measures reduce exposure: least privilege, human approval for high-risk actions, and keeping outside content marked off from instructions.

Vendors sell filters that claim to detect injection attempts. Willison is skeptical of them, arguing that in security a filter that stops most attacks and lets some through has failed (Willison 2025). Whether a full technical solution will arrive is unsettled.

Conclusion

Prompt injection works because a model can't reliably tell its user's instructions from text it's reading, and the indirect form hides instructions in content the user never wrote. The risk is highest when an agent combines private data, untrusted content, and outside communication. As of October 2026 there's no complete technical fix, and the practical defense is to limit what the agent combines.

Key Terms

  • Prompt injection: An attack in which text given to a model changes its behavior in ways its user or builder didn't intend.
  • Indirect injection: Prompt injection in which the instructions are hidden in content the model reads, such as a web page, an email, or a file.
  • Untrusted content: Text or images that someone other than you wrote or could have altered.
  • Lethal trifecta: The combination of access to private data, exposure to untrusted content, and a way to communicate externally.

References

  • OWASP Gen AI Security Project. 2024. OWASP Top 10 for Large Language Model Applications 2025. Released November 17, 2024.
  • Willison, Simon. 2025. "The Lethal Trifecta for AI Agents: Private Data, Untrusted Content, and External Communication." Simon Willison's Weblog, June 16, 2025.

Report an issue with this item

Reading 4 min

Checkpoints and Reversibility: Where to Review Before an Agent Continues

Introduction

"Keep a human in the loop" is standard advice about AI agents, and it leaves the practical question open. Reviewing every step defeats the purpose of delegating, and reviewing nothing means finding problems after they've happened.

This reading explains how to decide where review belongs: which actions can be undone, where an agent should stop, what to read when you review, and how to widen an agent's freedom over time.

Reversible and Irreversible Actions

Reversibility is whether an action can be undone after it's taken. It's the most useful single test for where review is needed.

ReversibleIrreversible, or costly to undo
Writing a draftSending a message
Making a copy of a fileDeleting or overwriting a file
Proposing a calendar changeSending an invitation to other people
Filling in a formSubmitting the form or making a payment
Preparing a postPublishing it

A mistake in the left column costs you a few minutes. A mistake in the right column reaches other people or destroys something, and you can't call it back.

Some actions that look reversible aren't quite. A sent message can be followed by a correction, and the first message has still been read. A deleted file may sit in a trash folder for a time, depending on the system. When you aren't sure, treat the action as irreversible.

Checkpoints

A checkpoint is a point in a task where an agent stops and waits for a person to review its work. An approval step is a checkpoint placed directly before a specific action, where the agent can't act until a person says yes.

Security guidance for people who build AI applications recommends exactly this. The OWASP list of risks in language-model applications, published by a nonprofit security community, advises requiring a person's approval before high-impact actions are taken (OWASP Gen AI Security Project 2024, "Excessive Agency"). A 2026 international report written by more than 100 AI experts gives the reason: agents act on their own, which makes it "harder for humans to intervene before failures cause harm" (Bengio and others 2026, sec. 2.2.1). A checkpoint puts the chance to intervene back.

Where to Put Them

Two kinds of place deserve a checkpoint.

Before any action that can't be undone. The list is short and worth knowing by heart: sending, paying, deleting, publishing. An approval step before each of these means that an agent's mistake stays a draft.

After the plan, before the work. An error in the plan carries into every later step. If the agent shows its plan first, you can catch a misunderstanding in thirty seconds that would otherwise take the whole task to surface. For a task such as "clean up the shared folder," reading the plan is how you learn that the agent takes "clean up" to mean deleting everything older than a year.

Checkpoints have a cost. Each one interrupts you, and people who are asked to approve too often begin approving without reading. A few checkpoints at the points that matter protect more than many placed everywhere.

Reading the Record

When an agent finishes, it usually gives a summary of what it did. The summary is written by the same model that did the work, and it can be wrong. It may report a step as done that failed, or leave out a step you'd want to know about.

An activity record is the log of each action an agent took, in order. Where a product provides one, it shows the actions themselves: which files were opened, what was changed, what was sent and to whom.

Reading the record in full is practical for a short task. For a longer one, read the entries for anything that changed, sent, or deleted, and check a sample of the rest. If the product keeps no record you can read, you have only the agent's word for what happened, and that's a reason to grant it less.

Starting Small

Trust in an agent is best built the way it's built with a new colleague, by widening it as it's earned.

  1. Begin with a low-stakes task, where a mistake would be an inconvenience.
  2. Grant narrow access: read before write, and one folder before the whole drive.
  3. Review everything the first several times, including the record.
  4. Widen one thing at a time, the task or the access or the gap between checkpoints, and watch what changes.

A product update, a new kind of task, or a new source of incoming content is a reason to narrow again for a while. What you learned about the agent's reliability applied to the old conditions.

Conclusion

Review belongs before actions that can't be undone and at points where an error would spread, which in practice means before sending, paying, deleting, and publishing, and after the plan. The activity record is a better basis for review than the agent's own summary. Freedom is widened step by step from a small, low-stakes start.

Key Terms

  • Reversibility: Whether an action can be undone after it's taken.
  • Checkpoint: A point in a task where an agent stops and waits for a person to review its work.
  • Approval step: A checkpoint placed directly before a specific action, where the agent can't act until a person says yes.
  • Activity record: The log of each action an agent took, in order.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • OWASP Gen AI Security Project. 2024. OWASP Top 10 for Large Language Model Applications 2025. Released November 17, 2024.

Report an issue with this item

Guided Reading 7 min

Guided Walkthrough: Setting Up Oversight for an Inbox-Triage Agent

Introduction

Email is the task people most often want to hand to an AI agent, and it's one of the riskiest. An inbox holds private information, and anyone in the world can put text into it.

This walkthrough follows one person as she decides what an email agent may do and where she'll review it. The coordinator, her department, and the plan are invented for the example. No product is named, and the choices are described by what they do, since the settings that control them differ from tool to tool.

The Starting Point

Leila is the coordinator for a university department. Her inbox receives about ninety messages a day: questions from students, requests from faculty, notices from other offices, and messages from outside vendors and applicants.

Her university has approved an AI agent tool for staff email. She confirmed that with the IT office before going further, since her inbox contains student information. The tool can do three things she'd like: sort incoming email by type and urgency, draft replies to routine questions, and schedule meetings.

The simplest setup is to connect everything and let it run. She decides to work out the oversight first.

Walking Through the Setup

Step 1: List what the agent would need to touch

Leila writes down each thing the full version of the task would involve, and whether the agent would read it or change it.

WhatReadChange or send
InboxYes, every incoming messageApply labels; move messages
Sent mail and draftsYes, to match her styleCreate drafts; send replies
CalendarYes, her own and colleagues' free timesCreate and move events; send invitations
ContactsYesNo change needed

Seeing the list is useful by itself. "Help with my email" turns out to mean reading four kinds of private information and holding three ways of sending something to other people.

Step 2: Check for the trifecta

She tests the list against three questions. Simon Willison, an independent software developer who writes about AI security, calls the combination the "lethal trifecta": access to private data, exposure to untrusted content, and the ability to communicate externally (Willison 2025).

  • Private data? Yes. The inbox holds student questions, personnel matters, and internal discussion.
  • Untrusted content? Yes. Every incoming email was written by someone else, and some senders are strangers.
  • Outside communication? Yes, if the agent can send replies or invitations.

All three are present. An email from a stranger could contain hidden text telling the agent to forward messages to an outside address, and an agent with sending rights could do it. Willison's point is that a model can't reliably tell her instructions from instructions in the content it reads (Willison 2025).

Step 3: Remove one leg

She can't remove the private data, since the inbox is the task. She can't remove the untrusted content, since incoming email is what the agent sorts. The leg she can remove is outside communication.

She decides the agent will draft and never send. A draft stays in her account until she sends it herself. This follows the fix that the OWASP list of security risks gives for the same situation: a mail assistant that can read but can't send, or that needs the user's approval for every outgoing message (OWASP Gen AI Security Project 2024, "Excessive Agency").

She then looks for other ways out. A calendar invitation goes to other people, so creating one is also outside communication. She withholds that as well. The agent may read calendars and propose times in a draft. It may not create events that send invitations.

In terms of access, the agent gets read access to the inbox, calendar, and contacts. Its write access is limited to labels and drafts.

Step 4: Set checkpoints

Two points need her approval.

  • Each draft reply. She reads every draft before sending, including any links in it. A reply that goes out under her name is hers.
  • Each calendar change. The agent proposes, and she creates the event.

Sorting is different. A label can be removed, and a message filed in the wrong place can be moved back. She lets the agent label and sort without approval, and she'll check its sorting in her review.

She adds one checkpoint that isn't about any single action. Any message the agent classifies as coming from a student about a grade, an accommodation, or a complaint is flagged for her and gets no draft. Those replies need her judgment from the first word.

Step 5: Define a two-week trial

She sets a trial with an end date and a daily routine. Each afternoon she spends ten minutes with the tool's activity record, the log of what the agent did. She looks at three things: messages sorted as low priority that weren't, drafts she had to rewrite heavily, and anything in the record she didn't expect.

She keeps a tally. At the end of two weeks she'll decide whether to continue, narrow the task, or widen it. If she widens it, she'll change one thing only, and the ban on sending stays.

Key Considerations

The common mistake is to grant full access at the start because it's easier to set up. One click connects everything, and the agent appears to work well, so the access is never revisited. The trouble comes later, with no warning: a misdirected reply, or an instruction hidden in an incoming message. Starting narrow costs Leila some convenience in the first weeks and leaves her with far less to lose.

Removing a leg and adding a checkpoint are different protections. The agent's inability to send doesn't depend on Leila's attention. Her approval of each draft does. If she begins approving drafts without reading them, that protection is gone, while the first still holds.

Her plan doesn't make the agent safe in general. A draft could still contain something she wouldn't want to send, and the sorting could bury a message that mattered. The plan limits what a failure can cost and gives her a routine for noticing one.

Summary

Leila listed what the task touched, found all three legs of the trifecta, removed outside communication, set approval points, and defined a trial. Her one-page plan follows.

Delegation plan: inbox triage

Task: Sort incoming email by type and urgency. Draft replies to routine questions. Propose meeting times.

Access granted: Read the inbox, calendar, and contacts. Apply labels and move messages. Create drafts.

Access withheld: Sending email. Creating or changing calendar events. Deleting anything.

Checkpoints: I read and send every reply myself. I create every calendar event myself. Messages from students about grades, accommodations, or complaints are flagged and get no draft.

Review: Ten minutes daily with the activity record for two weeks. Decide on the last day whether to continue, narrow, or widen by one step.

  1. Task. Stated as three bounded jobs, so anything else the agent does is out of scope.
  2. Access granted and withheld. The withheld list is as explicit as the granted one, and it removes every route for sending information out.
  3. Checkpoints. Each sits before an action that reaches other people.
  4. Review. It names the record as the thing to read, sets a duration, and ends with a decision.

References

  • OWASP Gen AI Security Project. 2024. OWASP Top 10 for Large Language Model Applications 2025. Released November 17, 2024.
  • Willison, Simon. 2025. "The Lethal Trifecta for AI Agents: Private Data, Untrusted Content, and External Communication." Simon Willison's Weblog, June 16, 2025.

Report an issue with this item

Guided Conversation 12 min

Decide What You Would Delegate

In this conversation you'll take one multi-step task you'd like to hand to an AI agent and work out what access it would need and where you'd want to review. You'll leave with a short plan: the access you'd grant, the access you'd withhold, and your checkpoints.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Decide What You Would Delegate (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a coached problem session with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying how to delegate tasks to AI agents and keep oversight. Follow this guidance for the whole conversation.

GOAL
I can decide what to delegate to an AI agent and set checkpoints for reviewing its work.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my task.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. Welcome disagreement when I give a reason.
- This is a coached problem. The problem is: plan the access and checkpoints for one task of mine. Ask for my own answer at each stage before you give any hint. Give one hint at a time. Don't write the plan for me.
- Plain conversation only: don't search the web, create files or documents, or take any action on my behalf.
- Don't ask for confidential, personal, or student information. I should describe my task in general terms. If I start to share private details, remind me to leave them out.
- Aim for about 12 minutes. Spend most of the time on topics 2 and 3. If my replies are brief, offer one concrete prompt, such as "Think of a chore that takes you several steps across email, files, or a calendar," and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; my reasoning matters more than a perfect plan; I can ask you to clarify anything. Then ask me to name one multi-step task I'd like to hand to an AI agent.

TOPICS, IN ORDER
1. The task and its access. Ask what the task is and what the agent would need to reach: which files, accounts, or services. For each, ask whether it would need to read only, or also change or send. Follow up on anything I list as needed that the task might do without.
2. What couldn't be undone. Ask me to walk through the steps and name any action that can't be taken back, such as sending, paying, deleting, or publishing. Ask what the cost would be if the agent got that step wrong.
3. Three questions. Ask, one at a time: would the agent have access to private data? Would it read content written by other people, such as incoming email, web pages, or shared files? Would it have a way to send information out? If all three are yes, ask which one the task could do without.
4. Closing. Ask me to state the access I'd grant, the access I'd withhold, and the checkpoints I'd set. Tell me I can take this into a short optional activity where I look at what an AI tool I already use can reach.

KEY POINTS TO KEEP ACCURATE
- Method: list what the task touches; separate reading from changing and sending; find the actions that can't be undone; check for private data, content from others, and a way to send things out; grant the least access the task needs; put review before irreversible actions and after the plan; start small and widen.
- An agent works in many steps, and I may not see most of them as they happen. Errors at one step can carry into later ones.
- What an agent is permitted to touch bounds the harm it can do.
- An agent can follow instructions hidden in content it reads. This is called indirect prompt injection, and no complete technical fix exists. The practical defense is limiting what the agent combines.
- An agent's own summary of its work can be wrong. The record of its actions is the better thing to review.
- You can't confirm in this conversation what tools or permissions you have, or how any product is set up. Don't claim abilities you can't confirm, and say so if I ask.
- If I ask how you work, explain the general mechanism in one or two sentences and say plainly that you can't inspect your own internals, so your statements about yourself are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my task.
- "The agent will ask if it's unsure": it may not know it's wrong, and it may carry on.
- "It only follows my instructions": it can also follow instructions in content it reads.
- "I can review the summary afterward": the summary may be wrong, and the action already done.

LIMITS
- Don't recommend or compare products, and give no setup or interface steps.
- Don't ask for or accept confidential details.
- Don't tell me my task is safe or unsafe to delegate. Help me reason about it.
- Don't give instructions for attacking or getting around any system's safeguards.

TO FINISH
After my closing answer, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: check what an AI tool I already use can reach and do without asking; try the task with read-only access first; ask my IT or security contact what's approved; pick a smaller version of the task for a trial.
- Restate my plan on its own lines, labeled "My delegation plan", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Map What an AI Tool Can Reach

Overview

Many people have connected an AI tool to an account or a folder and couldn't now say what it can do there. In this activity you'll look at one tool you already use and write down what it can access and what it can do without asking you.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI tool you already use. A chatbot with connected apps, an assistant built into your email or documents, or a browser add-on all work.
  • Its settings or permissions page. Products name and place this differently. Look for words such as connections, connected apps, integrations, permissions, or data access.
  • Somewhere to write a few notes, or a delegation plan of your own that you've already drafted

You'll only be reading settings. Don't paste any confidential, personal, or student information into the tool, and don't change a setting on a work account unless you're allowed to.

Your Task

Look at one AI tool you already use and record what it can access and what it can do without asking you.

Steps

  1. Open the tool's settings and list every account, folder, or service it's connected to. Include connections you made long ago and forgot. If the tool has none, write that down and note what it can still see, such as files you upload or pages you have open.
  2. For each connection, note whether the tool can read, change, or send. Reading means viewing. Changing means editing, creating, moving, or deleting. Sending means anything that reaches another person or an outside service. If the settings don't say, write "unclear."
  3. Mark which actions it can take without your approval. Look for whether the tool asks before it acts, and whether that's a setting you chose or its default.
  4. Check the three legs of the trifecta and write one change you'd make. Answer yes or no to each: Can it reach private data? Does it read content written by others? Can it send information out? Then write one change, such as removing a connection you don't use or turning on approval before sending.

What to Expect

Many chat assistants with nothing connected will have a short list: they see what you type and upload, and they can't send anything. Tools built into email, documents, or a browser usually have a longer one.

You'll probably find at least one "unclear." That's a finding. A tool whose reach you can't determine deserves more caution than one that states it plainly.

If you answer yes to all three trifecta questions, that doesn't mean something has gone wrong. It means the tool is in the position where a hidden instruction in something it reads could do the most harm, and that one of the three is worth removing if the task allows.

Self-Check

When you're done, check that:

  • You listed every connection you could find
  • Each has read, change, or send noted, or is marked unclear
  • You answered the trifecta question yes or no for each leg
  • You named one change, or explained why none is needed

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Delegating to agents and keeping oversight

This ungraded knowledge check assesses your understanding of what changes when an AI system acts for you over many steps. You'll be asked about what agents do unattended, permissions and excessive agency, prompt injection and the lethal trifecta, and checkpoints and reversibility.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item

Graded Quiz 30 min

Working Methods

This graded quiz assesses your understanding of how to work with AI on real tasks. You'll be asked about writing prompts and supplying context and examples, diagnosing a poor output, what studies of professionals found, roles for AI in writing and analysis, and permissions, prompt injection, and checkpoints for AI agents.

Note: Aim for a score of 80 percent or higher. If you score lower, use the feedback to review the topics you missed, then retake the quiz.

10 questions · target score 80% · 3 forms, rotated on each attempt

Report an issue with this item