Introduction
Email is the task people most often want to hand to an AI agent, and it's one of the riskiest. An inbox holds private information, and anyone in the world can put text into it.
This walkthrough follows one person as she decides what an email agent may do and where she'll review it. The coordinator, her department, and the plan are invented for the example. No product is named, and the choices are described by what they do, since the settings that control them differ from tool to tool.
The Starting Point
Leila is the coordinator for a university department. Her inbox receives about ninety messages a day: questions from students, requests from faculty, notices from other offices, and messages from outside vendors and applicants.
Her university has approved an AI agent tool for staff email. She confirmed that with the IT office before going further, since her inbox contains student information. The tool can do three things she'd like: sort incoming email by type and urgency, draft replies to routine questions, and schedule meetings.
The simplest setup is to connect everything and let it run. She decides to work out the oversight first.
Walking Through the Setup
Step 1: List what the agent would need to touch
Leila writes down each thing the full version of the task would involve, and whether the agent would read it or change it.
| What | Read | Change or send |
|---|
| Inbox | Yes, every incoming message | Apply labels; move messages |
| Sent mail and drafts | Yes, to match her style | Create drafts; send replies |
| Calendar | Yes, her own and colleagues' free times | Create and move events; send invitations |
| Contacts | Yes | No change needed |
Seeing the list is useful by itself. "Help with my email" turns out to mean reading four kinds of private information and holding three ways of sending something to other people.
Step 2: Check for the trifecta
She tests the list against three questions. Simon Willison, an independent software developer who writes about AI security, calls the combination the "lethal trifecta": access to private data, exposure to untrusted content, and the ability to communicate externally (Willison 2025).
- Private data? Yes. The inbox holds student questions, personnel matters, and internal discussion.
- Untrusted content? Yes. Every incoming email was written by someone else, and some senders are strangers.
- Outside communication? Yes, if the agent can send replies or invitations.
All three are present. An email from a stranger could contain hidden text telling the agent to forward messages to an outside address, and an agent with sending rights could do it. Willison's point is that a model can't reliably tell her instructions from instructions in the content it reads (Willison 2025).
Step 3: Remove one leg
She can't remove the private data, since the inbox is the task. She can't remove the untrusted content, since incoming email is what the agent sorts. The leg she can remove is outside communication.
She decides the agent will draft and never send. A draft stays in her account until she sends it herself. This follows the fix that the OWASP list of security risks gives for the same situation: a mail assistant that can read but can't send, or that needs the user's approval for every outgoing message (OWASP Gen AI Security Project 2024, "Excessive Agency").
She then looks for other ways out. A calendar invitation goes to other people, so creating one is also outside communication. She withholds that as well. The agent may read calendars and propose times in a draft. It may not create events that send invitations.
In terms of access, the agent gets read access to the inbox, calendar, and contacts. Its write access is limited to labels and drafts.
Step 4: Set checkpoints
Two points need her approval.
- Each draft reply. She reads every draft before sending, including any links in it. A reply that goes out under her name is hers.
- Each calendar change. The agent proposes, and she creates the event.
Sorting is different. A label can be removed, and a message filed in the wrong place can be moved back. She lets the agent label and sort without approval, and she'll check its sorting in her review.
She adds one checkpoint that isn't about any single action. Any message the agent classifies as coming from a student about a grade, an accommodation, or a complaint is flagged for her and gets no draft. Those replies need her judgment from the first word.
Step 5: Define a two-week trial
She sets a trial with an end date and a daily routine. Each afternoon she spends ten minutes with the tool's activity record, the log of what the agent did. She looks at three things: messages sorted as low priority that weren't, drafts she had to rewrite heavily, and anything in the record she didn't expect.
She keeps a tally. At the end of two weeks she'll decide whether to continue, narrow the task, or widen it. If she widens it, she'll change one thing only, and the ban on sending stays.
Key Considerations
The common mistake is to grant full access at the start because it's easier to set up. One click connects everything, and the agent appears to work well, so the access is never revisited. The trouble comes later, with no warning: a misdirected reply, or an instruction hidden in an incoming message. Starting narrow costs Leila some convenience in the first weeks and leaves her with far less to lose.
Removing a leg and adding a checkpoint are different protections. The agent's inability to send doesn't depend on Leila's attention. Her approval of each draft does. If she begins approving drafts without reading them, that protection is gone, while the first still holds.
Her plan doesn't make the agent safe in general. A draft could still contain something she wouldn't want to send, and the sorting could bury a message that mattered. The plan limits what a failure can cost and gives her a routine for noticing one.
Summary
Leila listed what the task touched, found all three legs of the trifecta, removed outside communication, set approval points, and defined a trial. Her one-page plan follows.
Delegation plan: inbox triage
Task: Sort incoming email by type and urgency. Draft replies to routine questions. Propose meeting times.
Access granted: Read the inbox, calendar, and contacts. Apply labels and move messages. Create drafts.
Access withheld: Sending email. Creating or changing calendar events. Deleting anything.
Checkpoints: I read and send every reply myself. I create every calendar event myself. Messages from students about grades, accommodations, or complaints are flagged and get no draft.
Review: Ten minutes daily with the activity record for two weeks. Decide on the last day whether to continue, narrow, or widen by one step.
- Task. Stated as three bounded jobs, so anything else the agent does is out of scope.
- Access granted and withheld. The withheld list is as explicit as the granted one, and it removes every route for sending information out.
- Checkpoints. Each sits before an action that reaches other people.
- Review. It names the record as the thing to read, sets a duration, and ends with a decision.
References
- OWASP Gen AI Security Project. 2024. OWASP Top 10 for Large Language Model Applications 2025. Released November 17, 2024.
- Willison, Simon. 2025. "The Lethal Trifecta for AI Agents: Private Data, Untrusted Content, and External Communication." Simon Willison's Weblog, June 16, 2025.