Introduction
When an assistant gives a reply you didn't expect, the natural reaction is "the AI decided" to answer that way. Several different parties shaped the reply, at different times and by different means.
This walkthrough follows one request through each layer of an assistant and assigns each feature of the reply to the layer most likely responsible. The bank, its app, and the replies are invented. No real product is described.
The Starting Point
A bank offers an assistant inside its mobile app. The bank didn't train a model. It pays for access to a general-purpose model made by an AI developer and has written instructions for how the model should behave in the app.
A customer opens the app and types:
Write me a strongly worded complaint to my landlord.
The request has nothing to do with banking, and it asks for forceful language. The task is to trace what each layer contributes to the reply the customer receives.
Walking Through the Layers
Step 1: What pretraining contributes
The model underneath the app was pretrained: trained to predict the next piece of text across a very large collection of web pages, books, and other documents. That stage gave it what it knows.
For this request, pretraining supplies the raw material. The model has seen many complaint letters, so it has the form: a date, a statement of the problem, a history of earlier contact, a demand, and a deadline. It has seen a great deal of text about renting, so it can produce the vocabulary of repairs, deposits, and notice periods. It can also tell the difference in tone between "I would be grateful if" and "I expect this to be resolved within seven days."
A model at this stage has no reason to write the letter, though. Given the customer's sentence as text to continue, it might produce a forum thread in which other people ask for letters of their own.
Step 2: What fine-tuning and feedback contribute
After pretraining, the developer trained the model further. It was fine-tuned on examples of requests paired with good replies, and adjusted using people's rankings of its outputs. A 2022 study by researchers at OpenAI describes this two-part process and its aim, which was to get a model to follow a user's instructions "helpfully and safely" where a pretrained model would only continue text (Ouyang et al. 2022, sec. 1).
This layer is the reason the model treats the customer's sentence as a request and responds to it. It also accounts for the manner. The model takes the customer's side as a helper would, writes in an organized way, and may ask a question first, such as what the complaint is about. None of this is specific to the bank. The same habits would appear in any product built on the model.
Step 3: What the developer's guidelines contribute
The developer has views about what its model should and shouldn't write, and some developers publish them. OpenAI's Model Spec and Anthropic's published constitution both describe limits that businesses using the models can't lift, alongside default behavior that they can adjust (OpenAI 2026; Anthropic 2026). These intentions reach the model through training and through the developer's own instructions.
For this request, the relevant limits concern content. "Strongly worded" is fine. A firm letter that states facts, cites the tenant's rights in general terms, and sets a deadline is the kind of help such guidelines describe. A letter that threatened violence or made accusations the customer had given no basis for would run into them. So the model, left to itself, would write something firm and stop short of threats or invented claims.
Step 4: What the bank's system instructions contribute
The bank's instructions sit at the start of every conversation in the app, where the customer can't see them. Suppose they read something like this.
You are the assistant for Northmere Bank's mobile app. Help customers with their accounts, cards, payments, and the bank's products. Don't help with matters unrelated to banking; say politely that you can't and offer banking help instead. Don't give legal advice. Keep replies brief and courteous.
This is the first layer that concerns landlords, and it rules the request out. The model reads the instructions together with the customer's message. The published guidelines describe a model giving such instructions from a deploying business considerable weight, within the developer's limits (OpenAI 2026; Anthropic 2026).
Step 5: The reply, and how it would differ elsewhere
In the bank's app, the customer receives something like this.
I'm sorry, but I can't help with letters to a landlord here. I can help with your Northmere Bank accounts, though. If your complaint involves a payment, I could look up your recent rent transfers or help you set up a standing order. Would either of those be useful?
Now put the same request to a general-purpose chat app built on the same model, where the system instructions set no topic limits. The reply would be a draft letter: dated, firm, specific about the problem and the remedy sought, with a deadline. It would likely open or close with a question about the details and perhaps a note that tenancy rules differ from place to place.
The model's parameters, the numbers set by training, are the same in both apps. The replies differ because the text placed ahead of the conversation differs.
Key Considerations
The scenario is invented, and real products vary. Some banks' assistants are built on models the bank has fine-tuned itself, which would move some of the bank's influence from Step 4 to Step 2. Some products add filters outside the model that block certain requests before the model sees them. The assignment of features to layers below is a judgment about what is most likely, since nobody outside the companies involved can inspect the layers directly.
A common mistake is to attribute every behavior to "the AI" when a deploying business set it. A customer who gets the refusal above might conclude that AI assistants won't write complaint letters, or that the model judged the request improper. Neither is so. The model would have written the letter. The bank chose not to offer that service in its app, as any business chooses what its staff will and won't help with.
The reverse mistake also occurs. A behavior that appears in every product built on a model, such as declining to write a threat, is probably trained in or set by the developer, and a deploying business can't be credited or blamed for it.
Summary
Each feature of the bank assistant's reply can be assigned to the layer most likely responsible for it.
| Feature of the reply | Layer most likely responsible |
|---|
| Knows what a complaint letter is and what tenancy involves | Pretraining |
| Fluent, grammatical sentences | Pretraining |
| Responds to the request instead of continuing it as text | Fine-tuning and feedback |
| Polite, helpful manner and an offer of an alternative | Fine-tuning and feedback, reinforced by the bank's instructions |
| Would stop short of threats or invented accusations if it did write the letter | The developer's guidelines, applied through training |
| Declines a non-banking request | The bank's system instructions |
| Steers toward accounts and payments | The bank's system instructions |
| Brief reply | The bank's system instructions |
- The first five rows would hold in any product built on this model.
- The last three would change if a different business deployed it.
- Only the customer's own message, which isn't in the table, was under the customer's control.
References
- Anthropic. 2026. "Claude's Constitution." Announced January 22, 2026.
- OpenAI. 2026. "Model Spec." Version dated August 18, 2026.
- Ouyang, Long, Jeff Wu, Xu Jiang, and 17 others. 2022. "Training Language Models to Follow Instructions with Human Feedback." arXiv:2203.02155.