A comparison of the methods that constrain agent behavior, and why only some of them hold up once the model itself is compromised
In our piece on agent architecture, we drew a hard line between the local agent (a thin client with no judgment) and the brain (a stateless model that decides what happens next). That split explains how agents work. It doesn't say anything about what they're allowed to do. Those are different questions, and the gap between them is where most of the damage in agentic systems actually happens.
An agent that can call a tool is able to call it. Whether it should, in this context, for this user, on this data, is a separate decision, and it's tempting to assume the model itself is where that decision belongs. Our post on prompt injection already showed why that assumption fails: if the model reads an instruction from a tool result, it has no structural way to know whether that instruction came from the user or from an attacker. A control that only works by convincing the model to behave isn't a control. It's a suggestion.
So the real engineering question is: where do you put a rule that holds regardless of what the model decides? There are several credible answers, and they sit at different layers of the stack. Here's a comparison of the main ones, what each actually stops, and where each one runs out of road.
Method 1: Prompt-level instructions
The most common and weakest control is telling the model what not to do: system prompt rules like "never delete files without confirmation" or "don't send data to external domains." This is worth doing, and it's not nothing. Well-trained models follow these instructions the large majority of the time.
But this control lives entirely inside the probabilistic layer. It's enforced by the same component that reads untrusted tool results as part of its context, which means it's exactly the layer that prompt injection targets. An instruction embedded in a webpage or a file the agent reads competes for influence with the system prompt on equal footing: both are just tokens. Prompt-level rules raise the bar for an attacker. They don't set a floor.
Method 2: MCP tool annotations
The Model Context Protocol gives server authors a way to declare behavioral properties of each tool directly in its schema: readOnlyHint, destructiveHint, idempotentHint, and openWorldHint. A tool tagged destructiveHint: true is telling any client that calls it: this can cause damage that isn't easily undone.
{
"name": "delete_customer_record",
"description": "Permanently removes a customer record from the CRM",
"annotations": {
"readOnlyHint": false,
"destructiveHint": true,
"idempotentHint": false,
"openWorldHint": true
}
}
Clients and governance layers can use these hints to decide how cautiously to treat a given call, and the spec defaults to the most pessimistic posture when a tool ships with no annotations at all: unannotated means treat it as destructive, non-idempotent, and reaching outside the system, until proven otherwise. That's a sensible default, but the mechanism has a specific and openly acknowledged weakness: annotations are self-declared by the server. Nothing forces a malicious or careless server to label a destructive tool honestly, and the spec itself says clients should not blindly trust these hints when the server isn't trusted. Annotations are a risk vocabulary, not an enforcement mechanism. They tell a downstream system what question to ask. They don't answer it.
Method 3: Human-in-the-loop confirmation
MCP's own guidance is direct on this point: there should always be a human in the loop with the ability to deny a tool invocation, and clients are expected to surface a confirmation prompt before executing anything consequential. The protocol formalizes this through elicitation, which lets a server pause mid-task and request explicit input or approval from the user before continuing, for example a migration tool asking for final confirmation before it touches a production database.
This is a real control, not a hint: if implemented correctly, no destructive action fires until a human clicks approve, and that approval doesn't depend on the model behaving well. Its limits are practical rather than architectural. Confirmation dialogs don't scale to an agent making hundreds of calls in a background job with nobody watching, and the more often a human is asked to approve something routine, the faster they learn to click through without reading it. A control that depends on sustained human attention degrades in exactly the situations, high tool-call volume, low supervision, where you need it most.
Method 4: OAuth 2.1 scopes at the authorization layer
Our post on the MCP handshake covered the mechanics of the protocol's optional OAuth 2.1 authorization layer for remote servers. The relevant point here is what scopes actually buy you: authentication confirms who's connecting, and scopes on the issued token determine what that connection is allowed to do, enforced at the protocol boundary rather than inside the model's reasoning.
POST /oauth/token
grant_type=client_credentials
scope=tools:read tools:invoice:write
# Token returned is only valid for read tools
# and the specific invoice-writing tool. A call
# to delete_customer_record is rejected by the
# authorization server before it reaches the tool.
This is a meaningfully different kind of control than the previous two, because it doesn't route through the model or the client's judgment at all. If a token only carries tools:read, no amount of clever prompting or injected instruction gets a destructive call to execute, because the resource server checks the token's scope and rejects anything outside it, full stop. The limitation is granularity and adoption: scopes are typically defined per tool or per tool category, set up in advance by whoever configures the integration, and the ecosystem is still standardizing on how fine-grained and how dynamic these scopes can get. A scope answers "is this class of action permitted at all." It doesn't evaluate the specific call in context.
Method 5: Policy gateways
This is the layer built specifically to answer the question OAuth scopes leave open: given that an action is permitted in principle, should this particular call execute right now. A policy gateway sits inline between the agent and its tools, intercepts every tool call before it runs, evaluates it against a declarative policy, and returns a verdict: allow, block, modify, or escalate to a human.
# Example policy rule (policy-as-code)
rule "block_bulk_delete_outside_business_hours" {
when: tool == "delete_customer_record"
and call_count_last_hour > 5
and current_hour not in [9..17]
then: deny
reason: "Bulk deletes outside business hours require manual approval"
}
The important architectural property here is that the decision is deterministic and sits outside the model entirely. There's no LLM in the enforcement path, so a prompt injection that successfully convinces the brain to attempt a bulk delete at 2 a.m. still hits a policy engine that doesn't care what the model was convinced of. It only cares whether the call matches a rule. This is the layer most directly addressing what our DLP post called the deeper problem: controls that live downstream of the decision, watching aggregate network traffic, can't see enough to act in time. A gateway sitting inline with the tool call itself sees the actual request before it executes, which is the only place a stop like this can still work. The tradeoff is operational: someone has to write and maintain the policies, and a policy that's too narrow misses new attack patterns while one that's too broad reintroduces the friction of Method 3.
Method 6: Execution sandboxing
Sandboxing takes a different angle entirely: instead of deciding whether a call should happen, it limits the damage if a call that shouldn't have happened, happens anyway. Running agent-invoked code or tools inside an isolated environment, whether that's container namespaces and cgroups, seccomp filters restricting which system calls are even available, or a microVM boundary with its own guest kernel, means a compromised or misused tool execution is contained to a disposable environment rather than the host system.
This is a genuinely different kind of control from everything above it: it doesn't try to catch the bad call, it just makes sure the bad call can't reach anything that matters. That makes it a good complement to the other methods rather than a replacement for any of them. Its blind spot is specific: sandboxing constrains what a process can touch on disk and, to a point, on the local network, but an agent that's allowed to make outbound API calls to a legitimate destination doesn't need to escape the sandbox to cause harm; it just needs to send data somewhere it shouldn't, through a channel the sandbox was never designed to police. That's the exact scenario our DLP post described: the traffic looks like a normal, successful, allowed request. A sandbox that isn't paired with egress controls will happily let that request through.
Comparing the methods
Read across that table and a pattern falls out. The methods that survive a compromised model are exactly the ones that don't ask the model anything: they check a token, evaluate a rule, or physically wall off the execution environment, all before or independent of whatever the brain decided. The methods that don't survive are the ones that rely on the model's own judgment, whether that's a system prompt it might be talked out of or a self-reported annotation it has no way to verify.
What this means in practice
None of these methods is a replacement for the others, and none of them is sufficient alone. A reasonable layered setup looks like this: OAuth scopes define the outer boundary of what's even possible for a given integration to call, a policy gateway makes the contextual, per-call decision about whether a permitted action should actually run right now, human confirmation is reserved for the genuinely ambiguous or high-consequence cases that a static policy can't anticipate, and sandboxing catches whatever gets through anyway so a bad call doesn't become a bad outcome. Tool annotations and prompt-level instructions still earn their place in this stack, not as the control itself but as the signal that tells the gateway and the human what to be careful about.
The throughline from this whole series holds here too. An agent's brain will do what the context in front of it makes plausible, and that context isn't fully trustworthy by construction. The methods worth building on are the ones that don't need the brain to get it right.
Further reading:
- Model Context Protocol, Tools
- Model Context Protocol, Understanding Authorization in MCP
- Model Context Protocol Blog, Tool Annotations as Risk Vocabulary: What Hints Can and Can't Do
This piece is part of our ongoing research into how agentic AI systems are actually built and where their behavior comes from.

