When Agent Failures Are Boundary Failures

An agent can produce a fluent, plausible answer and still fail because the surrounding system never defined what it may access, infer, change, or finalize. Before replacing a model or rewriting a prompt, inspect those boundaries.

An agent returns a confident answer. The answer sounds reasonable, but the evidence does not support it.

The first reaction is often to blame the model: improve the prompt, add examples, or move to a more capable model. Those changes may help. But they will not repair a system that never defined the agent’s authority.

A useful first question is not only, “Why did the model say that?” It is also:

What was the model allowed to see, conclude, do, and finalize?

I use four boundaries to diagnose this class of failure: access, inference, action, and finalization. This is a practical framework for the article, not a universal taxonomy.

Access: what may enter the model’s context?

The model can choose what information would be useful. It should not decide what information it is authorized to receive.

That distinction matters in retrieval systems. An agent may generate a search query, but application code should still enforce the user identity, tenant or account scope, allowed sources, result limits, and returned citation identifiers.

A prompt that says “only access this customer’s records” is an instruction. It is not an access-control boundary.

OWASP’s guidance on excessive agency recommends limiting tool functionality and permissions, executing actions in the user’s context, and implementing authorization in downstream systems rather than relying on the model to decide whether an action is allowed. NIST defines least privilege similarly: give a user, or a process acting for one, only the access needed for the assigned task.

The practical rule is simple:

  • Let the agent request context.
  • Let the backend decide whether that request is allowed.
  • Return only the approved evidence.

A safe empty result is part of this contract. If nothing authorized and relevant matches, the system should return no evidence rather than broaden the search until something plausible appears.

Inference: what may the model conclude from that context?

Authorized access does not make every conclusion valid.

A system should distinguish between:

  • facts directly supported by a source;
  • conclusions derived from multiple sources;
  • unresolved gaps that require more evidence;
  • speculation that should not be presented as fact.

This boundary can be enforced with structured outputs, citation requirements, evidence-to-claim checks, and explicit unresolved states. The model still performs reasoning, but the workflow makes the reasoning inspectable.

Without that distinction, retrieval can create false confidence. The agent cites a real document, but the cited passage does not support the claim it made.

Action: what may the agent change?

Reading a record, drafting an update, and applying that update are different capabilities.

OWASP describes excessive agency as a combination of excessive functionality, excessive permissions, or excessive autonomy. That framing is useful because the fix is architectural, not rhetorical. Give the agent narrow tools with explicit contracts instead of a generic command surface. Separate read operations from writes. Validate arguments outside the model. Require idempotency, audit records, and approval for consequential side effects.

A stronger model does not make an over-permissioned tool safe.

Finalization: who makes the result official?

“Human in the loop” is too vague unless the workflow defines where the loop occurs.

A review boundary should be an explicit state transition:

draft → review required → approved or rejected → finalized

The system should record what the reviewer saw, what evidence supported the draft, and what changed before approval. High-impact actions should remain pending until the authorized reviewer acts.

This makes human review part of the operating model rather than a disclaimer placed below an autonomous workflow.

A synthetic example

Consider a support agent asked to answer a customer’s policy question.

The agent requests relevant policy material. The backend checks the customer account, filters the permitted document set, and returns approved passages with stable citation IDs.

Two outcomes are valid:

  1. Supported result: the evidence addresses the question. The agent drafts an answer, links each consequential claim to a passage, and sends the draft to review.
  2. Safe empty result: no approved passage answers the question. The agent preserves the gap and routes it for review instead of searching outside the authorized scope or inventing a policy.

In both cases, the model chooses a useful next step. It does not choose its own authority.

User request
    ↓
Model requests context
    ↓
Backend enforces identity, scope, source, and limits
    ↓
Approved evidence — or an explicit empty result
    ↓
Model produces a cited draft or unresolved gap
    ↓
Authorized review
    ↓
Finalized output

Diagnose the failure before changing the model

Failure class Example First place to inspect
Model Misreads clear, sufficient evidence Prompt, examples, model capability, evaluation set
Tool Returns malformed, stale, or incomplete data Tool contract, validation, freshness, retries
Boundary Exposes the wrong data or permits an unauthorized action Identity, authorization, scope, permissions, approval gates
Workflow Loses state, retries a side effect, or bypasses review State transitions, idempotency, recovery, orchestration

Model and prompt improvements still matter when the task is well-scoped, the evidence is sufficient, the tools behave correctly, and the model still reasons poorly. But those conditions should be demonstrated rather than assumed.

The broader lesson is that agent reliability is not produced by the model alone. It emerges from the contract between probabilistic reasoning and deterministic control.

I ran into this while building Folium’s small chart-review flow: the meaningful problems were configuration, serialization, citation validation, and persisted review state, not a need for a more elaborate prompt. The project is evidence for the point, not the point itself.

Before upgrading the model, name the boundary that failed.

Sources