AI Agents
AI Agent Guardrails: How to Keep Autonomous Agents Reliable in Production
Demos are easy. Production is where agents fail in expensive, embarrassing ways. Here's the guardrail architecture we build into every agent before it touches a real customer or system.
James Holt
Head of AI Engineering

Every AI agent works in a demo. The gap between a demo and a production system that runs unattended, on real customer data, thousands of times a day, is almost entirely about guardrails — not model quality.
We covered the basic agent loop (observe, reason, act, reflect) in an earlier post. This one goes deeper into the part that actually determines whether an agent is safe to ship: the guardrail layer that sits around that loop.
Why guardrails matter more than model choice
Swapping in a better model rarely fixes a reliability problem. Most production incidents we've seen come from one of three places:
- The agent took an action it shouldn't have had permission to take
- The agent got stuck in a loop or spiraled on an edge case it had never seen
- Nobody noticed the failure until a customer complained
None of these are solved by a smarter model. They're solved by architecture.
The four guardrail layers
1. Permission boundaries
Every agent should have an explicit allowlist of actions it can take autonomously, and a separate list of actions that require human approval — regardless of how confident the agent is.
In practice, this means:
- Read actions (querying a CRM, checking an order status) run freely.
- Low-risk write actions (updating a ticket status, sending an internal Slack alert) run freely but are logged.
- High-risk write actions (issuing a refund, sending a customer-facing email, canceling a subscription) require either a hard dollar/impact threshold or a human-in-the-loop approval step.
The threshold isn't a one-time decision — it should move as you build confidence in the agent's judgment on a given task.
2. Evaluation sets, not vibes
Before any change ships — a new prompt, a new tool, a new model version — it runs against a fixed bank of test cases with known-correct outcomes. This is the single highest-leverage practice for agent reliability, and the one teams skip most often because it feels like overhead early on.
A good eval set includes:
- Common, everyday cases (most of the volume)
- Known edge cases that have broken the agent before
- Adversarial inputs — what happens if a user tries to manipulate the agent into an action it shouldn't take
If a change doesn't improve the eval score, it doesn't ship, no matter how good it looked in manual testing.
3. Bounded autonomy, not open-ended loops
Open-ended agents that keep reasoning and acting until they "feel done" are the most common source of runaway behavior — extra API calls, repeated actions, spiraling cost. Production agents need explicit stopping conditions:
- A maximum number of loop iterations per task
- A maximum number of tool calls per session
- A hard timeout, after which the task escalates to a human instead of continuing indefinitely
Bounded autonomy makes failure predictable. An agent that stops and escalates after 5 failed attempts is reliable in a way an agent that "tries harder" indefinitely never will be.
4. Monitoring and audit trails
Every action an agent takes — every tool call, every decision, every piece of reasoning that led to it — gets logged in a way a human can review after the fact. This isn't optional instrumentation; it's what makes the first three layers auditable and improvable over time.
At minimum, this means being able to answer, for any single agent run:
- What did it observe?
- What did it decide, and why?
- What action did it take, and what was the result?
- Did it hit a guardrail, and what happened when it did?
Without this, you can't debug a failure, and you can't build the next eval case from it.
What this looks like in a real deployment
For a support agent handling inbound tickets, this typically means: the agent can look up order history and draft a response freely, it can send a reply to routine questions without approval, but refunds above a set threshold, anything mentioning legal or safety concerns, and any customer who's already escalated twice get routed straight to a human — with the agent's draft attached as a starting point, not a decision.
That's the difference between "AI handles support" as a slogan and an agent that resolves a meaningful share of tickets without anyone worrying what it might do unsupervised.
The takeaway
An agent's reliability isn't a property of the model underneath it. It's a property of the guardrails, evaluation process, and monitoring built around it. Teams that treat those as the real engineering work — not an afterthought once the demo works — are the ones running agents in production that nobody has to babysit.
James Holt
Head of AI Engineering
James leads AI agent architecture at RexorAI, with a background in applied ML and production LLM systems.


