AI Engineering · August 27, 2026 · 4 min read
Guardrails Are Not a Prompt: What Actually Keeps an AI Agent in Line
TL;DR
"Add some guardrails" usually means one line in the system prompt telling the model to behave. That's a request, not a guardrail. A real guardrail is separate code that checks input before it reaches the model and checks output before it reaches the user. It doesn't care what the model was told or how convincingly it argues otherwise.
The Problem With Prompt-Only Safety
A system prompt is context, not control. The model reads it the same way it reads everything else, as tokens competing with the rest of the conversation for influence over the next output. Enough back and forth, a clever rephrase, or a long enough context window and the instruction gets diluted or ignored. That's just how autoregressive generation works. If the only thing stopping a bad output is a sentence that says "never do X," you don't have a guardrail. You have a suggestion the model can talk itself out of.
The fix is moving the check outside the model entirely.
What a Guardrail Actually Is
A guardrail is deterministic code, or a small purpose-built model, sitting on either side of the main LLM call.
- Input guardrails run before the prompt reaches the model. Reject or rewrite the request if it's off-topic, contains an injection attempt, or asks for something the system was never scoped to handle.
- Output guardrails run after the model responds, before the user sees it. Catch hallucinated claims, leaked system instructions, unsafe content, or an answer that drifted outside what the system actually knows.
Neither depends on the model policing itself. Both keep working even if the model gets jailbroken, because the check happens outside its own reasoning loop.
Guardrails I'd Actually Build Into an Agent
- Scope enforcement: an input classifier catches and rejects out-of-scope questions before they hit the expensive main call.
- Grounding checks: an answer's claims have to trace back to retrieved context or tool output before it ships. No source, no claim. It downgrades to "couldn't confirm this" instead.
- Injection detection: user input, scraped pages, and tool output all get treated as untrusted text, not instructions. A check flags anything trying to redirect the agent's behavior.
- PII and secret leakage checks: output gets scanned for credentials, internal system details, or personal data before it goes out, regardless of what the model meant to do.
- Cost and rate boundaries: a request-size cap or a hard limit on tool-call loops stops a runaway agent chain before it burns budget or time.
Guardrails That Stop Tool Calls, Not Just Text
Input and output guardrails check text going into or out of the model. Tool-call guardrails are a different thing entirely. They sit between the model deciding to call a tool and the tool actually running. The model can propose delete_database() or send_email(), but proposing a call isn't the same as the system letting it execute.
- Tool allowlisting: an agent only gets access to the tools its task actually needs, nothing outside its scope.
- Argument validation: a tool that takes a URL only accepts URLs on an approved domain list, a tool that takes an amount only accepts values inside a sane range.
- Human-in-the-loop confirmation: irreversible actions (payments, deletions, sending something external) pause for a real approval instead of firing automatically.
- Context-aware call checks: since tool inputs are often built from scraped or user-supplied content, a separate check should confirm the call still makes sense given where that content came from, not just that the syntax is valid.
This matters more in agentic systems than in a plain chatbot, because a chatbot's worst case is a bad sentence. An agent's worst case is a bad action.
Where This Fits in an Agent Pipeline
In a graph-based agent like LangGraph, guardrails belong in the graph as nodes, not bolted on at the end. An input-guard node sits right after entry and can short-circuit the whole run before retrieval or generation even happens. An output-guard node sits right before the response goes out and can force a retry, a fallback answer, or an honest "not found" instead of letting a bad response through. Making them real nodes keeps the enforcement point visible in the architecture instead of buried in a wrapper function.
The Part People Skip: Evals
A guardrail nobody has tested against real failure cases is just code you're hoping works. Run it against a set of known bad inputs and known bad outputs and check it fires every time. That's the only way to know a grounding check actually catches ungrounded claims or an injection filter actually catches injections. Guardrails without evals are decoration.
The Real Takeaway
A well-worded system prompt is hope, not a guardrail. Guardrails are the parts of the system the model doesn't control, the code that runs no matter what the model decided to say. Build those first. Worry about the prompt after.