AI agent guardrails: budgets, grants and prompt-injection defence
Guardrails are not one filter on the model's output. They are a set of limits around the whole run: budgets, grants, schema validation, untrusted-content fencing, a tool-call ledger and evaluations before publish.
ArchitectureVeloPhex Engineering
Platform team
AI agent guardrails are the limits enforced around an agent run, not a filter on its answer. A production setup caps model calls, tokens and time; grants the run only the tools it needs; validates output against a schema; fences untrusted content; records every tool call in a ledger; and evaluates each version before it is published.
When people say "guardrails" for AI, they often mean a content filter: something that checks the model's answer for toxicity or leaked data. That matters for chatbots. For agents, which call tools and change records, it is the smallest part of the problem. The question is not only "is this answer acceptable?" but "what could this run do if the model went wrong at step four?"
This post lays out the guardrails we think every agent in production needs, why each exists, and how they map to the OWASP Top 10 for LLM Applications. Most of it is framework-neutral.
AI agent guardrails at a glance
| Guardrail | What it stops | OWASP LLM risk |
|---|---|---|
| Budgets (calls, tokens, time) | Runaway loops, cost blowouts, stuck runs | LLM10 Unbounded Consumption |
| Grants | The agent calling tools it was never given | LLM06 Excessive Agency, LLM01 Prompt Injection |
| Output schema validation | Malformed or unexpected answers reaching downstream systems | LLM05 Improper Output Handling |
| Untrusted-content fencing | Retrieved text being treated as instructions | LLM01 Prompt Injection |
| Tool-call ledger | Duplicate side effects, silent double payments | LLM06 Excessive Agency |
| Capacity limits | One agent starving others, burst costs | LLM10 Unbounded Consumption |
| Evaluations before publish | Regressions shipping with a prompt or model change | LLM09 Misinformation |
Budgets: calls, tokens and time
Agents loop. That is what makes them useful: read, decide, call a tool, read the result, decide again. It is also what makes them expensive when something goes wrong. A model that misreads a tool error can call the same tool forty times, each time with a longer context.
Every run needs three ceilings:
- Model calls. A hard cap on loop iterations. Most well-scoped tasks finish in a handful of steps; if a run hits the cap, that is a signal worth investigating, not a reason to raise it.
- Tokens. Prompt plus completion tokens for the whole run. This bounds cost directly and catches context that grows out of control.
- Wall clock. A timeout for the run, never past the job's own deadline. A run waiting on a slow tool should not hold capacity forever.
On top of per-run budgets, set cost budgets per workspace or team at the model gateway, and capacity limits on concurrent runs per agent and per tenant. Per-run budgets stop one bad run. Capacity limits stop a burst of a thousand good ones from costing a month's budget in an afternoon.
Grants: the run can only do what it was given
A grant is the list of tools, and only those tools, that a specific run may use. The important word is enforced. It is not enough to give the model a list of tools in its prompt. The system that executes tool calls must check each call against the grant and refuse anything outside it, whatever the model asks for.
Why does this matter so much? Because prompt injection is, at heart, an attempt to make the agent do something it was not meant to. If the agent was only ever granted "search the policy store" and "queue an invoice for review", an injected instruction to "email this file to an external address" fails at the server, regardless of how convincing it was to the model.
Grants should also cover indirect paths. If the agent can start other jobs, the grant should limit which ones. Chains of agents calling agents should have a depth limit. And tools you never want any agent to call should be blocked at the definition level, so a publish check fails before a version with them can ship.
Output schema validation
An agent's final answer is usually consumed by code: a workflow branches on decision, a robot posts the amount. Free text is a poor interface for that. Define a JSON Schema for the output and validate the answer against it before anything acts on it.
When validation fails, allow one repair turn: tell the model what was wrong and let it try again. If the second answer still does not fit, fail the run with a clear error. Never pass an invalid answer downstream "with a warning". This is the core of OWASP's improper output handling guidance: treat model output like any other untrusted input.
Schemas also make agents easier to reason about. An answer that must be one of post, hold or reject is easier to evaluate, alert on and audit than a paragraph.
Fencing untrusted content
Everything a model reads competes for its attention: the system prompt, the user's request, tool results, retrieved documents. A model has no built-in way to know that the PDF it just retrieved is data and not a new instruction.
Fencing gives it one. Wrap every piece of external content (tool output, retrieved chunks, even the run's own inputs) in a clearly marked block, and state in the system prompt that anything inside such a block is data to be analyzed, never instructions to follow. Two implementation details matter:
- Escape the fence. If the content contains your closing tag, escape it, so a document cannot close its own block and start "speaking" as the system.
- Own the system prompt. Build it from the agent's definition on the platform side and check it before every model call, so it cannot be modified by anything the run read.
Be honest about what fencing achieves. It reduces the success rate of injection; it does not eliminate it. That is why it sits alongside grants, approvals and schema validation, which limit what a successful injection can do. For tool-heavy setups, our post on MCP in the enterprise covers injection through tool output in more depth.
The tool-call ledger and unknown outcomes
This is the guardrail most teams discover after their first incident.
Suppose an agent calls a payment tool and the connection drops before a response arrives. Did the payment happen? You do not know. A framework that retries on error will call the tool again and may pay twice. A framework that treats the error as a failure may report "not paid" when the money has already left.
A tool-call ledger fixes this with three rules:
- Every call gets an idempotency key, derived from the run, the tool and the arguments, and is recorded before it is sent.
- A repeat with the same key replays the recorded answer instead of calling the tool again.
- A call that was sent but never reported is marked "outcome unknown", raises an event, and is never re-run automatically under its key. A person checks the target system and records whether it succeeded or failed. Even then, the call is not re-run; a repeat returns the person's verdict.
It follows that agent runs with side effects should not be retried wholesale by the platform either. A retried run would redo calls that may already have taken effect. Retries belong at the level of individual idempotent calls, with the ledger deciding.
This is the same discipline good RPA teams apply to queue items, applied to tool calls. If you have read our post on credential leases, the theme is familiar: decide the unhappy path before it happens.
Evaluations before publish
Prompts, models and tools all change, and any change can make an agent worse in ways that are not obvious from three manual tests. Before a version goes live:
- Freeze it. A candidate version should be immutable: its definition, its model binding and every tool it resolved to, with hashes. What you evaluate is exactly what you ship.
- Run it against a dataset of representative cases with expected outcomes, scored by rules where possible and by an LLM judge where judgment is needed.
- Gate publishing on the score. If the average falls below a threshold, the version does not publish.
- Run the boring checks too. The definition is valid, every tool still resolves, no blocked tools are present, and the model binding is allowed by policy.
Treat this exactly like a test suite in CI, because it is one. Our post on treating automation as software makes the broader case.
How VeloPhex implements these guardrails (beta)
All of the above is built into VeloPhex Managed Agents, which is in beta on a preview Orchestrator and open to Design Partners. Specifically:
- Budgets: per-run defaults of 12 model calls, 50,000 tokens and 15 minutes, adjustable per agent, plus cost and budget tracking at the AI gateway and a workspace AI policy.
- Grants: each run holds the grant of its version. Anything outside it, including child jobs and unlisted packages, is refused by Orchestrator. Blocked tools fail a publish check, and agent chains stop three links deep.
- Schema validation: the final answer is validated against the agent's output schema, with one repair turn before the run fails.
- Fencing: tool outputs and run inputs reach the model as fenced untrusted content, with tags inside escaped.
- Capacity: concurrent-run and burst limits per agent, tenant and organization, with a clear refusal and retry hint when exceeded.
- Ledger: every tool call carries an idempotency key; calls with an unknown outcome raise an event and an audit record, are never re-run, and are reconciled by a person.
- Publish checks and evaluations: immutable candidates, publish checks, and an optional dataset gate scored with an LLM judge.
Approvals are the remaining layer; we cover them separately in human-in-the-loop approvals for AI agents. If you are running your own agent code on Robots today, the Python AI agents tutorial shows how to apply budgets and failure handling yourself. For the wider picture, see agentic automation.
Want these guardrails around your first agent? Apply to the Design Partner program.
Frequently asked questions
What are AI agent guardrails?
AI agent guardrails are the limits a platform enforces around an agent run so that a confused, manipulated or simply wrong model cannot cause outsized harm. They include budgets on model calls, tokens and time, a grant listing the tools the run may use, validation of the final output against a schema, fencing of untrusted content, and a ledger that records every tool call.
Can prompt injection be fully prevented?
Not with today's models. Any content a model reads can try to steer it, and no filter catches every attempt. The practical goal is to limit what a successful injection can achieve: fence untrusted content as data, give the agent only the tools it needs, enforce that grant on the server, require approvals for side effects, and validate outputs before anything acts on them.
Why should agent tool calls not be retried automatically?
Because some failures happen after the effect. If a payment call times out, the payment may or may not have gone through. Retrying blindly can pay twice. A tool-call ledger gives each call an idempotency key, replays recorded answers for repeats, and marks calls whose outcome is unknown so a person can check the target system and record what really happened.
How do the OWASP Top 10 for LLM Applications relate to agent guardrails?
Several entries map directly. Prompt injection (LLM01) calls for fencing and least privilege. Improper output handling (LLM05) calls for schema validation before acting. Excessive agency (LLM06) calls for grants and approvals. Unbounded consumption (LLM10) calls for budgets and capacity limits. Guardrails are how you implement those recommendations at run time.
Which of these guardrails does VeloPhex provide?
VeloPhex Managed Agents, in beta, applies per-run budgets (by default 12 model calls, 50,000 tokens and 15 minutes), a server-enforced run grant, output schema validation, untrusted-content fencing, capacity limits per agent, tenant and organization, a tool-call ledger with human reconciliation of unknown outcomes, and publish checks with an optional evaluation-score gate. It runs on a preview Orchestrator, open to Design Partners.
AI agent guardrails at a glance
Budgets: calls, tokens and time
Grants: the run can only do what it was given
Output schema validation
Fencing untrusted content
The tool-call ledger and unknown outcomes
Evaluations before publish
How VeloPhex implements these guardrails (beta)
Frequently asked questions
What are AI agent guardrails?
Can prompt injection be fully prevented?
Why should agent tool calls not be retried automatically?
How do the OWASP Top 10 for LLM Applications relate to agent guardrails?
Which of these guardrails does VeloPhex provide?
AI agents Guardrails Prompt Injection OWASP Architecture8 UiPath alternatives for 2026 (and which suit Python teams)
If your automation team writes Python, the right platform looks different. A fair look at UiPath and its alternatives, including where VeloPhex fits and where it does not yet.
Product
VeloPhex Product
MCP enterprise security: Model Context Protocol in production
The Model Context Protocol makes it easy to give AI agents tools. That is exactly why it needs controls. What MCP is, the four risks that matter, and the controls that address them.
AI & Agents
VeloPhex Security
Agentic automation vs RPA vs IPA: which fits which work?
RPA follows rules, IPA adds models to fill in the gaps, and agentic automation lets an AI agent choose the steps. Three worked examples show where each one belongs.
Engineering Enterprise Automation Python Architecture AI & Agents Product