AI Agent Output Guardrails: How to Stop API Key Leaks Before Your Agent Posts
Table of Contents
Chapter 18 of Agentic Design Patterns covers guardrails and safety. The core idea, in one line: an agent's input has to be screened for injection, and its output has to be checked before it goes out, so nothing harmful or anything that should not leave the system walks straight out the door. Input guardrails stop what bad actors feed in. Output guardrails stop what the agent sends out.
When we read the words "output guardrails", we could not laugh. This is exactly what our company sells. UltraProbe is our open-source AI security scanner, output guardrails are one of its signature scan vectors, and we use it to scan customers' agents for whether they leak things they should not. Then, on the day we ran an adversarial audit on ourselves, we found something embarrassing: our own posting agent had been running naked the whole time.
A company that sells security, with its own agent running naked
We have an entire automated posting pipeline. The agent generates content and POSTs it straight to Discord and social platforms, with no egress check anywhere in between. Whatever it produces goes out.
Think through what that means. If a prompt ever accidentally pulls environment variables into the context, the agent is fully capable of posting a real API key, the boss's real name, or the skeleton of the model's own system prompt, untouched, to a public channel. An OpenRouter key starting with sk-or-v1 pasted into Discord is an invitation for someone else to burn through our credits.
On one side we sell "we check whether your agent's output leaks". On the other, our own agent had never been checked that way. It is not that we did not know how. We had not eaten our own dog food first.
The guardrail we added: the hard part is not blocking, it is how much to block
The logic of an egress guardrail fits in one sentence: before content goes out, run it through a check first. The real difficulty is what to block, and how far to go.
Our posting pipeline has hard requirements. Discord posts like the pre-market and post-market updates are time-sensitive, and wrongly killing one normal post hurts more than letting one suspicious post slip through. So the rules are split into two tiers. Real keys get hard-blocked: sk-or-v1, AIza, sk-ant, private key headers, Bearer tokens. None of these can ever appear in a normal public post, so a match means the post is not sent. The criteria are rigid enough that the false-positive rate is close to 0. Suspicious content gets a soft warning: the model spitting out its system prompt skeleton, canned refusal phrases, the boss's real name, Simplified Chinese characters. These might be a problem or might be a misfire, so they are only logged, not blocked. That gives observability without risking a broken posting run. That only counts if someone reads the log, and nobody was reading our guardrail log.
This is the most central trade-off in guardrail design: hard blocking lowers the miss rate and raises the false-positive rate. In our scenario we would rather miss than wrongly kill, so only things that "can never be normal content" earn a hard block. Even when the checker itself is missing, we choose to let posts through rather than block them.
Change the scenario and the trade-off flips completely. If what the agent sends out is a money transfer or a command that deletes a database, you should rather wrongly kill than miss: block the wrong thing a hundred times rather than let one bad one through. Guardrails are not better the stricter they are. They have to be aimed at whichever hurts more, the cost of sending the wrong thing or the cost of blocking the wrong thing.
Spend a few minutes checking your own agent
01 Check whether there is any check at all before your agent sends something out Go through every place that sends content outward: posting, email, message replies, webhook calls. Follow the content from generation to send and see whether any function checks it in between. If none does, it is running naked.
# List every outbound send call, then check one by one whether it passes through a guard first
grep -rnE 'requests.post|curl.*-d|sendMessage|webhook|POST ' your-agent/ \
| grep -viE 'guard|sanitize|redact|filter'
Red flag: the search turns up a pile of send actions, and not one of them has a check in front of it.
02 Feed it a fake key and see whether it blocks Do not guess, test it. Put a string that looks like a real key into the content about to be sent, run your pipeline once, and see whether it stops. If it cannot stop this, it will not stop the real leak either.
# Push a fake key into the egress check; expect it to be blocked (non-zero exit)
echo 'sk-or-v1-AAAABBBBCCCCDDDDEEEEFFFF1234 出貨' | your-guard; echo "exit=$?"
# exit=0 = let through = your guardrail cannot see keys
Red flag: you put an obvious key in, and the exit code is still 0.
03 Separate what should be hard-blocked from what should only get a soft warning List your check rules and ask one question of each: when this matches, do I refuse to send, or just write a log line? Only rules with rigid criteria and a false-positive rate close to 0 (real keys, private keys) earn a hard block. Fuzzy ones (suspicious wording, style issues) should be soft warnings, or you will kill ten good posts to stop one dirty one.
# See whether your guardrail has tiers, or treats everything the same as block
grep -iE 'block|warn|hard|soft|exit 1|log_only' your-guard.*
Red flag: every rule gets the same treatment (all hard-block or all pass), with no tiers based on the cost of a false positive.
Guardrail health checklist
Run this against your own agent:
- Is there an egress check before your agent sends any content outward?
- Have you actually tested it with a fake key, and does it really block it?
- For your hard-block rules, is the false-positive rate really close to 0, or could they kill normal content?
- Are fuzzy signals soft warnings written to a log, or have they been dragged into hard blocks too?
- Have you worked out whether your scenario should rather miss than wrongly kill, or rather wrongly kill than miss?
- When the checker itself goes down, does it fail open or fail closed, and does that fit your scenario?
- Are blocks and warnings logged, so you can look them up afterwards?
This chapter in four lines
- If you sell security, make yourself secure first. Your own agent running naked while it sends content out is more embarrassing than a customer's scan turning up problems.
- An egress guardrail is not better the stricter it is. The core is a trade-off: hard blocking lowers the miss rate and raises the false-positive rate, and your scenario decides which cost hurts more.
- Only things that "can never appear in normal content" earn a hard block (real keys), with a false-positive rate close to 0. Anything fuzzy gets a soft warning, so you do not kill ten good posts to stop one dirty one.
- Eat your own dog food first. Whatever you check for other people, run it on yourself first.
Source location: ~/.openclaw/scripts/lib/output-guardrail.py (the main egress check: 5 HARD real-key rules + SOFT warnings + Simplified Chinese detection), output-guardrail.sh (the bash interface that posting scripts source; it lets posts through when the checker is missing). Both blocks and warnings land in ~/.openclaw/logs/guardrail.log, which is quiet most of the time.
This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet chapter by chapter, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.
FAQ
What is an output (egress) guardrail for an AI agent?
Input guardrails stop what bad actors feed in. Output guardrails stop what the agent sends out. The logic of an egress guardrail fits in one sentence: before content goes out, run it through a check first, so nothing harmful or anything that should not leave the system walks straight out the door, such as a real API key, the boss's real name, or the skeleton of the model's own system prompt.
How do I test whether my AI agent can leak an API key?
First go through every place that sends content outward (posting, email, message replies, webhook calls) and see whether any function checks the content between generation and send. Then do not guess: put a string that looks like a real key into the content about to be sent and run your pipeline once. If the exit code is still 0, it was let through, and it will not stop the real leak either.
Which guardrail rules should hard-block and which should only warn?
Only things that can never appear in normal content, with rigid criteria and a false-positive rate close to 0, earn a hard block: real keys such as sk-or-v1, AIza and sk-ant, private key headers, Bearer tokens. Fuzzy signals (a system prompt skeleton, canned refusal phrases, the boss's real name, Simplified Chinese characters) should be soft warnings written to a log, or you will kill ten good posts to stop one dirty one.
Are stricter AI agent guardrails always better?
No. Hard blocking lowers the miss rate and raises the false-positive rate, so it depends on which hurts more, the cost of sending the wrong thing or the cost of blocking the wrong thing. A time-sensitive posting pipeline would rather miss than wrongly kill, which is why it lets posts through even when the checker itself is missing. If the agent sends out a money transfer or a command that deletes a database, you should rather wrongly kill than miss.