AI Agent Tool Integrations Rot Silently: How to Catch Tools That Return Nothing
Table of Contents
Chapter 5 of Agentic Design Patterns covers tool use. The core idea, in one line: an agent should not rely only on the bit of static knowledge in its head. It has to be able to reach out and call external tools (look up news, fetch a web page, run a calculation) and feed the live information it gets back into its own reasoning.
We read it and nodded. Our automated posting agent has a whole research step before it generates a post: it calls blogwatcher to pull RSS news, calls hn-trending to grab what is hot on Hacker News, and calls summarize to summarise a related article. The output of the three tools is stitched together into "today's real material" and fed to the model. Tool use, implemented. Then we ran an adversarial audit on ourselves. The truth of this chapter: for months, that research step got 0 results every single time. The posts never had a single bite of real news from start to finish, and nobody noticed.
One wrong flag, and months with nobody noticing
We dug out the tool call. It looked like this:
BLOG_NEWS=$(blogwatcher scan --json 2>/dev/null | head -c 2000 || true)
The problem is --json. blogwatcher has no such flag. Its command for fetching data is blogwatcher scan -s, listing the results takes blogwatcher articles, and there has never been anything called scan --json. So on every call, blogwatcher printed an unknown option error and failed.
But look at the rest of that line: 2>/dev/null throws the error message into a black hole, || true makes the failed command pretend it succeeded, the pipeline carries on as usual, and BLOG_NEWS gets an empty string. The guard check on the next line looks like this:
if [ -n "$BLOG_NEWS" ] && [ "$BLOG_NEWS" != "[]" ]; then
RESEARCH_CONTEXT="..." # stuff the news into the material
log "Research: blogwatcher returned data"
fi
The empty string fails -n, so the whole block is skipped and waved through as "no news today, normal". That returned data log line did not print once between going live and being caught. And the absence of a log line looks exactly the same as a system running quietly.
First, spend a few minutes checking your own tool integrations
01 Run the tool call by hand exactly as written, and this time do not swallow stderr
Take the call from your code, flags and all, paste it into a terminal and run it by hand, but this time do not throw away the error messages. A silent failure is silent precisely because 2>/dev/null wiped out the complaint that "this flag does not exist".
# Run the exact flag your code calls, by hand, and let stderr print
blogwatcher scan --json
# Check the official help again to confirm the flag / subcommand still exists today
blogwatcher --help
Red flag: running it by hand prints unknown option, the usage text, or nothing at all, and your program treats that as normal and carries on.
02 grep the tool call sites and see whether the program waves an empty value through as normal The dangerous combination is "swallow stderr + treat failure as true + skip quietly on empty". Pull those out and follow what happens when the value is empty.
# Find calls that swallow stderr and treat failure as success
grep -nE '2>/dev/null.*\|\| true' your-script.sh
# Then check whether the captured variable, when empty, is waved through as "no data today"
grep -nE 'if \[ -n|if \[ -z|!= "\[\]"' your-script.sh
Red flag: the tool call swallows stderr, the variable is quietly skipped the moment it is empty, and not a single line stops or raises an alert because "this time there were 0 results".
03 Is there "alert when the output is empty", not just "log when there is data" Logging a line only on success means failure is completely silent. You need a positive assertion: if the count sits at 0 for a long time, someone should be woken up.
# Do you only leave a trace when there is data, or does failure make noise too
grep -nE 'log|WARN|ALERT|returned' your-script.sh
# Harsher: count how many times data was retrieved; if it stays at 0 for a long time, the well has run dry
grep -c 'returned data' ~/.openclaw/logs/autopost.log
Red flag: a log line is written only on success and failure is dead quiet; the long-run count of "got data" is 0, and nobody has ever gone back to look at that 0.
What we actually changed
We changed that line to commands the tool actually has: first blogwatcher scan -s to fetch a round, then blogwatcher articles to list the results, and we left a blunt comment for the next person: "blogwatcher has no --json, stop typing it." The same hole existed, identically, in three sister scripts (mindthread, probe and advisor), and they were patched in the same batch. As of 2026-07-06, the research step really carries news into generation every time, and the log has finally started printing returned data.
It is worth being clear about what this was not: the tool was not broken, blogwatcher had been working fine all along, the script did not crash, the timer fired on schedule every day, and the log recorded the task as complete every day. On the surface everything was healthy. The only broken thing was a flag nobody had checked again, paired with 2>/dev/null || true wiping the evidence clean, paired with a guard check that treated "empty" as "normal". Each of the three is harmless on its own. Put together, they made a hole that stayed silent for months. A tool integration can also be fake the other way round, with real output but only one hand-typed run and nothing scheduling it; see how to tell a one-off run from a real integration.
Tool use checkup list
Run through it against your own system:
- For every external tool call, do the flags / subcommands still exist when you run them by hand today
- Is any
2>/dev/nullor|| trueswallowing the tool's error messages - When a tool returns an empty value, does the program stop and raise the alarm, or carry on as "no data today"
- Is there a check that alerts when the output is empty, not just "log when there is data"
- Over the long run, has the count of "the tool got data" been stuck at 0 with nobody looking
- When a tool you depend on ships a new version or changes its commands, will anything tell you the old call has stopped working
Four lines from this chapter
- Tools rot silently. A flag gets renamed, a subcommand is removed, an API changes format: your program will not throw an error, it will just quietly get nothing back.
- An empty value is not a normal value. Waving "the tool returned 0 results" through as "no data today" hands silent failure a free pass.
2>/dev/nulland|| truehelp a broken integration destroy the evidence. To verify a tool, first let stderr out and run it once by hand.- Success has to leave a trace; it is not enough to alert only on failure. A "got data" count stuck at 0 is more honest than any error message.
Source location: ~/.openclaw/scripts/moltbook-autopost.sh (the research step, with three tools: blogwatcher / hn-trending / summarize), moltbook-autopost-mindthread.sh / -probe.sh / -advisor.sh (the same scan --json hole, fixed in the same batch); the tools themselves live in /usr/local/bin/{blogwatcher,hn-trending,summarize}. The broken original call is still preserved in the *.bak-* backups: blogwatcher scan --json.
This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.
FAQ
Why does an AI agent's tool call return nothing without throwing an error?
Because the error message gets swallowed and the empty value gets treated as normal. In our case the tool call used a flag that does not exist (--json), so the tool printed an unknown option error and failed on every call. But 2>/dev/null threw the error message into a black hole, || true made the failed command pretend it succeeded, the variable got an empty string, and the guard check on the next line waved the empty value through as 'no news today, normal'. Each of the three is harmless on its own. Put together, they made a hole that stayed silent for months.
How can a broken tool integration go unnoticed for months?
Because on the surface everything looks healthy: the tool is not broken, the script does not crash, the timer fires on schedule every day, and the log records the task as complete every day. The only broken thing was a flag nobody had checked again. The program logged a line only when it got data, and the absence of a log line looks exactly the same as a system running quietly.
How do I check whether my agent's tool calls are failing silently?
Three steps. First, paste the tool call from your code into a terminal, flags and all, run it by hand without swallowing stderr, and check --help again to confirm the flags and subcommands still exist today. Second, grep the call sites for the combination of swallowing stderr, treating failure as true and skipping quietly on empty, and follow what happens when the variable is empty. Third, check whether anything alerts when the output is empty, not just logs when there is data, and count how many times the log says data was retrieved to see whether it has been stuck at 0.
What should a program do when a tool returns zero results?
Do not wave 'the tool returned 0 results' through as 'no data today', because an empty value is not a normal value. You need a positive assertion: alert when the output is empty, and if the count sits at 0 for a long time, someone should be woken up. Success has to leave a trace, not only failure, and a 'got data' count stuck at 0 is more honest than any error message.