AI AgentExploration and DiscoveryProactive ExplorationBuildInPublicSolo Business

How to Audit an AI Agent's Exploration System: Catch LLM Errors Saved as Insights and Stop Temporary Failures from Burning Material

· 9 min read
Lobster Fleet · Pattern Audit · Part 22 of 25
Table of Contents
  1. Half alive, and the living half bites
  2. The other half lost its pulse long ago
  3. Spend a few minutes checking your own exploration system
  4. How we actually fixed it
  5. Exploration and discovery checklist
  6. This chapter in four sentences

A few months ago, in We Built a Self-Learning AI Sales System in 48 Hours, I described "100 targeted cold emails per day" and "auto-blacklisting bounced domains" as an exploration engine that evolves on its own. This post lays out what it looks like now: that outreach channel has an 82% bounce rate and has effectively been stopped for a long time, and the other half, the research chain that is still alive, is writing LLM error messages into its notes as insights.

Chapter 21 of Agentic Design Patterns covers exploration and discovery. The core idea, in one line: an agent cannot just wait passively for instructions. It has to go out and dig up new material and new opportunities, pulling the outside world in and turning it into something it can use. An agent that does not explore only ever sees the small slice you feed it.

We read it and nodded. We have two legs here: a research-chain that scans blogs and HN twice a day, reads new articles and produces research notes on "what this means for us"; and an outreach/hunter pipeline that digs up prospects and reaches out proactively. Proactive exploration, implemented. Then we ran an adversarial audit on ourselves and found this exploration system was half dead: one half was still mining real insights, the other half had an 82% bounce rate and had long since stopped, and the living half was quietly hurting itself in two ways.

Half alive, and the living half bites

Start with the living half. research-chain is still running today (2026-07-06), and today it still produced notes with real insight, not empty shells. But it is broken in two nasty places.

First: it writes LLM error messages into its notes as "analysis". The last step of the chain asks an LLM to analyse the article and return 2 to 3 insights. When that call fails, what comes back is not an insight but an error string, and the script accepts it as is and writes it into RESEARCH-NOTES.md. What the audit dug up: on 6/27 the notes held API_KEY_INVALID, on 7/3 they held run out of credits. The downstream posting script reads these as material, which amounts to passing "the API key is invalid" around as a business insight.

Second, and worse: it marks articles as read permanently. To avoid processing anything twice, the chain records every URL it has looked at in a seen list. But after the analysis fails and it gets an error message back, it marks the URL as read anyway. So a single temporary billing or API error costs a good article for good: it will never be analysed again, because the system thinks it has already been processed. An error that lives for five seconds causes a permanent loss.

The most dangerous failure in an exploration system is not coming up empty. It is finding something, storing the error as an insight, and burning the raw material on the way out, so you believe the well has run dry.

The other half lost its pulse long ago

Now the dead half. The outreach/hunter pipeline claimed to be "proactively exploring for customers". The audit pulled the numbers: an 82% bounce rate. That is not a low reply rate. The channel itself is dead: eight in ten emails sent never even reach the recipient's inbox, and it had long since been paused. The value of exploration is pulling outside things in, and a channel with an 82% bounce rate pulls nothing in. It just pretends, every day, to reach outward. On the surface "we do proactive exploration"; laid out in full, it is half dead: one half mines real insights but bites itself, the other half is completely dead and still sitting on the shelf.

Spend a few minutes checking your own exploration system

01 Is the "analysis" you saved actually an LLM error message? Go through your exploration notes or material files and scan them by keyword for common LLM error strings. The last step of an exploration chain usually asks an LLM to produce insights. Once that call fails, what comes back is an error message, and most scripts do not validate the format and take it as is. Your insight library may be holding a few run out of credits right now.

# Scan exploration notes for LLM error messages that slipped in
grep -niE 'API_KEY_INVALID|run out of credits|quota exceeded|401 Unauthorized|429 Too Many|INVALID_ARGUMENT|no valid API' RESEARCH-NOTES.md

Red flag: any match at all. Your exploration output has error messages mixed in, and downstream will use them as material.

02 Can a temporary failure cause a permanent loss? Find every irreversible action in your exploration flow: marking as read, blacklisting, marking as processed, deleting from a queue. Then ask the question that matters: does this action happen after success, or does it happen whether the step succeeds or fails? If the failure path also marks items as read, one temporary API error will permanently burn a good piece of material.

# Check whether mark-as-read / mark-as-processed actions also run on the failure path
grep -nB3 -iE 'seen\.push|mark.*read|blacklist|processed|dedupe|\.delete\(' your-explore.sh

Red flag: writes to the seen list or blacklist still execute in the "analysis failed" branch. Irreversible actions must happen only after success is confirmed.

03 Is your outreach / exploration channel still open today? For exploration to pull outside things in, you need a live channel. Pull the recent performance numbers of your outreach channel: bounce rate, reply rate, number of records returned by the API. A channel with an 82% bounce rate is no different from one that is switched off, except that it is still on the shelf, letting you believe you are exploring outward.

# Pull the outreach channel's recent bounce / failure rate; don't just check whether it is running
grep -ciE 'bounce|550|undeliverable|rejected' outreach.log
grep -c 'sent' outreach.log   # divide the two numbers and you have your real bounce rate

Red flag: a bounce rate above 20%, or you cannot say at all who this channel actually reached most recently.

How we actually fixed it

On 2026-07-04 the fix went in right after the analysis step of research-chain: a deterministic validation gate. If the analysis result is empty, or matches any known LLM error message by keyword (API_KEY_INVALID, run out of credits, quota exceeded, 429 and so on), the chain does not write the note and does not mark the URL as read; it leaves it for a retry on the next run. One cut closes two holes: error messages cannot get into the insight library, and a good article can no longer be burned permanently by a single temporary error. As for the dead outreach half, we did not pretend it was still alive. We paused it outright, to be reopened once there is a live channel.

Exploration and discovery checklist

Run this against your own exploration or discovery system:

  • Has the output of the last step in your exploration chain been checked to be "analysis" rather than "an LLM error message"?
  • Are empty responses and error strings kept out of your notes, or written into the material library anyway?
  • Are irreversible actions such as marking as read, blacklisting and marking as processed done only after success, or also on the failure path?
  • Can a single temporary API error cause a good piece of material to be skipped permanently and never retried?
  • Can you say what the recent bounce rate or reply rate of your outreach / exploration channel is?
  • Has your dead exploration channel been honestly paused, or is it still on the shelf pretending to reach outward?

This chapter in four sentences

  • The most dangerous failure in exploration is not coming up empty; it is storing an LLM error message as an insight. Validate the output format: an error string is not analysis.
  • A temporary failure should not cause a permanent loss. Irreversible actions such as marking as read or blacklisting may only happen after success is confirmed; otherwise an error that lives for five seconds burns a good article.
  • Exploration needs a live channel. An outreach pipe with an 82% bounce rate pulls nothing in. It is not exploring, it is pretending to reach outward.
  • A half-dead exploration system is more deceptive than a fully dead one. The living half makes you believe the whole system is healthy, when in fact it is mining insights with one hand and biting itself with the other.

Source locations: ~/.openclaw/scripts/research-chain.sh (the exploration chain, including the error detection gate added on 2026-07-04: empty responses and error messages are never written to the notes, never marked as read, and are retried on the next run); exploration notes land in ~/.openclaw/workspace/RESEARCH-NOTES.md, and the seen list is at ~/.openclaw/data/research-seen-urls.json; the outreach/hunter half (82% bounce rate) has been paused.

This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet chapter by chapter, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.

FAQ

Why do LLM error messages end up saved as analysis in exploration notes?

The last step of an exploration chain usually asks an LLM to analyse an article and produce insights. When that call fails, what comes back is not an insight but an error string, and most scripts do not validate the format and write it into the notes as is. Our audit found API_KEY_INVALID from 6/27 and run out of credits from 7/3 in the notes, and the downstream posting script reads these as material.

How do I check whether LLM error messages have slipped into my exploration notes?

Go through your exploration notes or material files and scan them by keyword for common LLM error strings such as API_KEY_INVALID, run out of credits, quota exceeded, 401 Unauthorized, 429 Too Many and INVALID_ARGUMENT. Any match at all is a red flag: your exploration output has error messages mixed in, and downstream will use them as material.

Why can a single temporary API error cause a permanent loss?

Because actions such as marking as read, blacklisting, marking as processed and deleting from a queue are irreversible. If the chain marks a URL as read even after the analysis fails and returns an error message, that article will never be analysed again, because the system thinks it has already been processed. Irreversible actions must happen only after success is confirmed.

How do you stop error messages from being written to notes without burning good articles?

Add a deterministic validation gate right after the analysis step: if the analysis result is empty, or matches any known LLM error message by keyword, do not write the note and do not mark the URL as read; leave it for a retry on the next run. One cut closes two holes: error messages cannot get into the insight library, and a good article can no longer be burned permanently by a single temporary error.

How can I tell whether an outreach or exploration channel is actually dead?

Do not just check whether it is running; pull its recent numbers: bounce rate, reply rate, number of records returned by the API. A bounce rate above 20%, or not being able to say who the channel actually reached most recently, is a red flag. A channel with an 82% bounce rate is no different from one that is switched off, and our response was to pause it outright until there is a live channel.

Lobster Fleet · Pattern Audit · Part 22 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.