AI AgentAdversarial AuditBuildInPublicSolo BusinessAgentic Design Patterns

Assume You Are Lying: How to Audit an Entire AI Agent Fleet

· 6 min read
Lobster Fleet · Pattern Audit · Part 1 of 25
Table of Contents
  1. The patient chart
  2. Why we are publishing this
  3. Do it yourself: run the audit on your own system
  4. Self-audit checklist
  5. This post in four sentences

We read Antonio Gulli's Agentic Design Patterns and used it to build out the AI agent fleet of a one-person company. 21 design patterns, implemented one by one against the book: routing, reflection, memory management, multi-agent collaboration, resource awareness, evaluation and monitoring. When that checklist was done, it felt good. Every box was ticked.

Then we did something most tutorials never do: we assumed every one of those ticks was false and went through them chapter by chapter.

The audit had one rule: trust no self-reported claim. A document saying "implemented" is not evidence. We wanted the last time it actually ran, the output it actually produced, the record of it actually being called. If we could not find those, it was an empty shell.

The result was ugly. Of the 14 chapters we claimed as "implemented", not one was completely clean.

The patient chart

Here is that same system, laid open.

Flagships die. Chapter 6, Planning: agent-meeting.py, the agent we were proudest of, turned out under audit to have been dead for three months. No timer was calling it; all that was left was an archived empty shell, hanging there for show.

Numbers get inflated. Chapter 3, Parallelization: externally we said "six brands running in parallel". In reality only one brand gateway was alive; the other five were empty shells left behind after the migration.

Names lie. Chapter 20, Prioritization, sounds clever. The audit showed it was only static scheduling plus rotation plus rate limiting, with not a single line doing any ranking by value. The only scoring script had no timer at all and never ran.

Monitoring goes quiet exactly when you need it most. Chapter 11: we had installed a whole suite of alerts, yet the two most painful bills both happened while monitoring was in place, because the thresholds were set so loose that the money slipped out underneath them and no alarm went off.

Even the evaluation system hallucinates. Chapter 19: we built an LLM judge to catch wrong answers, and the judge marked a system that had answered correctly as wrong, then pushed a CRITICAL alert to scold it.

The most expensive chapter was Chapter 16. Of the money the gateway burned in one week, 98.8% went to rereading its history again and again; actual output accounted for 1.2%.

Why we are publishing this

The usual move, when you find out half of what you built is dead, is to fix it quietly and then only talk about the fixed version. Nine out of ten AI agent tutorials out there are written that way: clean architecture diagrams, perfect flows, everything apparently running to plan.

We decided to do the opposite and turn the audit report itself into the teaching material. The reason is practical: this is what real agent systems look like. Half the claimed features are dead, the flagship is there for show, monitoring stays silent during a disaster, and the learning loop runs diligently every night and learns 0. Look at any automated system that has been running for a few months without someone watching it every day, and eight times out of ten it looks just like this. Pretending it is perfect does nothing for the people who are stepping into the same holes right now.

An honest audit report is more useful than a pretty architecture diagram.

Do it yourself: run the audit on your own system

You can apply this audit as it is. The core fits in one sentence: trust no self-reported claim, accept only evidence of execution.

01 Assume every "done" is false List every feature in your system that you believe is working, then treat each one as a hypothesis to be proven, not a known fact. Only with that mindset can the audit actually be carried through.

02 Check when it last actually ran Flagship features are the ones most likely to die without a sound. Check whether their timer still exists and how long ago the last log entry was written.

# When did this service last actually run
systemctl --user list-timers | grep 你的服務
ls -la --time-style=long-iso 你的服務.log   # check when the log was last updated

Red flag: a flagship feature's log stops weeks or even months ago.

03 Check whether it actually produces anything A timer calling it does not mean it produces anything. Verify "is it running" and "is what it produces right" separately.

# Running does not mean producing: count the real output
grep -c 'produced\|posted\|wrote' 你的服務.log

Red flag: plenty of runs, and real output close to 0.

04 Reconcile the numbers you tell people Take the scale you describe to others (how many run in parallel, how many agents, how much has been processed) and go back and count, one by one, how many are actually alive. Red flag: the number you tell people is higher than the number you can actually count as alive.

Self-audit checklist

  • For every feature marked "done", do you have evidence of the last time it actually ran?
  • For your flagship features, how long ago was the last log entry?
  • For services whose timer is firing, are they really producing output, or just running empty?
  • For the numbers you tell people, have you gone back and counted how many are actually alive?
  • Are your monitoring thresholds so loose that a disaster can slip out underneath them?
  • Is there a flagship feature you think is still in use that is in fact already dead?

This post in four sentences

  • Trust no self-reported claim, accept only evidence of execution. The document saying it is done does not count; the log decides.
  • Flagship features are the most likely to die silently, because nobody questions the flagship.
  • A timer that is firing does not mean output is being produced. Verify "running" and "running correctly" separately.
  • An honest audit report is more useful than a pretty architecture diagram. That is also how this whole series is written.

This is the opening post of the Agentic Design Patterns × Lobster Fleet series. Over the next 21 posts, we lay the audit open chapter by chapter: what the book says, what we thought we had done, what the audit exposed instead, and finally how we fixed it or admitted we have not fixed it yet. The credibility of this series comes from our willingness to publicly prove ourselves wrong.

FAQ

How do you audit your own AI agent system?

The core fits in one sentence: trust no self-reported claim, accept only evidence of execution. First assume every 'done' is false and treat it as a hypothesis to be proven. Then check when it last actually ran and whether it actually produces anything. Finally, reconcile the numbers you tell people by counting, one by one, how many are actually alive.

What does an audit find after building an agent fleet from Agentic Design Patterns?

In our case, of the 14 chapters we claimed as implemented, not one was completely clean. agent-meeting.py, the agent we were proudest of, had been dead for three months with no timer calling it; we said six brands were running in parallel while only one brand gateway was alive; the LLM judge marked a system that had answered correctly as wrong; and 98.8% of the money the gateway burned in one week went to rereading its history again and again.

How do you check when an automated service last actually ran?

Check whether its timer still exists and how long ago the last log entry was written. You can find the timer with systemctl --user list-timers, then check when the log file was last updated. Red flag: a flagship feature's log stops weeks or even months ago.

If a timer is firing, does that mean the agent is working?

No. A timer calling it does not mean it produces anything. Verify 'is it running' and 'is what it produces right' separately, for example by counting the real output in the log. Red flag: plenty of runs, and real output close to 0.

Why do things still go wrong when monitoring alerts are in place?

The thresholds may be too loose. We had installed a whole suite of alerts, yet the two most painful bills both happened while monitoring was in place, because the thresholds were set so loose that the money slipped out underneath them and no alarm went off. When you audit, check whether your monitoring thresholds are so loose that a disaster can slip out underneath them.

Lobster Fleet · Pattern Audit · Part 1 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.