Multi-Agent Collaboration Audit: How to Check How Many Agents Are Really Running and What They Cost
Table of Contents
A few months ago, in Multi-Agent Architecture: How to Keep 4 AIs from Fighting, I wrote up the whole six-layer collaboration mechanism as a success story: a group of agents dividing the work seamlessly and handing off to one another on their own. This post goes back and counts properly. Publicly I inflated it into six agents running in parallel, an audit found only four actually alive, and today, to save money, I personally shut down the only agent that ever came to collect from that relay pipeline.
Chapter 7 of Agentic Design Patterns covers multi-agent collaboration. The core idea, in one line: complex work should not be stuffed into one all-purpose agent. It should be split across several agents, each with its own specialty, passing results to one another and picking up where the other left off.
We read it and felt very pleased with ourselves. We had the full set: four roles, security, content, community and finance, each posting on its own, with a content relay pipeline laid between them. When A finishes a post it drops a lead for B, and B picks it up when it wakes. Multi-agent collaboration, implemented. Then we ran an adversarial audit on ourselves. The truth of this chapter has three layers, each more embarrassing than the last.
The pipeline is real, the numbers were inflated, and we shut down the recipient today
The good part first: the relay pipeline is not decoration. relay-queue.json holds six messages. After the security agent posted that 90% of LLM applications have vulnerabilities, it dropped one to the main agent reminding it to pick up the business angle; after the finance agent posted, it dropped a strategy lead. The most recent message was consumed yesterday at 12:03 p.m. Looking back today, the pipeline really is still flowing.
Now the embarrassing part. Publicly we talked about six brand agents running in parallel. An audit count found only four actually alive (security, content, community, finance). There really are six folders in the profile directory, but two of them have a last-modified time stuck three weeks ago: empty shells left behind after the migration and never cleaned up. The interesting part is that the relay pipeline's own code says 4-agent fleet from start to finish. The code was honest. It was our mouths that turned four into six.
The harshest layer is today. Every one of those six messages has the same recipient: the main agent. And the main agent's gateway burns $17 a week (it is the central character of the bill in Chapter 16). Today, to save money, we shut it down outright. So the pipeline is still there and the senders are still stuffing messages into it, but the only agent that ever came to collect them was shut down by our own hand. The most expensive part of multi-agent collaboration was never the relaying. It is feeding the mouths that talk back.
Spend a few minutes checking your own multi-agent collaboration
01 Check how many agents are actually alive, against how many you claim Go through the directory where your agent configs live and look at each one's last-modified time. The cold ones are shells that were never cleaned up after a migration or rewrite, yet you still count them in your parallelism figure.
# List every agent profile; the cold ones (mtime stuck far in the past) are empty shells
ls -lt ~/hermes-bench/data/profiles/
# Or count the gateways that are actually still active, not the config files
systemctl --user list-units --state=active | grep -i gateway
Red flag: the directory has six agents, but half of them were last modified weeks ago. Half of the parallelism you claim is dead. How to count live agents at the process level and reconcile that count against your public figure is covered separately in the parallelization chapter.
02 Check whether the relay channel between agents still moves today, and who is still receiving A relay pipeline most easily turns into a one-way dead letter box: the senders keep diligently stuffing it while nobody on the receiving end has read it for ages. Look at when it was last consumed, then look at who all the messages flow to.
# When the relay queue was last "consumed", not when it was last written to
jq -r 'map(.consumedAt)|sort|last' ~/.openclaw/data/relay-queue.json
# Who are all the messages addressed to? Everything sent to one agent = switch that one off and the whole line breaks
jq -r '.[].to' ~/.openclaw/data/relay-queue.json | sort | uniq -c
Red flag: the channel still moves, but every message flows to the same recipient, and that recipient happens to be your most expensive agent, the one you most want to shut down. You think it is mesh collaboration; really it is one central node straining to hold everything up.
03 Check whether each agent costs more than it produces, and shut it down if it does This item ties straight back to Chapter 16. Put what each agent burned this week next to what it actually produced. Collaboration sounds beautiful, but every additional agent that talks back is one more bill.
# Compare each agent's spend this week (grouped and summed from the cost log)
grep '"agent"' ~/.openclaw/logs/llm-cost.log | jq -sr 'group_by(.agent)[]|"\(.[0].agent) $\(map(.cost)|add)"'
# Then look at the gateway that burns the most: is its output worth it
systemctl --user status hermes-gateway.service | grep Active
Red flag: one agent burns far more per week than the value it produces, and you keep it only because it is a node on the architecture diagram.
What we actually did today
No parameter tuning. We simply stopped the main agent's gateway and masked it. It cost $17 a week, and almost all of its output was rereading the conversation history again and again, which is not worth that price. We left the relay pipeline alone and let it sit: a quiet JSON file costs almost nothing, and if the day comes when we really need a central agent to pull things together, we will turn it back on. Keeping the senders and shutting off the most expensive receiver is not breaking collaboration. It is admitting that one of the four agents is not currently worth keeping on to talk back. In the end, multi-agent collaboration is the same thing as resource awareness: ruthless pruning.
Multi-agent collaboration checklist
Run it against your own system:
- Does the number of agents in your config match the ones actually still running, or does it include shells left over from a migration
- Does the parallelism you claim publicly match the agent count written in your code
- When was the relay channel between agents last consumed, or has it long since become a one-way dead letter box
- Do all messages flow to one central agent, so that switching that one off breaks the whole line
- Is what each agent burned this week worth its output
- Is there an expensive agent kept around only because it is a node on the architecture diagram
Four lines from this chapter
- A real relay pipeline does not mean the collaboration is real. Count the agents that are actually alive and actually receiving, not the nodes on the architecture diagram.
- Public parallelism figures inflate. If the code says four, do not say six; one audit count and it falls apart.
- Mesh collaboration is often fake, with one central agent straining underneath to hold it all up. Switch that center off and the whole line breaks.
- The most expensive part of multi-agent collaboration is not the relaying; it is feeding the mouths that talk back. Shutting down an agent that costs more than it produces is nothing to be ashamed of. That is the discipline collaboration should have.
Source location: ~/.openclaw/scripts/team-context.sh (injects teammates' leads into the posting prompt, with 4-agent fleet hardcoded in the code), content-relay.sh (add/get/consume for the cross-agent relay pipeline); the relay queue's actual file is ~/.openclaw/data/relay-queue.json (all six messages addressed to main, last consumed on 2026-07-05); the main agent gateway that was shut down is hermes-gateway.service (masked to inactive today, saving $17 a week).
This is part of the Agentic Design Patterns × Lobster Fleet series. We follow the book to systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.
FAQ
How do I check how many agents in a multi-agent system are actually still alive?
Go through the directory where your agent configs live and look at each profile's last-modified time: the ones stuck far in the past are shells that were never cleaned up after a migration or rewrite. You can also count the gateways that are actually still active with systemctl --user list-units --state=active, rather than counting config files. If the directory has six agents but half of them were last modified weeks ago, half of the parallelism you claim is dead.
How do I know whether the relay pipeline between agents is still working?
Check when the relay queue was last consumed, not when it was last written to, then check who all the messages are addressed to. A relay pipeline most easily turns into a one-way dead letter box: the senders keep diligently stuffing it while nobody on the receiving end reads it.
What is the risk when every message flows to the same agent?
It means what looks like mesh collaboration is really one central node straining to hold everything up: switch that agent off and the whole line breaks. All six messages in our relay queue were addressed to the main agent, whose gateway burned $17 a week. After we shut it down to save money, the senders kept stuffing messages in, but the only agent that ever came to collect them was gone.
How do I decide whether an agent should be shut down?
Put what each agent burned this week next to what it actually produced. If one burns far more than the value it produces, and you keep it only because it is a node on the architecture diagram, shut it down. Our main agent cost $17 a week and almost all of its output was rereading the conversation history, so we stopped its gateway and masked it.
Should you remove the relay pipeline after shutting down the central agent?
We did not. We left it alone: a quiet JSON file costs almost nothing, and if the day comes when we really need a central agent to pull things together, we will turn it back on. Keeping the senders and shutting off the most expensive receiver is not breaking collaboration. It is admitting that one of the four agents is not currently worth keeping on to talk back.