Inter-Agent Communication Audit: How to Check Which A2A Channels Are Still Alive
Table of Contents
A few months ago, in Multi-Agent Coordination Architecture: How to Keep 4 AI Agents from Stepping on Each Other, I wrote up the message pipelines between our agents as an elegant mechanism from which collaboration grew on its own. This post checks, one by one, whether the two communication channels I claimed actually work. The file relay is real and still flowing. The other one, gateway IPC on port 18789, has been at zero connections for a long time, yet it is still written, intact, in the live config file and in a whole row of backups, looking every bit as credible as a live channel.
Chapter 15 of Agentic Design Patterns covers Inter-Agent Communication (A2A). The core idea, in one line: the value of a multi-agent system is not how smart each agent is, but whether the messages passed between them actually arrive. Once the channel breaks, a group of agents degrades into a set of independent scripts talking past each other.
After reading it, we were confident: we have two of these. One is the file relay: when an agent finishes a post, it drops a summary into a queue, and the other agents read it the next time they wake up. The other is gateway IPC, running on local port 18789, letting the fleet route to one another through a single gateway. Two channels, A2A implemented. Then we ran an adversarial audit on ourselves. The truth for this chapter: of the two, one had long been at zero connections.
One is still flowing, one is a dead pipe
Start with the file relay. Open its queue file, relay-queue.json, and it is alive: mindthread handed main the engagement numbers for "6,500 followers, zero manual posts" as case-study material, probe handed main the finding that "90% of LLM applications have vulnerabilities" to pick up the business angle, and advisor passed over a personal finance post to get a sales perspective. The latest consumedAt is yesterday, and the four agents (main, mindthread, probe, advisor) really are handing off to one another. This one is real.
Now the gateway IPC. The audit found no process listening on port 18789, and zero established connections. The reason is not hard to see: after the migration to OpenRouter, every script calls OpenRouter directly, and not a single route passes through that local gateway anymore. The last time that bus carried a message was before the migration.
The most glaring part is not that it died. It is that it is still claimed everywhere after dying. port: 18789 is still written, intact, in the gateway block of the live openclaw.json, and allowedOrigins still lists http://localhost:18789. The backups are worse: a whole row of openclaw.json.bak.* files, each repeating the same 18789. Anyone who opens the config file, or any agent that reads this section into its context, will naturally assume the fleet has an IPC bus running. In reality it is a dead pipe with zero connections. Half of the communication mechanisms you claim are dead, and they died without a sound.
Spend a few minutes verifying each channel you claim is alive
01 For every port on your architecture diagram, ask "is anyone still listening?" Every IPC / gateway channel you claim, in documents or out loud, lay them out and verify them one at a time. If nothing is listening and established connections are zero, the channel is dead, however complete the config file looks.
# Is any process listening on this port (no output = nobody listening)
ss -ltnp | grep 18789
# Number of established connections on this port (0 = nobody using it)
ss -tnp | grep -c 18789
# Without ss, use netstat: netstat -ltnp | grep 18789
Red flag: a communication port you name in your architecture shows nobody listening in ss, zero established connections, and you cannot remember the last day you "confirmed it was alive".
02 A relay file existing does not mean the relay is still flowing A queue file sitting there only proves it was once created. It does not prove anything goes in or out today. Check when it was last consumed, and whether a pile of items is stuck with nobody picking them up.
# Time of the last consumption, should be recent
jq -r 'map(select(.consumed))|max_by(.consumedAt)|.consumedAt' relay-queue.json
# Number of items not yet consumed, piling up
jq '[.[]|select(.consumed==false)]|length' relay-queue.json
Red flag: the file is there, but the last consumedAt is several days old; or unconsumed items keep piling up and never get read. The relay exists in name but has actually stopped flowing.
03 After a migration or architecture change, have the dead mechanisms been cleaned out of config and claims? Retiring a channel does not end when you switch it off. Its name and port will linger in config files, documents and backups, so that you keep claiming, time after time, a mechanism nobody has connected to in ages. Dig it out and see how many places it is still scattered across.
# Is the dead port still in the live config files (backups excluded)
grep -rn 18789 ~/.openclaw --include='*.json' | grep -v '\.bak'
# How many backups are still repeating it
grep -rl 18789 ~/.openclaw/*.bak* | wc -l
Red flag: a mechanism already at zero connections is still written in gateway.port of the live config file and still echoing through a whole row of .bak files. The config file says it exists, reality says it is dead, and the two do not match.
What our audit turned up
Two channels: one holds up under verification, one does not. The file relay is backed by real timestamps in relay-queue.json, and the four agents were still handing off yesterday. The gateway IPC is reduced to a number in a config file: nobody listening on port 18789, zero connections, and not a single message carried since the migration to OpenRouter. It is not broken. The whole route was bypassed, and nobody went back to delete it from what we claim.
It is worth being clear about what this is "not": it is not that the relay died too, and it is not that the agents have no communication at all. The truth is subtler and more common. Of the two channels claimed, one is real and one is fake, and the fake one, because it is still printed in the config file, looks every bit as credible as the real one. This is exactly where multi-agent systems most easily fool themselves: a channel's "claim" and a channel's "connection count" are two different things. You think you are describing a live bus, when you are actually describing a pipe nobody has connected to in three months.
Inter-agent communication checklist
Run this against your own multi-agent system:
- For every A2A channel you claim, can you state, one by one, "how it is verified alive today"?
- For every IPC / gateway port, does ss show anything listening, and how many established connections are there?
- Is the last
consumedAton the relay / queue file recent, or has it stopped flowing for days? - After a migration or architecture change, have the retired old channels been deleted from the live config files?
- How many documents and backups still carry the dead mechanism's name / port, waiting for you to claim it again?
- Can you say when each channel "last actually carried a message"?
Four lines from this chapter
- The value of multiple agents is in the channels, not the count. Before claiming how many communication channels you have, verify that each one still works today.
- A channel's "claim" and its "connection count" are two different things. A port written in a config file does not mean anyone is listening on it.
- A relay file existing does not mean the relay is flowing. Check the last
consumedAt; a queue that has stopped flowing for days is as good as dead. - Retiring a channel means clearing out its config, documentation and verbal claims together. Otherwise you will keep claiming a dead mechanism with zero connections, and it will look every bit as credible as a live one.
Source location: ~/.openclaw/scripts/team-context.sh (injects the relay into the prompt), content-relay.sh (relay add/get/consume/clean); the queue lives at ~/.openclaw/data/relay-queue.json (four agents, last consumed 2026-07-05); the retired gateway IPC port 18789 is still written in the gateway block of ~/.openclaw/openclaw.json and in a whole row of openclaw.json.bak.* files, measured at zero connections.
This is part of the Agentic Design Patterns × Lobster Fleet series. We follow the book to systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.
FAQ
How do I check whether an agent IPC or gateway port is still in use?
Check every port on your architecture diagram with ss, one at a time. ss -ltnp piped to grep for the port shows whether any process is listening (no output means nobody is listening), and ss -tnp piped to grep -c counts the established connections (0 means nobody is using it). Without ss, use netstat. If nothing is listening and established connections are zero, the channel is dead, however complete the config file looks.
If the relay queue file still exists, is the relay between agents still working?
Not necessarily. A queue file sitting there only proves it was once created, not that anything goes in or out today. Check whether the last consumedAt is recent and whether unconsumed items keep piling up. If the last consumedAt is several days old, or items pile up and never get read, the relay exists in name but has stopped flowing.
How do I clean out an old communication mechanism after a migration?
Retiring a channel does not end when you switch it off: its name and port linger in config files, documents and backups. Grep the live config files (excluding .bak backups) to see whether the dead port is still there, then count how many backups still repeat it. Clear out the config, documentation and verbal claims together, or you will keep claiming a dead mechanism with zero connections.
What is inter-agent communication (A2A), and why does a broken channel matter?
Chapter 15 of Agentic Design Patterns covers Inter-Agent Communication (A2A). The value of a multi-agent system is not how smart each agent is, but whether the messages passed between them actually arrive. Once the channel breaks, a group of agents degrades into a set of independent scripts talking past each other.
Why is a gateway port in the config file not proof that the channel is alive?
A channel's claim and its connection count are two different things. After our migration to OpenRouter, every script called OpenRouter directly, so port 18789 had no process listening and zero connections, yet it was still written in the live openclaw.json and in a whole row of backups. A port written in a config file does not mean anyone is listening on it.