AI AgentHuman-in-the-LoopMonitoring and AlertingBuildInPublicSolo Business

Full Monitoring, Nobody Reading It: How to Make Human-in-the-Loop Actually Work for AI Agents

· 9 min read
Lobster Fleet · Pattern Audit · Part 14 of 25
Table of Contents
  1. Monitoring everywhere, and no human in the loop
  2. Another irony: alerts that flood people until they go numb
  3. First, spend a few minutes checking your own human-in-the-loop
  4. What we actually changed
  5. Human-in-the-loop checkup list
  6. Four lines from this chapter

Chapter 13 of Agentic Design Patterns covers Human-in-the-Loop. The core idea in one line: not every decision should be automated, and a human should stay at the points where a machine should not make the call alone. But the real bar in this chapter is hidden somewhere easy to skip: the interface you prepared for the "human", will a human actually read it?

We read it and nodded. We had soft-warning logs from the guardrails, semantic heartbeat scores, error classification, and a cost record for every call, all dutifully written to files, waiting for someone to look. Human-in-the-loop, implemented. Then we ran an adversarial audit on ourselves and found something embarrassing: the machines were keeping diligent records, and not a single person had ever sat down and read those logs.

Monitoring everywhere, and no human in the loop

Open ~/.openclaw/logs and there are four or five files in it: guardrail.log for output the guardrails let through but flagged as suspicious, semantic-quality.jsonl for semantic quality, llm-errors.log for LLM errors, and llm-cost.log for the cost of every call. Every one of them is being written. Not one of them is being opened.

This is the exact opposite of human-in-the-loop. The book wants you to put a human at the points where intervention is needed. We had collected all the signals, but never built an interface a person would be willing to open on their own. The guardrails catching suspicious output, the semantic heartbeat losing points, the bill burning: these are all moments that call for human judgment, yet that "human" was never called to the scene. A signal buried in a pile of files nobody reads is no different from having no signal.

So we built review-digest: a read-only, on-demand review panel that does not push to TG (Telegram). It gathers those four or five scattered logs into one page, with hard blocks marked red, soft warnings marked yellow, semantic score drops marked red, and cost ranked to show the top five burners. It deliberately does not push notifications, because a human-in-the-loop interface is not "one more notification". It is "a place a person is willing to open on their own".

Another irony: alerts that flood people until they go numb

If the first pit is that nobody reads the signals, the second pit is the exact opposite: so many signals that nobody wants to read them.

We have a Claude Stop hook, tg-mirror-check.py, built with good intentions. You ask Claude something from TG, Claude answers in the terminal but forgets to mirror the answer back to TG, and with nothing showing on your phone, the delivery silently drops. This hook checks once at the end of every turn and reminds you when it catches one.

The problem is that it lives in ~/.claude/hooks, which every session shares. It fires at the end of every Claude session, including terminal coding sessions that have nothing to do with TG. Those sessions read the same shared TG log and, over a message that does not belong to them, push the boss a line saying "you just forgot to mirror that". Pure cross-session noise. Worse, at first it had no dedup: the same unmirrored message was re-sent every turn and every loop tick, and the boss got the same alert roughly every 5 minutes.

An alert built to pull a human into the loop ended up flooding that human into numbness. By the third time, people skip it automatically, and they skip the one that truly matters along with the rest. This is more dangerous than having no alert, because you think someone is watching.

First, spend a few minutes checking your own human-in-the-loop

01 Count: how many are "machines keeping records", and how many will "a human actually read" The first breach in human-in-the-loop is signals that get collected diligently while nobody actually goes to look at them. First count how many logs or alert channels are quietly writing, then honestly answer how many of them you actually opened this week.

# How many places are quietly writing logs / pushing alerts
ls ~/.openclaw/logs/*.log ~/.openclaw/logs/*.jsonl 2>/dev/null | wc -l
# Of these, how many did you actually sit down and read this week? (Answer honestly, usually 0)

Red flag: you can count a whole row of logs being written, but cannot say when you last actually sat down to read one. Machines keeping records does not mean a human is in the loop.

02 Is your review interface built for people to read, or written and never touched A log written to disk that nobody opens does not count as human-in-the-loop. Check whether you have an interface that "gathers scattered signals into one page a person is willing to pull on their own", or whether the signals are spread across eight files, each doing its own thing. While you are at it, see how far apart that log's last read and last write are.

# How far apart this monitoring log's last read and last write are
stat -c '讀 %x/寫 %y' ~/.openclaw/logs/guardrail.log
# Read far older than write = constantly written, nobody reading (if atime is disabled, ask yourself whether you have a review panel instead)

Red flag: you have no review panel that gathers everything into one page and is pulled on demand. The signals sit in separate files, and to see them you have to open eight at once.

03 Do your alerts repeat and flood Flooding is a slow poison for human-in-the-loop. Check whether your alerts "fire every time they trigger" or "dedup and fire only on a state transition (OFF→ON)", then look at how many times the same alert actually repeats in the log. Also check whether a shared hook fires in places where it should stay quiet.

# Does the alert script dedup / fire only on an OFF→ON transition
grep -niE 'dedup|already.?alert|last.?state|OFF.*ON|轉態|transition' your-alert.*
# How many times the same alert line was sent in the log
sort alerts.log | uniq -c | sort -rn | head

Red flag: no dedup or transition check anywhere, the same alert repeated dozens of times in the log, or a shared hook firing in contexts that have nothing to do with it.

What we actually changed

Two cuts. The first cut gathered the scattered signals into one page, review-digest: read-only, run on demand, zero noise, one command that pulls up guardrails, semantics, errors and cost so you can read them all at once, deliberately not pushing to TG. The second cut, on 2026-07-04, fixed the flooding hook. We added dedup so the same missed mirror only alerts once, then switched TG push off entirely and made it log-only, so the signal stays in missed-mirrors.log instead of flooding the boss.

It is worth being clear about what it is not: review-digest is genuinely built and runs. It is not an empty shell. The embarrassing part is not the tool. It is the reason the tool had to exist: we first piled up a mountain of monitoring nobody read, and then that hook over-reminded until people did not want to read at all. The two pits are two sides of the same coin. Signals nobody reads, and signals so numerous nobody wants to read them, both leave human-in-the-loop existing in name only.

Human-in-the-loop checkup list

Run it against your own system:

  • How many logs / alerts are quietly writing, and how many of them you actually read this week
  • Whether there is a review interface that gathers scattered signals into one page a person is willing to pull on their own
  • Whether that interface is a zero-noise panel that "waits for a person to look", or a notification stream that "keeps pushing"
  • Whether alerts are deduplicated, and fire only on a state transition (OFF→ON)
  • Whether a shared hook fires in sessions / contexts that have nothing to do with it
  • How many times the same alert repeats in the log, and whether people have already started skipping it automatically

Four lines from this chapter

  • The bar for human-in-the-loop is not "is there an interface". It is "will a person actually open that interface". Logs that fill up and never get read mean there is no human in the loop.
  • Monitoring is for machines to record, and the review panel is for humans to read. Gather scattered signals into one page, pulled on demand, with zero noise, and people will actually look.
  • Flooding is a slow poison for human-in-the-loop. By the third time the same alert arrives, people start skipping it automatically, and they skip the one that truly matters along with it. Dedup, and fire only on state transitions.
  • A shared alert will speak up where it should stay silent. Before sending, ask one question: should this person receive this one, right now?

Source locations: ~/.openclaw/scripts/review-digest.sh (a read-only, on-demand, zero-noise human review panel that gathers Ch18 guardrails / Ch19 semantics / Ch12 errors / Ch16 cost, and only brings in AI for root cause with --diagnose), ~/.claude/hooks/tg-mirror-check.py (detects missed TG mirroring; on 2026-07-04 it gained dedup and had cross-session push switched off in favour of log-only, with the signal landing in missed-mirrors.log).

This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented has been verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publicly prove ourselves wrong.

FAQ

What is human-in-the-loop for AI agents?

Human-in-the-loop is the subject of Chapter 13 of Agentic Design Patterns. The core idea: not every decision should be automated, and a human should stay at the points where a machine should not make the call alone. The real bar is whether a human will actually open the interface you prepared for them. Logs that fill up and never get read mean there is no human in the loop.

My AI agent's monitoring logs pile up and nobody reads them. What should I do?

Build a review panel that gathers the scattered signals into one page, is pulled on demand, and makes zero noise. Our review-digest is read-only, runs on demand and does not push to TG (Telegram). It gathers those four or five guardrail, semantic, error and cost logs into one page: hard blocks in red, soft warnings in yellow, semantic score drops in red, and cost ranked to show the top five burners. It deliberately does not push notifications, because a human-in-the-loop interface is not one more notification but a place a person is willing to open on their own.

How can I check whether anyone is actually reading my monitoring?

Count how many logs or alert channels are quietly writing, then honestly answer how many of them you actually opened this week (usually 0). You can also use stat to see how far apart a log's last read and last write are: a read far older than the write means it is constantly written and nobody is reading. If atime is disabled, ask yourself whether you have a review panel instead.

How do I stop repeated alerts from flooding people?

Dedup, and fire only on a state transition (OFF→ON). By the third time the same alert arrives, people start skipping it automatically, and they skip the one that truly matters along with it. That is more dangerous than having no alert, because you think someone is watching. On 2026-07-04 we added dedup to our flooding hook so the same missed mirror alerts only once, then switched TG push off and made it log-only.

Why does a shared hook turn into alert noise?

A hook in a location every session shares fires at the end of every session, including sessions that have nothing to do with it. Our tg-mirror-check.py lives in ~/.claude/hooks, so terminal coding sessions unrelated to TG read the same shared TG log and pushed the boss alerts over messages that did not belong to them. With no dedup at first, the boss got the same alert roughly every 5 minutes. Before sending, ask one question: should this person receive this one, right now?

Lobster Fleet · Pattern Audit · Part 14 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.