AI AgentMulti-Step PlanningFlagship FeaturesBuildInPublicSolo Business

How to Check Whether Your Flagship AI Agent Still Runs: Ours Had Been Dead for 3 Months

· 7 min read
Lobster Fleet · Pattern Audit · Part 7 of 25
Table of Contents
  1. Opened Up, the Flagship Was an Empty Shell Kept for Show
  2. Spend a Few Minutes Checking Your Own Flagship Features
  3. Flagship Feature Health Checklist
  4. This Chapter in Four Sentences

A few months ago, in We Made 4 AI Agents Talk to Each Other on Discord, I wrote up "they hold their own meetings every morning, review their own data, make their own recommendations" as something happening in my setup every day. This post goes back to check the last run time of that flagship meeting agent: it has not held a single meeting in three months, and I had assumed all along that it was meeting every day. Code that is not broken is not the same as code that is still alive.

Chapter 6 of Agentic Design Patterns covers planning. The core idea, in one line: faced with a complex task, an agent should first break it into an ordered multi-step plan and then execute it step by step, rather than forcing the whole thing through in one go.

After reading it, we thought we had an ace for this chapter. The fleet has a script called agent-meeting.py in which four roles hold a meeting every day: they read fleet data, read yesterday's meeting memory, read cross-agent signals, then produce the day's discussion, extract action directives and write them back to the signal bus as tomorrow's plan. It was the script we most liked to show off, a living showcase of multi-step planning. Then we ran an adversarial audit on ourselves. The embarrassment of this chapter: this flagship had not run in 3 months, and we had assumed all along that it was holding a meeting every day.

Opened Up, the Flagship Was an Empty Shell Kept for Show

The first cut went at "when did it last actually run". The script file itself stopped at 2026-04-03 and has not been touched since. That could still be explained as "written once and never needed changing", so we went through what it writes itself: the meeting memory file meeting_memory has 14 entries in total, and the date on the last one is also 2026-04-03. It was not that nobody bothered to edit it. After that day, it did not hold a single meeting. As of today, that is a full 3 months.

The second cut went deeper: why did it die? Because nothing was calling it at all. We scanned all 23 systemd timers: not one named meeting, not one named agent. Then we grepped the whole fleet directory for "agent-meeting". Apart from the script itself, the only hit was a wrapper sitting in _archive/ whose entire contents are one line, exec python3 agent-meeting.py. A shell in an archive folder, wrapped around another shell that nobody calls. No timer, no service, no cron: the last time this "meets every day" agent was triggered was three months ago, when someone ran it once by hand.

It is worth being clear about what it is not: the code is not broken, it runs when started, the Gemini API key is still there, and the Discord webhook still works. If you ran it by hand today, it would hold a perfectly tidy meeting as usual. What broke is not whether it can run; it is that no wire at all is pulling it to run. This is the most dangerous state for a flagship artifact: it sits there fully intact, and the picture in your head is frozen on the day it launched.

Spend a Few Minutes Checking Your Own Flagship Features

01 When did it last actually run Do not ask "can it run"; ask "what day did it last run". Check the script file's modification time, then the date of the last entry in the data file it writes, and look at both together.

# How long ago the script was last touched (a dead script stops on a particular day)
stat -c "%y  %n" ~/.openclaw/scripts/你的招牌.py
# Harder evidence: the date of the last entry in the data file it produces
python3 -c "import json;m=json.load(open('.../data.json'))['records'];print(m[-1]['date'])"

Red flag: the last entry stopped weeks or months ago, and you are still telling people it "runs every day".

02 Is any timer / cron actually calling it Whether it can run and whether anything tells it to run are two different things. List every schedule, search it by script name, then grep globally to see who references it.

# Is any schedule actually triggering it
systemctl --user list-timers --all | grep -i 你的招牌
crontab -l | grep -i 你的招牌
# Global search for references: apart from itself, who else calls it
grep -rIl "你的招牌" ~/.openclaw/

Red flag: list-timers returns nothing; grep finds only the script itself, plus a wrapper lying in _archive/ or old/. No trigger source = it will never wake up on its own.

03 Does the step order it claims match the order it actually runs The point of planning is order. Read the script from top to bottom and mark which line "decides what to do" and which line "gathers the evidence". Our autopost claims publicly to "research first, then pick the topic". Opened up, the topic-selection line uses a clock modulo (POST_SLOT % 題庫數, the slot number modulo the size of the topic pool) to lock the topic in first. The research step comes after, and even which URL to fetch is chosen according to the topic already decided. The order is reversed: it does not research and then decide what to write, it decides what to write and then fetches material to fill it in.

# Find which line "decides the topic" and which "does research", and see which comes first
grep -nE '選題|pillar|PILLAR=|research|summarize' 你的招牌.sh

Red flag: the line that decides what to do sits before the line that gathers the evidence. Your "plan first, then execute" is really "decide first, do the homework afterwards".

Flagship Feature Health Checklist

Run through it against the agent or script you are proudest of:

  • Can you say what day it last actually ran (not "it can run", but "it last ran")
  • Is any timer / cron / service triggering it, or can it only be kicked off by hand
  • Grep its name globally: apart from itself, who references it, and are those references live or lying in _archive/
  • Is the date on the last entry in the data file it produces fresh
  • Does the step order it claims (research first, then pick the topic) match the actual top-to-bottom order in the source code
  • Is the impression you give when introducing it to others something you verified this week, or a snapshot from the day it launched

This Chapter in Four Sentences

  • A flagship artifact may just be there for show. Code that is not broken does not mean it is still alive; the question to ask is "what day did it last actually run".
  • Whether it can run and whether anything is telling it to run are two different things. With no timer and no cron, even the most polished agent is just a shell waiting for someone to kick it off by hand.
  • Reconcile the claimed step order against the source code. Deciding first and backfilling the research is not planning; it is running backwards while describing it as running forwards.
  • Your impression of a flagship feature most easily freezes on its launch day. Every so often, verify its trigger source and last run time yourself, and do not let memory stand in for the current state.

Source locations: ~/.openclaw/scripts/agent-meeting.py (the flagship planning agent, last run 2026-04-03, no timer triggering it); the only thing that references it is ~/.openclaw/scripts/_archive/discord-agent-chat.sh (a wrapper in the archive folder). Evidence of the reversed order is in ~/.openclaw/scripts/moltbook-autopost.sh (the pillar is decided first by clock modulo, research comes after).

This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented has been verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publicly prove ourselves wrong.

FAQ

How do you tell whether an AI agent or scheduled script is still running?

Do not ask 'can it run'; ask 'what day did it last run'. Check the script file's modification time, then the date of the last entry in the data file it writes, and look at both together. If the last entry stopped weeks or months ago while you are still telling people it runs every day, that is a red flag.

Why would an agent stop running when its code is not broken?

Whether it can run and whether anything tells it to run are two different things. Our flagship meeting agent's code was not broken and it still held a meeting when run by hand, but none of the 23 systemd timers triggered it, and the only reference to it was a wrapper in an archive folder. With no timer, no service and no cron, it will never wake up on its own.

How do you check whether any timer or cron job triggers a script?

Search systemctl --user list-timers --all and crontab -l for the script name, then grep the whole directory for the name to see who else references it. If list-timers returns nothing and grep finds only the script itself plus a wrapper lying in _archive/ or old/, there is no trigger source.

How do you verify that an agent really plans first and then executes?

Read the script from top to bottom and mark which line decides what to do and which line gathers the evidence. If the deciding line sits before the evidence-gathering line, your 'plan first, then execute' is really 'decide first, do the homework afterwards'. Our autopost claimed to research first and then pick the topic, but the topic was locked in first by a clock modulo and the research came after.

What is the planning pattern in Agentic Design Patterns?

Chapter 6 covers planning. The core idea, in one line: faced with a complex task, an agent should first break it into an ordered multi-step plan and then execute it step by step, rather than forcing the whole thing through in one go.

Lobster Fleet · Pattern Audit · Part 7 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.