AI AgentLearning LoopsSelf-ReflectionBuildInPublicSolo Business

How to Tell Whether an AI Agent's Learning Loop Is Actually Learning or Just Spinning

· 9 min read
Lobster Fleet · Pattern Audit · Part 10 of 25
Table of Contents
  1. All three loops are running, but the ledger is empty
  2. Spend a few minutes checking your own learning loops
  3. The root cause we found digging through the logs
  4. We did not pretend to fix it
  5. Learning loop health checklist
  6. This chapter in four sentences

A few months ago, in We Built a Self-Learning AI Sales System in 48 Hours, I presented "it learns from its own results and gets smarter every day" as the most valuable part of the whole system. This post takes that line apart: my three learning loops run on schedule every night, yet the ledger is empty when you open it, and for three weeks straight they produced the exact same 5 sentences. A timer that runs is not a loop that learns.

Chapter 9 of Agentic Design Patterns covers learning and adaptation. The core idea, in one line: an agent has to be able to learn from the results of what it has done, then adjust its behaviour the next time. Instead of starting from zero every time, yesterday's results change how it works today.

When we finished reading it, we thought we were already ahead on this one. We have three learning loops: dream-cycle reflects in advance late at night while the GPU is idle, daily-reflect runs a fleet-wide strategy review every night at 23:00, and lobster-learn distils the reflections into behaviour changes every week. All three systemd timers fire on time, and the logs do not miss a line. It looks incredibly diligent. Then we ran an adversarial audit on ourselves. The truth of this chapter: the timers move every night, and what they learn is 0.

All three loops are running, but the ledger is empty

Start with the data file the first loop consumes. Every night daily-reflect asks itself "which post beat its target, which fell short, and what topic should each agent go after tomorrow". Every one of those questions needs post performance data to answer. The performance file it reads looks like this:

$ cat ~/.openclaw/data/post-performance.json
[]

An empty array. Further downstream, dream-cycle reads "the best-performing posts" every night as a style reference, and what it reads is empty too; the performance report it produces, laid out in full, looks like this:

## Summary
- Total posts tracked: 0
- Average score: 0

## Actionable Insights
- High-score titles tend to: [agent will fill this after reading]
- Best posting time: [agent will analyze created_at patterns]

Those [agent will fill this after reading] entries are placeholders, and not once since launch have they been filled. Filling them requires post results first, and the results are always 0.

The most glaring part is the self-reflection in dream-cycle. Every night it reads the system logs and sums up the patterns in the problems. I laid out the whole reflection file: from 6/18 to 7/06, every night produced the same 5-sentence passage: "The most frequent problem is that the Gateway failed to start", "The time pattern is every night at 22:00 and in the afternoon at 17:00", "Check the Gateway status regularly", "Increase the self-healing frequency". Three weeks in a row, nothing of substance changed. It is not learning. It is repeating yesterday's version of itself.

Spend a few minutes checking your own learning loops

01 Check whether real signal is entering the loop The first cell of a learning loop is always data. Open the data file it reads. If it is an empty [] or 0 records, however polished the reflection that follows, it is talking to thin air.

# Count how many records the learning loop's input data file actually has
jq 'length' ~/.openclaw/data/post-performance.json
# Or check whether the report it produces says "tracked: 0"
grep -iE 'tracked: 0|total: 0|\[.*will fill' report.md

Red flag: the data file is [], the report says tracked: 0, or it contains placeholders that have never been filled. The loop is reflecting on an empty set.

02 Diff this week's learning output against last week's to see whether it changed If it is really learning, the output changes as new data comes in. Put this week's reflection next to last week's and diff them. If they are nearly identical, it is not learning, it is repeating.

# See how many distinct sentences the whole reflection file actually contains
grep -vE '^\s*$|^##' dream-reflections.md | sort | uniq -c | sort -rn | head
# Or diff the two most recent outputs directly
diff <(sed -n '/上週那段/,/結束/p' f) <(sed -n '/這週那段/,/結束/p' f)

Red flag: three consecutive outputs are highly similar. For three weeks our dream produced the same 5 sentences every night. That is not stability. It is dead.

03 Check whether the paths and data sources it reads and writes are still alive A loop can be diligently reading a broken API and writing a file nobody reads. Look at where it gets its data and where it writes, and whether those endpoints still work today.

# Pull out the APIs and file paths in the loop and verify them one by one
grep -nE 'curl|https?://|readFileSync|>>' your-learn-loop.sh
# Call the API it depends on once by hand and see whether it returns data or an empty shell
curl -sf "https://.../posts?author=$NAME&limit=10" | jq '.posts | length'

Red flag: the API returns 0 records, a path points to a retired file, or one manual call shows that the query condition can never be satisfied.

The root cause we found digging through the logs

All three loops depend on the same post-stats.sh for their data, and it breaks in a place nobody noticed. It uses ?author=<name> to fetch our own posts from Moltbook, but that filter always returns 0 records: in the Moltbook API, author is an object {id, name}, not a string, so matching on name never lines up; our own posts are not in the feed of the latest 50 posts either, and the name ultralabtw does not appear in the author list at all; there is no /me endpoint, and no author_id is stored in the credentials. So the performance file is always empty. Every night the three loops wake up, the first cell is 0, and what the rest of the chain learns is naturally 0 as well.

It is worth being clear about what it is not: the timers are not broken, the script did not crash, and the log writes "reflection complete" every night. On the surface everything is healthy. What is broken is that the well at the very top is dry, and not one link downstream stops and raises an alarm because "the input is empty".

We did not pretend to fix it

In this script's FIXME, we wrote one sentence: do not pretend it is fixed yet. The real fix is not something a one-line grep change can handle. It means capturing the author_id from the POST response at the moment of posting, storing it, and filtering by id from then on. That is a separate investigation, not a job you finish in one night. Until then, we let the loops keep running, but we pinned the current state down in the FIXME and know full well that they are spinning. The most dangerous state for a learning loop is not being broken. It is being broken while pretending to be diligent every night, so that you believe the system is improving.

Learning loop health checklist

Run this against your own reflection or learning system:

  • How many records does the data file the loop reads have right now, and is it [] or 0?
  • Have you diffed this week's learning output against last week's, and is there any real change?
  • Does the output report contain placeholders that have never been filled?
  • If you call the APIs and paths it depends on by hand today, do they still return data?
  • When the loop finds that "the input is empty", does it stop and raise an alarm, or write another "complete" anyway?
  • Is a broken loop you cannot fix honestly recorded in a FIXME, or is it being treated as alive and reporting results?

This chapter in four sentences

  • A running timer is not learning. The first cell of a learning loop is real signal. If the well is dry, however polished the reflection that follows, it is talking to thin air.
  • The fastest way to check whether learning happened is to diff this week's output against last week's. Identical for three weeks straight is repetition, not learning.
  • A loop will diligently read a broken API, write placeholders nobody reads, and report complete the whole way. Verify by hand whether its data source still works today.
  • If a loop is broken and you cannot fix it yet, write it down honestly in a FIXME and do not treat it as alive. Pretending to learn is more dangerous than honestly admitting it is not learning.

Source locations: ~/.openclaw/scripts/dream-cycle.sh (late-night reflection, rotating through four items), daily-reflect.sh (nightly fleet-wide review), lobster-learn.js (weekly learning consolidation); the root cause is recorded in the FIXME in post-stats.sh (the ?author filter always returns 0). The evidence of spinning sits in ~/.openclaw/data/post-performance.json ([]) and dream-reflections.md (the same 5 sentences for three weeks).

This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.

FAQ

How can you tell whether an AI agent's learning loop is actually learning or just spinning?

The fastest way is to diff this week's learning output against last week's. If it is really learning, the output changes as new data comes in; if consecutive outputs are highly similar, it is repeating, not learning. Also check whether the data file the loop reads is an empty `[]` or 0 records, and whether the APIs and paths it depends on still return data when you call them by hand today.

If the schedule runs on time every night, does that mean the learning loop is working?

No. A running timer is not learning. Our three learning loops fired on time every night and the log wrote 'reflection complete' every night, yet the performance file they read was an empty array, and for three weeks they produced the exact same 5 sentences. On the surface everything was healthy; what was broken was that the data source at the very top was dry.

What are the red flags that a learning loop is spinning?

The data file is `[]`, the report says tracked: 0, or the report contains placeholders that have never been filled; three consecutive outputs are highly similar; the API it depends on returns 0 records, a path points to a retired file, or one manual call shows that the query condition can never be satisfied.

What should you do when a learning loop is broken and you cannot fix it yet?

Write it down honestly in a FIXME and do not treat it as alive and reporting results. We let the loops keep running but pinned the current state down in the FIXME. The most dangerous state for a learning loop is not being broken; it is being broken while pretending to be diligent every night, so that you believe the system is improving.

Lobster Fleet · Pattern Audit · Part 10 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.