AI AgentSelf-ReflectionReflectionHallucinationBuildInPublicSolo Business

When AI Self-Reflection Fabricates Numbers: How to Guard LLM Rewrites Against Fake Stats

· 8 min read
Lobster Fleet · Pattern Audit · Part 5 of 25
Table of Contents
  1. The draft really did get better, but a McKinsey line appeared
  2. Spend a few minutes checking your own reflection
  3. The problem was one missing sentence in the prompt
  4. Reflection checkup checklist
  5. Four takeaways from this chapter

A few months ago, in Automated AI Content Doesn't Have to Be Junk: Three Quality Gates in Practice, I described "if it is not good enough, let the AI rewrite it itself" as the best-value quality improvement. This post adds the other side that post left out: reflection really did turn a bad draft into a better one, but in that same version it casually invented a McKinsey figure that never existed. A fake number hidden in an improved draft is more dangerous than a bad draft left unimproved.

Chapter 4 of Agentic Design Patterns covers Reflection. The core idea in one line: after an agent produces a first version, do not ship it straight away; have it act as its own reviewer and score it first, and if it falls short, have it rewrite a version before sending. Getting the model to look back at its own work is the cheapest kind of quality improvement there is.

After reading it, we thought we had this one covered. Our posting script already had an inline quality gate, and we extracted it into reflect.sh, which any content script can source: pass in a draft and it returns either the original as is (good enough) or a self-rewritten version (not good enough). Implementation done. Then we took it out and tested it on a bad draft. The truth of this chapter has two sides, and the second one is chilling: reflection really did improve the draft, and in the very version it improved, it casually made up a number that never existed.

The draft really did get better, but a McKinsey line appeared

What we fed in was a deliberately bad draft: loose, no hook, low information density. reflect gave it a score below the threshold and returned REWRITE plus a rewritten version. The rewrite really was better: a tighter hook, denser sentences, exactly the kind of improvement we wanted.

The problem was hidden in one sentence. The rewrite said "according to McKinsey research, teams that adopt AI see a 40% efficiency gain". That number was not in the original draft, and it was not in the context we gave it either; we never mentioned McKinsey at any point. It was an authoritative-sounding citation that reflection produced out of thin air to make the draft "more credible".

The most ironic part: reflect.sh's default scoring criteria explicitly list "credibility". To score high on credibility, it fabricated a fake authoritative source. In the same rewrite, it made the draft read more like the truth and made the draft less true. Reflection rewrites the structure, and along the way it touches the facts. It rewrote the structure beautifully; it had no obligation to stay faithful to the facts, because we never told it to.

Spend a few minutes checking your own reflection

01 Does your reflection prompt explicitly require "use only numbers already in the original draft"? Pull up the part of the prompt where you ask the model to rewrite, and look for one hard constraint: it must not add any number, percentage, organisation name or source that the original draft does not have. If you cannot find one, it means you only told it to "improve it" and "strengthen credibility", without closing the door on it inventing a number to look credible.

# Does the reflection prompt contain a hard constraint like "do not add numbers/sources yourself"?
grep -niE '不.*新增|不准.*(捏|杜撰|編)|只.*用.*原稿|do not (invent|add)' reflect.sh

Red flag: the whole prompt only asks for things like "improve it" and "make it more credible", with not a single sentence forbidding new facts.

02 When the rewrite comes back, check whether it contains numbers the original draft does not Pull the numbers out of the original draft and the rewrite and compare them. Numbers that appear only in the rewrite and not in the original are suspected hallucinations, especially a percentage paired with an organisation name.

# Pull the numbers from both sides and find the ones that appear only in the rewrite
grep -oE '[0-9]+(\.[0-9]+)?%?' draft.txt   | sort -u > before.nums
grep -oE '[0-9]+(\.[0-9]+)?%?' rewrite.txt | sort -u > after.nums
comm -13 before.nums after.nums   # numbers only in the rewrite = suspected fabrication

Red flag: the rewrite has numbers the original draft does not, and you cannot find their source when you go back through the original context.

03 The more "professional" a number sounds, the more it needs a dedicated check McKinsey, Gartner, IDC, or "research shows" followed by a neat percentage: this is exactly the shape LLM hallucinations love most, because it looks the most real, is the easiest to cite, and is the least likely to be challenged on the spot. Pull these combinations out separately and check them sentence by sentence.

# Target the "organisation/research + percentage" combination, the most real-looking and most often fabricated
grep -nE '(麥肯錫|McKinsey|Gartner|IDC|研究顯示|報告指出).{0,20}[0-9]+%' rewrite.txt

Red flag: an authoritative number that is very easy to cite appears, and its source is not in any context you provided. The more professional it sounds, the more suspicious it is.

The problem was one missing sentence in the prompt

The rewrite instruction in reflect.sh only says: if the score is below the threshold, "output REWRITE, then give the improved full version". The scoring criteria include "credibility", but there is not one sentence requiring it to "only rewrite existing content, and not add any number or source the original draft does not have". Without that hard constraint, the model will quite naturally add a professional-sounding number to push up credibility. And the function returns the rewrite as is, with no step in between that checks whether the rewrite contains numbers the original draft lacks. The fail-open design means that when it breaks, it does not block the draft, and it also means that when it fabricates a number, it does not block the draft either.

The fix has two layers. First layer: add one airtight sentence to the prompt: only rewrite existing content; do not add any number, percentage, organisation name or source that the original draft does not have. Second layer: after the rewrite comes back, run a number diff, and if the rewrite contains a number the original draft does not, reject it and fall back to the original draft. Deterministic comparison stands in front of the LLM's gut feeling, the same trick we learned in the judge chapter.

Reflection checkup checklist

Run through this on your own self-reflection or self-rewrite:

  • Does the reflection prompt have a hard constraint forbidding it from adding numbers and sources the original draft does not have?
  • When the rewrite comes back, do you check whether it contains numbers the original draft does not?
  • For sentences that pair an organisation name with a percentage, have you gone back to the original context and checked each source?
  • Your scoring criteria include "credibility", but is credibility earned by making the draft look more real, or by making it closer to the truth?
  • Can you tell apart reflection failing (fail-open, returning the original draft) and reflection fabricating a fake number and returning a "better draft"?

Four takeaways from this chapter

  • Reflection improves the structure, and it can also damage the facts. What it rewrites is the prose; what it may touch is numbers you never authorised it to handle.
  • A reflection prompt missing one sentence, "do not add numbers yourself", tacitly allows it to invent one to look credible. Credibility cannot be pieced together from hallucinations.
  • The more professional-sounding and easy to cite a number is (McKinsey 40%, Gartner, "research shows"), the more it has the shape hallucinations love, and the more you need to go back to the original and check it.
  • Do not trust a rewrite as soon as it comes back; put a deterministic number diff in front of the LLM's gut feeling. One fake number hidden in an improved draft is more dangerous than a bad draft left unimproved.

Source location: ~/.openclaw/scripts/lib/reflect.sh (the reflection primitive reflect_content; the rewrite instruction lacks a hard constraint forbidding new numbers, and the rewrite is returned as is without any number comparison). It depends on the premium generation capability in ollama-helper.sh.

This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet chapter by chapter, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.

FAQ

What is self-reflection (Reflection) in an AI agent?

Chapter 4 of Agentic Design Patterns covers Reflection: after an agent produces a first version, do not ship it straight away; have it act as its own reviewer and score it first, and if it falls short, have it rewrite a version before sending. Getting the model to look back at its own work is the cheapest kind of quality improvement there is.

Why does an AI make up numbers when it rewrites its own draft?

In our test, reflection improved a bad draft but added the line 'according to McKinsey research, teams that adopt AI see a 40% efficiency gain', a number that was in neither the original draft nor the context we gave it. The scoring criteria included 'credibility', but the rewrite instruction had no sentence requiring it to only rewrite existing content and not add any number or source the original lacks, so the model quite naturally added a professional-sounding number to push up credibility.

How do you stop an AI rewrite from inventing fake statistics?

The fix has two layers. First, add one airtight sentence to the reflection prompt: only rewrite existing content; do not add any number, percentage, organisation name or source that the original draft does not have. Second, run a number diff after the rewrite comes back, and if the rewrite contains a number the original draft does not, reject it and fall back to the original, so a deterministic comparison stands in front of the LLM's gut feeling.

How can I check whether the numbers in an AI rewrite were made up?

Pull the numbers out of the original draft and the rewrite and compare them; numbers that appear only in the rewrite are suspected hallucinations. Pay special attention to McKinsey, Gartner, IDC or 'research shows' followed by a neat percentage, the shape LLM hallucinations love most, and pull these out separately and go back to the original context to check each source.

Lobster Fleet · Pattern Audit · Part 5 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.