LLM-as-Judge Pitfalls: How Our Judge Hallucinated and Marked a Correct Answer Wrong
Table of Contents
A few months ago, in Automated AI Content Doesn't Have to Be Junk: Three Quality Gates in Practice, I called cross-agent Peer Review my favourite mechanism: one AI reviews another AI's output. This post takes apart its blind spot. When the judge itself hallucinates, it will confidently mark a system that answered correctly as wrong, and push a CRITICAL alert to scold it. The judge needs to be evaluated too.
Chapter 19 of Agentic Design Patterns covers evaluation and monitoring. The core idea in one line: an agent being alive does not mean it answered correctly. systemctl is-active only checks whether something is running; it cannot check whether what it produces is right. The book's signature technique is LLM-as-Judge: use an LLM to score output, aimed specifically at the kind of failure that fools is-active. The service is still alive, but the answer is wrong, hallucinated, or contains zero useful content.
After reading it, we built a "semantic heartbeat". Every day it sends a question with a known answer to a key service, and uses the cheap Gemini Flash as a judge to score the reply from 1 to 5. Only when the score drops below 3 does it push an alert to Telegram; otherwise it stays quiet. The implementation was finished and we were pleased with it. Then the judge itself hallucinated. The irony of this chapter is that the evaluation system turned into the very disease it was meant to catch.
The day the judge hallucinated
The question used as the regression test was "What is AVS?", and the reference answer is AVS = AI Visibility Score (a composite score of SEO + AEO + AAO). This question has a history: our Q&A system once hallucinated AVS as Agent Verification System, which we later fixed by adding knowledge anchors, so using it as the daily checkup question made perfect sense.
That day the lobster answered correctly. Its reply said AI Visibility Score in plain words. But the judge gave it a low score, ruled it a hallucination, and pushed a CRITICAL-level alert to Telegram, the gist being that the lobster's quality had dropped and it had answered AVS wrong.
The problem was the judge itself. The cheap flash model has its own world knowledge: in IT and payments, the most common meaning of AVS is Address Verification System (address verification for credit cards). The judge compared the lobster's answer against the definition in its own head, found they did not match, and ruled that the lobster had hallucinated. In truth the judge was the one hallucinating: it marked a system that answered correctly as wrong, then escalated to CRITICAL to scold it.
We built this system precisely to catch "active but wrong". The first thing to fall into that trap was the judge itself: it was active, it was wrong, and it confidently pushed an alert. The evaluation system itself needs to be evaluated. We later ran into the same thing twice on our legal Q&A: a regression suite that stayed green for weeks because it was looser than production, and a benchmark whose empty catch quietly dragged the score down.
Spend a few minutes checking your own judge
01 Give the judge a batch of known correct answers and regression-test the judge first Before you let a judge go live monitoring anything else, feed the judge itself a batch of questions whose reference answers you already know, and see whether it rules on them correctly. Send in an obviously correct answer. If the judge marks it wrong, what you have caught is not a problem in the system under test; it is a problem in the judge.
# Take an answer you know is correct, deliberately send it to the judge, and see whether it misjudges
echo 'Q=AVS 是什麼 | A=AVS 是 AI Visibility Score' | your-judge --score
# Expect a 5; a low score = the judge is broken, not the answer
Red flag: the judge marks an answer you know is correct as wrong. That means your scoring criteria lean too heavily on the LLM's own world knowledge.
02 For questions with a reference answer, use deterministic keyword matching, not the LLM's gut feeling If a question has an authoritative keyword that must appear, judge it by string matching, not by asking an LLM. If the keyword is there, the answer is right. This is deterministic and will not flip because of whatever world knowledge the judge happens to bring that day. Only when the keyword is missing do you fall back to the LLM judge, and only then might the answer really be wrong. While you are at it, pin the LLM's prompt to the reference answer you provide, and state explicitly "do not use what you know about this acronym from other fields". That sentence closes exactly the gap through which it applied Address Verification to AVS.
# If the authoritative keyword is present, PASS directly and bypass the flaky judge
if printf '%s' "$ANS" | grep -qiE 'AI Visibility Score|可見度'; then
SCORE=5 # deterministic criterion
else
SCORE=$(llm_judge "$ANS") # ask the LLM only when the keyword is missing
fi
Red flag: your score is decided entirely by a single LLM call, with no deterministic checkpoint in between.
03 Log the judge's verdicts and spot-check them too Write every score the judge gives, along with its reasoning, to a jsonl file, and pull a few entries to review every so often. What you are verifying is not the score of the system under test, but whether the judge itself rules accurately. The few anomalous scores are usually not a broken system; they are a broken judge.
# Persist every verdict so you can sample it later
echo "{\"probe\":\"AVS\",\"score\":$SCORE}" >> semantic-quality.jsonl
Red flag: the judge only leaves a trace when it rules something a failure, so you never know whether its everyday verdicts are accurate.
How we actually fixed it
The fix on 2026-07-05 went into the order of the criteria: if the answer contains the authoritative keyword "AI Visibility Score" or "可見度" (Chinese for "visibility"), it always passes with a score of 5, and the flash judge is never called. The judge is kept only for cases where not even the keyword appears, which is the only time the answer might really be wrong. The deterministic criterion stands in front, and the LLM's gut feeling moves to the back as a fallback. The CRITICAL pushed that same day was false. The principle of putting the deterministic check in front applies to Chapter 4's self-rewrites as well: when a rewrite contains numbers the original draft lacks, a number diff sends it back to the original.
Judge checkup checklist
Run through this on your own evaluation system:
- Before the judge went live, did you regression-test it with a batch of known correct answers?
- For questions with a reference answer, do you use deterministic keyword matching, or leave everything to the LLM's gut feeling?
- Is the LLM judge's prompt pinned to the reference answer you provide, forbidding it from using its own world knowledge?
- Is every verdict the judge makes (score plus reasoning) logged for spot checks?
- Can a CRITICAL-level alert be triggered by a single LLM judgement, or is there a deterministic checkpoint in between?
Four takeaways from this chapter
- An LLM used as a judge can also be wrong and can also hallucinate. The evaluation system itself needs to be evaluated.
- For checks with a known correct answer, use deterministic keyword matching; do not hand all the scoring to another LLM's gut feeling.
- The cheaper the judge, the more easily its world knowledge overrides the reference answer you gave it, and the more you need deterministic criteria to pin it down.
- A judge that marks an obviously correct answer as wrong is more dangerous than one that misses a wrong answer. It will confidently push a CRITICAL alert to scold a system that did nothing wrong.
Source locations: ~/.openclaw/scripts/semantic-heartbeat.sh (the semantic heartbeat, including the keyword short-circuit fix from 2026-07-05); every verdict the judge makes lands in ~/.openclaw/logs/semantic-quality.jsonl.
This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet chapter by chapter, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.
FAQ
Can an LLM used as a judge (LLM-as-Judge) get it wrong?
Yes. An LLM judge can hallucinate too. We used the cheap Gemini Flash as a judge and asked 'What is AVS?' every day. The lobster answered correctly, and its reply said AI Visibility Score in plain words. But in the judge's own world knowledge the most common meaning of AVS is Address Verification System (address verification for credit cards), so it compared the answer against that definition, ruled that the lobster had hallucinated, and pushed a CRITICAL-level alert to Telegram. The evaluation system itself needs to be evaluated.
How do I test whether an LLM judge rules accurately?
Before you let the judge go live monitoring anything else, feed the judge itself a batch of questions whose reference answers you already know, and see whether it rules on them correctly. Send in an obviously correct answer. If the judge marks it wrong, what you have caught is not a problem in the system under test; it is a problem in the judge. Once it is live, also write every score the judge gives, along with its reasoning, to a jsonl file and pull a few entries to review every so often.
For questions with a reference answer, should I use an LLM to score or keyword matching?
Use deterministic keyword matching. If a question has an authoritative keyword that must appear, judge it by string matching: if the keyword is there, the answer is right, and that will not flip because of whatever world knowledge the judge happens to bring that day. Only when the keyword is missing do you fall back to the LLM judge. While you are at it, pin the LLM's prompt to the reference answer you provide, and state explicitly 'do not use what you know about this acronym from other fields'.
Why does a cheap LLM judge need deterministic criteria even more?
The cheaper the judge, the more easily its world knowledge overrides the reference answer you gave it, and the more you need deterministic criteria to pin it down. A judge that marks an obviously correct answer as wrong is more dangerous than one that misses a wrong answer: it will confidently push a CRITICAL alert to scold a system that did nothing wrong. So check whether a CRITICAL-level alert can be triggered by a single LLM judgement, or whether there is a deterministic checkpoint in between.