RAG's Most Dangerous Failure Is Confidently Citing the Wrong Authority: How to Prevent It
Table of Contents
Chapter 14 of Agentic Design Patterns covers knowledge retrieval, RAG. The core idea in one sentence: do not let the model answer from its own memory. First retrieve real content from a knowledge base, inject it into the prompt, and have the model speak from what is actually there.
We read it, nodded, and ticked a box in our heads: we have this. When someone asks a question in the community's #ask-agent channel, a script first searches the knowledge base for blog passages, scores them, and puts the three passages that score high enough into the prompt. Then a layer of "term definition" anchors pins authoritative definitions at the highest priority, forcing the model to copy them instead of making things up. RAG, implemented, with citation links included. Then we ran an adversarial audit on ourselves. The truth in this chapter has two layers. The first is embarrassing. The second is dangerous.
The first layer: for 7 weeks, the only ones asking questions were us
The script is designed respectably. It wakes up every minute, grabs the latest human questions in #ask-agent, and answers them with RAG before the slower main lobster agent gets there. It also has a --test dry-run mode: feed it a question, it runs the full retrieval and generation path and prints the answer, which is handy for smoke testing.
The problem is that for 7 weeks, the only questions it ever answered were the few we fed it ourselves through --test. Not one real person asked a question in #ask-agent. Lay the log out and it is a solid sheet of [TEST] OK, without a single ✓ Answered line for an answer actually POSTed.
It wakes up on time every day, passes its own smoke test every day, and writes "all normal" in the log every day. It looks perfectly alive. But is a race-to-answer system that can only answer the exam questions it set for itself really alive? Running is not the same as serving.
The second layer: it cites the wrong authority more confidently than anyone
This is where RAG is truly dangerous, and it is the point of this chapter.
People usually assume what RAG fears most is not finding anything: retrieval misses and the model cannot answer. We actually got this path right. When the score falls below the threshold (we set it at 8), it takes the fallback and says honestly "our blog has not written a dedicated post on this yet", instead of forcing an answer. Not finding anything is a safe failure.
What actually bites is at the other end. RAG's signature move is putting retrieved authoritative definitions at the highest priority in the prompt, with a line saying "do not use general knowledge, do not infer on your own, this is the reference". The design is meant to suppress hallucination, but it backfires: if the authority you put in is wrong, the model will state the wrong thing with maximum confidence, because you personally ordered it not to doubt.
Our anchor layer did pass along a wrong number once, writing one product's capability as 25 when the official figure was 12. At that moment RAG was not fixing hallucination. It was stamping a wrong answer with "authoritative, do not doubt". A miss falls back to honesty. A wrong authority does not: it lies to you with complete certainty. That is far worse than finding nothing.
(Correction, 2026-09-27: this sentence has the two numbers reversed. The anchor actually said 12, and the official figure is 25. UltraProbe has had 25 defense vectors since 2026-06-15 (12 LLM-era + 13 agent/ASI-era); 12 was an older figure. The anchor was changed to 25 on 2026-09-26.)
First, spend a few minutes checking your own RAG
01 Check every definition and number your RAG injects as "authoritative", one by one RAG puts a batch of definitions marked "highest priority, do not doubt" into the prompt. How long has it been since anyone checked that batch? Pull them out and compare every number against the current reality. When retrieval misses, the model falls back to honesty. When an authority is wrong, the model copies it with doubled confidence. A wrong authority is far more dangerous than a miss.
# Pull out every definition / number injected as "authoritative" and check each one is still true
grep -nE '[0-9]+ ?(篇|帳號|向量|%|分)|Score|=' lib/rag-concept-anchors.js
cat ~/.openclaw/workspace/PUBLIC-FACTS.md
Red flag: you have a layer of authoritative definitions that "must not be overridden", but you cannot say when they were last checked one by one.
02 Is it answering real people's questions, or only the exam questions it set itself
Separate "running" from "serving". Put two numbers side by side: how many answers it has actually POSTed, and how many times it has run the --test smoke test. If the real column is 0, the feature is just practising with itself every day.
# Actually answered (POST) vs only setting its own exam (--test)
grep -c '✓ Answered' ~/.openclaw/logs/rag-answerer.log
grep -c '\[TEST\]' ~/.openclaw/logs/rag-answerer.log
jq '.answered | length' ~/.openclaw/data/rag-answerer-state.json
Red flag: the log is all [TEST] and the real answer count is 0. A feature with no real traffic does not count as alive.
03 When retrieval misses, does it honestly say it does not know, or does it force content on the model and let it make things up Give retrieval a score threshold. Below it, take an honest fallback that says plainly "we have not written about this yet", instead of stuffing irrelevant low-score passages into the prompt and letting the model make something up. Not finding anything should be the safest failure. Do not turn it into a breeding ground for hallucination. There is also a partial miss the threshold cannot see: when rephrasing the question pushes the key source out of range, other passages still fill every slot and the answer can flip.
# Which path runs below the threshold: an honest fallback, or stuff it in anyway
grep -nE 'MIN_SCORE|hasMatch|topScore|沒寫過' discord-rag-answerer.js
Red flag: no score threshold, or low-score retrieval passages still get put into the prompt and the model is told to answer anyway.
What we actually fixed
Two things, taken separately. The cut for the wrong authority has been made: we changed that capability number in the anchors back to the official 12, aligned with our public figures, so it no longer stamps authority on a wrong answer. The low-score fallback path was right from the start, and it stays.
(Correction, 2026-09-27: "changed back to the official 12" above is also wrong. The official figure has been 25 since 2026-06-15, so changing it to 12 did not align the anchor with our public figures and did not stop it stamping authority on a wrong answer. The anchor was changed from 12 to 25 on 2026-09-26.)
What we did not fix is traffic. We did not inject fake questions to make the numbers look good. Zero real questions in 7 weeks is zero. We let it keep running and keep logging, knowing full well it is only a race-to-answer desk that passes its own smoke test and has never played a real game. RAG's two most dangerous diseases: one is confidently citing a wrong authority. The other is quieter: a Q&A system nobody asks, which always looks healthy.
RAG checkup list
Run it against your own retrieval system:
- When the definitions and numbers you inject as "authoritative" were last checked one by one
- When that "must not be overridden" layer of authoritative content is wrong, whether anything will notice, or whether the model will just copy it
- Whether retrieval has a score threshold, and whether below it you get an honest fallback or forced stuffing
- Whether you can tell apart how many real people's questions it has actually answered and how many times it has run
--test - Whether you treat a feature with 7 weeks of zero real traffic as alive, or admit it is just practising with itself
- Whether the numbers in your public figures (the PUBLIC-FACTS kind) have been reconciled against old numbers in the blog, with it settled which one is authoritative
Four lines from this chapter
- The most dangerous thing in RAG is not finding nothing. It is confidently citing a wrong authoritative definition. A miss falls back to honesty; a wrong authority lies with complete certainty.
- Every definition and number you inject as "highest priority, do not doubt" needs someone checking it regularly. The more firmly you order the model not to doubt, the more fatal a wrong entry becomes.
- A retrieval miss should be the safest failure. Set a score threshold and have low scores honestly say "I do not know". Do not turn a miss into a breeding ground for forced answers.
- Running is not the same as serving. A Q&A system that can only answer its own exam questions looks permanently healthy, and has in fact never played a real game.
Source locations: ~/.openclaw/scripts/discord-rag-answerer.js (retrieval + injection + webhook race-to-answer, including the --test dry-run mode and the MIN_SCORE=8 threshold), ~/.openclaw/scripts/lib/rag-concept-anchors.js (the highest-priority "term definition" anchors, with the capability number changed to the current official 25 defense vectors on 2026-09-26); state and logs live in ~/.openclaw/data/rag-answerer-state.json (the answered list) and ~/.openclaw/logs/rag-answerer.log (where real answers and [TEST] are told apart).
This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented has been verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publicly prove ourselves wrong.
FAQ
What is the most dangerous failure in RAG?
It is not a retrieval miss. It is confidently citing a wrong authoritative definition. When retrieval misses, the model falls back to honesty. But RAG puts authoritative definitions at the highest priority in the prompt and orders the model not to doubt them, so if the authority you put in is wrong, the model will state the wrong thing with maximum confidence.
How do you stop RAG from citing a wrong authoritative definition?
Pull out every definition and number your RAG injects as 'authoritative', check each one against the current reality, and have someone check them regularly. The more firmly you order the model not to doubt, the more fatal a wrong entry becomes. The red flag: you have a layer of authoritative definitions that 'must not be overridden', but you cannot say when they were last checked one by one.
What should RAG do when retrieval finds nothing relevant?
Give retrieval a score threshold (we set it at 8). Below it, take an honest fallback that says plainly 'we have not written about this yet', instead of stuffing irrelevant low-score passages into the prompt and letting the model make something up. Not finding anything should be the safest failure, not a breeding ground for hallucination.
How can you tell whether a RAG Q&A system is actually serving real users?
Separate running from serving: put the number of answers it has actually POSTed next to the number of times it has run the --test smoke test. If the real column is 0, the feature is just practising with itself every day. Our own answerer went 7 weeks without a single real person asking a question.