Defined but Never Run: How to Find the Features in Your Codebase With Zero Callers
Table of Contents
Chapter 17 of Agentic Design Patterns covers reasoning techniques. The core idea in one sentence: complex problems cannot be answered in one step, so the model should lay out its reasoning and work through it in several steps (Chain-of-Thought and similar techniques), and some tasks should even go to a dedicated reasoning model, such as deepseek-r1, which takes its time and closes in on the root cause step by step.
We read it and nodded. We have a deepseek-r1 "reasoning tier": the function is called ollama_generate_reasoning, and its comment sounds very sure of itself: "for multi-step diagnosis (self-heal, root-cause)", meaning it is there for self-healing and root-cause analysis. Reasoning techniques, implemented, and with the model the industry agrees reasons best. Then we ran an adversarial audit on ourselves. The truth of this chapter: from the day that tier was defined to the day of the audit, it had not run a single time.
Alive in the docs, dead in execution
The first step of the audit was to grep the function name across all of $HOME. Result: apart from the one line that defines it, zero hits. No script, no timer, no self-healing flow had ever called it. It was not broken, and it did not fail now and then. It had simply never been woken up.
The comment was the most glaring part. "for self-heal, root-cause" reads like a capability that is up and running, yet nothing in the whole system that does self-healing calls it, and nothing that does root-cause analysis calls it. That description was not out of date. It had been fiction from day one. The function picked the right model, set the right parameters, and wired up a tidy fallback chain; the only thing never connected was the line between it and the system. A correct function with no caller is no different from a piece of fake documentation. A tier that is defined but has never run is a lie written into the docs.
First spend a few minutes checking your own "defined but never run" features
01 Grep the whole codebase and count how many real callers the feature has The line that defines it does not count. Grep the function name across the whole project, filter out the definition line, and whatever remains are the callers. Zero means a feature that exists only in the docs.
# Count how many places call a function, excluding its definition
grep -rn 'ollama_generate_reasoning' ~/.openclaw | grep -v '() {'
# 0 lines returned = zero callers; this feature has never run
Red flag: grep only finds the definition line, and nothing else. It is not broken. It has never been called.
02 If it claims to be "for X", go into X and look for it Comments love to say "for diagnosis / self-heal / scheduling". So grep for the name inside those scripts. If you cannot find it, the "for X" was wishful thinking, and X does not even know it exists.
# It claims to be for "diagnosis / self-healing", so check whether those scripts actually call it
grep -rln 'diagnose\|self.heal\|root.cause' ~/.openclaw/scripts
# For each script that comes back, grep again for the reasoning function name
Red flag: not one of the scripts for its claimed purpose actually calls it. The purpose is written in a comment, not wired into the code.
03 If it has really run, the logs will hold its fingerprints; search the logs for its tracks A feature that has actually been executed leaves traces somewhere: the cost log has the model it used, the error log has its caller tag, the output directory has what it produced. If you search every log and cannot find its name, it has not emitted a single token from launch to today.
# This tier uses deepseek-r1; does the cost log show any record of it running?
grep -c 'deepseek' ~/.openclaw/logs/llm-cost.log
# Does its caller tag appear in any log?
grep -rl 'reasoning\|diagnose' ~/.openclaw/logs
Red flag: no fingerprints anywhere in the cost log, the error log, or the output directory. Alive in the docs, dead in execution.
The reverse also holds: a fingerprint alone does not prove it is wired in. One output, a timestamp stuck on a single day, and nothing scheduling or calling it afterwards means it was run by hand once, not integrated.
How we actually fixed it
The moment grep returned zero callers, there were only two paths: delete the fake "for diagnosis" description and admit the tier was dead, or connect a real caller so the description became true. We chose the second.
We added a --diagnose switch to our existing review-digest panel. Normally the panel is purely read-only with zero noise. When you run review-digest.sh 7 --diagnose, it pulls the past seven days of guardrail hard blocks, low semantic scores and LLM errors into a single anomaly list, feeds it to deepseek-r1 for multi-step root-cause analysis, and asks for 3 to 5 points on what to fix first. That is what a reasoning tier is for: a multi-cause diagnosis that cannot be answered in one step and can only be worked out by connecting several signals, handed to a model that lays out its reasoning. Each run costs money, so it is opt-in and off by default. Only from that moment did the comment stop being a lie.
It is worth being clear about what this was not: the function itself had no bug, the model choice was right, and deepseek-r1 really does suit this kind of multi-step reasoning. What was broken was never the implementation. It was the line between the function and the system, which had never been connected. An empty shell throws no error, does not crash, and still passes tsc. It just sits there quietly and lets the docs tell a lie on its behalf about something it cannot do.
Empty-shell feature checklist
Run this against your own codebase:
- For every "implemented" feature, have you grepped the whole project to confirm it has real callers (remembering to filter out the definition line)?
- For anything whose comment says "for X", does X actually call it?
- If it has really run, can you find its fingerprints in the logs, cost records or output files?
- Is there a function or tier that is beautifully defined, with the right parameters, but has zero callers?
- Are dead features honestly deleted or marked, or left in place for the docs to lie about?
- After connecting a real caller, did you actually run it once to confirm it really works?
This chapter in four sentences
- Being defined is not the same as running. A function with zero callers is a lie in the docs, not a feature of the system.
- If a comment says "for X", go into X and grep for it. If it is not there, that sentence is fiction.
- A feature that has really run leaves fingerprints in the logs. If you search every log and cannot find its name, it has never emitted a single token.
- There are two ways to fix an empty shell: delete it honestly, or connect a real caller so the docs become true. Keeping it around to lie for you is the worst, third option.
Source location: ~/.openclaw/scripts/ollama-helper.sh (ollama_generate_reasoning, the deepseek-r1 reasoning tier, which had zero callers before the 2026-07-04 audit); the real caller is the --diagnose switch in ~/.openclaw/scripts/review-digest.sh (opt-in and costs money; it feeds the anomaly signals from the past N days to the reasoning model for multi-step root-cause analysis).
This is part of the Agentic Design Patterns × Lobster Fleet series. We systematise a solo company's AI agent fleet chapter by chapter, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.
FAQ
When should a task go to a reasoning model such as deepseek-r1?
The core of chapter 17 of Agentic Design Patterns: complex problems cannot be answered in one step, so the model should lay out its reasoning and work through it in several steps (Chain-of-Thought and similar techniques), and some tasks should even go to a dedicated reasoning model, such as deepseek-r1, which takes its time and closes in on the root cause step by step. A multi-cause diagnosis that cannot be answered in one step and can only be worked out by connecting several signals is exactly what a reasoning model is for.
How do I check whether a function has any real callers?
Grep the function name across the whole project and filter out the definition line; whatever remains are the callers. The line that defines it does not count. If grep only finds the definition line and nothing else, the feature exists only in the docs: it is not broken, it has never been called.
A comment says a feature is 'for X'. How do I confirm X actually uses it?
Go into the scripts for its claimed purpose and grep for the function name. If you cannot find it, the 'for X' was wishful thinking, and X does not even know it exists. The purpose is written in a comment, not wired into the code.
How can logs tell me whether a feature has ever actually run?
A feature that has actually been executed leaves traces somewhere: the cost log has the model it used, the error log has its caller tag, the output directory has what it produced. If you search every log and cannot find its name, it has not emitted a single token from launch to today.
Should I delete a function with zero callers or keep it?
There are only two paths: delete the fake description and admit it is dead, or connect a real caller so the description becomes true. Keeping it around to lie for you is the worst, third option. After connecting a real caller, actually run it once to confirm it really works. We chose the second path: we added a --diagnose switch to our existing review-digest panel that feeds the past seven days of anomaly signals to deepseek-r1 for multi-step root-cause analysis, and because each run costs money, it is opt-in and off by default.