The Model Routing Cost Trap: Does Your LLM Fallback Chain Quietly Upgrade to the Most Expensive Model?
Table of Contents
Chapter 2 of Agentic Design Patterns covers routing. The core idea in one sentence: choose the model by task difficulty, use cheap models for simple tasks and bring in expensive ones only for hard tasks, instead of reaching for the biggest model every time.
We read it and nodded. Our LLM helper has three tiers, each with its own fallback chain, all with cheap models as the primary, and the default tier led by the free gemma. Choosing the model by difficulty: done. Then we ran an adversarial audit on ourselves. The truth of this chapter is not on the normal path. It is in the fallback: the router correctly picked a cheap primary, yet let the degradation on failure climb toward the most expensive model.
When the free tier gets rate limited, it upgrades itself to the most expensive model
The old default chain looked like this: gemma:free → gemini-2.5-flash → claude-haiku. The logic at design time was simple: use the free one first, and fall back only if it does not work. The problem is that "falling back" lands on the most expensive model in the whole chain, Claude.
Rate limiting is what free models do most. Once the free tier is blocked, the router steps down one slot as designed, and step by step it ends up at claude-haiku, roughly 10 times the price of Gemini flash. It also gets worse the busier you are: the more requests and the denser the rate limiting, the more requests get pushed into the most expensive slot. It is not choosing a model by difficulty. It is pushing each request toward the expensive end based on whether the front of the chain is jammed. A rate limit is a transient error, not a dead model, so it should back off and retry the same model first; see how to classify errors before you retry.
The deadliest part is that we could not see it. The config clearly selected the free gemma, so we naturally assumed we were paying $0. Which model actually ran, no log entry could tell us. That is how $8 of Claude burned quietly, and we only found it by accident while going through the dashboard. Routing picks the "intent", the bill records the "actual", and between the two sat an entire fallback chain nobody was watching.
First, spend a few minutes checking your own fallback chain
01 Lay out your model chain and price every slot Pull out the primary plus every fallback, label each slot with its unit price, and read from top to bottom. A healthy chain gets cheaper as it falls back, or at least stays flat. It should not get more expensive.
# Pull the model chain and fallbacks out of your config
grep -niE 'model|fallback' config.* .env
# Label each model with its unit price and see whether the chain moves toward cheaper or pricier
Red flag: any slot in the chain whose next model costs more than the one before it. The moment the free tier is rate limited, you start paying for the most expensive slot.
02 Check whether you log which model each call actually used The config selects the intent. When rate limiting kicks in, the model that actually takes over could be any one at the tail of the chain. Without per-call model logging, you are blind.
# Does the cost log record the model per call? Group by model to see actual usage
grep '"model"' llm-cost.log | jq -r .model | sort | uniq -c | sort -rn
# If this log does not exist at all, that is the problem itself
Red flag: you cannot answer "what percentage of requests in the past day actually ran on the most expensive model".
03 Make the primary fail on purpose and see where it degrades Set the primary to a model that will fail, force the fallback to take over, and check the log for which model actually stepped in. The direction it falls is the direction your bill moves under high load.
# Deliberately break the primary, see who takes over and which model the cost log records
OPENROUTER_MODEL='does/not-exist' your-script "測試一句"
tail -1 llm-cost.log | jq '{model, cost}'
Red flag: after the primary fails, the model that takes over costs more than the primary. Your degradation runs toward the expensive end.
What we actually changed
Two cuts on 2026-07-04. The first cut removed every Claude from every chain: the default tier became gemini-2.5-flash → gemini-2.0-flash-001, premium switched from claude-sonnet to gemini-2.5-pro, and the reasoning tier's fallback also dropped claude-sonnet. None of the three tiers can fall to Claude anymore; degradation only moves toward cheaper or equal cost, never upward. The second cut added _llm_log_cost, which pulls tokens from the usage field of every response, estimates the cost, and writes each call to llm-cost.log. From then on, which service burned how much on which model is always visible, with no more waiting for the dashboard to spring a surprise. That $8 could only burn because this one line of logging did not exist.
Routing checkup checklist
Run this against your own model routing:
- Price every slot in the fallback chain: does it get cheaper or more expensive as it falls back?
- When rate limited, does degradation move toward cheaper or more expensive models?
- Is the model each call actually used recorded in a log, call by call?
- Can you answer "what percentage of requests in the past day ran on the most expensive model"?
- Does the cheap model your config selects match the model you actually pay for on the bill?
- When the free tier (the one most prone to rate limiting) fails, is the model that takes over the most expensive one in the whole chain?
This chapter in four sentences
- Routing is not only about choosing the primary; it also has to control where the fallback drops. The direction of degradation is the direction your bill moves under high load.
- Free models rate limit the most. If the fallback climbs toward expensive models, you pay the most exactly when you are busiest.
- The config records intent, the bill records reality. Without per-call model logging, you do not know which model you are burning.
- Routing saves money not by how good the normal path looks, but by the failure path never quietly upgrading to the most expensive model.
Source location: ~/.openclaw/scripts/ollama-helper.sh (the three-tier routing and fallback chains; on 2026-07-04 every Claude was removed and the default changed to gemini-2.5-flash → gemini-2.0-flash-001), _llm_log_cost (same file, per-call cost logging written to ~/.openclaw/logs/llm-cost.log).
This is part of the Agentic Design Patterns × Lobster Fleet series. We follow the book to systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.
FAQ
Why can an LLM routing fallback make costs go up?
Because the fallback chain can be ordered toward more expensive models. Our old default chain was gemma:free → gemini-2.5-flash → claude-haiku. Once the free tier was rate limited, the router stepped down one slot at a time and ended up at claude-haiku, roughly 10 times the price of Gemini flash. The more requests and the denser the rate limiting, the more requests got pushed into the most expensive slot.
How do I check whether my model fallback chain gets more expensive as it falls back?
Pull out the primary plus every fallback, label each slot with its unit price, and read from top to bottom. A healthy chain gets cheaper as it falls back, or at least stays flat. It should not get more expensive. The red flag is any slot whose next model costs more than the one before it.
How do I know which model each LLM call actually used?
Log it per call. The config selects the intent; when rate limiting kicks in, the model that actually takes over could be any one at the tail of the chain. We added _llm_log_cost, which pulls tokens from the usage field of every response, estimates the cost, and writes each call to llm-cost.log, so actual usage can be grouped by model. If you cannot answer what percentage of requests in the past day actually ran on the most expensive model, you are not logging it.
How do I test where the fallback degrades when the primary model fails?
Set the primary to a model that will fail on purpose, force the fallback to take over, and check the log for which model actually stepped in. The direction it falls is the direction your bill moves under high load. If the model that takes over costs more than the primary, your degradation runs toward the expensive end.
How do you fix a fallback chain that climbs toward the most expensive model?
We made two cuts on 2026-07-04. The first removed every Claude from every chain and changed the default tier to gemini-2.5-flash → gemini-2.0-flash-001, so none of the three tiers can fall to Claude anymore and degradation only moves toward cheaper or equal cost. The second added per-call cost logging, so which service burned how much on which model is always visible.