AI AgentLLM CostOpenRouterHermesBuildInPublicSolo Business

How to Check Whether Your AI Agent Is Quietly Burning Money

· 8 min read
Lobster Fleet · Pattern Audit · Part 17 of 25
Table of Contents
  1. What the ledger looked like once we opened it
  2. Spend three minutes checking your own agent first
  3. What we found in our logs
  4. Fix it: three cuts
  5. The last cut, and the hardest
  6. Resource-awareness checklist
  7. This chapter in four sentences

A few months ago, in AI Agent Token Optimization in Practice, I wrote that we had audited our four agents, cut their waste by 40%, and called context the biggest hidden cost. This chapter proves that post wrong. The main gateway burned $17 in one week, and 98.8% of that money went to the very disease I had claimed was cured: making the agent reread the conversation history over and over.

Chapter 16 of Agentic Design Patterns covers resource-aware optimization. The core idea in one sentence: every cost an agent incurs must be visible, and the choice of model should follow the difficulty of the task, because bigger is not always better.

We read it, nodded, and ticked a box in our heads: we already do this. We have cost alerts, a three-tier model fallback, and a spend monitoring script. Resource awareness, implemented. Then we ran an adversarial audit on ourselves. What this chapter turned up is the most painful cut in the whole series.

What the ledger looked like once we opened it

After the migration to OpenRouter, the main gateway burned $17 in a week. That does not sound like much until you open the ledger it keeps for itself: 702K input tokens, in exchange for only 8K output tokens.

98.8% of the money went to rereading conversation history again and again. The part that actually produced content was 1.2%.

Type a single word, "huh?", and this agent packs 5,600 tokens into the request. The same stretch of history gets reread up to 60 times. It is not thinking, it is ruminating. This is not the first time, either. In June we took an NT$1,954 (New Taiwan dollars) Gemini bill from the same disease: when nobody watches the volume, it grows on its own.

Spend three minutes checking your own agent first

01 Check the idle ratio: input tokens versus output tokens A healthy ratio sits roughly between 1:1 and 3:1. If it shoots up to 88:1 like ours did, the agent is ruminating on its history. The Activity page in the OpenRouter dashboard shows this directly; if you keep your own records, sum the prompt_tokens and completion_tokens of every call in your log and compare them.

# Sum the in/out totals for the most recent day from your own cost log
grep '"usage"' cost.log | jq -s 'map(.prompt_tokens)|add', 'map(.completion_tokens)|add'

Red flag: input is more than 10 times output. You are paying to reread history.

02 Check whether max_tokens has a cap Open your agent's config and look for max_tokens. If it is unset, or set to the model's limit (tens of thousands), you have handed the model a blank cheque, and every request reserves the maximum allowance. Set it to the output length you actually need: 500 to 1000 is enough for a chat bot, 2000 to 4000 for long-form writing.

grep -niE 'max_tokens|max_output' config.yaml .env

Red flag: not found, null, or greater than 8192.

03 Check for a crash-and-restart loop A service that keeps crashing and keeps getting revived by the system reruns a full cycle, and burns money again, every time it comes back. Look at how many times your service has been restarted; an absurdly large number means something is crashing silently.

systemctl --user show your-service-name -p NRestarts
# or search the journal for a crash loop
journalctl --user -u your-service-name | grep -c 'Started'

Red flag: dozens to hundreds of restarts within a day or two.

What we found in our logs

We ran all three checks on ourselves, and every one of them hit. OpenRouter was returning HTTP 402, with a blunt message:

HTTP 402 This request requires more credits, or fewer max_tokens.
You requested up to 65535 tokens, but can only afford 15650.

The problem is that 65,535: it is the model's limit, which applies when max_tokens has no cap. Every request first reserved an allowance of 65,000 tokens. Once the balance hit zero, every request was bounced with a 402 and the gateway crashed on the spot. And because it runs under systemd with automatic restart, every crash got it revived. In twelve days it came back to life 403 times, reconnecting, reloading and rerunning each time. Why an error that can never succeed kept getting retried is a separate chapter: how to classify errors before you retry.

It is worth being clear about what it was not: the cron schedule was empty and the task board was empty, so there was no runaway time bomb. It was one uncapped max_tokens line, combined with pay-per-use billing, plus a crash loop that kept resurrecting itself. Even a cheap model cannot survive burning money like that.

Fix it: three cuts

Match these config changes to the problems you just found. These are the three cuts we actually made; you can copy them directly and then tune the numbers to your own needs.

# Before: a blank cheque
max_tokens: null        # equals the model limit, 65535
max_turns: 60           # cap on back-and-forth rounds per task
history_limit: 400      # number of history messages sent with each request

# After: set to what you actually need
max_tokens: 4096        # enough for a chat bot
max_turns: 20
history_limit: 80

Then do this, or you have not changed anything: restart, then check the logs to confirm the new values actually took effect. We fell into exactly this trap. max_turns in the config had clearly been changed to 20, yet the log showed the agent actually running 90 rounds. The cap was being overridden by an environment variable in .env, and the config was never read at all.

# After restarting, pull the values actually in effect from the log; do not trust the config file
grep -iE 'max_iterations|max_turns|budget' gateway.log | tail -3

The last cut, and the hardest

It is not a parameter tweak. It is shutting down the main agent outright. An agent that spends more than it produces is better turned off than left running to ruminate every day and resurrect after every crash. Taken to its conclusion, resource-aware optimization is not tuning every agent to be as cheap as possible. It is admitting that some agents are not worth running right now. That main agent was also the only one that collected messages from our inter-agent relay pipeline; what shutting it down left behind, and how to audit your own multi-agent setup, is in our multi-agent collaboration audit.

Resource-awareness checklist

Run this against your own system:

  • Are the input / output tokens of every API call recorded in a log?
  • Is max_tokens set, and is it the length you need or the model's limit?
  • Do conversation turns and history length each have a cap?
  • Is your cost alert threshold low enough? (Our $17 week slipped in under a $20 threshold; we later lowered it to $5.)
  • After changing a setting, did you go back to the log to confirm it took effect, instead of trusting only the config file?
  • Is there an agent that spends more than it produces and should really be shut down?

This chapter in four sentences

  • A cost you cannot see is a bill that has not exploded yet. The spend on every call must be visible.
  • Cap everything. max_tokens, turns, history: leave any one of them uncapped and it is a blank cheque.
  • Verify that the change took effect. If the config changed but the log never picked it up, nothing changed.
  • There is no shame in shutting down an agent that spends more than it produces. That is what resource awareness means.

Source location: ~/.openclaw/scripts/ollama-helper.sh (cost logging), llm-cost-report.sh, openrouter-spend-watch.sh (the alert threshold has been lowered from $20/week to $5/week), ~/hermes-bench/data/config.yaml (where the three cuts live).

This is part of the Agentic Design Patterns × Lobster Fleet series. We systematise a solo company's AI agent fleet chapter by chapter, then run an adversarial audit on ourselves. Every chapter we claim to have implemented gets verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publish our own failures.

FAQ

How do I tell whether my AI agent is idling and burning money?

Look at the ratio of input tokens to output tokens. A healthy ratio sits roughly between 1:1 and 3:1; if input is more than 10 times output, you are paying to reread conversation history. The Activity page in the OpenRouter dashboard shows this directly, and if you keep your own records, sum the prompt_tokens and completion_tokens of every call in your log and compare them.

What should I set max_tokens to?

Set it to the output length you actually need: 500 to 1000 is enough for a chat bot, 2000 to 4000 for long-form writing. Leaving it unset, or setting it to the model's limit, hands the model a blank cheque, and every request reserves the maximum allowance. Not found, null, or greater than 8192 are all red flags.

What happens if max_tokens has no cap?

Every request first reserves an allowance up to the model's limit, which in our case was 65,535 tokens. Once the balance hit zero, OpenRouter bounced every request with HTTP 402 and the gateway crashed on the spot. Because it ran under systemd with automatic restart, it came back to life 403 times in twelve days, reconnecting, reloading and rerunning each time.

Why did my agent config change not take effect?

Something else may be overriding it. We changed max_turns to 20 in the config, yet the log showed the agent actually running 90 rounds, because an environment variable in .env overrode the cap and the config was never read at all. After changing a setting, restart, then pull the values actually in effect from the log instead of trusting the config file.

How do I check whether a service is stuck in a crash-and-restart loop?

Look at how many times the service has been restarted: systemctl --user show with -p NRestarts gives the count, or count how often 'Started' appears in the journal. Dozens to hundreds of restarts within a day or two means something is crashing silently, and every revival reruns a full cycle and burns money again.

Lobster Fleet · Pattern Audit · Part 17 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.