AI AgentException HandlingRetry LogicBuildInPublicSolo Business

How to Classify Errors Before You Retry: What to Back Off On and What to Stop Immediately

· 7 min read
Lobster Fleet · Pattern Audit · Part 13 of 25
Table of Contents
  1. The traps we hit before we got classification right
  2. Spend a few minutes checking your own error handling
  3. What we actually changed
  4. Exception handling checkup list
  5. This chapter in four sentences

Chapter 12 of Agentic Design Patterns covers exception handling and recovery. The core idea in one sentence: errors are not all the same thing, so classify them first. Only transient errors are worth a backoff and retry; permanent errors should be stopped on the spot. Forcing retries on an error that will never get better is not resilience. It turns one failure into a storm.

We read it and nodded. Our gateway had long since rewritten _openrouter_call: 402 and 401 stop immediately with no retry, only 429 and 5xx back off, and the backoff has a cap. Exception handling, implemented. Then we ran an adversarial audit on ourselves and turned up two things: an incident that really did drive the gateway into a crash loop, and an alert re-reporting bug where the same unchanged state bit us 4 times in one day.

The traps we hit before we got classification right

The earliest version of _openrouter_call fired off a single curl and never looked at the HTTP status that came back, so it lumped two completely different kinds of failure into one:

Transient errors such as rate limits and 5xx were treated as "this model is down", and it jumped straight to the next fallback model. The next model in the fallback chain is often more expensive, so a momentary rate limit quietly upgraded you to a pricier model.

Account-level permanent errors, a 402 for no credit or a 401 for a wrong key, were also replayed across the whole fallback chain. The problem is that every model on the chain uses the same key and the same balance. If model A returns 402, models B and C will return 402 too. Replaying the whole chain on a 402 is a pointless storm that is guaranteed to fail.

This is exactly where the real incident grew. On the day the credit hit zero, max_tokens was still set to 65535, and every request was bounced with a 402 as soon as it went out. The gateway ran under systemd with automatic restart, so it crashed, got a 402, crashed again, and was revived again: at the process level too, a permanent error was treated as one that could simply be run again. This is the same crash loop as in Chapter 16, but that chapter counted how much money it burned. The question for this chapter is: why would an error that can never succeed be retried by the system over and over? The answer is that it was never classified in the first place.

Spend a few minutes checking your own error handling

01 Do you actually classify errors, or do all errors take the same path? Open the part of your code that calls an LLM or external API and see whether it routes by HTTP status. The errors to retry (429, 5xx, connection timeouts) and the errors to stop on immediately (402, 401, 403) need separate paths. The most dangerous case is putting 402 into the retry loop as well, which can only produce a storm.

# Does your error handling read the HTTP status and route on it, or treat every error the same
grep -nE 'http_code|status|case .*in|402|401|429|50[0-9]' your-llm-call.sh

Red flag: the whole section has a single error path (retry everything, or switch model for everything) and never reads a status code from start to finish.

02 Is retry capped? Both the count and the backoff need a ceiling Find your retry loop and confirm that both things have a ceiling: the number of retries is capped, and so is the backoff interval. Miss either one and a single failure snowballs into a crash loop on its own.

# Find the retry loop: is there a cap on attempts, and is the backoff capped
grep -nE 'max_attempts|max_retries|backoff|sleep|while ' your-llm-call.sh

Red flag: a while loop with no attempt counter, a backoff that keeps doubling (×2) with no cap, or relying on systemd to restart whatever crashed without thinking, with no circuit breaker in between.

03 Do your alerts have dedup? If the state has not changed, do not report it again Re-reporting a state that has not changed is not monitoring, it is spam. Check whether your alert remembers the previous state and fires only at the moment things go from healthy to failing, repeating at most once per interval while the failure persists.

# Does the alert keep state and fire only on a transition; count how many times the same one fires in a day
grep -nE 'state|stamp|last_fire|OFF|ON|轉態' your-alert.sh
grep -c 'crash-loop' ~/.openclaw/logs/tg-alert.log   # how many entries per day for the same problem

Red flag: the same unchanged condition shows up in the log several times a day. Our re-report bug bit us 4 times in one day.

What we actually changed

_openrouter_call now routes on http_code as the first thing it does when a response comes back. 2xx returns normally. 401 and 402 are judged permanent: it writes a log line and stops the whole chain on the spot, with no retry and no model switch, because switching gives the same 402. 429, 500, 502, 503, 504, and 000 for a connection timeout are judged transient: it backs off and retries the same model at intervals of 1, 2 and 4 seconds, up to 3 times, and gives up on that model only once those are used up. Everything else, such as 400, 403 and 404, is judged a model-level error that will fail the same way on retry: no retry, and the layer above moves straight on to the next fallback. Bounded, backed off, no storm.

On the alert side we added hysteresis. Each condition writes a state stamp and pushes once, only at the moment it goes from healthy to failing; while the failure persists it repeats at most every 60 minutes, instead of biting us every 5 minutes. The same crash loop now fires once, not 4 times.

Exception handling checkup list

Run through it against your own system:

  • Does your error handling classify by HTTP status first, or do all errors take the same path
  • Do permanent errors like 402 and 401 stop immediately, or do they also get retried and switched to another model
  • Could your model-switching fallback amplify an error that "fails for the whole account" into a storm across the entire chain
  • Is retry capped, with both the count and the backoff interval bounded
  • Do services that rely on systemd/PM2 auto-restart have a circuit breaker, or are they revived without thinking every time they crash
  • Do your alerts have dedup, so an unchanged state is not re-reported, and how many times a day does the same problem fire

This chapter in four sentences

  • Errors are not all the same. Classify first: only transient errors get backoff and retry; permanent errors like 402 and 401 must stop immediately.
  • Retrying a permanent error is not resilience, it is a storm. With a 402 on the same key, switching models, replaying the chain or restarting the process all end the same way.
  • Backoff must have a cap. Both the count and the interval need a ceiling, or a single failure snowballs into a crash loop on its own.
  • Alerts should fire only when the state has really changed. Re-reporting an unchanged state is not monitoring, it is spam.

Source locations: ~/.openclaw/scripts/ollama-helper.sh (_openrouter_call: 401/402 permanent errors stop immediately, 429/5xx transient errors back off 1→2→4 seconds with a cap of 3 attempts, other model-level errors are not retried and move on to the next model); alert dedup lives in ~/.openclaw/scripts/lib/hard-alert-watchdog.sh (each condition writes a state stamp, pushes only on the OFF→ON transition, then repeats at most every 60 minutes).

This is part of the Agentic Design Patterns × Lobster Fleet series. Following the book, we systematise a solo company's AI agent fleet, then run an adversarial audit on ourselves. Every chapter we claim to have implemented has been verified again, and the investigation and the fix are written up as steps you can run. The credibility of this series comes from our willingness to publicly prove ourselves wrong.

FAQ

Which API errors should you retry, and which should stop immediately?

Only transient errors are worth a backoff and retry, such as 429, 5xx (500, 502, 503, 504) and connection timeouts. Account-level permanent errors like 402 (no credit) and 401 (wrong key) should stop immediately, with no retry and no model switch. Model-level errors such as 400, 403 and 404 fail the same way on retry, so skip the retry and move straight on to the next fallback.

Why shouldn't you retry a 402 error or switch models?

A 402 is an account-level permanent error for no credit. Every model on the fallback chain uses the same key and the same balance, so if model A returns 402, models B and C will return 402 too. Switching models, replaying the chain or restarting the process all end the same way: a pointless storm that is guaranteed to fail.

Why is jumping to a fallback model on a rate limit a bad idea?

Rate limits and 5xx are transient errors, not a sign that the model is down. The next model in the fallback chain is often more expensive, so treating a momentary rate limit as a dead model and switching quietly upgrades you to a pricier one. Transient errors should back off and retry the same model instead.

How should you cap retry backoff?

Both the retry count and the backoff interval need a ceiling. Miss either one and a single failure snowballs into a crash loop on its own. Our gateway backs off and retries the same model at intervals of 1, 2 and 4 seconds, up to 3 times, and gives up on that model only once those are used up. Services that rely on systemd or PM2 auto-restart also need a circuit breaker, instead of being revived without thinking every time they crash.

How do you dedup an alert that keeps firing for the same problem?

Have the alert remember the previous state and fire only at the moment things go from healthy to failing, repeating at most once per interval while the failure persists. We have each condition write a state stamp, push once only on that transition, then repeat at most every 60 minutes. The same crash loop now fires once, not 4 times.

Lobster Fleet · Pattern Audit · Part 13 of 25

Weekly AI Automation Playbook

No fluff — just templates, SOPs, and technical breakdowns you can use right away.

Join the Solo Lab Community

Free resource packs, daily build logs, and AI agents you can talk to. A community for solo devs who build with AI.

Want to try it yourself?

UltraProbe is free and needs no sign-up. One scan tells you whether Google and AI engines can find your site.