Skip to content

Back to Articles

5 min read ·

Your fallback is lying to you

A fallback that nobody hears about is worse than a crash. The crash, at least, is honest.

This month I changed what the model client returns in a teammate's email pipelines. Two steps there lean on an LLM through OpenRouter: classifying incoming mail, and working out which commercial invoice belongs to which house waybill. Both have a regex path for when the model can't answer.

The client had the habit most model wrappers pick up. chat_completion returned None for four different reasons: no API key, an HTTP error, a transport failure, a malformed body. Four different problems went in, one shrug came out, and every caller read the shrug as "use the deterministic path."

Fine. Use it.

But a caller cannot be loud about a fallback it cannot detect. I call this polite failure. Each layer catches the error and hands something harmless to the layer above, which does the same, until the only thing that reaches a human is a green tick on a dashboard. It gets waved through as defensive programming. It's lying, and the cheapest kind, because return None costs one line and the truth costs a type.

I'm not against the fallback. When the model can't answer, the regex keeps the mail moving, and I kept the regex. But a regex out of its depth still returns something, and something looks a lot like an answer. A silent fallback is the restaurant that runs out of paneer and brings you tofu without a word.

So every fallback now gets a name, a type and an alarm.

The client returns a typed LLMResult with a degraded_reason, and fallback is a required argument, so no caller reaches the regex path without writing down, at the call site, that it means to. A 200 is not an answer either. A body with an error field in it, a choice whose finish_reason is "error", or an empty choices list is degraded, and the result says so.

The alarm is one P1 per reason per hour, not one per call, because alerting per call is indistinguishable from alerting on none of them. A 403 is moderation, and it doesn't page. Neither does a 400 or a 404, which is our bug, not the provider's, so it goes to the logs marked degraded instead of to a pager.

The bill fails the same way. The gateway costs about a dollar a day, and a prepaid balance makes no noise on its way to zero. At zero, every call is an HTTP error, which the client now names and pages on, and a page that arrives after the money is gone is late. That dollar is a production dependency, so a cron job reads the balance at 09:00 and 15:00 IST and warns at 80% of the limit.

Inngest takes polite failure at its word. A function that returns an error dict instead of raising is, to the platform, a completed run, and a completed run burns its 24-hour idempotency key. The work never happens, and for a day the platform refuses to try it again. The rule can't stop at the model client.

A run that does no work must not report success.


Polite failure doesn't need a model. A browser SDK that ignored my region parameter and handed me a session anyway, in Oregon, had the same manners. Logs have their own version, and it's called truncation. Cut an Outlook message id to 30 characters and every email in a shared folder gets the same name, because the first 30 characters of the id are the folder's GUID. Nothing throws. The logs tell you, calmly, that a whole folder is one email.

Truncation is not identity.

The trace id now comes from the work: a SHA-256 of the origin and the natural key, cut to 16 characters, across about 20 entry points. Yes, that's a cut too. The first 16 characters of a hash are not a folder name. The id is derived, never generated, so if the same email ever runs twice, both runs carry the same trace without asking each other.

The part I trust most is the test. Context variables don't cross client.send, so every emitter carries the trace by hand, and anything carried by hand gets forgotten. Under the old suite you could delete with_trace( from all 38 emitters and stay green. The new one parses the source and goes red if any emitter drops the trace.


That leaves the loud kind, the function that dies outright, and the question of who hears it. The Inngest SDK's answer is a per-function on_failure hook, which makes alerting opt-in, and opt-in alerting covers exactly the functions somebody remembered. But that hook is itself just a generated consumer of inngest/function.failed, filtered to one function, and the platform emits that event for every function. One consumer, written once, covers all of them, including the ones nobody has written yet. It retries three times. The generated handlers retry zero times, and a Slack blip would eat the alert.

Then I got clever. On top of per-key suppression and a cap on how many alerts one burst can send, I wrote more code to remember which channels each alert had already reached. Across three review rounds each fix created the next finding, until a recheck showed that my code could drop a real alert and report it as sent. Polite failure, in the code I wrote to end it.

I deleted the mechanism. Slack decides whether an alert reached a human.

Open your model client and find the return None. Count the things it can mean.