Post-mortem: the AI spending cap and the fallback that said nothing
· 4 min · post-mortem, ai, openrouter, osservabilità
In July 2026 the model provider key hit its spending cap. The brief generator silently fell back to a canned text and the chat answered 502. How the zero-token rows gave it away.
Timeline
Between 16 and 19 July 2026, for about two days, every AI function on the site was down: the brief generator, the chat, the Seantral assistant and the triage of improvement requests. Recovery came from raising the key’s spending cap in the provider’s dashboard.
Cause
Model calls go through OpenRouter with a key that carries a dollar spending cap. The cap was low, it was reached, and from then on every call was rejected. The cap itself was a sound choice; the problem was what happened next.
When the model does not answer, the brief function returns fallback content flagged _fallback: true: three generic questions instead of the analysis. The chat answered with a 502 and a polite message. Both are designed not to break the page for its user, and together they hid the outage from whoever runs the site.
How it was spotted
From behaviour: a brief with only three generic questions, no analysis. No alarm fired. The diagnosis took two steps. Call the brief’s analyze action with the public key and find _fallback: true in the response. Then read the latest rows of the usage table, ai_usage_events, by date: the log writes a row even when the call fails, and the rows with zero input and output tokens were the rejected calls. The first row with tokens above zero, going backwards, marks the last moment the AI worked.
One false lead cost time: the Supabase management API returns a digest of function secrets, not their value. Trying that string as a key against the provider gives authentication errors that look like a broken key, and the key was fine.
Fix
The key’s spending cap was raised in the provider’s dashboard, and calls resumed. No code change was needed for the recovery.
Ratchet added
Here the ratchet is incomplete, and that needs saying. What remains is the two-minute diagnosis procedure written in the project notes, and the rule that a secret is never checked by reading it from the management API. The missing tooth is an alarm when the usage table piles up consecutive zero-token rows: that is the signal that revealed the outage, and today a person still reads it. Until that alarm exists, the same incident would come back the same way.