The short version
Treat your LLM provider like any dependency that will eventually fail. Set strict timeouts, stop calling a provider that is clearly down with a circuit breaker, and keep a second route ready, whether that is another region, another provider or an honest degraded mode. Retries alone will not save you from an outage that lasts six hours.
What happened on 29 September
According to a postmortem write up by Artur Markus, Azure OpenAI Service, Foundry Agent Service, Foundry Models and Cognitive Services in the Sweden Central region failed for 5 hours and 55 minutes on 29 September 2026, from 10:03 to 15:58 UTC. Customers saw intermittent request failures, higher latency and HTTP 5XX errors against model and data plane APIs. Every cloud and model provider has incidents at some point, so this is not a post about blame. It is about how to engineer for the day it happens to you.
The same write up mentions a separate networking incident the next day across several regions, but I am not drawing conclusions from it because those details have not been confirmed. For the confirmed record, the official Azure status history is the place to look, rather than anyone’s summary, including mine.
The details of this incident matter less than its shape. A model API in one region was unusable for most of a working day. If your product calls an LLM on the critical path, that is your outage too, and the question is what your code does during hour three.
Why retries are the wrong tool here
Retries are built for blips: a dropped connection, a single overloaded node. Against a six hour incident they do harm. Every user request waits through several failed attempts, your workers pile up, your queue grows, and when the provider recovers it receives a flood.
The first step is to stop treating all errors the same. I sort them into three groups, because each one has a different correct reaction.
- A rate limit (HTTP 429) means not now. Reschedule the work instead of burning attempts, as described in Do Not Retry an LLM Rate Limit, Reschedule It.
- A client error (most other 4XX) means your request is wrong. Retrying changes nothing. Fix the request.
- An outage signal, meaning a timeout, a connection reset or a 5XX, means the provider is unhealthy. This is the group that needs a breaker and a fallback.
Timeouts come first
Without a timeout, an LLM call can hang for minutes and hold a connection, a worker and a user. Set one on every call, and choose it from what a healthy call looks like for your use case. A short classification should answer in a few seconds. A long generation needs a longer limit, and when you stream it is worth watching time to first token separately, because that is the number that tells you the provider is struggling.
const isOutage = (error) =>
error.name === 'TimeoutError' ||
error.code === 'ECONNRESET' ||
(error.status >= 500 && error.status < 600)
export async function callProvider(provider, messages) {
const signal = AbortSignal.timeout(provider.timeoutMs)
return provider.chat(messages, { signal })
}A circuit breaker in twenty lines
A circuit breaker remembers that a dependency is failing and stops sending it traffic for a while. That protects your users from waiting on a dead service and protects the service from a retry storm when it comes back.
export class CircuitBreaker {
constructor({ threshold = 5, coolDownMs = 30000 } = {}) {
this.threshold = threshold
this.coolDownMs = coolDownMs
this.failures = 0
this.openUntil = 0
}
allows(now = Date.now()) {
return now >= this.openUntil
}
success() {
this.failures = 0
this.openUntil = 0
}
failure(now = Date.now()) {
this.failures += 1
if (this.failures >= this.threshold) {
this.openUntil = now + this.coolDownMs
}
}
}After the cool down, allows returns true again and the next request acts as a probe. If it fails, the failure count is still above the threshold, so the breaker opens again straight away. If it succeeds, everything resets. That is the whole idea, and it is enough for most services.
A fallback chain
With timeouts and a breaker per provider, the fallback is a loop over an ordered list. Each entry has a name, a chat function, a timeout and its own breaker.
export async function askWithFallback(providers, messages) {
let lastError
for (const provider of providers) {
if (!provider.breaker.allows()) continue
try {
const reply = await callProvider(provider, messages)
provider.breaker.success()
return { reply, provider: provider.name }
} catch (error) {
if (!isOutage(error)) throw error
provider.breaker.failure()
lastError = error
}
}
throw lastError ?? new Error('Every provider is unavailable')
}Two cautions. A second region protects you from a regional failure like Sweden Central, but not from a problem that spans the provider, so for the critical path a second provider is the stronger choice. And a fallback model is not a drop in copy. Prompts that work well on one model can behave differently on another, so keep a small set of real examples and run it against every provider you list. A fallback you have never tested is a hope, not a plan.
This is easier when your code never imports a vendor SDK directly. If each provider is an adapter behind one small interface, adding a second is an afternoon of work. The same design is used in the booking assistant described in An AI Booking Assistant That Cannot Double Book.
Decide what failure looks like for the user
When every route is down, the worst outcome is a spinner that never ends. Decide in advance. Background work can go into a queue and run later. A chat can say plainly that the assistant is unavailable and offer another way to reach you. Something that has a cached answer can serve it, clearly marked.
The Copper Larder chatbot takes this approach: when the model is down, rate limited or missing its key, every path still returns a warm, on brand message and a callback card, so a visitor can still leave a number. The design is covered in the Copper Larder case study.
Test the failure before it finds you
You do not need a real outage to rehearse one. Write a fake provider that returns a 503, another that hangs past its timeout, and a healthy one, then assert how your code behaves.
import test from 'node:test'
import assert from 'node:assert'
const broken = {
name: 'broken',
timeoutMs: 100,
breaker: new CircuitBreaker({ threshold: 1 }),
chat: async () => { throw Object.assign(new Error('down'), { status: 503 }) },
}
const healthy = {
name: 'healthy',
timeoutMs: 100,
breaker: new CircuitBreaker(),
chat: async () => 'ok',
}
test('falls back when the first provider is down', async () => {
const result = await askWithFallback([broken, healthy], [])
assert.equal(result.provider, 'healthy')
assert.equal(broken.breaker.allows(), false)
})Run something like this in CI, and once in a while rehearse it for real by pointing a staging environment at a provider that always fails. You will find the one forgotten call that has no timeout.
A short checklist
- Every LLM call has a timeout, and the limit matches what a healthy call looks like.
- Rate limits, client errors and outage signals are handled differently.
- Each provider has a circuit breaker, and the order of providers is explicit.
- The fallback has been tested against your own prompts, not just reached.
- There is a written answer to what the user sees when everything is down.