Your AI Feature Has a Single Point of Failure
If your product calls one model from one provider, then your uptime is capped by theirs, and theirs is not a number you control. Every major provider has had partial and full outages. They publish status pages, they apologize, and they fix it. None of that helps the user who is staring at a spinner inside your app.
Founders usually discover this on a bad day. A customer demo stalls, support tickets pile up, and someone in Slack asks whether we can "just switch to the other one." Switching is possible, but not in ten minutes unless you built for it before the incident. That is what LLM provider failover means in practice: the work you do on a quiet Tuesday so the loud Thursday is boring.
What Actually Fails
Full outages get the headlines, but they are the smallest part of the problem. The failures that hurt are quieter. Here are the four you will meet.
Hard outages
The API returns 500s or 503s, or nothing at all. This is the easy case to detect and the rarest. It usually lasts minutes to a couple of hours, and it tends to hit a single model or region before it hits everything.
Rate limits and capacity errors
The 429 is the failure you will see most. Sometimes it is your own quota, because a batch job ate the budget that your chat feature needed. Sometimes it is the provider shedding load during a busy period, and you get "overloaded" responses even though you are well under your limits. The fix for the first is on your side. The fix for the second is a second route.
Latency spikes
This is the sneaky one. The API is technically up, but a call that normally takes two seconds now takes forty. Nothing returns an error, so your retry logic never fires, your request queue backs up, and your servers run out of connections. A slow dependency is often worse than a dead one because dead ones fail fast.
Model deprecations and silent changes
Providers retire model versions on a schedule, usually with months of notice, and the notice goes to an email address nobody reads. When the date passes, calls start failing. There is a related quieter version: a provider updates a model alias, the behavior shifts, and your carefully tuned prompt starts producing slightly different output. You will not see an error. You will see a support ticket three weeks later.
Each needs a different response. A timeout fixes the latency spike. A second provider covers the outage and the capacity error. Pinned versions and a calendar reminder cover deprecation.
Timeouts and Retries Come First
Before you add a second provider, fix the basics. Most teams we meet have default timeouts from their HTTP client, which are often 60 seconds or no limit at all, and no retry policy beyond whatever the SDK does on its own. That is the first thing to change, and it costs an afternoon.
Set timeouts by use case
A chat reply that streams can tolerate a long total time as long as the first token arrives quickly. So set two limits: a short one for time to first token (5 to 10 seconds is a fair starting point) and a longer one for the whole response. A background summarization job can wait a minute or more. A classification call inside a form submit should give up after a few seconds. Pick numbers from your own latency data, not from a blog post, including this one. Look at your p95 and set the timeout at two to three times that.
Retry with backoff and jitter
Retry on 429, 500, 502, 503 and timeouts. Do not retry on 400 or 401, because the request itself is wrong and sending it again changes nothing. Use exponential backoff with random jitter: wait about one second, then two, then four, each with a random offset. Without jitter, every client retries at the same instant and you create a second spike right as the provider recovers. If the response carries a retry-after header, respect it.
Cap the attempts at two or three. A request that has failed three times is not going to succeed on the fourth, and your user left a while ago.
Use a circuit breaker
If the last twenty calls to a provider mostly failed, stop sending the twenty-first. A circuit breaker trips after a failure threshold, rejects calls instantly for a cooldown period (30 to 60 seconds is typical), then lets a trickle through to test recovery. It protects your own servers from piling up on a dead dependency, and it is also the signal that tells your failover logic when to switch routes.
Make requests safe to repeat
If a model call triggers a tool that sends an email or charges a card, a retry may do it twice. Attach an idempotency key to the action.
Adding a Second Provider Through a Gateway
Once timeouts and retries are in place, a second provider is the real answer to outages and capacity errors. You do not want to wire two vendor SDKs into your application code. You want one internal interface, and a gateway behind it that knows how to route and fall back. We cover the broader pattern in our piece on AI gateway architecture, rate limiting and caching. Here is the failover slice of it.
The options fall into a few groups:
- LiteLLM. An open source proxy and Python library that puts an OpenAI-style interface in front of a long list of providers. You host it yourself, which means you own its uptime, but you also own the config. Fallback lists, retries and cooldowns are first-class settings. This is the pick I would make for teams with someone comfortable running a small service.
- OpenRouter. A hosted router with one API key and one bill across many models, and it can fall back between providers that serve the same model. It is the fastest way to get a second route live, often in a day. The tradeoff is an extra hop, a platform fee on top of model prices that you should check against current terms, and one more vendor in the path of your data.
- Vercel AI Gateway. A good fit if you already deploy on Vercel and use their AI SDK. Setup is minimal and fallback ordering is built in. It is less attractive if your backend lives somewhere else.
- Portkey. A hosted or self-hosted gateway with stronger observability, guardrails and config-driven routing. It suits teams that want logs, budgets and fallback rules in one place and are willing to learn another dashboard.
Pick one and stop shopping. Under ten engineers, go with OpenRouter or Vercel AI Gateway if speed matters most, and LiteLLM if you want control and already run containers. A gateway can also be a single point of failure. Run at least two LiteLLM instances behind a load balancer, or accept that a hosted gateway's uptime is now part of your math.
The Prompt Portability Catch
Here is what the gateway marketing leaves out. Routing a request to a different provider is easy. Getting an equally good answer back is not. A prompt tuned against one model does not transfer cleanly to another, even when the API shape is identical.
The differences that bite:
- Instruction following. One model obeys a long system prompt closely. Another drifts after a few turns. Your tone, format and refusal rules may all shift.
- Structured output. JSON mode, schema enforcement and tool calling work differently across vendors. A response your parser handled fine from model A may arrive wrapped in a code fence or with a trailing sentence from model B.
- Tool calling. Function schemas, parallel calls and error handling vary. If your agent loop depends on one vendor's quirks, the fallback model may loop or stall.
- Safety behavior. A request one model answers may be refused by the other, and the refusal text will not match your UI.
The practical answer is to treat the fallback model as a product you test, not a switch you flip. Keep an evaluation set of 50 to 200 real prompts with expected properties, and run it against both models whenever you change a prompt. Score on what your product needs: valid JSON, correct tool selection, tone, length. When the fallback fails on a category of prompts, write a model-specific prompt variant for it and store variants alongside the primary prompt in your config. A fallback that scores 85 percent of the primary is perfectly acceptable for an outage. A fallback you have never run is a gamble.
Decide which features get a fallback at all. Your chat assistant probably does. A heavily tuned extraction pipeline probably does not, and its degraded mode should be something other than a different model.
Degraded Modes Users Can Live With
Failover to a second model is one tool. It does not help when both providers are slow, when your budget is gone, or when the fallback cannot do the job. Design a ladder of degraded modes, ordered from least to most noticeable, and decide in advance when each one applies.
Cached answers
If a large share of your traffic is repeat or near-repeat questions, a cache gives you an answer with no model call at all. Exact-match caching is trivial and safe. Semantic caching, which matches similar questions by embedding, saves more but can return a wrong answer to a question that only looks similar, so use a high similarity threshold and keep it away from anything personalized. In an outage, serving a day-old answer with a small "this may be out of date" note beats an error.
A smaller or different model
When your primary is overloaded, a smaller model from the same provider is often still up, because capacity is tracked per model. We go deeper on tiers in our guide to model routing and LLM cost optimization. The same routing table you build to save money doubles as your failover map. Quality drops, so show it honestly: shorter answers, fewer features, no long document analysis.
A queue and a promise
Many AI tasks do not need an instant answer. Report generation, document processing, email drafting and enrichment can all go into a job queue and complete when the provider recovers. Tell the user "We will email you when this is ready" and mean it. This is the cheapest degraded mode to build if you already use a queue like SQS, BullMQ or Cloud Tasks, and it turns an outage into a delay.
A human fallback
For high-stakes flows such as support escalation, medical intake, compliance review or anything with money moving, the right fallback is a person. Route the request to a support inbox or an ops queue with the context attached. It is slow and it costs labor, but it keeps the promise your product made.
An honest message
The last rung is a clear error. Say what is not working, what still works, and when to check back. Users forgive "our AI features are temporarily limited" faster than a frozen screen.
Test It With Fault Injection
Failover code that has never run in anger does not work. Untested fallback paths almost always have a bug. The fix is to break things on purpose, in staging first and then carefully in production.
A practical test plan:
- Return fake errors. Put a mock server or a proxy in front of the provider in staging and make it return 429, 500 and 503 at a configurable rate. Confirm that retries fire, backoff looks right in the logs, and the circuit breaker trips.
- Add latency. Make the mock hold responses for 30 seconds. Confirm timeouts fire, connections are released and your app stays responsive for other users. This test finds more bugs than any other.
- Kill the primary route. Blackhole the primary provider's hostname and verify the gateway moves to the secondary within your target time. Measure how long it took.
- Send malformed output. Have the mock return truncated JSON, an empty string and a refusal. Your parser should degrade, not crash.
- Run a game day. Once a quarter, pick a time, tell the team, and switch the primary off in production for a small slice of traffic or an internal-only tenant. Watch the dashboards, time the recovery and write down what surprised you.
A local proxy such as Toxiproxy or a feature flag that forces an error path is plenty. Someone should see the degraded UI with their own eyes before a customer does. If you have never run an incident, our walkthrough on handling a production outage covers the human side: who owns communication, what to write in the status update and how to run the review afterward.
Status Monitoring That Warns You First
Provider status pages are useful but slow. They tend to confirm problems after your own users have already felt them, and they often report "degraded performance" for issues that break your specific workload. Do not rely on them as your alarm.
Measure from your side instead. Track these per provider and per model:
- Error rate by status code, so a rise in 429s looks different from a rise in 500s.
- Time to first token and total latency, at p50 and p95.
- Fallback rate: the share of requests served by something other than the primary route. This is the number that tells you the primary is hurting before it fully fails.
- Output sanity checks, such as the share of responses that fail JSON validation. A quiet model change shows up here first.
Alert on the trend, not only the threshold. A page for "error rate above 5 percent for three minutes" is good. A second, lower-priority alert for "fallback rate doubled over the last hour" catches the slow burn. Send the first to whoever is on call and the second to a channel.
Subscribe to the status pages anyway, and route deprecation notices to a shared inbox with a real owner. Put every model retirement date on a team calendar with a reminder 60 days out. Moving to a new model version is a normal prompt-and-eval project. Doing it the week before a hard cutoff is a fire drill.
What Redundancy Costs and When to Skip It
Failover is not free, and anyone who says it is has not done it. The costs fall into four buckets.
- Engineering time. Timeouts, retries, a circuit breaker and a degraded UI are about a week of work for one engineer on a simple feature. Adding a gateway and a second provider with an eval suite is more like two to four weeks, depending on how many prompts and tools you have.
- Ongoing maintenance. Every prompt change now needs checking against two models, plus a day or two per quarter for the game day.
- Infrastructure. A hosted gateway may add a platform fee or a small markup. A self-hosted one costs a couple of small instances. Either is usually tens to a few hundred dollars a month, small next to your model spend.
- Model spend. A standby provider costs nothing while idle on pay-as-you-go pricing. Reserved or committed capacity is different, so check contract terms before you commit to two vendors.
The cheaper half of the work, timeouts, retries, caching and an honest error screen, is worth doing for almost everyone. The expensive half, the second provider with tested prompts, depends on what an outage costs you.
When not to bother
Skip the second provider if your AI feature is a nice extra and the product works without it. Skip it if you are pre-launch or still testing whether anyone wants the feature, since your prompts will change weekly and an eval suite for two models will slow you down. Skip it if the work is asynchronous and a queue plus retries already turns every outage into a delay. And skip it if a few hours of downtime a year is cheaper than a month of engineering, which is true for many internal tools.
Build it if AI is the product, if your customers have contractual uptime expectations, if the feature sits in a revenue or safety path, or if you have already had one painful incident. A rough test: if you would refund customers for a four-hour outage, you can afford to build the fallback.
If you are not sure where your product sits on that line, book a free strategy call.