Skip to content
Back to the blog

Updated September 29, 2026

API resilience: surviving slow, unreliable providers

A results screen with six cards: five complete and one empty, in place and aligned with the others

If your product calls third-party APIs to build its responses, your availability is all of theirs multiplied together. With ten providers at 99.5% each, a system that waits for all of them drops to about 95%. The job is to respond well when they fail, because they will.

When the problem shows up

A provider going down cleanly is the easy case: it’s detected in milliseconds and handled.

The problem shows up when a provider gets slow. That’s when everything else falls over, and the mechanism is always the same:

  1. A provider that used to respond in 200 ms starts taking 8 seconds.
  2. Your requests to that provider stop completing, and each one holds a thread or a connection.
  3. The shared connection pool runs dry.
  4. Requests that have nothing to do with that provider start failing.
  5. Your system is down because of a secondary feature almost nobody uses.

Step 4 is the one that surprises people the first time. A slow provider exhausts a shared resource and takes the rest down with it, well beyond its own feature.

With an 8 s wait and 50 connections, 7 requests per second to the slow provider are enough to use them all up.

Why retrying makes things worse

The instinctive reaction to a failure is to retry. And uncontrolled retrying is what turns a slowdown into an outage.

If a provider is slow because it’s overloaded, sending it every request three times triples exactly the load that overloaded it. The result is that a provider that would have recovered in two minutes doesn’t, because you, and every other customer that’s retrying too, keep it down.

Retries are needed, but under three conditions that are almost never all in place:

  • Only on transient failures. A validation error retried is still a validation error.
  • With increasing backoff and random jitter. Without the jitter, all your processes retry at once and the spike repeats identically every few seconds.
  • With a cap. Two or three, plus an overall budget: if the retry rate crosses a threshold, stop retrying, even if each individual request would otherwise qualify.

The four patterns

1 · Latency budget

Before writing any code, decide how much time the user has. If the response has to arrive in 2 seconds and you talk to three providers, none of them can have a 2-second timeout.

The budget is split from the top down, and each call gets its share. A timeout copied from the provider’s documentation example, or left at the library’s default, means there’s no budget at all.

No infinite timeouts, anywhere. It’s the cheapest rule to apply and the one most often broken.

2 · Circuit breaker

When a provider racks up failures, you stop calling it for a while. Requests fail immediately, without using up connections or waiting, and every so often a single request tests whether it’s back.

It has three states: closed (all normal), open (no calls, fail fast) and half-open (test with one). The implementation is in any library; the thresholds are what matter: how many failures open the circuit, how long it stays open and how many successes close it. Those numbers come from measuring. Copying them from somewhere else doesn’t work.

The real benefit goes beyond protecting the provider: it frees your own resources so the rest of your product keeps working.

3 · Bulkheads

The name comes from a ship’s hull: if one compartment floods, the ship stays afloat.

In practice, each provider gets its own connection pool and its own concurrency limit. If the slow provider is allotted ten connections, it can exhaust ten and no more. The rest of the system never notices.

It’s the pattern that solves step 4 of the mechanism above, and the one most often forgotten because you don’t notice it’s missing until the day you need it.

The slow provider's compartment still fills up. The difference is that now it only switches off its own feature, and the rest of the product carries on.

4 · Graceful degradation

It’s a product decision: what gets shown when a piece is missing.

The options, in order of preference:

  • Serve what you have. Eight providers out of ten responded: show those eight and say that some results are missing. It’s almost always better than showing nothing.
  • Serve cached data and flag it. A price from ten minutes ago is usually worth more than an error.
  • Hide the feature. If the part that’s failing is secondary, let it disappear rather than break the page.
  • Fail clearly. The last resort, and only when serving something incomplete would be worse than serving nothing. A payment, for example.

This decision belongs to whoever knows the business, and it has to be made in advance, because at three o’clock on a Tuesday afternoon there’s no time left to make it.

What it looks like in production

Three things you need to be able to see, or the four patterns above are just an act of faith:

  • Latency per provider, in percentiles. The average always lies. What matters is p95 and p99, because the worst 5% of experiences are the ones that generate complaints.
  • Error rate and circuit state per provider. If a circuit has been open for twenty minutes, you need to know while it’s happening.
  • How many responses were partial. It’s the metric nobody sets up, and the one that really tells you whether your product is delivering what it promises.

And one alert worth having: silent degradation. A system serving partial results without complaining can run half-broken for weeks without anyone noticing, because technically it returns 200.

What still fails even if you do all this

  • The provider that responds successfully with wrong data. No resilience pattern detects a 200 with garbage in it. You have to validate the response, and almost nobody does.
  • The unannounced change. A field that changes type, a different date format. Other people’s API contracts are less stable than their documentation suggests.
  • Rate limits you didn’t know about. Many providers don’t publish their real one, and you find out the day your product grows.
  • Correlated outages. If three providers run on the same cloud, they aren’t three independent risks.

When this is over-engineering

If you call a single provider, off the user’s path, and a failure only delays a background job, don’t build any of this. Set a reasonable timeout, one retry with increasing backoff and a dead-letter queue. Then move on.

The practical threshold is two conditions at once: the call is on the user’s path and there’s more than one external dependency. With a single dependency and no user waiting, the timeout and the retry cover almost all of the risk for very little effort.

What’s always worth doing, from the very first integration: set an explicit timeout. There’s no case in which waiting indefinitely on a third party is the right decision.


Nimboo builds whole products, including the integrations the product needs to work: payment provider, email, calendar, e-signature. How all of this gets decided in the design phase is in custom SaaS, and the price of each phase is on the pricing page.

← Back to the blog Tell us about your project →