API Integration

Wrangling Third-Party APIs: How to Handle Rate Limits and Downtime Gracefully

Build integrations that slow down deliberately, recover safely, and tell users what is happening when a provider struggles.

3 min read

A third-party API is part of your application's behavior even though you do not control its deployment. When it slows down, customers experience your product as slow. When it rejects requests, your retries can either support recovery or make the situation worse.

Reliability starts with deciding what each integration is allowed to delay, repeat, or temporarily omit. A payment submission and a cached weather lookup should not share the same retry policy simply because both use HTTP.

Give each operation a time budget

Set a deadline for the whole user operation, then allocate smaller limits to downstream requests. If a screen must answer promptly, spending its entire budget on the first attempt leaves no room for recovery or a useful fallback. A timeout should release waiting resources and produce a state the caller can handle.

Distinguish concurrency from request rate. Ten workers can still exceed a provider's quota if each sends requests quickly. A queue with a bounded worker pool protects local resources, while a shared rate limiter coordinates quota usage across instances. Decide whether limits apply per account, API key, tenant, or application, because placing the counter at the wrong scope makes apparently correct limits ineffective.

Respect the provider's recovery signal

HTTP 429 indicates rate limiting. A provider may include Retry-After to express a waiting period, either as a number of seconds or an HTTP date. Parse both forms and inspect the specific provider's contract. A missing or invalid header needs a bounded fallback policy, not an immediate retry loop.

If the requested wait exceeds the remaining operation budget, stop the synchronous attempt and schedule later work only when that is valid for the product. Do not shorten the provider's requested delay just to keep the interface busy. Explain to the caller whether the request was rejected, deferred, or accepted for background processing.

Retry only when repetition is safe

Exponential backoff with jitter spreads retries over time instead of sending every waiting client back at once. Bound the number of attempts and total elapsed time. Keep one clear owner for retries; nested retries in an SDK, service wrapper, and job runner can multiply traffic unexpectedly.

A timeout does not prove that a write failed. The provider may have completed the operation before the response disappeared. For operations that create orders or charges, use the provider's documented idempotency mechanism or reconcile by a durable operation reference. If neither exists, surface uncertainty for review instead of blindly repeating the write. Retrying a validation error with identical input is equally unhelpful: nothing relevant has changed.

Design a degraded mode before an outage

A circuit breaker temporarily stops calls to a failing dependency and later permits controlled recovery probes. It is useful only when the product has a defined response while the circuit is open. That might be an explicitly dated cached value, a disabled optional feature, or a queued task with a visible status. A fabricated success is never a useful fallback.

Track outcome categories, provider latency, queue age, retry counts, and the number of operations awaiting reconciliation. Alert on customer impact rather than every isolated failed attempt. Exercise a provider outage in a development or staging environment: confirm that queues stay bounded, the interface remains understandable, and recovery does not create a sudden wave of duplicate writes. A resilient integration behaves intentionally during the bad minutes, not just during a successful demo.