Retries can turn one slow dependency into an API outage
I was recently debugging a new relic slow distributed trace across 3 different services in my stack.
Client → API → billing service → tax service.
Each of the three callers allows one initial attempt and two retries. If failures propagate through the entire chain and every caller exhausts its attempts, the tax service receives 3 × 3 × 3 = 27 calls for one user request.
One user request can trigger 27 calls to a struggling dependency when retries multiply across three layers.
Each retry policy looks reasonable in isolation. Together, they multiply the work reaching a struggling dependency.
The CAUSE of the failure matters here.
If a request fails because of a brief network interruption, another attempt may succeed. If it fails because the dependency is overloaded, another attempt adds work while that dependency is already falling behind.
As responses slow down, requests hold connections and other resources for longer. Shared capacity like thread workers or db connection pools can fill up, allowing the problem to spread to other API operations.
We can analyse this by separating original requests from retry attempts. Stable user traffic alongside rising downstream attempts is a reason to inspect the retry chain.
Trace where retries happen, including inside client libraries. Then choose which layer owns retries for that operation and limit the additional attempts it can generate. Consider adding jitter to your retry attempts. This makes the total retry behaviour easier to reason about when a dependency slows down.
Also when reviewing a retry policy, follow one request through the full chain and count the downstream attempts it can produce. That is the load the dependency may face while trying to recover.