Uber published a post this week on retry storms that is worth your time if you operate anything with more than two services in a call chain. The interesting part is not the math, though the math is good. It is the admission that the industry-standard fixes, backoff, jitter, retry budgets, all manage the volume of retries without asking the only question that actually matters: whose fault is this error?

The amplification problem in one formula

When every hop in a dependency chain retries failed requests, the numbers multiply. Uber models it as R^d x N, where R is the per-hop retry count, d is the call depth and N is the baseline traffic. With one retry at every hop, a service sitting three levels deep serves eight times its normal request load the moment something below it errors. Five levels deep and it gets ugly fast. Google’s SRE book warns about the same effect: frontend, backend and client each retrying three times turns one user action into 64 attempts against the database.

The standard mitigation is a retry budget. Cap retries at, say, 10 percent of traffic per hop, and the formula becomes (1+B)^d, which caps that depth-3 node at about 1.33x baseline instead of 8x. Budgets work. They also share a blind spot: a budget cannot tell the difference between an error this service caused and an error this service is merely passing along from somewhere further down. So during an outage, every upstream service keeps dutifully spending its budget to re-ask a question that a dying service has no capacity to answer.

Error ownership

Uber’s fix is an idea they call error ownership, built into their shared retry middleware and a Service Dependency Analysis system. When a service fails, the error propagates up the chain with a claim attached. Each service inspects the failure: if the error originated in code it owns, it claims it; if it is just relaying, it leaves the claim unclaimed. The retry middleware only retries claimed errors. An unclaimed propagated error stops dead at the first hop that recognizes it did not cause the problem.

This requires infrastructure most teams don’t have: correlation between a service’s inbound failures and its outbound failures, plus claim context carried hop by hop. That is the honest caveat about the whole approach. It was not a config change, it was a mesh-wide mechanism.

What it bought them

The numbers come from a real incident. On November 18, 2025, a Core Entity service more than five levels deep in the call chain failed. Simple retry budgets would have pushed a 46 to 135 percent traffic increase onto the degraded service. With error ownership, Uber measured up to 200,000 spurious requests per affected caller stopped, and 9.5 million mesh-wide. Across their user-facing APIs, the maximum retry storm radius dropped from 25 to 3, and the average from 20 to 2.

There is a subtlety here worth stealing even without the mesh. The team found that with a 10 percent retry budget, retries only improve perceived availability while the callee’s base availability stays near 90 percent or above. At 99 percent base availability, retries get you to 99.99. At 80 percent, you land at 88, because the budget starves most of the retries that would have helped. Below some floor, retrying is a waste of a budget you should be saving. Most retry implementations never check.

One risk the post addresses: coincidental errors. An internal failure that happens to coincide with a fail-open dependency failure can get misattributed, suppressing retries that would have fixed the real problem. Uber’s analysis blames downstreams first and remembers failure patterns, keeping the worst-case misclassification near 2 percent of high-failure edges. Good enough, apparently, but it is the kind of number you want to know about before copying the design.

What you can do without Uber’s mesh

The primitives everyone should already have are well documented. The AWS SDK retry behavior uses exponential backoff with full jitter, a 50 ms base delay for transient errors and 1,000 ms for throttling, a 20-second cap, and a token-bucket quota that makes clients fail fast once sustained failures pass roughly 22 percent of requests. Google SRE guidance adds per-request retry caps and server-wide budgets, around 60 retries per minute per process in their example. Google’s Cloud team has also documented what happens without jitter: fixed-interval retries during a 15-minute outage lock all client load together and demand 15x to 20x normal capacity at recovery.

On top of those, Uber’s pattern can be approximated with nothing fancier than a response header. Propagate an error-claim marker on error responses. Each service claims errors it caused and leaves pass-through errors unclaimed. Retry middleware retries only claimed errors, while still guaranteeing at least one retry somewhere along the chain so availability does not silently drop. Pair the whole thing with circuit breakers and per-call timeouts, and test the behavior under injected load in staging, because the first real test should not be an outage.

The one-line takeaway from the post: a retry is a bet that the second attempt will land somewhere healthier than the first. Spending that bet on an error you didn’t cause is how one failed service takes its neighbors down with it.

Leave a Reply

Your email address will not be published. Required fields are marked *