Retries, Timeouts, Backpressure & Load Shedding
The Slow Recommendation Service That Took Down Checkout
Test your architecture intuition: Pitch a 7-axis solution, survive two aggressive reviewer objections, and inspect the staff-level Teacher Gold Answer.
Part 0. Start here
The problem: an optional widget took the home page down
A streaming service's home page shows rows of titles, and each title shows its star rating. The ratings come from a separate ratings service, and they are optional: the page works fine without stars. One evening ratings slows from 30 ms to 5 seconds per call, and the home page stops loading for everyone. The numbers below come from step 2.3 of the Netflix case study: one API instance serves 3,333 home-page requests a second, each request calls ratings once, and all requests share one pool of 1,000 threads.
ratings healthy | ratings slow | |
|---|---|---|
| How long one call holds a thread | 0.03 s | 5 s |
| Threads needed (requests a second × seconds held) | 3,333 × 0.03 ≈ 100 | 3,333 × 5 = 16,665 |
| The pool | 1,000: one tenth in use | 1,000: full after 1,000 ÷ 3,333 ≈ 0.3 s |
Once the pool is full, every new request waits for a thread, including the ones that never call ratings. An optional star rating has taken the whole page down.
Two more facts make it worse:
- Retries multiply. If five layers each make 3 attempts, the bottom layer sees 3 × 3 × 3 × 3 × 3 = 243 calls for one user action (AWS Builders' Library).
- The outage can outlive its cause. In an example from Bronson and colleagues' paper on metastable failures (HotOS 2021), a database is fast only below 300 queries a second, and it normally receives 280. A 10-second network outage at the switch makes clients time out and retry. After the network is back, the retries hold the load at 560 a second, and goodput stays at zero until the load drops under 150 a second or the retries drop under 20 a second.
ratings is healthy again ten seconds later, but the home page is still down five minutes after. Why? Would raising the timeout, or doubling the API servers, bring it back? What would you change?
The big picture
Synthesizing vector architecture diagram...
What to notice: three layers each make 3 attempts. Product pages never call stock, but they run on the same 400 threads as the cart views and checkouts that do. The red database is the only thing that gets slow.
What you'll be able to do after this page
- Explain with Little's law why a slow dependency takes down callers that don't use it, and tell goodput from throughput (Part 1).
- Set a timeout from measured latency, and send each attempt's own deadline down the call chain (Part 2).
- Decide which errors to retry, where, and how many times, and count the retries you didn't configure (Part 3).
- Spread retries and periodic work with jitter, and say what jitter does not fix (Part 4).
- Cap retries with a budget, and hedge a slow replica without adding load in an overload (Part 5).
- Let a dependency refuse work early and cheaply, find its real queue, and put priority and an age limit there (Part 6).
- Configure a circuit breaker per operation, and predict how it behaves on a dependency that is slow rather than dead (Part 7).
- Size a bulkhead with Little's law, and know when it saves you and when it is only a backstop (Part 8).
- Degrade the optional part of a page instead of failing the whole page (Part 9).
- Diagnose a failure that outlives its trigger, and get out of it (Part 10).
- Trace one checkout through every layer, with the deadline and the guard at each hop (Part 11).
- Map every guard to AWS, count the retries AWS makes for you, and name the look-alikes (Part 12).
You may have arrived from a step that relies on this: step 2.3 of the Netflix case study (one slow service took down the home page), step 3.1 of the Shopify case study (a slow payment provider ties up all the workers), step 3.4 of the rate limiter loop (the backend is slow, but nobody is over their limit), step 2.6 of the job scheduler loop (a target is slow and we're making it worse) or step 2.5 of the chat loop (6.7 million phones reconnect at once). This page is the "why" behind all five, and behind the retry, timeout and shedding lines in 36 of the 37 interview loops.
Part 1. Why slow is worse than down
A dependency that is down fails fast: the caller gets an error in a millisecond and moves on. A dependency that is slow holds everything that waits for it. This Part names the one law behind that, the one measurement that tells a healthy system from a busy one, and the example the rest of the page follows.
Little's law
It holds for any stable system: a thread pool, a connection pool, a queue, a whole service. The hook used it twice: 3,333 a second × 0.03 s ≈ 100 threads in use, and 3,333 × 5 s = 16,665 threads needed. A call that gets 100 times slower needs 100 times the threads. When the pool runs out, every request waits for a thread, whatever it was going to do.
Two consequences run through the whole page:
- A thread (or a connection, or a buffer) is held for the whole wait, retries and backoff included.
- A longer timeout makes this worse, not better. It raises the time each request spends inside, so the same arrival rate needs more threads.
Goodput is not throughput
Throughput counts everything a system finishes. Goodput counts only the requests answered successfully within the caller's deadline. The difference is waste: a query that finishes after the user has closed the page, or a reply to an attempt the caller has already retried. During an overload a service can run at 100% throughput and 0% goodput. That is exactly what happens below.
The example we follow
A real company has hundreds of services; that is too many to watch. So one small story runs through the whole page: one slow dependency turns a retry storm into an outage. Every trace on this page was produced by running a private reference simulation of this setup (apps, four shop-api instances with their threads, stock with its workers, a database with its connections, and every guard as a setting), with two random seeds, not worked out by hand. The example's numbers are smaller than the hook's so that the whole fleet fits on one line.
Basketly sells groceries online. Its shop-api serves product pages, carts and checkouts. It calls prices (always healthy here) and stock. stock answers every request with one query to its own database. On Friday at 20:00:00 the database's storage has a 20-second hiccup: every query takes ten times longer. At 20:00:20 it is healthy again.
| Setting | Our example | At real scale |
|---|---|---|
| Traffic (random arrivals) | Product pages 800/s (call prices only); cart views 1,200/s (prices + a stock read; the page can hide its stock badges); checkouts 60/s (prices + a stock reserve, a write with an idempotency key). Total 2,060 actions/s | Netflix: 3,333 requests/s per API instance; Shopify: 850 payments/s at peak |
shop-api | 4 instances × 100 threads = 400. A request holds its thread for its whole life. prices takes 10 ms. stock calls arrive at 1,260 ÷ 4 = 315/s per instance | Netflix: 1,000 threads per instance; Shopify: 100 hosts × 30 workers |
stock | 4 instances, 128 workers in total; one database query per request | |
stock's database | 8 connections; a query takes 4 ms on average (random): capacity 8 ÷ 0.004 s = 2,000 queries/s. Normal load 1,260/s, which is 63% | RDS Proxy caps connections with MaxConnectionsPercent |
| The trigger | 20:00:00.000 to 20:00:20.000: queries average 40 ms, so capacity is 8 ÷ 0.040 s = 200/s | Bronson: a 10 s network outage at the switch |
| C0: as shipped | Apps: 2 s per attempt, 3 attempts, retry at once after a timeout or an error. shop-api → stock: 1 s timeout, 3 attempts, no backoff, retries a 503, one shared thread pool, an unbounded first-in-first-out (FIFO) queue, no deadlines. stock → database: 5 s timeout, 3 attempts | Common library defaults (Part 2) |
Each Part adds one guard to the list above and replays the same 30 seconds: C1 timeouts and per-attempt deadlines (Part 2), C2 one retry layer (Part 3), C3 backoff with jitter (Part 4), C4 a retry budget (Part 5), C5 stock protects itself (Part 6), C6 circuit breakers (Part 7), C7 a bulkhead (Part 8) and C8 a brownout (Part 9). Each setting keeps all the guards before it.
The story in eleven beats: the site falls over and stays down (Part 1); bounded waits (Part 2); one place to retry (Part 3); jitter (Part 4); a retry budget (Part 5); stock protects itself (Part 6); breakers (Part 7); bulkheads (Part 8); brownouts (Part 9); why it stayed down (Part 10); one checkout, end to end (Part 11). Each Part shows only its own slice of events; the full table is in Part 13.
Words we use
| Word | Meaning here |
|---|---|
| Action | One thing a user does: open a product page, view the cart, check out. An action may send several attempts |
| Attempt | One try of one call. A retry is a second (or third) attempt |
| Timeout | How long a caller waits for one attempt, a duration |
| Deadline | The moment after which nobody wants the answer. Sent between services as the time left, a duration |
| Goodput | Actions answered successfully within the caller's deadline, per second |
| Degraded | Answered, but without an optional part (a cart without stock badges) |
| Shed | Refused on purpose, early and cheaply, with 503 and Retry-After |
| Recovered | Goodput back to at least 90% of the 2,060 actions a second offered, and staying there for 5 s |
| For nobody | Work finished after the caller who asked for it stopped waiting |
The same 30 seconds, as shipped (C0)
Snapshot W1, 19:59:55: a normal Friday evening (seed 1; rates are per second)
| Tier | What it's doing | Its limit |
|---|---|---|
| Apps | 2,065 attempts | |
shop-api | 28 threads busy, nothing queued | 400 threads |
stock | 1,240 calls in; nobody waiting for a connection | 128 workers, 8 connections |
| Database | 1,240 queries finished, 0 for nobody | 2,000 a second |
| Goodput | products 822, cart views 1,184, checkouts 56 |
Every tier is far below its limit. Little's law predicts the threads: 800 × 0.010 s + 1,260 × (0.010 + about 0.0042) s ≈ 8 + 18 = 26 busy on average (the run, time-averaged over 19:59:46 to 19:59:54: about 26; 28 at this instant).
| # | Time | Event (seeds 1 and 2) |
|---|---|---|
| 1 | 19:59:50 | Baseline: about 2,060 actions/s; about 26 of 400 shop-api threads busy; the database finishes about 1,250 queries/s, all for callers still waiting |
| 2 | 20:00:00.000 | The hiccup: queries take 40 ms; capacity falls to 200/s |
| 3 | 20:00:00.38 to 20:00:00.40 | All 400 threads busy. Cart views and checkouts arrive at 1,260/s and most now hold a thread for at least 1 s; only about 200/s can finish. The pool fills in about 400 ÷ (1,260 − 200) ≈ 0.38 s. Product pages, which never call stock, queue behind them |
| 4 | 20:00:02 onward | App attempts time out at 2 s and retry at once: about 5,800 to 5,900 attempts/s on average over 20:00:02 to 20:00:19, and 6,180 once every action is retrying, for 2,060 actions/s |
| 5 | 20:00:02 | Goodput 0 in all three classes. The database still finishes about 200 queries/s, and all of them are for nobody |
| 6 | 20:00:20.000 | The database is healthy again |
| 7 | 20:00:20 → 20:05:00 | Goodput still 0. The database now finishes about 2,000 queries/s, all for nobody. shop-api's queue holds 104,000 to 106,000 requests at 20:00:20 and grows by about 2,900 a second: 403,000 at 20:02 and 923,000 at 20:05 |
Snapshot W2, 20:00:01: one second in (seed 1)
| Tier | What it's doing |
|---|---|
| Apps | 2,094 attempts (no retries yet: the first timeouts come at 20:00:02) |
shop-api | 400 of 400 threads busy; 1,192 requests queued behind them, product pages among them |
stock | 400 calls in; all 128 workers busy, 120 waiting for one of the 8 connections |
| Database | 183 queries finished, 163 of them for nobody already |
| Goodput | products 13, cart views 19, checkouts 1 |
Why it can't drain
Why doesn't the site recover at 20:00:20, when the database is fast again? One inequality says it. The apps now send about 3 × 2,060 ≈ 6,180 attempts a second (each action times out twice and tries three times). How many can shop-api finish? It is limited by the database: 2,000 queries a second, and 1,260 of every 2,060 requests (61%) need one. So it can finish at most 2,000 ÷ (1,260 ÷ 2,060) ≈ 3,270 requests a second. 6,180 > 3,270: work arrives faster than it can be finished, so the queue grows by the difference whatever caused it. And every request that reaches the front has waited far longer than 2 s, so its app has already given up: the database runs at 100% on work nobody wants.
The database is healthy at 20:00:20. Why is goodput still zero at 20:00:25?
Snapshot W3, 20:00:25: the trigger is gone, the outage isn't
Synthesizing vector architecture diagram...
What to notice: the line that starts near 2,060 is goodput. It falls to zero within two seconds of the hiccup and stays there after the database recovers at second 20. The line that starts at zero is database work finished for nobody: about 200 a second during the hiccup, then about 2,000 a second, the database's full capacity, once it is healthy. The database is busy and useless. (Seed 1; the axis is not evenly spaced after second 5.)
What to remember from Part 1
- In flight = rate × time: a call that slows 100× needs 100× the threads, and when they run out everything waits.
- Throughput isn't goodput: work finished after the caller has left is waste.
- If callers send more than can be finished, the backlog never drains, whatever caused it.
Part 2. Timeouts and deadlines
Every call in C0 already had a timeout. They were simply wrong: too long, and set without looking at each other.
| Hop | Timeout per attempt | Attempts | Longest the caller can be kept waiting |
|---|---|---|---|
App → shop-api | 2 s | 3 | 6 s |
shop-api → stock | 1 s | 3 | 3 s |
stock → database | 5 s | 3 | 15 s |
These timeouts are inverted: each callee is allowed to work longer than its caller will wait. shop-api may spend 3 s on an app attempt that is abandoned after 2 s, and stock may spend 15 s on a call that shop-api drops after 1 s. In this trace the database layer's 5 s never fires (a query waits about 0.6 s for a connection: 120 waiting ÷ 8 connections × 40 ms; under 0.8 s at the longest in the run), so the damage comes from the other two hops: stock finishing queries for calls shop-api has dropped, and shop-api serving a queue of requests whose apps have left.
A timeout from percentiles
A timeout should come from the dependency's measured latency, not from a guess. The method below is the AWS Builders' Library's ("Timeouts, retries, and backoff with jitter"):
textCHOOSE A TIMEOUT (per dependency; re-measure after every deploy) pick the false-timeout rate you accept -> 0.1% of healthy calls read that percentile of healthy latency -> P99.9 pad it when P99.9 is close to P50, and for network distance -> about 2 x P99.9 put connection setup and TLS inside the timer, or connect before the process takes traffic keep a separate, short connect timeout
Our simulation measured healthy stock calls, as shop-api sees them, over 60 s:
| Percentile | P50 | P95 | P99 | P99.9 |
|---|---|---|---|---|
Healthy stock call (seeds 1 and 2) | 3.0 ms | 12.2 ms | 18.7 ms | 28 ms (27.8 and 28.3) |
So C1 sets the shop-api → stock timeout to 60 ms, about 2 × P99.9: at most 0.1% of healthy calls could time out, and padding keeps a small latency rise from turning into a wave of timeouts.
Put connection setup inside the timer. The Builders' Library tells of a dependency called with a timeout of around 20 ms that timed out only right after deployments. The timer included setting up a new secure connection, which took longer than 20 ms, so new servers timed out on their first calls. The lasting fix was to open the connections when the process started, before it took traffic.
Connection timeout vs request timeout. A connection timeout bounds the TCP and TLS handshake: a host that doesn't answer at all should be given up on quickly. A request timeout bounds the wait for the answer and is the one set from percentiles.
Timeout vs deadline
- A timeout is a duration for one attempt: "wait 60 ms for this call".
- A deadline is a point in time for the whole request: "nobody wants this answer after 20:00:07.032".
Between machines the deadline travels as the time left, a duration, because clocks on different machines disagree. gRPC does exactly this: it turns the deadline into a timeout and subtracts the time already spent at each hop. (Timeouts that protect safety rather than capacity, such as leases, are a different job: see Leases, Fencing Tokens & Distributed Locks, Part 3 there.)
Each attempt's own deadline, cancelled when abandoned
C1 adds one rule at every hop: send the deadline of the attempt you are making, which is min(this attempt's timeout, the request's time left), minus a small margin. stock therefore hears "60 ms", never the whole 2 s. When the caller gives up on an attempt, it cancels it (an HTTP/2 stream reset or a gRPC cancellation), so the callee can stop. And every queue checks the deadline when it takes work off the queue.
Synthesizing vector architecture diagram...
What to notice: stock hears 60 ms, the deadline of the attempt shop-api is making, not the app's 2 s. When shop-api gives up, it cancels, and stock frees the connection instead of finishing a query nobody will read.
textSHOP-API CALLING STOCK (C1 onward) left = request.deadline - now if left <= 0: stop, the app has already given up budget = min(60 ms, left) - margin send the call with "deadline: budget" as a duration, not a clock time wait up to budget; if no reply: cancel the attempt, count a timeout STOCK, AT EVERY QUEUE on arrival: expires_at = my clock now + the deadline received at dequeue: if now >= expires_at: drop it, never start the work run the query with statement timeout = expires_at - now on cancel from the caller: cancel the query, free the connection
Check the sum at every hop: the attempts, their timeouts and the backoff must fit inside the deadline above. For shop-api from C3 on: 2 attempts × 60 ms plus at most 25 ms of backoff is about 145 ms, well inside the app's 2.0 s.
The same 30 seconds, with bounded waits (C1)
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 8 | 19:59:50 | C1 | Healthy stock calls: P50 3.0 ms, P99 18.7 ms, P99.9 about 28 ms, so the timeout is 60 ms |
| 9 | 20:00:00 → 20:00:20 | C1 | Threads no longer fill with 1 s waits: product pages about 723 of 800 a second get through. But the retry storm is still there: stock receives about 6,850 calls a second for the 1,260 it needs, 5.4×, because the apps and shop-api still retry at once, three times each. Cart views about 190 of 1,200, checkouts about 9 to 11 of 60 |
| 10 | 20:00:20 | C1 | Recovered at once: goodput is about 2,830 to 2,850 in the second starting 20:00:20 (new arrivals plus queued requests that still had time). shop-api's queue, about 8,200 at 20:00:20, is empty by 20:00:27 to 20:00:28. The database finishes 0 queries for nobody: stale work is dropped at dequeue, not served |
Side row 10z: passing the whole deadline down, with the cancel lost. Same C1, but each hop sends the request's whole time left (up to 2 s) rather than the attempt's 60 ms, and the cancel never reaches stock (a proxy or client library that drops it). With the cancel working, sending the whole time left changes nothing here: the run recovers at 20:00:20, exactly like C1.
| Setting | Work for nobody during the hiccup | Goodput after 20:00:20 |
|---|---|---|
| C1, the attempt's deadline sent (cancel working or lost) | 0 queries a second | Recovered at 20:00:20; queue empty at 20:00:27 to 20:00:28 |
| C1, the whole time left sent, cancel working | 0 queries a second | Recovered at 20:00:20, like C1 |
| C1, the whole time left sent, cancel lost | About 200 queries a second (all the database can do) | Never recovers in the 40 s after the hiccup; shop-api's queue about 8,600 at 20:00:20 and growing |
Nobody tells stock that shop-api gave up, and it believes it has up to 2 s, so it keeps working on attempts abandoned after 60 ms, and the stale-work loop of C0 comes back. The attempt's own deadline is the backstop for a lost cancel: it stops the work when the caller stops waiting, whether or not the cancel arrives. Envoy gets this right: its x-envoy-expected-rq-timeout-ms header carries the per-try timeout when one is set (the route timeout otherwise, or when hedging on per-try timeouts), never more than the route's time left.
Defaults that are too long
| Default | Value | What to do |
|---|---|---|
| gRPC call deadline | None: a call can wait forever | Always set one |
AWS Step Functions Task TimeoutSeconds | 99,999,999 s (over three years) if not set | Set it, and HeartbeatSeconds for long tasks |
Amazon RDS Proxy ConnectionBorrowTimeout | 120 s (allowed 0 to 300 s) | Set it below your callers' deadline |
| Amazon CloudFront origin | Response timeout 30 s; connection attempts 3 (timeout 10 s each); on a response timeout it tries a GET or HEAD again, up to the connection attempts | Lower both, or set a response completion timeout that caps the whole wait |
| Amazon API Gateway REST integration | 50 ms to 29 s (29 s by default); HTTP APIs 30 s | Behind it, a Lambda function may run up to 900 s. Set the function's timeout below the integration's, or it keeps working after the client got a 504: the AWS version of an inverted timeout |
Why not 50 ms, and why not longer
This is the second question of the drill The Slow Recommendation Service That Took Down Checkout, on our numbers:
- Too short. With a 5 ms timeout on a healthy day, every call slower than 5 ms fails: about 31% of healthy
stockcalls (seeds 1 and 2: 30.7% and 30.5%), because P50 is 3.0 ms and P95 is 12.2 ms. Each false timeout becomes a retry, so a healthy dependency gets about a third more load, and some users fail for no reason. - Too long. With a 5 s timeout, a hanging
stockneeds 1,260 calls a second × 5 s = 6,300 threads, andshop-apihas 400.
A timeout is set from the healthy tail. "The dependency is plainly down" is a different signal, handled by a breaker or a concurrency limit (Parts 7 and 8).
Every call already had a timeout in C0, and C1 also passes a deadline down. Why does passing the whole 2 s down, when the cancel gets lost (side row 10z), still leave the site down?
What to remember from Part 2
- Set each timeout from the dependency's measured latency (the percentile of the false-timeout rate you accept, padded), never a guess.
- Send each attempt's own deadline down as a duration, and cancel the attempt when you give up.
- Every queue drops work whose deadline has passed.
Part 3. Retries: what, where and how many
In C0, one tap on "view cart" can become many database queries. The app tries 3 times; for each app attempt, shop-api tries stock 3 times; for each of those, stock tries its database 3 times.
Synthesizing vector architecture diagram...
What to notice: the attempts multiply, 3, then 9, then 27. This is the worst case: every attempt at every layer fails. In the C0 trace the database layer never times out, so the real product is smaller; the multiplication is still why the apps send about 6,180 attempts a second for 2,060 actions.
Which errors to retry
| Class | Examples | Retry? |
|---|---|---|
| Transient | A timeout, a connection reset, 500, 502, 503, 504 | Yes, with backoff, if the call is safe to repeat |
| Throttle | 429; throttling errors that come as 4xx, such as DynamoDB's ProvisionedThroughputExceededException and ThrottlingException (both HTTP 400) | Yes, after a longer wait; honour Retry-After |
| Permanent | 400, 401, 403, 404, 422, a schema or validation error | Never: the same request fails the same way |
| "Overloaded, don't retry" | A 503 from a service that is shedding on purpose, or an explicit flag | Not in this layer; pass it up |
The AWS SDKs use the same three families (transient, throttling, non-retryable), and a throttle is not always a 429: DynamoDB's throttles are 400s that are "OK to retry".
Side effects: the same key on every retry
A checkout's reserve is a write. If its attempt times out, the reserve may or may not have happened: a timeout is an unknown, not a failure. Retry a side effect only with the same idempotency key, so the callee can return the first result instead of acting twice, and never fail over to a different provider on a timeout. Idempotency & Effectively-Once Processing owns the rules (Part 1 there for unknowns, Part 4 for how long to remember a key and the late original).
Event 11y. In C0, during the hiccup, about 3.8 checkouts a second get their reserve committed more than once by the simulator's database (about 7.5 extra commits a second): shop-api gave up on its attempt after 1 s and sent another, but stock was never told and ran both. (The app's own retries land behind the queue and commit a second time after 20:00:20.) With the idempotency key, stock returns the first result to the second. From C1 on, the run shows none, because an abandoned attempt is cancelled before it commits. The simulator's network adds no delay; in a real one a reply can still be lost after the commit, so the key stays.
When the server says when
| Hint | Where | Meaning |
|---|---|---|
Retry-After: 30 or Retry-After: <HTTP date> | HTTP, with 503 or 429 (RFC 9110; RFC 6585 says a 429 may include it) | Wait at least this long |
x-amz-retry-after | AWS service responses, in milliseconds | The SDKs use it, clamped to [computed backoff, backoff + 5,000 ms] |
grpc-retry-pushback-ms | gRPC response metadata | Wait this many ms; a negative or unreadable value means "don't retry" |
Where to retry: one layer
Attempts multiply across layers: three layers of 3 attempts is 27 at the bottom, and the Builders' Library's example of five layers is 3⁵ = 243. The rule, from the Google SRE book: retry only in the layer immediately above the one that is failing, and let every other layer pass the failure up. When a service is overloaded, it says so with an "overloaded, don't retry" error, and its callers don't retry it.
textSHOP-API CALLING STOCK: the only layer that retries (C2) for attempt in 1, 2: reply = call stock with this attempt's deadline if reply is success: return it if reply is 503 "overloaded" or a permanent error: stop for a reserve: keep the same idempotency key return 503 + Retry-After to the app, marked "don't retry" STOCK CALLING ITS DATABASE: one attempt, no retry APP: retry once, only if it got no answer at all; on a 503 wait for Retry-After
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 11 | 20:00:00 → 20:00:20 | C2 | The worst case per app attempt falls from 3 × 3 × 3 = 27 queries to 2 (the app's one retry needs a lost answer, which doesn't happen here). Measured: about 2,410 stock calls a second, 1.9× the 1,260 needed, down from 5.4×. Product pages about 797 of 800, cart views about 190, checkouts about 9. Nothing queues in shop-api; recovered at 20:00:20 |
| 11x | side row | C0 copy | A sidecar proxy that also retries failed connections twice (3 tries) adds a fourth layer nobody configured: 3 × 3 × 3 × 3 = 81 per tap |
Hidden retry layers
The layers you configure are not the only ones. Count these too:
| Layer | What it does by default |
|---|---|
| Amazon CloudFront | Tries a GET or HEAD again after an origin response timeout, up to its connection attempts (3 by default); never POST, PUT, PATCH, DELETE or OPTIONS |
| Amazon ECS Service Connect proxy | Retries a failed connection twice; the second attempt avoids the host of the previous one |
| AWS SDKs, standard mode | 3 attempts in total (DynamoDB clients 4 in the updated behaviour) |
| AWS SDKs, legacy mode | Varies by SDK, with no standard retry quota. Java, Python, Ruby, PHP, C++ and AWS CLI version 1 (version 2 already defaults to standard) default to legacy until November 2026; Python's legacy mode makes 5 attempts (10 for DynamoDB) |
| AWS Lambda, asynchronous invocations | 2 retries after a function error (after 1 minute, then 2); throttles and system errors retried for up to 6 hours |
AWS Step Functions Retry | 3 retries (MaxAttempts) when a Retry rule matches |
| gRPC with a retry policy | Up to maxAttempts (values above 5 count as 5) |
Your app, your API and your database driver each try 3 times. How many queries can one tap send to a dead database?
What to remember from Part 3
- Retry only what can succeed later and is safe to repeat; a side effect retries with the same key.
- Attempts multiply across layers: retry in one place and pass "don't retry" up.
- Honour
Retry-After, and count the retries your proxies, SDKs and managed services make.
Part 4. Backoff and jitter
At 20:00:20.000 the database recovers, and in one simulated instant 1,200 cart views fail at once. If each retries after a fixed 1 s, all 1,200 retries land together at 20:00:21.000: a second, self-made spike.
Synthesizing vector architecture diagram...
What to notice: with a fixed 1 s wait (the tall bar), all 1,200 retries fall in the one 10 ms slot starting at 1.0 s. With full jitter over 0 to 1 s (the short bars), each 10 ms slot gets about 12 (seed 1: never fewer than 5 or more than 25). One slot in ten is drawn.
Backoff, then jitter
Backoff makes each retry wait longer than the last, so a failing dependency gets fewer retries per second. Capped exponential backoff doubles the wait each time up to a cap. But clients that failed together still retry together. Jitter adds randomness so they don't. The AWS SDKs, and C3, use full jitter:
C3 uses base 25 ms in shop-api (so the first retry waits 0 to 25 ms, inside the 2 s request) and base 1 s in the app.
| Kind | Wait before retry n | AWS's simulation (100 clients contending) |
|---|---|---|
| None | min(cap, base × 2ⁿ) | "The clear loser": most work and most time |
| Full | random(0, min(cap, base × 2ⁿ)) | Least work; slightly more time than decorrelated |
| Equal | half of min(cap, base × 2ⁿ), plus random(0, the other half) | "The loser" among the jittered ones: slightly more work than full, much longer |
| Decorrelated | min(cap, random(base, 3 × the previous wait)) | Slightly less time than full, more work |
What jitter does and doesn't do
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 12 | 20:00:00 → 20:00:20 | C3 | About 2,440 stock calls a second, the same as C2's 2,410. Jitter spread the retries in time; it didn't remove any. On C0's settings, jitter alone never recovers (the run is still at zero goodput 40 s after the hiccup) |
Both are true, and they don't contradict each other. Jitter doesn't change how many retries failing clients may send: that is set by the attempt cap. When the dependency fails every call (as here, for 20 s), every allowed retry is sent, jittered or not. Under contention (clients competing for something that succeeds for some of them), jitter lets attempts succeed sooner, so fewer retries are needed: in the AWS Architecture Blog's simulation ("Exponential Backoff And Jitter", 2015) with 100 contending clients, jitter cut the number of calls by more than half.
Periodic work: a stable offset, not a fresh random wait
Replay P runs on a copy: 200 warehouse agents each re-read 50 stock counts every 60 s, 10,000 reads a minute, against a database with 740 reads a second to spare.
Synthesizing vector architecture diagram...
What to notice: the line that sits at 740 for the first 13.5 s of every minute, then drops to zero, is every agent starting at :00: 10,000 reads queue for 10,000 ÷ 740 ≈ 13.5 s. The line that never goes above 350 is stable offsets: each agent starts at hash(its ID) mod 60 s, the same second every minute, averaging 167 reads a second (seed 1; offsets are hashed, so some seconds get a few more agents).
| Schedule | Database load | One agent's gap between syncs |
|---|---|---|
| All at :00 | 10,000 reads in the first second; 13.5 s of queueing every minute | 60 s |
| A fresh random offset each minute | Spread out (busiest second 450 to 550 reads, seeds 1 and 2) | Anywhere from 0 to 120 s: in the run, 2.9 to 118.8 s; 12.5% of gaps are longer than 90 s |
| A stable offset: hash(agent ID) mod 60 s | 167 reads a second on average; busiest second 350 | Always 60 s |
The sync agents already add a random wait each minute. Why do some stores' stock counts go two minutes without a refresh?
Check your library's default
| Library or service | Backoff by default |
|---|---|
| AWS SDKs, standard mode | Full jitter, capped at 20 s; in the updated behaviour, base 50 ms for transient errors and 1,000 ms for throttles (DynamoDB 25 ms) |
| Envoy | Fully jittered exponential backoff from 25 ms, capped at 250 ms (10 × the base); when configured to follow reset headers such as Retry-After (rate_limited_retry_back_off), it waits random(interval, 1.5 × interval) |
| gRPC (retry design A6) | ±20% jitter around each exponential step, in the Go, Java and C-core implementations (the original design used full jitter) |
AWS Step Functions Retry | JitterStrategy is NONE unless you set FULL; IntervalSeconds 1, BackoffRate 2.0 |
resilience4j Retry | 3 attempts with a fixed 500 ms wait: no backoff, no jitter |
What to remember from Part 4
- Backoff lowers the rate; jitter breaks the lockstep; neither lowers how many retries failing clients may send.
- Give periodic work a stable offset per client, not a fresh random wait each period.
- Check your library's default: several don't jitter.
Part 5. Retry budgets and hedged requests
C3 still sends 1.9× the calls stock needs, because in a 20-second failure nearly every request uses its second attempt. A per-request cap ("at most 2 attempts") can't tell a blip, where retrying is cheap and helpful, from a real failure, where every retry is extra load. A budget can.
A cap per request vs a budget per client
- Per request: "at most N attempts". When the dependency fails everything, the load is N×. The SRE book's per-request budget of 3 attempts gives "a threefold increase" in the worst case.
- Per client: "retries at most 10% of my requests". When the dependency fails everything, the load is about 1.1×, and a blip that fails a few requests is still retried. The SRE book reports exactly this: a 10% per-client retry ratio reduces the growth "to just 1.1x".
textRETRY BUDGET, ONE PER SHOP-API INSTANCE (C4) tokens = 10 the bucket's size, the burst allowance on every first attempt: tokens = min(10, tokens + 0.1) before any retry: if tokens < 1: don't retry, fail now else tokens = tokens - 1 so retries stay near 10% of first attempts, plus a burst of 10
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 13 | 20:00:00 → 20:00:20 | C3 → C4 | stock calls fall from about 2,440 to about 1,380 a second (1.1×); the budget refuses about 950 retries a second |
Real budgets
| Where | The budget |
|---|---|
| AWS SDKs, standard mode | A retry quota: a token bucket of 500 tokens per client instance, "not shared across processes or hosts". In the updated behaviour (opt in now with AWS_NEW_RETRIES_2026=true, the default from November 2026) a transient retry costs 14 tokens (before: 5, and 10 for a timeout) and a throttling retry 5; a first-try success refunds 1. It drains once more than about 22% (transient) or 32% (throttling) of requests keep failing. Adaptive mode adds a client-side rate limiter that can delay even first attempts, recommended only for a client that talks to one resource |
| Envoy | A retry budget: budget_percent 20% of active requests, with min_retry_concurrency 3 |
| gRPC retry throttling | Per server name: maxTokens and tokenRatio; a failure costs 1, a success earns tokenRatio; no retries (or hedges) while tokens are at or below maxTokens ÷ 2 |
| Google SRE book | At most 3 attempts per request, and retries under 10% of requests per client |
Budgets live in each client
Every one of these budgets is local: per process, per client instance or per proxy. The fleet's total is the sum.
Side row 13a. On a copy of C4 with 8 shop-api instances instead of 4, the run shows about the same 1,380 calls a second and 950 refusals. Our bucket earns per request, so the fleet's retries stay near 10% of traffic, however many instances share it. What does grow with the fleet is everything fixed per client: the burst (8 × 10 tokens instead of 4 × 10), the SDK's 500 tokens per client (about 35 transient retries each at 14 tokens: a thousand clients hold about 35,000 retries of burst), and any "N retries allowed" floor per proxy. Size those from the dependency's capacity ÷ the number of callers, or let the dependency protect itself (Part 6).
Hedging the tail
A hedged request is a retry sent early, for a different problem: one slow replica among healthy ones.
Synthesizing vector architecture diagram...
What to notice: the second call goes out only when the first is already slower than 95% of calls, to a different instance, for a read. The first answer wins and the loser is cancelled.
Replay H (on a copy, with a healthy database): each stock call independently stalls 250 ms with probability 1%, as if one instance were slow.
| P95 | P99 | P99.9 | Extra calls | |
|---|---|---|---|---|
| No hedging | 13.1 ms | 250 ms | 259 ms | 0 |
| Hedge after the P95 (13.1 ms) | 17 ms | 24 ms | 5% | |
| 60 ms timeout and one retry, no hedging | 30 to 50 ms | 70 ms | 1% |
The first row has no timeout. With C1's 60 ms timeout and one retry instead, every stalled call becomes a timeout and a retry: P99.9 about 70 ms (P99 sits right on the 1% of stalled calls, so it swings between 30 and 50 ms by seed). Hedging after the P95 limits the extra load to about 5%, the same number as The Tail at Scale (Dean and Barroso), whose Google benchmark cut the 99.9th-percentile latency of reading 1,000 keys from 1,800 ms to 74 ms with just 2% more requests after a 10 ms hedge delay. (Their "tied requests" go further: both copies are sent, and each can cancel the other once it starts.)
The rules: hedge only idempotent reads, only after the P95, inside the same retry budget, and cancel the loser. During the hiccup the slowness is in the shared database, so every instance is slow and a hedge only adds load; with the budget empty, no hedges go out. A slow replica in a quorum read is the same case: Replication, Quorums & Read-Your-Writes.
C4 retries at 1.1×. Why is the database still far beyond its 200 a second?
What to remember from Part 5
- Cap retries as a share of traffic, so a failing dependency sees about 1.1×, not 2× or 3×.
- Budgets live in each client: fixed per-client allowances grow with the fleet.
- A hedge is a retry sent early: read-only, after the P95, inside the budget, loser cancelled.
Part 6. Backpressure: stockprotects itself
C4 still sends about 1,380 calls a second to a database that can do 200. Every caller setting so far has shaved repeats, but the 1,260 calls that are really needed are still 6× what the database can serve. Only stock knows how busy it is, so stock must be able to say no.
What backpressure is
A queue turns overload into delay: work waits instead of failing, and the backlog outlives the trigger (Part 1). Backpressure is the opposite reflex: the slow side's signal makes the fast side slow down or stop, instead of buffering without limit. The loops use it at every scale: the news feed's fanout workers slow down when cache latency rises (step 2.3 of the news feed loop), and the matching engine rejects a new order rather than queue a stale one (step 2.7 of the stock exchange loop). For a request/response service, the signal is a fast refusal: 503 with Retry-After.
Refuse early and cheaply
A refusal must cost far less than a success, or refusing becomes the new overload. So stock counts the requests inside it and refuses at the door, before any real work: before TLS and authentication where it can (the chat loop admits connections at TCP accept, before TLS: step 2.5 of the chat loop), and always before the database. C5's limit is 64 requests inside stock, and 80 for reserves, so checkouts still get in when views are refused.
Find the real queue
Where do requests actually wait? stock has 128 workers, and at most 80 requests are ever admitted, so a worker is always free: the worker queue never forms. The only real queue is the wait for one of the 8 database connections. So the other two tools act there:
- Bound its age (the CoDel variant in Maurer's "Fail at Scale", ACM Queue 2015): once the queue has not been empty for 100 ms (it is standing), drop every request that has waited more than 5 ms. A short burst still queues; a standing queue can't grow old.
- Serve newest first while it stands (adaptive LIFO, from the same article): the newest request is the one whose caller is most likely still waiting.
- Priority where requests wait: a reserve waiting for a connection goes before any view. Priority only at the admission limit helps little: it decides who gets in, not who gets a connection, and once inside, views and reserves would wait in the same line. (Google's SRE book gives requests one of four criticalities, carried with the request to every backend, for the same purpose.)
How many waits are even useful? Little's law again: during the hiccup the database finishes 200 queries a second, and a caller waits at most 60 ms, so only about 200 × 0.06 = 12 requests can usefully wait. A fixed limit of 64 is far above that; it helps here only because priority and the age limit act in the connection queue. An adaptive limit (Part 8) would find a number near 12 on its own.
Synthesizing vector architecture diagram...
What to notice: the refusals come from two places, the admission check at the door and the age limit in the connection queue, and both go back to shop-api as a fast 503: that arrow is the backpressure. Priority and the age limit sit in the connection queue, the only place where requests actually wait.
textSTOCK (C5) on a call: limit = 80 for a reserve, else 64 if calls inside stock >= limit: reply 503 + Retry-After before any real work take a worker and join the connection wait queue when a connection frees: if the queue has been non-empty for 100 ms: reply 503 to every waiting call older than 5 ms pick from the newest end last in, first out else: pick from the oldest end first in, first out a waiting reserve is picked before any view if its deadline has passed: drop it else run the query with statement timeout = its time left
The same 30 seconds, with stockprotecting itself (C5)
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 14 | 20:00:00 → 20:00:20 | C5 | stock refuses about 165 calls a second at admission (162 to 167) and about 705 more (701 to 711) that waited over 5 ms in the standing connection queue, all with 503 + Retry-After, which shop-api doesn't retry. Reserves go first: checkouts rise from about 9 to about 44 a second (43.7 and 43.9); cart views about 157 in full; product pages about 793. Recovered at 20:00:20 |
| 14a | 20:00:00 → 20:00:01 | C4 vs C5 | Queries started in the first second: 1,300 to 1,400 (C4) vs 550 to 585 (C5). Queries finished for a caller still waiting: about 190 in both (the database can't do more than about 200). In C4 most started queries had already waited most of their 60 ms and were cancelled soon after starting. Checkouts served in that second: 15 to 16 (C4) vs 43 to 55 (C5) |
Side row 17x: stock protecting itself from as-shipped callers. The same stock guards on C0, whose apps and shop-api retry everything at once, including 503s:
| C0 | C0 + stock's limit, priority and age limit | |
|---|---|---|
Calls reaching stock | About 400 a second | About 8,400 to 10,300 a second (at most 1,260 × 9 = 11,340) |
| Refused at admission / by the age limit | 0 / 0 | About 7,300 to 9,500 / about 550 to 650 a second |
| Checkouts | 0 | 56 to 65 a second |
| Cart views in full / product pages | 0 / 0 | About 140 to 520 (it depends on how fast the callers retry) / about 800 |
| After 20:00:20 | Never recovers | Recovered at 20:00:20 |
Server-side protection works even against badly behaved callers, for one reason: each refusal costs microseconds, so a checkout refused at the door soon gets through on one of its many retries. If a refusal cost as much as a query (TLS, authentication, a database read), those 7,300 refusals a second would be the new overload. That is why the refusal must be early.
Will it ever drain?
A backlog empties only at what you can finish minus what keeps arriving. With the retries of C0, arrivals (6,180 a second) exceed what can be finished (3,270), so it never empties (Part 1).
Side row 17y. A copy of C0 with no retries anywhere: arrivals are 2,060 a second, below 3,270. At 20:00:20 shop-api holds about 28,000 queued requests, and stock about 4,100 stale calls of its own. stock works those off first (by 20:00:25); then shop-api's queue falls from 28,700 at 20:00:25 to 3,900 at 20:00:45, about 1,240 a second, close to the 3,270 − 2,060 ≈ 1,210 the arithmetic predicts. It is empty at 20:00:48 to 20:00:49. The Builders' Library ("Avoiding insurmountable queue backlogs") puts it in one line: an hour-long outage needs double the capacity for another hour to catch up.
Buffers you pull from (an SQS queue, a stream) are backpressure by design: the consumer takes work at its own pace, and the producer never waits. The danger moves to the backlog's age, so alarm on the age of the oldest message, not only on how many there are. Consumer lag on a change stream is the same backlog (Change Streams & the Transactional Outbox). Visibility timeouts, redelivery and dead-letter queues belong to a later loop primitive, Queues & Delivery Semantics.
Queues in load balancers. The Classic Load Balancer queued excess requests in a surge queue of up to 1,024; a request that waited there reached the server with no idea how old it was. The Builders' Library ("Using load shedding to avoid overload") recommends failing fast instead of queueing, and says the Application Load Balancer rejects excess traffic.
503, 429 or backpressure?
503 + Retry-After (shedding) | 429 + Retry-After (rate limiting) | Backpressure in a pipeline | |
|---|---|---|---|
| What it protects | The service, from itself: it is overloaded, whoever is sending | The service, from one client: that client sent too much | The consumer, from a faster producer |
| Decided by | The service's health (in flight, queue age, latency) | The client's usage against its quota | The consumer's pace (pull, credits, a full buffer) |
| The caller should | Not retry in this layer; wait at least Retry-After | Slow down that client; wait at least Retry-After | Nothing: it is slowed down or pulled from |
| Owned by | This page | A later loop primitive, Rate Limiting Algorithms; per-key limits are in Sharding, Hot Keys & Rebalancing (Parts 3 and 4 there) | This page (buffers) and Queues & Delivery Semantics |
Lowering every tenant's rate limit is not overload protection: when the backend is slow and nobody is over their limit (step 3.4 of the rate limiter loop), only shedding by the service's own health helps.
C5 lets most views in, and the database still finishes only about 200 queries a second. Why do checkouts go from 9 to 44 a second?
What to remember from Part 6
- The dependency must be able to refuse work, early and cheaply, with
503andRetry-After. - Find where requests actually wait, and put priority and an age limit there.
- A buffer only helps if the consumer catches up: work out the drain time, and watch the backlog's age.
Part 7. Circuit breakers
With C5, about 157 cart views a second get their stock badges and more than 1,000 get an error, while shop-api keeps sending about 1,400 calls a second that mostly fail. A circuit breaker notices that a dependency is plainly failing, stops calling it for a while, and answers at once instead: with a fallback, or with a fast error.
The three states
Synthesizing vector architecture diagram...
What to notice: every arrow uses the example's numbers (C6). The note is the flap of event 15f: the probes test the dependency at a load far lighter than the real one.
textBREAKER, one per shop-api instance and per operation: stock.view, stock.reserve (C6) before a call: if OPEN and 5 s have passed: go HALF_OPEN (0 probes sent, 0 passed) if OPEN: refuse now -> fallback for a view, fast 503 for a reserve if HALF_OPEN: if 3 probes already sent: refuse now else send this call as a probe, with a 100 ms timeout after a call: a probe: failure -> OPEN; success -> if 3 passed: CLOSED, clear the window in CLOSED: add the outcome to the last-20 window if the window is full and at least half failed: OPEN
Count window or time window
A breaker decides on recent calls. A count window looks at the last N calls: C6 uses the last 20 per instance and operation, as resilience4j does by default (COUNT_BASED, though its default size is 100). A time window looks at the last T seconds. With a sudden total failure, a time window full of healthy history is slow to trip: a 10 s window holds about 315 × 10 ≈ 3,000 healthy calls per instance, and must see failures outnumber them before it opens, so the time to open is roughly the window × the threshold. A window of the last 20 calls trips within the first second. (Side row 15t in Part 13 runs the 10 s time window: about 345 stock calls a second get through during the hiccup instead of about 80.)
The same 30 seconds, with breakers (C6)
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 15 | during 20:00:00 | C6 | All 4 view breakers open within the first second. stock calls fall to about 80 a second; checkouts about 47 (46.5 and 47.0). Cart views in full: only about 6 a second. Nothing else is served for views yet, because a refused view is still an error: there is no fallback until Part 9 |
| 15f | 20:00:05, 20:00:10, 20:00:15 | C6 | The flap. After each 5 s open period, 3 probes pass: a query averaging 40 ms beats the 100 ms probe timeout about 92% of the time (e^(−100 ÷ 40) ≈ 8% don't). The breakers close, the full view load returns (stock calls jump from about 60 to 125 to 230 in those seconds), and they open again within the second. Now and then a reserve breaker on one instance opens too, for several seconds |
| 16 | 20:00:20 → 20:00:21 | C6 | Probes pass on the healthy database; the view breakers close during 20:00:20; goodput is back at 20:00:21. A reserve breaker that opened late stays open for its full 5 s, so checkouts can lag for a few more seconds |
Probes and their timeout
The probe's timeout must sit above the healthy P99, or a recovered dependency can never pass it. Shopify's Semian names it half_open_resource_timeout, and the half-open call is a real request: for payments, a real buyer's payment, so a failed probe is an unknown to resolve like any other timeout (step 3.1 of the Shopify case study; the loop's fact-check caught a probe timeout set below the healthy 1.5 s call). The same step shows the price of probing: while a dependency stays down, each open circuit lets one real request wait up to h (the probe timeout) every open period e, about h ÷ e of a thread per failing resource. Shopify's published example multiplies that across 42 Redis instances, all down, on a 2-thread worker: "an extra utilization of 263%" with a 0.25 s half-open timeout and a 2 s error_timeout (0.25 × 42 ÷ (2 × 2) ≈ 2.63), and 4% after moving to a 50 ms half-open timeout and a 30 s error_timeout.
| Row | Change on C6 | Result (seeds 1 and 2) |
|---|---|---|
| 15x | One breaker for both operations | When views trip it, checkouts are blocked too: about 0.7 checkouts a second instead of 47 |
| 15y | Probe timeout 5 ms, below the healthy P99 of 18.7 ms | A probe round passes only if all 3 calls beat 5 ms: about 0.69³ ≈ 1 in 3 even on a healthy database. Goodput returns only at 20:00:50 and 20:01:00, 29 and 39 s later than with 100 ms probes |
Key it by what fails separately
Key a breaker by dependency and operation, and further (tenant, country, route) wherever failures differ: Shopify keys its payment circuits by the merchant's country, so one country's outage doesn't close checkout elsewhere. Our view breaker may open; the reserve breaker must not open because of views. Decide in advance what each open breaker returns: a fallback for an optional call (Part 9), a fast 503 for a required one. And remember the state is per process: Semian's is explicitly not shared across workers, so each of our 4 instances needs its own evidence before it opens.
Per host, per dependency, or a cap?
| Outlier ejection (per host) | Circuit breaker (per dependency and operation) | Concurrency cap (a bulkhead; Envoy's "circuit breaking") | |
|---|---|---|---|
| Question it answers | Is this one host bad? | Is this whole dependency failing? | How much of me may this dependency hold at once? |
| Signal | Errors from one host in a row (Envoy: 5 consecutive 5xx) | The failure or slow-call rate over recent calls | Calls in flight |
| Action | Stop sending to that host for a while (Envoy: 30 s × the number of ejections, at most 10% of hosts) | Refuse every call; probe later | Refuse calls over the cap, at once |
| Helps when | One bad host among healthy ones | The dependency is plainly failing | The dependency is slow, for any reason |
| Doesn't help when | Every host is slow (our hiccup: one shared database) | It is slow but alive: it flaps | The cap is set too high for the timeout (Part 8) |
| State | Per proxy or client | Per process | Per process |
Breaker or a retry budget only, on equal terms
The AWS Builders' Library prefers limiting retries locally with a token bucket over breakers: breakers add a mode that is hard to test and can add significant time to recovery. On this slow-but-alive database, with everything else equal (both runs have stock's protections and the brownout fallback of Part 9, and neither has the bulkhead):
| Measured 20:00:02 to 20:00:19 (seeds 1 to 5) | No breakers (side row 19b: C5 + brownout) | Breakers (C6 + the brownout, no bulkhead) |
|---|---|---|
Calls reaching stock | About 1,390 a second | About 75 a second |
| Cart views in full / without badges | About 155 / about 1,047 | About 6 / about 1,193 |
| Checkouts | About 44 (43.0 to 46.2) | About 43 (36.9 to 47.0) |
| Back to normal | 20:00:20 | 20:00:21 for the page; checkouts on an instance whose reserve breaker opened late stay refused until its 5 s open time ends (up to about 20:00:25) |
| Flapping | None | Every 5 s |
Over five seeds, checkouts are about the same with and without breakers (about 44 without; 43 to 47 with, depending on details: 43 as above, 45 if stock's 503s don't count as breaker failures, 47 in C8 with the bulkhead). Adding the bulkhead without breakers is worse for checkouts (about 34), because one gate shared by both operations starves reserves when nothing stops the views. The breakers' real gain is 95% less load on a struggling dependency; their cost is about 150 cart views a second shown without badges that could have had them, a second of slower recovery and a mode that flaps. So: a breaker where a fallback exists or the dependency is plainly down; a retry budget always.
Drill 23, the second question (The Slow Recommendation Service That Took Down Checkout): why not 50 ms timeouts everywhere instead of a breaker? A timeout is set from the healthy tail; below it, healthy calls fail and turn into retries (Part 2). And a timeout still sends every call: each one to a dead dependency costs a thread for the full timeout and adds load. A breaker stops sending, gives the dependency room to recover, and answers the caller at once. You need both: the timeout bounds one call, the breaker bounds the calls you keep making.
The views breaker is open. Why must checkouts have their own, and why does the views breaker keep reopening every 5 seconds?
What to remember from Part 7
- A breaker stops calling what is plainly failing and gives it room to recover.
- Key it by what can fail separately, trip it on recent calls, and decide each one's fallback or fast refusal.
- On a slow-but-alive dependency it flaps; the half-open probe's timeout must exceed a healthy call's P99.
Part 8. Bulkheads and concurrency limits
Go back to event 3: in C0, product pages waited because cart views and checkouts held all 400 threads. The pool was shared. A bulkhead (named after a ship's watertight walls) caps how many threads, connections or requests one dependency may hold, so a slow dependency can sink only its own compartment.
Synthesizing vector architecture diagram...
What to notice: inside the "shop-api instance: 100 threads" panel, only calls to stock pass the gate of 12; calls to prices go straight through. However slow stock gets, it can hold at most 12 of the 100 threads, and product pages always have the rest.
Size a cap with Little's law
Size = calls a second × the P99 latency × headroom. For stock: 315 calls a second per instance × 0.0187 s × 2 ≈ 11.8, so 12 per instance (48 across the fleet). With C1's 60 ms timeout, even if every call waited the full timeout, an instance would hold at most 315 × 0.06 ≈ 19 calls. A timeout bounds each call; the cap bounds how many.
A backstop, and when it saves you
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 17 | 20:00:00 → 20:00:20 | C7 | With a correct 60 ms timeout, and the breakers already cutting the load, the cap of 12 refuses only about 6 to 8 calls a second. Checkouts about 44 to 50. On the main line it is a backstop |
| 16x | side row | C0 + a cap of 40, and of 60 | The cap alone, on the as-shipped settings where the timeout is 1 s and wrong. With 40 per instance: product pages about 800 a second, cart views about 188, checkouts about 9, the cap refusing about 3,370 calls a second; recovered at 20:00:20. With 60 per instance: product pages still about 800, but no cart view or checkout succeeds during the hiccup; recovered at 20:00:21 |
Why does 40 work so well and 60 not? With 1 s timeouts, 40 × 4 = 160 calls in flight, served at the slow database's 200 a second, take about 160 ÷ 200 = 0.8 s each (Little's law): inside the 1 s timeout. At 60 per instance, 240 in flight take about 1.2 s each: every call waits past its 1 s, and all of them fail. Either way the cap protects the other traffic (product pages) and ends the outage at 20:00:20, which C0 never does. That is a bulkhead's job: the backstop for a timeout that turns out to be wrong.
(In the recommended order, the bulkhead comes after the breakers. Placed before them, on C5, a cap of 12 refuses about 290 to 300 calls a second, but checkouts fall from about 44 to about 33, because the one cap is shared by views and reserves. A cap per operation avoids that.)
Pools and semaphores
A semaphore caps how many of the caller's own threads may be inside calls to one dependency. A separate thread pool per dependency does the same with its own threads, and lets the caller walk away from a blocked call, at the price of more threads and context switches. Neither can end a call that is already blocked: that is what the timeout beside it is for.
Caps multiply with the fleet
A cap is per instance, so the fleet's total is cap × instances. Replay A scales shop-api from 4 to 8 instances at 20:00:10, the moment an autoscaler might react. On C7 it changes nothing visible (the breakers keep 2 to 4 calls in flight; recovered at 20:00:21), but the fleet-wide cap has doubled from 48 to 96 at exactly the moment stock is weakest. On C0 it doubles the threads waiting on stock (400 → 800): at 20:00:15 about 5,200 calls are inside stock, against 3,100 to 3,200 without the extra instances, and goodput stays at zero. Size caps from the dependency's capacity ÷ the number of callers, and let the dependency enforce its own limit too (Part 6).
Limits that adapt
A fixed cap goes stale as deploys and scaling change the system. Two adaptive forms, one sentence each:
- At the client, AIMD (additive increase, multiplicative decrease, as in TCP): add 1 to the limit after each healthy interval; halve it on timeouts,
429/503or latency above a target, at most once per interval. The job scheduler does this per target (step 2.6 of the job scheduler loop). - At the server, latency-based: Netflix's concurrency-limits library estimates the queue from how far current latency is above the minimum (its Vegas limit estimates the queue as limit × (1 − minRTT ÷ sampleRTT), adding 1 to the limit when that queue is small and subtracting 1 when it is large) and refuses the excess at once. For our
stockit would cut the limit sharply during the hiccup; the page doesn't simulate where it settles.
| Fixed cap (C7) | Client-side AIMD | Server-side latency-based limit | |
|---|---|---|---|
| Where it runs | Each caller | Each caller | The dependency |
| Signal | Calls in flight | Timeouts, 429/503, latency above a target | Current latency vs the minimum seen |
| On trouble | Refuses over the cap | Halves the limit, at most once per interval | Lowers the limit as the queue estimate grows |
| When healthy | Nothing changes | Adds 1 per healthy interval | Adds 1 while the queue estimate is small |
| Goes stale when | Deploys, scaling or latency change | Rarely; it keeps probing | Rarely; it keeps probing |
| Costs | Tuning by hand; × the fleet | Oscillation; per caller, so × the fleet | Needs a clean latency signal |
| In the loops | Shopify's Semian tickets | The job scheduler, per target (step 2.6) | The Netflix API (step 2.3) |
Managed bulkheads on AWS: Lambda reserved concurrency is both the maximum and the minimum concurrency of one function, carved out of the account's Regional pool (1,000 by default, shared by every function); an SQS event source's maximum concurrency caps how many invocations that one queue can drive, and should be set no higher than the function's reserved concurrency.
Drill 23, the first question (The Slow Recommendation Service That Took Down Checkout): why does a gateway run out of threads when an optional service slows from 30 ms to 5 s? Little's law: threads in use = rate × time held. At 30 ms the calls hold a few threads; at 5 s they need about 167 times as many, and in one shared pool they take every thread, so requests that never call the slow service wait too. A per-dependency cap keeps the rest of the pool free.
Why did a 40-call cap per instance keep product pages up even with the as-shipped timeouts, and why does a 12-call cap barely act in C7?
What to remember from Part 8
- Cap what each dependency may hold, sized by rate × latency × headroom.
- It's the backstop for a timeout that turns out to be wrong.
- A per-instance cap multiplies with the fleet; let the dependency set its own limit too.
Part 9. Brownouts and fallbacks
After C7, about 1,190 cart views a second fail fast: the breaker is open or the gate is full, and the view has nowhere to go. But a cart view needs its items and prices; the stock badges are optional. A brownout drops the optional part and serves the rest.
Synthesizing vector architecture diagram...
What to notice: the same five kinds of failure lead to two different answers, decided in advance per operation: a degraded cart page for the optional call, a fast, honest refusal for the required one.
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 19 | 20:00:00 → 20:00:20 | C8 | Cart views served without badges: about 1,180 to 1,200 a second, plus about 6 in full. Product pages about 793 to 804. Checkouts 44 to 50 of 60 a second (43.5 and 49.9; 42.7 to 51.8 over seeds 3 to 5). Back to normal at 20:00:21 for the page; checkouts can lag a few more seconds where a reserve breaker opened late (seed 1: 37, 49, 37 and 34 a second at 20:00:20 to 20:00:23) |
Snapshot W4, 20:00:05: in the hiccup, at a flap (C8, seed 1; rates are per second)
| Tier | What it's doing |
|---|---|
| Apps | 2,000 attempts, no storm |
shop-api | 18 of 400 threads busy, nothing queued; view breakers open on all 4 instances (they flap in this second); 1,105 calls refused by open breakers, 22 by the gate, 17 retries refused by the budget |
stock | 116 calls in, 12 dropped by the age limit, none refused at admission |
| Database | 104 queries started, 71 finished, 0 for nobody |
| Goodput | products 763, cart views 32 in full and 1,150 without badges, checkouts 39 |
Snapshot W5, 20:00:21: recovered (C8, seed 1)
| Tier | What it's doing |
|---|---|
| Apps | 2,021 attempts |
shop-api | 25 threads busy, nothing queued; view breakers closed; one reserve breaker still open (9 checkouts refused) |
stock | 1,171 calls in, nothing refused |
| Database | 1,168 queries finished, 0 for nobody |
| Goodput | products 840, cart views 1,119 in full, checkouts 49 |
Three kinds of fallback
Netflix's API (step 2.3 of the Netflix case study) names three, decided per call in advance:
| Kind | What it returns | Our example |
|---|---|---|
| Custom | A default or a cached value | "In stock" for items that were in stock an hour ago (with a staleness limit) |
| Fail silent | Nothing: the part is left out | The cart without badges |
| Fail fast | An error, at once | The checkout's 503 + Retry-After |
Fallbacks that never run don't work. Exercise them on purpose (step 2.5 of the Netflix case study: failures surprise us in production). Two things are never shed: the load balancer's health checks, because a server that ignores them is taken out of service just when every server is needed; and emergency paths, which the ride-sharing loop exempts from shedding (step 3.5 of the ride-sharing loop).
Why checkouts top out near 50
Even with no queue at all, a reserve attempt has 60 ms, and during the hiccup a query takes 40 ms on average, with an exponential spread. The chance that one query takes longer than 60 ms is e^(−60 ÷ 40) = e^(−1.5) ≈ 22%. The run agrees: 22% and 23% of reserve attempts timed out. A second attempt helps (at best 1 − 0.22² ≈ 95% of checkouts would succeed), but the age limit, the shared gate and an occasional open reserve breaker take their share, so about 44 to 50 of 60 get through. To do better, give the checkout a longer attempt timeout (a checkout is worth more than a badge), or let fewer views compete for the 8 connections.
Shedding and autoscaling
Shedding hides the signal autoscaling waits for. The Builders' Library warns that a service shedding at a CPU target keeps its CPU at that target, so CPU-based scaling never fires; and fast refusals make the median latency look excellent while successful requests are slow. So scale and alarm on offered load, shed counts, breaker opens and budget refusals, and report latency of successes separately. Shedding also doesn't replace headroom: Amazon services keep enough spare capacity to handle an Availability Zone failure without adding capacity, and shedding can hide how close a fleet is to that limit.
Which is better for the customer at 20:00:05: a cart without stock badges, or an error?
What to remember from Part 9
- Degrade optional parts before failing required ones.
- Decide each call's fallback in advance, and test it.
- Shedding hides load from CPU-based scaling: scale and alarm on offered load and shed counts.
Part 10. Why it stayed down: metastable failures
Look at snapshot W3 again: at 20:00:25 the database is healthy, busy at full capacity, and serving nobody. The trigger lasted 20 seconds; C0's outage would have lasted until someone intervened. This Part names that kind of failure, the loops that sustain it, and the ways out.
Stable, vulnerable, metastable
Bronson and colleagues ("Metastable Failures in Distributed Systems", HotOS 2021) describe three states:
Synthesizing vector architecture diagram...
What to notice: nothing goes wrong in the vulnerable state; it only needs a trigger. And there is no arrow from Metastable back to Vulnerable when the trigger ends: the only way out is a push from outside.
- Stable: even the worst amplification the system can produce (every retry, every stale request) still fits in its capacity.
- Vulnerable: normal load is fine, but the amplified load would not fit. Basketly on Friday evening is vulnerable: 2,060 actions a second is fine, but C0's 3 × 2,060 = 6,180 attempts a second is not.
- Metastable: a trigger pushed it over, and a sustaining effect keeps it there after the trigger is gone.
Systems run in the vulnerable state because it is cheaper: higher utilization, fewer servers. The job of the guards on this page is to make the amplification small enough, or breakable enough, that the vulnerable state is safe.
The loops that kept it down
- Retries (event 4): about 6,180 attempts a second against about 3,270 a second that can be finished.
- Stale work (event 21, below): even with no retries at all,
shop-apiserved its queue first in, first out, for callers who had left, and recovered only 25 s after the trigger ended. - A cache that can't refill, the look-aside cache example in Bronson's paper (its second, section 2.2): behind a look-aside cache with 90% hits, the database sees a tenth of the reads. Lose the cache while the system is vulnerable and the database sees ten times as many; the reads time out, nothing gets back into the cache, and the system stays down. Cold caches and stampedes belong to Caching & Invalidation (Parts 5 and 8 there); coalescing identical reads is in Sharding, Hot Keys & Rebalancing (Part 4 there).
| # | Time | Setting | Event (seeds 1 and 2) |
|---|---|---|---|
| 21 | side row | C0, no retries anywhere | The same hiccup with 1 attempt at every layer: goodput back at 20:00:45, shop-api's queue (about 28,000 at 20:00:20) empty at 20:00:48 to 20:00:49. Stale work alone costs 25 s; with retries (C0), it never ends |
Synthesizing vector architecture diagram...
What to notice: three lines, seed 1, every other second. The line near 2,060 all the way across is C8, counting cart views served without badges; the line that sits near 850 until second 20 and then joins it is C8 counting only full answers (the gap is the 1,190 degraded cart views a second). The line at zero from second 2 onward is C0: it never comes back.
Getting out
At 20:02:00, two minutes into C0's outage, with about 403,000 requests queued in shop-api, two operator actions were tried on copies:
| # | Action at 20:02:00 (C0) | What happened (seeds 1 and 2) |
|---|---|---|
| 22 | Shed 70% of new arrivals at the edge for 30 s | Not enough. The queue only falls from about 403,000 to about 363,000 by 20:02:30, and goodput stays at zero. The queued work is still stale, and the as-shipped apps retry a refused attempt at once, so about 1 − 0.7³ ≈ 66% of actions still get in |
| 22 | Flush the queues (restart the workers) | Goodput back in the same second: about 2,830 to 2,990 answered in 20:02:00 (new arrivals plus requests that still had time), and no queue left by 20:02:15. It held in all 8 of our seeds; an independent model relapsed in 1 of its 8, when the retry wave from the flushed attempts brought it back down within 13 s |
Getting out needs two things together: load below capacity, and the stale backlog thrown away. Shedding new arrivals does the first but not the second; a flush does the second, and the load then usually fits. It isn't guaranteed while the as-shipped retries remain: every flushed attempt times out 2 s later and is retried at once, and that wave alone can tip the site straight back. So flush and stop the retries (or switch on deadlines) together; and the next trigger will start it again if nothing else changes. Dropping expired work by deadline (C1) is the same flush, done continuously.
Why "add capacity" didn't help
Replay A (Part 8) doubled shop-api at 20:00:10: on C0 goodput stayed at zero, and about 60% more calls were waiting inside stock than without the extra instances. More callers put more pressure on the slow dependency, and capacity that arrives in minutes can't beat a storm that arrives in seconds. The loops' fact-checks keep catching "autoscale" as the fix (for example, step 2.6 of the mobile stock-trading loop, the opening bell).
Every guard on equal terms
What each guard is:
| Guard | Runs in | Bounds | Reacts to | Cost on a healthy day | Cost in overload | Misconfigured, it... | State | Needs idempotency? |
|---|---|---|---|---|---|---|---|---|
| Timeout + per-attempt deadline | Every caller and queue | Time per attempt and per request | The clock | False timeouts (≤ 0.1%) | Work cancelled mid-way | Too short: retries on a healthy day; too long: threads held; whole deadline sent and the cancel lost: stale work | Per request | No |
| One retry layer | The layer above the failure | Repeats across layers | Errors | Blips far from the failure aren't hidden | None | Hidden layers multiply anyway | None | Yes, for writes |
| Backoff + full jitter | Every retrier | When retries happen | Errors | Slower recovery for an unlucky client | None | No jitter: synchronized waves | Per request | Yes, for writes |
| Retry budget | Each client | Retries as a share of traffic | Its own failure rate | None (a full bucket) | A blip isn't retried | Fixed per-client sizes grow with the fleet | Per client | Yes, for writes |
| Hedging | Each client | Tail latency of one slow replica | Time past the P95 | About 5% more reads | Harmful unless inside the budget | On writes: duplicates | Per client | Reads only |
| Admission, priority, age limit | The dependency | Work it accepts, and which | Its own queue and age | None | Refuses while some capacity is left | Priority in a queue that never forms: no effect | Per server | No |
| Circuit breaker | Each client, per operation | Calls to a failing dependency | Recent failure rate | None | Flaps on slow-but-alive; slower recovery | Shared key: blocks what still works; low probe timeout: never closes | Per process | Probes may be real writes |
| Bulkhead (fixed cap) | Each client | Threads a dependency may hold | Calls in flight | Idle reserved threads | Refusals at the cap | Too big for the timeout: no protection; shared by operations: starves one | Per process | No |
| Adaptive limit | Client (AIMD) or server (latency) | Concurrency | Latency or errors | Small oscillation | Refusals | Needs a clean latency signal | Per process | No |
| Brownout | The caller with a fallback | What the user loses | A refused or failed call | None | A worse page | Fallback never tested: fails when needed | None | No |
What the run shows, each guard alone on the as-shipped settings (seed 1):
| Guard, alone on C0 | Prevents the metastable state? (on from the start) | Ends it once started? (switched on at 20:02, 403,000 queued) |
|---|---|---|
| Timeout 60 ms + per-attempt deadline, cancel, drop at dequeue | Yes: recovered at 20:00:20 | No within 80 s: the 403,000 requests already queued were sent without a deadline, so they are still served |
| One retry layer | No | No |
| Backoff + jitter | No | No |
| Retry budget | No | No |
stock's admission, priority and age limit | Yes: 20:00:20 | Yes, in 14 s |
| Breakers | Yes: 20:00:23 | No: at 20:02 stock is healthy, so the breakers never open; the loop lives in shop-api's own queue |
| Bulkhead of 12 (of 40) | Yes: 20:00:20 | Yes, in 13 s (22 s with a cap of 40) |
| Brownout | No: a view still holds its thread through three 1 s attempts before it degrades | No |
| Hedging, adaptive limits | Not simulated | Not simulated |
Four guards prevent it on their own; only two end it once it has started, because only they stop shop-api's threads from being consumed by the backlog. That is why the page stacks guards: each layer should be able to break the loop by itself.
The database was fine at 20:00:20. What kept the site down?
What to remember from Part 10
- The trigger ends; a loop (retries, stale queues, a cold cache) keeps the system down.
- Getting out needs load below capacity and the stale work thrown away.
- Design so each layer can break the loop on its own.
Part 11. End to end through the layers
With C8, every hop bounds time, repeats and concurrency. Follow one checkout and one cart view through all of them at 20:00:07, in the middle of the hiccup (seed 1, event 23).
Synthesizing vector architecture diagram...
What to notice: the deadline shrinks at every hop, 2.0 s at the app, 60 ms at stock, whatever is left at the database. The highlighted boxes, shop-api and stock, are the two services that refuse work: one on behalf of its callers, one on its own behalf.
The trace (seed 1, C8):
| Time | Cart view | Checkout |
|---|---|---|
| 20:00:07.0003 | App sends attempt 1 (deadline 20:00:09.000); a shop-api thread takes it at once | |
| 20:00:07.0103 | prices done; the view breaker is open: refused at once; served without badges, 10 ms after the tap | |
| 20:00:07.0321 | App sends attempt 1 (deadline 20:00:09.032); a thread takes it at once | |
| 20:00:07.0421 | prices done; attempt 1 to stock with a 60 ms deadline (1.99 s were left); the reserve goes straight to a connection, statement timeout 60 ms | |
| 20:00:07.0515 | The query finishes in 9.4 ms; checkout confirmed, 19.4 ms after the tap |
The deadline at each hop
| Hop | The deadline it receives | For this checkout |
|---|---|---|
App → edge → shop-api | The app attempt's 2.0 s | Until 20:00:09.032 |
shop-api → stock | min(the attempt's 60 ms, the time left) | 60 ms (1.99 s left) |
stock → database | Statement timeout = the attempt's time left when the query starts | 60 ms (it didn't wait for a connection) |
Which guard acts where
| Hop | Guards | What it refuses | What it costs |
|---|---|---|---|
| App | 2.0 s deadline per attempt; one retry, only with no answer; full jitter; honours Retry-After; the same idempotency key | Retrying a 503 | A user waits up to 2 s, then maybe one more try |
| Edge | Early refusal before TLS and authentication when overloaded; per-client rate limits (Rate Limiting Algorithms) | Excess, before any work | Refused users |
shop-api | Deadline check at dequeue; breaker per operation; gate of 12; 60 ms attempt with its own deadline, cancelled on give-up; budget of 10%; jittered retry with the same key; fallback | Calls to a plainly failing operation, over the gate, or retries over the budget | Degraded carts; a few false timeouts |
stock | Admission 64 (reserves 80); reserves first in the connection queue; age limit and newest first when the queue stands; deadline at dequeue | Work it can't finish in time | Refusals while some capacity is left |
| Database | Statement timeout = the time left; 8 connections | Queries past their deadline | Cancelled work |
Sizing for an Availability Zone loss
stock's 128 workers run in two Availability Zones, 64 in each. Lose one zone during the hiccup and 64 remain. On a copy of C8 with 64 workers, the run is unchanged: checkouts 44 to 50 a second, cart views degraded, back at 20:00:21. At most 80 calls are ever inside stock, and the 8 database connections, not the workers, are the limit; on a healthy day stock needs only about 1,260 × 0.005 ≈ 6 workers. With the admission limits also halved (32, reserves 40), checkouts are 47 to 50. So stock can lose a zone with views degraded and still serve every checkout the database can. The database's own zone failure is a different mechanism: see Replication, Quorums & Read-Your-Writes.
What to remember from Part 11
- Every hop bounds time, repeats and concurrency.
- Each attempt carries its own deadline, and an abandoned one is cancelled.
- The cheapest refusal is the earliest.
Part 12. On AWS
Every guard on this page exists somewhere in AWS, often switched on by default, sometimes with a default you should change. Three rules carry over: count every layer's retries, remember which limits are per client and which are per account, and know that you still choose the deadlines, the one retry layer, the priorities and the fallbacks.
Managed services that use it
| Service | What it provides | What AWS documents |
|---|---|---|
| Amazon API Gateway | Admission by token bucket; integration timeouts | An account-level throttle per Region of 10,000 requests a second with a burst of 5,000 (2,500 and 1,250 in some newer Regions), applied on a best-effort basis; excess gets 429 Too Many Requests; stage, method and usage-plan throttles on top. REST integration timeout 50 ms to 29 s (Regional and private APIs can go higher in exchange for a lower throttle quota); HTTP APIs 30 s. A Lambda function behind it can run up to 900 s: set its timeout below the integration's |
| Application Load Balancer | Rejects rather than queues; per-target concurrency caps; routing around a bad target | 503 when no target is registered or, with target optimizer, when no target is ready (an agent on each target enforces its maximum concurrency); 504 when a target doesn't answer within the idle timeout or the 10 s connection timeout. Automatic target weights: detection is always on and needs at least 3 healthy targets; mitigation (shifting traffic away) needs weighted random routing and anomaly mitigation turned on, and is off by default. Least outstanding requests routing; slow start of 30 to 900 s for new targets |
| Amazon CloudFront | Origin timeouts, and its own retries | Connection attempts 1 to 3 (default 3); connection timeout 1 to 10 s (default 10); origin response timeout default 30 s (1 to 120, more on request). After a response timeout it tries GET and HEAD again, never POST, PUT, PATCH, DELETE or OPTIONS. An optional response completion timeout caps the whole wait |
| AWS Lambda | Concurrency as a bulkhead; managed retries | Concurrency = requests a second × average duration (the docs' formula; it is Little's law). 1,000 concurrent executions per account per Region by default, shared by all functions; reserved concurrency is both a function's maximum and its minimum; throttles are 429. Asynchronous invocations: 2 retries after function errors by default (1 minute, then 2), set with MaximumRetryAttempts (0 to 2) and MaximumEventAgeInSeconds (up to 6 h); throttles and system errors retried for up to 6 h with backoff from 1 s to 5 minutes. SQS event sources: a maximum concurrency per source, set no higher than the reserved concurrency |
| Amazon SQS | A buffer the consumer pulls from | The consumer works at its own pace; alarm on the age of the oldest message. Visibility timeouts, redelivery and dead-letter queues belong to the Queues & Delivery Semantics loop primitive |
| AWS Step Functions | Declarative retries with backoff and jitter; task timeouts | Retry: IntervalSeconds 1, MaxAttempts 3, BackoffRate 2.0, MaxDelaySeconds, JitterStrategy FULL or NONE (default NONE). Task TimeoutSeconds defaults to 99,999,999; HeartbeatSeconds for long tasks; HTTP tasks are capped at 60 s |
| Amazon DynamoDB | Throttles the client must back off from | ProvisionedThroughputExceededException and ThrottlingException come back as HTTP 400 and are "OK to retry"; the SDKs retry them with backoff. Other 4xx errors need the request fixed |
| Amazon RDS Proxy | A bounded connection pool with a wait limit | MaxConnectionsPercent caps connections to the database; ConnectionBorrowTimeout defaults to 120 s (0 to 300 s): set it below your callers' deadline |
| Amazon ECS (Service Connect) | A proxy with timeouts, retries and outlier detection | Per-request timeout 15 s by default; retries a failed connection twice, the second attempt avoiding the previous host; outlier detection stops using a task after 5 or more failed connections in 30 s, for 30 to 300 s; round-robin routing. A hidden retry layer to count |
The AWS SDKs (not a service, so not a chip) are the retry layer most applications already have:
| Standard mode | Adaptive mode | Legacy mode | |
|---|---|---|---|
| Attempts | 3 in total (DynamoDB clients 4 in the updated behaviour) | As standard | Varies by SDK: Python makes 5 (10 for DynamoDB) |
| Backoff | Full jitter, capped at 20 s; updated behaviour: base 50 ms (transient), 1,000 ms (throttling), 25 ms for DynamoDB | As standard | Varies |
| Retry quota | 500 tokens, typically per client instance, not shared across processes or hosts; updated costs: 14 per transient retry (before: 5, and 10 for a timeout), 5 per throttling retry; 1 refunded per first-try success | As standard | None standardized |
| Extra | x-amz-retry-after (ms) honoured, clamped | A client-side rate limiter that can delay first attempts; for a client that talks to one resource, not a multi-tenant client | |
| Default today | Most SDKs | Opt in | Java, Python, Ruby, PHP, C++ and AWS CLI version 1 (version 2 already defaults to standard) until November 2026 |
The updated behaviour can be switched on now with AWS_NEW_RETRIES_2026=true and becomes the default in November 2026 (AWS, 20 May 2026), which also moves the legacy-mode SDKs to standard mode.
Running it yourself
| Option | What it is | Facts and sizing |
|---|---|---|
| Amazon EKS or Amazon EC2 with Envoy (as a sidecar or at the edge) | Timeouts, retries, retry budgets, concurrency caps and outlier ejection in the proxy | Envoy's "circuit breaking" is concurrency caps per upstream cluster and priority, not a breaker: max_connections, max_pending_requests and max_requests default to 1,024, max_retries to 3, and they are enforced by each proxy on its own. Retry budget budget_percent 20%, min_retry_concurrency 3. Outlier detection: 5 consecutive 5xx, 10 s interval, 30 s base ejection time multiplied by the number of ejections, at most 10% of hosts ejected. Route timeout 15 s; retries only if configured; backoff fully jittered from 25 ms, capped at 250 ms; per-try timeouts; hedging on a per-try timeout; x-envoy-expected-rq-timeout-ms carries the per-try timeout when one is set (the route timeout otherwise, or when hedging on per-try timeouts), never more than the route's time left |
| Libraries in the application (on EC2, ECS or EKS) | Breakers, bulkheads and retries in code | resilience4j breaker defaults: COUNT_BASED window of 100 calls, at least 100 calls, 50% failure threshold, 60 s open, 10 calls in half-open, slow-call threshold 60 s; its Retry: 3 attempts, fixed 500 ms. Semian (Ruby): per-process breaker state, bulkhead tickets through SysV semaphores, half_open_resource_timeout. Netflix concurrency-limits (Java): Vegas and gradient limits. Hystrix is in maintenance mode |
| Sizing in words | Caps from rate × P99 × headroom per instance, checked against the dependency's capacity ÷ the number of callers. The dependency's own limit sized to survive an Availability Zone loss with low-priority work shed. Alarms on shed counts, breaker opens, budget refusals and queue age, not only on CPU |
Look-alikes that are not this mechanism
| Look-alike | Why it looks like this | Why it isn't |
|---|---|---|
| EC2 Auto Scaling, Application Auto Scaling | "More capacity fixes overload" | Capacity arrives in minutes, the storm in seconds; more callers press harder on the slow dependency; shedding can hide the CPU signal it scales on (Parts 9 and 10) |
| Elastic Load Balancing and Route 53 health checks | "They take bad targets out" | They probe a cheap endpoint on a timer; a target that is slow on real requests usually still passes. Automatic target weights (above) is the closer relative, though it reacts to 5xx responses and connection failures, not to slowness |
| AWS WAF rate-based rules, API Gateway usage plans | "They refuse excess traffic" | Per-client limits on what a client sent, not on the service's health (Rate Limiting Algorithms) |
| DynamoDB adaptive capacity and on-demand mode | "No more throttling" | Capacity management; per-partition limits still apply (Sharding, Hot Keys & Rebalancing) |
| SQS dead-letter queues and redrive | "Failed work is set aside" | Handling for messages that keep failing (Queues & Delivery Semantics), not a retry budget |
| Amazon VPC Lattice | A service network in front of services | Health checks, round-robin routing, and fail-open when every target is unhealthy; its documentation describes no retry or outlier-detection settings |
| AWS App Mesh | An Envoy-based mesh with retries and outlier detection | End of support on 2026-09-30; use Envoy directly or ECS Service Connect |
| AWS Fault Injection Service | "Resilience" | It tests these guards (Multi-Region Failover, Part 10 there); it doesn't provide them |
| A market-wide trading "circuit breaker", an SMS "budget breaker" | The same word | A trading halt (step 3.6 of the stock exchange loop) and a spending cap (the bot-defense loop), not the call-level pattern |
What to remember from Part 12
- Every managed layer has its own timeouts and retries: count them.
- The SDK's retry quota is per client instance, and Lambda's concurrency is per account and Region, not per dependency.
- You still choose the deadlines, the one retry layer, the priorities and the fallbacks.
Part 13. What you've learned
Back to the home page
An optional star rating took a home page down because waiting held threads (Little's law: 3,333 a second × 5 s = 16,665 threads needed from a pool of 1,000) and failure multiplied work (retries at every layer, and answers finished for callers who had left). Basketly's 20-second hiccup did the same, and C0 never came back. With every guard in place (C8), the same 30 seconds kept product pages whole, served about 1,190 cart views a second without badges, confirmed 44 to 50 of 60 checkouts a second, and was back to normal one second after the database. Here is what each piece did:
- Timeouts from percentiles and per-attempt deadlines (Part 2) bounded how long anyone waited, and dropped stale work instead of serving it: C1 alone recovered at 20:00:20.
- One retry layer (Part 3) cut the calls reaching
stockfrom 5.4× to 1.9×; the idempotency key made the reserve's retries safe. - Jitter (Part 4) spread retries in time without removing any; stable offsets kept periodic work flat and fresh.
- A retry budget (Part 5) brought retries down to about 1.1×.
stockprotecting itself (Part 6) refused early and gave its 200 queries a second to reserves and to fresh requests: checkouts from 9 to 44 a second.- Breakers per operation (Part 7) cut the load on
stockby 95%, at the price of a flap every 5 s. - A bulkhead (Part 8) was a backstop on the main line, and the thing that saved the site when the timeout was wrong.
- A brownout (Part 9) turned 1,190 errors a second into carts without badges.
The snapshots, side by side
| Snapshot | Setting and time | What it showed |
|---|---|---|
| W1 | C0, 19:59:55 | Healthy: 28 threads busy, 1,240 queries a second, nothing for nobody |
| W2 | C0, 20:00:01 | 400 of 400 threads busy, 1,192 queued, 163 of 183 queries already for nobody |
| W3 | C0, 20:00:25 | Trigger gone; goodput 0; 128,309 queued; 1,941 queries a second, all for nobody |
| W4 | C8, 20:00:05 | Breakers flapping for views; 1,150 carts without badges; checkouts flowing; nothing for nobody |
| W5 | C8, 20:00:21 | Recovered: 1,119 full cart views, 25 threads busy |
| W6 | C0 vs C8, 20:00 to 20:01 | C0 at zero from 20:00:02 onward; C8 answering about 2,060 a second throughout, about 850 of them in full during the hiccup |
What it costs
- False timeouts on a healthy day (at most 0.1% of
stockcalls with a 60 ms timeout), each one a retry. - Requests refused while some capacity is left: admission limits, age limits, gates and budgets all refuse work that might have succeeded.
- A degraded page: 1,190 carts a second without badges for 20 s, and some customers who reach checkout and find an item gone.
- Tuning per dependency: timeouts, caps, windows and probe timeouts come from measurements, and go stale after deploys.
- Counters that don't know the fleet's size: breakers, budgets and caps are per process, so fixed allowances grow with every instance you add.
- Discipline: deadlines and cancellation at every hop, idempotency keys on every retried write, fallbacks that are exercised.
The whole story, event by event
Side rows and side replays are marked "(side)". Seeds 1 and 2 unless stated.
| # | Time | Setting | Event |
|---|---|---|---|
| 1 | 19:59:50 | C0 | Baseline: about 2,060 actions/s; about 26 of 400 threads busy; about 1,250 queries/s, none for nobody |
| 2 | 20:00:00.000 | C0 | The hiccup: queries take 40 ms; capacity 200/s |
| 3 | 20:00:00.38 to 20:00:00.40 | C0 | All 400 threads busy; product pages queue behind stock calls |
| 4 | 20:00:02 onward | C0 | Apps retry at once: about 5,800 to 5,900 attempts/s, then 6,180 |
| 5 | 20:00:02 | C0 | Goodput 0; the database's 200 queries/s are all for nobody |
| 6 | 20:00:20.000 | C0 | The database is healthy again |
| 7 | 20:00:20 → 20:05:00 | C0 | Goodput still 0; queue 104,000 to 106,000 at 20:00:20, +2,900/s: 403,000 at 20:02, 923,000 at 20:05; 6,180 attempts/s > 3,270 finishable |
| 8 | 19:59:50 | C1 | Healthy stock calls: P50 3.0 ms, P95 12.2, P99 18.7, P99.9 about 28 → timeout 60 ms |
| 9 | 20:00:00 → 20:00:20 | C1 | Product pages about 723/s; stock receives about 6,850 calls/s (5.4×) |
| 10 | 20:00:20 | C1 | Recovered at once; queue about 8,200, empty at 20:00:27 to 20:00:28; 0 queries for nobody |
| 10z | (side) | C1, whole deadline sent, cancel lost | About 200 queries/s for nobody; never recovers (with the cancel working: recovered at 20:00:20, like C1) |
| 10y | (side) | C1 copy | A 5 ms timeout fails about 31% of healthy calls; a 5 s timeout needs 1,260 × 5 = 6,300 threads |
| 11 | 20:00:00 → 20:00:20 | C2 | Worst case 27 → 2 queries per app attempt; about 2,410 stock calls/s (1.9×); recovered at 20:00:20 |
| 11x | (side) | C0 copy | A proxy retrying connections twice: 81 per tap |
| 11y | (side) | C0 | About 3.8 checkouts/s reserved more than once during the hiccup (7.5 extra commits/s, all from shop-api's own retries); the idempotency key returns the first result; none from C1 on |
| 12 | 20:00:00 → 20:00:20 | C3 | About 2,440 stock calls/s: spread, not removed. On C0, jitter alone never recovers |
| 12J | (side) | Replay J | 1,200 retries: all in one 10 ms slot with a fixed 1 s wait; about 12 per slot (5 to 25) with full jitter |
| 12P | (side) | Replay P | Aligned at :00: 10,000 reads, 13.5 s of queueing a minute. Fresh random offsets: gaps 2.9 to 118.8 s, 12.5% over 90 s. Stable offsets: 167 reads/s on average (busiest second 350), every gap 60 s |
| 13 | 20:00:00 → 20:00:20 | C4 | About 1,380 stock calls/s (1.1×); about 950 retries/s refused by the budget |
| 13a | (side) | C4, 8 instances | Still about 1,380 calls/s: a budget that earns per request scales with traffic; the burst and fixed per-client allowances scale with the fleet |
| 13H | (side) | Replay H | 1% of calls stall 250 ms: P99 250 ms and P99.9 259 ms without hedging; 17 ms and 24 ms hedged after the P95 (13.1 ms), for 5% more calls |
| 14 | 20:00:00 → 20:00:20 | C5 | About 165 refused/s at admission and 705 by the age limit; checkouts about 44/s; about 157 full cart views/s |
| 14a | 20:00:00 → 20:00:01 | C4 vs C5 | Queries started 1,300 to 1,400 vs 550 to 585; finished for a waiting caller about 190 in both; checkouts 15 to 16 vs 43 to 55 |
| 17x | (side) | C0 + stock's guards | About 8,400 to 10,300 calls/s reach stock, about 7,300 to 9,500 refused at admission (the range is how fast callers retry: at the same instant, or 1 to 5 ms later); checkouts 56 to 65/s; recovered at 20:00:20 |
| 17y | (side) | C0, no retries | 28,000 queued at 20:00:20 drain at about 1,240/s once stock has cleared its own 4,100 stale calls; empty at 20:00:48 to 20:00:49 |
| 15 | during 20:00:00 | C6 | All 4 view breakers open in the first second; stock calls about 80/s; checkouts about 47/s; full cart views about 6/s |
| 15f | 20:00:05, :10, :15 | C6 | The flap: probes pass, breakers close, the full load returns, they reopen |
| 16 | 20:00:20 → 20:00:21 | C6 | Breakers close; goodput back at 20:00:21 |
| 15x | (side) | C6, one shared breaker | Checkouts about 0.7/s |
| 15y | (side) | C6, 5 ms probes | Back only at 20:00:50 and 20:01:00 |
| 15t | (side) | C6, 10 s time window | About 345 stock calls/s get through instead of 80; full cart views 35 to 40/s |
| 17 | 20:00:00 → 20:00:20 | C7 | The gate of 12 refuses 6 to 8 calls/s: a backstop |
| 16x | (side) | C0 + gate of 40 / of 60 | 40: products about 800/s, recovered at 20:00:20. 60: no view or checkout during the hiccup, recovered at 20:00:21 |
| 16A | (side) | Replay A on C7 | 4 → 8 instances at 20:00:10: the fleet cap doubles (48 → 96); the breakers keep 2 to 4 in flight; recovered at 20:00:21 |
| A | (side) | Replay A on C0 | 4 → 8 instances: calls inside stock about 5,200 at 20:00:15 vs 3,100 to 3,200 without the scale-out; goodput stays 0 |
| 17b | (side) | C5 + gate of 12, no breakers | About 290 to 300 refused/s; checkouts fall to about 33/s (one gate shared by both operations) |
| 19 | 20:00:00 → 20:00:20 | C8 | About 1,190 cart views/s without badges; products about 800/s; checkouts 44 to 50/s (42.7 to 51.8 over seeds 3 to 5); 22 to 23% of reserve attempts time out; back at 20:00:21 |
| 19b | (side) | C5 + brownout, no breakers | About 157 full + 1,048 degraded cart views/s; checkouts about 44/s; back at 20:00:20 |
| 19x | (side) | text | Shedding at a CPU target keeps CPU low, so CPU-based scaling never fires |
| 20 | 20:00:00 → 20:01:00 | C0 vs C8 | Goodput per second (W6) |
| 21 | (side) | C0, no retries | Goodput back at 20:00:45; queue empty at 20:00:48 to 20:00:49 |
| 22 | 20:02:00 | C0 copies | Shedding 70% for 30 s: queue 403,000 → 363,000, still 0. A flush: back in the same second (8 of 8 seeds here; not guaranteed while the as-shipped retries remain) |
| 22s | 20:02:00 | C0 copies (seed 1) | Switched on in the metastable state: the gate of 12 ends it in 13 s (40: 22 s), stock's guards in 14 s; deadlines, breakers, the budget, jitter, one retry layer and the brownout don't end it |
| AZ | (side) | C8, 64 workers | Unchanged: checkouts 44 to 50/s; with the admission limits halved too, 47 to 50/s |
| 23 | 20:00:07 | C8 | A cart view degraded in 10 ms at an open breaker; a checkout confirmed in 19.4 ms with a 60 ms deadline at stock |
The cheat card
| Topic | Remember |
|---|---|
| Little's law | In flight = rate × time held: a call 100× slower needs 100× the threads |
| Goodput | Answered successfully within the caller's deadline; work for nobody is waste |
| Can it drain? | Only if what arrives (with retries) is less than what can be finished: 6,180 > 3,270 never drains |
| Timeout | The percentile of your false-timeout rate (0.1% → P99.9), padded; connection setup inside the timer |
| Deadline | Send each attempt's own: min(attempt timeout, time left), as a duration; cancel on give-up; drop expired work at every dequeue |
| Which errors | Transient and throttles, if safe to repeat; never 400/403/404/422; a side effect only with the same key |
| Where | One layer, right above the failure; others pass "overloaded, don't retry" up; count hidden layers |
| Backoff | Full jitter: random(0, min(cap, base × 2ⁿ)); it spreads retries, it doesn't remove them |
| Periodic work | A stable offset per client (hash of its ID mod the period), not a fresh random wait |
| Budget | Retries ≤ about 10% of requests per client: about 1.1× in a total failure; fixed per-client allowances grow with the fleet |
| Hedging | Reads only, after the P95, inside the budget, loser cancelled: about 5% more calls |
| Protect the dependency | Admission limit with 503 + Retry-After, early and cheap; priority and an age limit (5 ms once standing 100 ms) where requests really wait; newest first while it stands |
| Breaker | Per dependency and operation; count window; probe timeout above the healthy P99; flaps on slow-but-alive |
| Bulkhead | Rate × P99 × headroom per instance (315 × 0.0187 × 2 ≈ 12); a backstop for a wrong timeout; × the fleet |
| Brownout | Optional parts fail silent, required parts fail fast; decided in advance, tested |
| Metastable | Trigger + sustaining loop; out = load below capacity and stale work thrown away |
Failure checklist
- Is every timeout set from a measured percentile, with connection setup inside it or done before taking traffic?
- Does every hop send the deadline of the attempt it is making, cancel what it abandons, and drop expired work at dequeue?
- Is exactly one layer retrying each dependency, and have you counted the proxies, SDKs, CDNs and managed runtimes that also retry?
- Does every retried side effect carry the same idempotency key, and is a timeout treated as an unknown?
- Do all retries use jitter, and does periodic work use stable offsets?
- Is there a retry budget, and is any fixed per-client allowance sized for the whole fleet?
- Can each dependency refuse work before its expensive part, with
503andRetry-After, and do callers honour it? - Do priority and age limits sit in the queue where requests actually wait?
- Is each breaker keyed by operation, with a probe timeout above the healthy P99?
- Does every dependency have a cap sized by rate × P99 × headroom, and is it checked against the dependency's capacity ÷ the fleet?
- Does every optional call have a tested fallback, and are health checks and emergency paths exempt from shedding?
- Do alarms and scaling use offered load, shed counts, breaker opens and queue age, not only CPU and median latency?
Think-first drills
Drill 1. A request crosses 4 layers, each making 3 attempts, to a dependency that fails every call. How many calls reach it per user action? What if only the layer right above it retries (2 attempts) and the others pass "don't retry"? And with a 10% retry budget at that layer?
Drill 2. A dependency's healthy P50 is 8 ms, its P99 20 ms and its P99.9 40 ms. Your service handles 900 requests a second per instance, each calling it once. Pick a timeout for at most 0.1% false timeouts, and size a bulkhead by rate × P99 × 2. How many threads would a hang consume with a 5 s timeout and no bulkhead?
Drill 3. 300 devices sync every 60 s, 20 reads each; the database has 500 reads a second to spare. All start at :00. How long does the burst queue each minute, and what does a stable offset give? With a fresh random offset each minute, what's the longest gap between two syncs of one device?
Interview questions
| Question | Model answer |
|---|---|
| How do you choose a timeout, and why not 50 ms everywhere? | From the dependency's measured latency: choose the false-timeout rate you accept (say 0.1%), read that percentile (P99.9), pad it when it is close to the median, and include connection setup and TLS in the timer or connect before taking traffic. 50 ms everywhere fails healthy calls wherever the tail is longer, and every false timeout becomes a retry; a long timeout holds threads (Little's law). "Plainly down" is a different signal, for a breaker or a concurrency limit. Each hop sends its attempt's deadline down and cancels what it abandons. |
| One optional dependency slowed and the whole service went down. Explain it and fix it. | Little's law: threads in use = rate × time held. A call that goes from 30 ms to 5 s needs 167 times the threads; in one shared pool it takes them all, so requests that never call it wait too. Fix: a timeout from its latency, a per-dependency cap sized by rate × P99 × headroom, a breaker per operation, and a fallback that drops the optional part. The cap alone would have kept the rest of the pool free. |
| Retries made an outage worse. Design the retry policy. | Retry only transient errors and throttles, never permanent 4xx; side effects only with the same idempotency key, and treat a timeout as an unknown. Retry in one layer, right above the failure, and pass "overloaded, don't retry" up; count hidden layers (proxies, SDKs, CDNs). Capped exponential backoff with full jitter; honour Retry-After. A retry budget of about 10% per client, so a total failure sees about 1.1×; hedges draw from the same budget. |
| Walk through a circuit breaker's states. What can go wrong in half-open, and on a dependency that is slow rather than dead? | Closed: count outcomes over recent calls (a count window trips fast; a long time window holds healthy history and trips slowly); open when the failure rate passes the threshold. Open: refuse at once with a fallback or a fast error. After the open time, half-open: let a few probes through; successes close it, a failure reopens it. The probe's timeout must exceed the healthy P99, or it never closes (and a probe may be a real write: resolve it like any unknown). On a slow-but-alive dependency the light probes pass, the full load returns and it opens again: it flaps. Key breakers by operation, so a degradable read can't block a required write. |
| Load shedding, rate limiting and backpressure: which, when? | Shedding protects a service from overload whoever causes it: it refuses by its own health (in flight, queue age), early and cheaply, with 503 and Retry-After, and callers don't retry it in that layer. Rate limiting protects it from one client's excess, by that client's quota, with 429. Backpressure lets a slow consumer set the pace of a faster producer (pull, credits, bounded buffers). Lowering rate limits doesn't fix a slow backend; only shedding by health does. |
| The trigger is gone but the system is still down. What's happening, and how do you get out? | A metastable failure: the trigger pushed a vulnerable system over, and a sustaining loop keeps it there: retries that keep load above capacity, a queue of stale work served for callers who left, or a cache that can't refill. To get out, bring load below capacity and throw the stale work away (drop by deadline, flush queues, restart); shedding new arrivals alone isn't enough while the backlog stays. Adding capacity is too slow and adds callers. To stay out, give each layer a guard that breaks the loop alone: deadlines that drop stale work, a dependency that refuses early, caps on what each dependency may hold. |
Where to go next
- Idempotency & Effectively-Once Processing: keys, the retry horizon and resolving unknowns, which Part 3 only recaps (Parts 1 and 4 there).
- Sharding, Hot Keys & Rebalancing: per-key admission with
429andRetry-After, single-flight and coalescing (Parts 3 and 4 there); this page handles a whole dependency, that one a single hot key. - Multi-Region Failover: the Region-level version: client backoff after a routing flip (Part 5 there) and cells as a blast-radius unit (Part 9 there).
- Leases, Fencing Tokens & Distributed Locks: timeouts that protect safety rather than capacity (Part 3 there).
- Change Streams & the Transactional Outbox: consumer lag is a backlog with a drain time.
- Replication, Quorums & Read-Your-Writes: the slow replica that hedging works around.
- Caching & Invalidation: stampedes (Part 5 there) and the cold cache that can keep a database down (Part 8 there), the third loop of Part 10.
- Coming later: Queues & Delivery Semantics (visibility timeouts, redelivery, dead-letter queues) and Rate Limiting Algorithms (token buckets, sliding windows, distributed counters).
- Drill: The Slow Recommendation Service That Took Down Checkout, both questions answered in Parts 7 and 8.
- Background: Circuit Breaker, Bulkhead & Fault Tolerance Patterns, the broader primitive.
- Loops that rely on this page: the key-value store (R2.8 and R2.9), the rate limiter (steps 1.4, 1.5, 2.4, 3.4), the news feed (step 2.3), chat (steps 1.6, 2.5), notifications (steps 1.2, 2.1, 2.4, 3.3), the S3-like object store (steps 2.4, 3.6), search autocomplete (step 2.4), the web crawler (step 3.4), the job scheduler (steps 1.4, 2.3, 2.6), payments (steps 1.2, 2.4), the stock exchange (step 2.7), hotel reservations (step 2.2), ride-sharing (step 3.5), the offline-first news feed (step 2.5 and R2.8), the mobile stock-trading app (steps 2.6, 3.2), the paging library (R2.8), and the Netflix (steps 2.3, 2.5), Uber (step 1.3) and Shopify (steps 2.3, 3.1, 3.4) case studies.