Designing Fault-Tolerant Distributed Systems Cheat Sheet
Comprehensive resiliency cheat sheet covering Circuit Breakers, Bulkheads, Retries with Jitter, Dead-Letter Queues, and Graceful Degradation.
Comprehensive resiliency cheat sheet covering Circuit Breakers, Bulkheads, Retries with Jitter, Dead-Letter Queues, and Graceful Degradation.
1. The 6 Resiliency Pillars
Synthesizing vector architecture diagram...
2. Resiliency Patterns Cheat Sheet
Pattern 1: Circuit Breaker State Machine
Synthesizing vector architecture diagram...
Every breaker starts Closed: calls pass through normally while failures are counted. When the failure rate crosses the threshold (for example more than 50% over 10 s), it moves to Open, and every call fails immediately with a fallback, so a struggling downstream service gets no load at all and your threads are not stuck waiting on timeouts. After the reset timeout (for example 30 s), it moves to Half-Open and lets a few trial calls through: if they succeed it returns to Closed, and if any fails it goes straight back to Open for another full timeout. Half-Open is what makes recovery automatic, and it must let only a few calls through, otherwise the backlog floods a service that has only just recovered.
Pattern 2: Bulkhead Isolation
Isolate resources into discrete pools so that a failure in one non-critical domain cannot starve thread pools or memory in critical domains:
- Connection Pool Bulkhead: Separate database connection pools for read-heavy catalog searches vs. payment checkouts.
- Thread Pool Bulkhead: Isolate third-party webhook dispatchers from user-facing login threads.
Pattern 3: Exponential Backoff with Full Jitter
Never retry immediately on network failures; synchronized retries cause a "thundering herd" or "retry storm" that prevents struggling services from recovering.
pseudocodesleep_duration = random_between(0, min(max_backoff, base * 2 ^ attempt_count))
This formula decides how long a client waits after a failed attempt before it tries again. attempt_count is the number of attempts that have failed so far (1 after the first failure). base is the starting wait in milliseconds, chosen by the team that configures the client. base * 2 ^ attempt_count doubles the upper bound of the wait after every failure, so a service that is still failing receives retries less and less often. max_backoff is the largest wait in milliseconds the team allows, and min(...) keeps the upper bound from growing past it. random_between(0, upper bound) picks a random number of milliseconds between 0 and that bound; this random part is the "full jitter", and it spreads the retries of many clients that failed at the same moment over the whole window instead of sending them all at once. The result, sleep_duration, is the time in milliseconds the client sleeps before its next attempt; the formula changes nothing else.
Synthesizing vector architecture diagram...
What to notice: the diagram follows one request through three failed attempts, with base = 50 ms chosen as the example value and max_backoff large enough that it does not cap these three waits. After attempt 1 fails, attempt_count is 1, so the upper bound is 50 × 2^1 = 100 ms and the client sleeps a random time between 0 and 100 ms. After attempt 2 fails, the bound doubles to 50 × 2^2 = 200 ms, and after attempt 3 fails it doubles again to 50 × 2^3 = 400 ms. Attempt 4 either succeeds or, if the retry limit is reached, the request goes to a dead-letter queue (DLQ), a separate queue that holds requests that could not be processed. Because every client picks its own random sleep inside the window, clients that failed at the same moment retry at different moments instead of all at once.
3. Graceful Degradation Strategies
| Component Failure | Naive Failure Behavior | Graceful Degradation Strategy |
|---|---|---|
| Personalized Recommendations Down | Entire home page crashes with HTTP 500 error. | Fallback to cached static top-10 popular items list. |
| Real-Time Inventory Lock Service Down | User checkout fails immediately. | Accept order provisionally into an async reconciliation queue. |
| Search Autocomplete Inverted Index Down | Search bar input errors out. | Fallback to simple SQL LIKE 'query%' or client-side history. |
| Image Transcoding Worker Queue Jammed | User upload rejected. | Store raw original in S3; show placeholder while queue drains. |
4. Health Check Probing Matrix
| Probe Type | Frequency | Purpose | Action on Consecutive Failures |
|---|---|---|---|
| Liveness Probe | Every 5-10s | Checks if process is deadlocked or stuck in infinite loop. | Restart / kill container immediately. |
| Readiness Probe | Every 2-5s | Checks if service is warm and capable of accepting traffic. | Remove instance from Load Balancer target group. |
| Startup Probe | On Boot | Checks if slow initialization (JVM warmup, cache pre-fill) done. | Pause liveness/readiness evaluation until complete. |