Multi-Region Failover
The Booking That Existed in Frankfurt but Not in Virginia
Test your architecture intuition: Pitch a 7-axis solution, survive two aggressive reviewer objections, and inspect the staff-level Teacher Gold Answer.
Part 0. Start here
The problem: Frankfurt goes dark at 20:00
Seatly sells event tickets. It is Friday, 20:00, the busiest hour of the week. Frankfurt (eu-central-1) stops answering: the front doors, the app servers, the Aurora writer, everything. Frankfurt serves the EU checkout, which is 40% of Seatly's customers (our example's number). The status page promises "automatic failover".
Here is the setup Seatly had last year, the naive one:
- a Route 53 failover record for
api.seatly.eu, with health checks at the defaults: a check every 30 s, and 3 failed checks in a row to mark the endpoint unhealthy; - a DNS time to live (TTL) of 300 s (our number);
- a warm standby in Ireland (
eu-west-1) with 15 instances, which scales out by launching new ones; - an Aurora Global Database secondary in Ireland that nobody has promoted. It is a read-only copy until someone runs a failover.
Two AWS facts shape what happens next. An Aurora Global Database failover loses what replication hadn't shipped yet, which AWS says is "typically measured in seconds", and it takes minutes. And it starts only when someone runs it: nothing promotes the secondary on its own.
The budget: Seatly promises 99.99% availability. A 30-day month has 30 × 1,440 = 43,200 minutes, and 0.01% of that is 43,200 × 0.0001 = 4.32 minutes of allowed downtime a month.
What the naive setup does that night:
- At about 20:01:30 (3 checks × 30 s = 90 s, "about" because many checkers report), Route 53 marks Frankfurt unhealthy and starts answering with Ireland's address. Clients follow as their cached answers expire, within the 300 s TTL.
- By about 20:06:30 (6.5 minutes in), EU customers reach Ireland, where every write fails: the database there is still a read-only secondary.
- On-call notices at about 20:10 (our number) and runs the failover; Ireland takes writes at about 20:12. Meanwhile 15 instances carry the load of 60, 4× their capacity (60 ÷ 15). Auto Scaling launched new instances at about 20:02; Seatly's own measured launch-to-healthy time is about 23 minutes, so they are healthy at about 20:25.
- At about 20:14, Frankfurt's power returns for a moment. Clients still holding Frankfurt's address write there, and nothing tells Frankfurt it is no longer the writer: two writers now accept orders. On-call shuts Frankfurt down by 20:20.
EU customers can't reliably buy tickets until the new capacity is healthy at about 20:25: about 25 minutes. Weighted by the share of customers hit, 25 × 0.4 = 10 customer-weighted minutes, which is 10 ÷ 4.32 = 2.3 months of budget in one evening.
One more trap hides in the "fix" someone suggests afterwards: make Ireland's health check test a write, so DNS never sends users to a read-only Region. Now both records are unhealthy until promotion. When every record in the group is unhealthy, Route 53 treats them all as healthy (it fails open), so it answers with the primary record: Frankfurt's dead address.
Route 53 says automatic failover is on. Frankfurt goes dark at 20:00. List everything that must be true before an EU customer can buy a ticket in Ireland, and who decides each one.
The big picture
Synthesizing vector architecture diagram...
What to notice: Frankfurt is the home and the only writer. Ireland runs a small fleet and keeps a read-only copy that trails by 0.8 s. The probes that judge Frankfurt run in Ireland and Stockholm, never inside Frankfurt. Route 53's answer is decided by routing controls, not by editing records.
What you'll be able to do after this page
- Name the five decisions of a Region failover in order, say who owns each, and compute the orders at risk from the commit rate and the lag (Part 1).
- Decide that a Region is gone from outside it, from two places, with a business metric, and tell a network split from an outage (Part 2).
- Build the decision gate: what it checks, when a person must approve, and how ARC safety rules enforce it (Part 3).
- Fence the old writer and promote the new one, and say why Aurora's own fence is not enough (Part 4).
- Move users with pre-created routing controls, and add up the DNS tail and the recovery time (Part 5).
- Size the survivor for the moved load, check its quota and residency, and pick a rung on the disaster-recovery ladder on equal terms (Part 6).
- Count and resume the writes caught in the middle, and answer the double-booking drill (Part 7).
- Bring the old Region back as a replica, reconcile its tail, and fail back on purpose (Part 8).
- Keep most failures smaller than a Region with cells and shuffle sharding (Part 9).
- Rehearse failovers and spend an error budget in minutes (Part 10).
- Trace one checkout through every layer after the move (Part 11).
- Map all of it to AWS, and name the look-alikes that are not Region failover (Part 12).
You may have arrived from a step that relies on this: step 3.3 of the Netflix case study (move all traffic out of a Region, fast), step 3.4 of the Uber case study (a whole site goes down mid-trip), step 3.4 of the Google Drive loop (EU data stays in the EU, and a Region can fail), or step 3.5 of the payment loop (data must stay in-country, and a Region can fail). This page is the "why" behind all four, and behind about thirty other loop steps that move a Region.
Part 1. Five decisions, one home per key, and two numbers
A Region is gone. What has to happen, and in what order? Teams that skip the order get the naive night: users moved before the database can take writes, and an old writer nobody stopped. This Part names the order, the one design rule that makes a move possible at all, and the two numbers every failover is judged by.
Five decisions, in order
Synthesizing vector architecture diagram...
What to notice: every arrow has an owner and evidence. Only one arrow before Moved needs a person: the approval from Gated to Promoted; the clean-up arrows after it are operators' planned work. "Promoted" includes the lease takeover and the new epoch (Part 4); "Moved" is only the routing switch (Part 5).
| Decision | The question | Where on this page |
|---|---|---|
| 1. Is it the Region? | Is Frankfurt really gone, or is only our view of it broken? | Part 2 |
| 2. Do we move? | Is the move safe and legal, and is it worth what it loses? | Part 3 and Part 6 |
| 3. Is the standby safe to write? | Has the old writer stopped, and will anything it sends later be refused? | Part 4 |
| 4. Move the users | How do clients reach Ireland without the dead Region's help? | Part 5 |
| 5. Clean up | What happens to the orders caught in the middle, and to Frankfurt when it returns? | Parts 7 and 8 |
One home per key
The design rule that makes a Region move safe is one home per key. Every event Seatly sells has exactly one home Region, and only the home writes that event's seats and orders. US events are homed in us-east-1 and EU events in Frankfurt. A failover moves a home: Frankfurt's events get Ireland as their home, with a new epoch, and nothing else changes owner.
With one home per key, "who may sell seat 14C?" always has one answer. Without it (two Regions writing the same event), every seat is exposed to the window in which the two Regions disagree, all the time, not only during failover (Part 7). The home is recorded in one place a split can't move (Part 4), and a write for an event whose home is elsewhere is refused or forwarded, never quietly accepted. Sharding, Hot Keys & Rebalancing (Part 8 there) applies the same "moved" refusal to a key whose shard has moved; this page applies it to a whole Region.
Writes from far away pay for this: a customer in Madrid buying a ticket for a Frankfurt-homed event writes to Frankfurt, wherever the customer is.
RPO and RTO
Every failover is judged by two numbers:
- RPO (recovery point objective): how much acknowledged data you can lose. It is set by replication: whatever the standby hasn't received when the home dies is lost to the standby.
- RTO (recovery time objective): how long customers can't use the service. It is set by the timing chain of the five decisions (Part 5 adds it up).
Our first formula counts the data at risk:
For Seatly at 20:00: 200 orders a second × 0.8 s = 160 acknowledged orders that exist only in Frankfurt. This is the same rule the Replication, Quorums & Read-Your-Writes page derives for one database; this page takes what cross-Region replication promises from there and adds only what a Region move needs.
The lag is not a constant. When the link between the Regions breaks, the lag grows by one second every second, and so does the data at risk (Part 2 shows it). The lease the home holds and the epoch it writes with come from Leases, Fencing Tokens & Distributed Locks (Part 7 there); Part 4 here applies them to a whole Region.
Words we use
| Word | Meaning here |
|---|---|
| Home Region | The one Region allowed to write a key (an event, its seats and its orders) |
| Standby | A Region ready to become the home; warm means it runs a small fleet and can grow |
| Vantage point | A place that probes the home from outside: here, Ireland and Stockholm |
| CPS | Checkout starts per second, the business metric each Region reports every 10 s |
| Lease | The home's right to write, renewed every 2 s in a table stored in three Regions |
| Epoch | A number that only grows, bumped on every change of home; the new home's database checks it on every write |
| Gate | The decision to move: six checks, then an approval |
| Routing control | An on/off switch in ARC (Amazon Application Recovery Controller) that decides which Region's DNS records are answered |
| Lost tail | Writes the home acknowledged that the standby never received |
| Unknown | A request that was in flight with no reply when the home died |
| Failover / switchover | Promote a standby when the home is gone and lose the lag / move the home on purpose when both are healthy and lose nothing |
| Failback | Moving the home back to the original Region, later, on purpose |
The example we follow
A real company moves dozens of services between many Regions; that is too much to watch. So one small story runs through the whole page: Frankfurt goes dark, step by step. Every trace on this page was produced by running a private reference simulator of this setup (Regions, links that can be cut, canaries, a three-replica control table, the lease, a lagging database, routing controls, DNS caches, capacity and quota, clients and queues), not worked out by hand.
| Setting | Our example | At real scale |
|---|---|---|
| Regions | us-east-1 (home of US events); Frankfurt (home of EU events); Ireland (warm standby for Frankfurt); Stockholm (third copy of the control table, second set of canaries, no customer data); us-west-2 (target of the US backups) | Netflix runs three active Regions; the loops pair Regions inside one jurisdiction |
| Where EU data may live | EU events' personal data (orders, carts, sessions, seat maps, tickets) is stored and processed only in EU Regions. This is Seatly's business rule, not a legal claim | The notification loop: "EU never moves to the US" |
| EU load | 40% of Seatly's traffic; 200 order writes a second at the Friday peak | |
| Orders and seat holds | Aurora PostgreSQL Global Database: writer in Frankfurt, secondary in Ireland; replication lag 0.8 s at the moment of failure. The rds.global_db_rpo parameter left at its default | AWS: RPO "typically measured in seconds" |
| Carts and sessions | DynamoDB global table in multi-Region eventual consistency (MREC) mode, Frankfurt + Ireland | AWS: replication "typically within a second or less" |
| Seat maps and ticket PDFs | S3 buckets in Frankfurt and Ireland, Cross-Region Replication with Replication Time Control, behind a Multi-Region Access Point in active-passive mode (Frankfurt active) | |
| Control table | A DynamoDB global table in multi-Region strong consistency (MRSC) mode with three full replicas: Frankfurt, Ireland, Stockholm (all three are Regions MRSC supports; full replicas, not a witness, so Stockholm can read and write). It holds the item eu-writer = {home, epoch, ack_until, rvn} and each Region's CPS reports. No customer data | MRSC uses exactly three Regions |
| Writer lease | Frankfurt renews every 2 s, writing ack_until = its wall-clock send time + 10 s. It stops writing at its monotonic send reading + 10 s. A taker's condition: epoch = :seen AND ack_until < :taker_wall_now − 5 s; the 5 s margin must exceed the worst gap between Frankfurt's and Ireland's clocks (so each host stays within 2.5 s of true time) | The payment loop: a 10 s lease and a 5 s margin |
| Epoch | A writer_epoch row inside the Aurora database, read in every write transaction. Frankfurt holds epoch 7 | The outbox loop |
| Detection | Canaries in Ireland and Stockholm, attached to a VPC, probe Frankfurt's public checkout path every 10 s; a vantage point fails after 3 misses in a row. Each Region writes its CPS every 10 s to the control table; "down" means below its normal band, or missing, for 60 s. The rule runs separately in Ireland and in Stockholm. Nothing on the EU path reads us-east-1 | Netflix: stream starts per second, every 10 s |
| Gate | A person approves in 2 min, or pre-approved rules approve in 10 s. Above 1,000 orders at risk, the business owner must also sign | The Uber case study |
| Routing | An ARC cluster. Routing controls eu-central-1-on and eu-west-1-on sit behind the Route 53 failover records for api.seatly.eu. TTL 60 s; the app's retry backoff is capped at 30 s | |
| Capacity | Frankfurt: 60 instances of 8 vCPUs, 20 per Availability Zone (AZ), so losing an AZ leaves 40, the peak need. Ireland: 15 running plus a warm pool of 45 stopped instances, already initialized: 120 vCPUs running before the failover | Netflix: 402 Zuul instances, 102 of them dark |
| Ireland's vCPU quota | 640 vCPUs of running On-Demand instances (raised from 512 after last month's game day) | Quotas are per Region |
| Retry surge | Load × 1.5 for 5 minutes after users move | |
| Ireland's prerequisites | Encryption keys, secrets, TLS certificates and container images already present in Ireland | The KV loop: a key that lived only in the dead Region |
| App order journal | The app keeps each order's key and body until it has fetched the ticket; on 404 unknown order it re-sends under the same key | |
| Cells | 4 cells per EU Region; deploys go one cell at a time with a 1 h bake | The KV and message-queue loops |
| Client timeout | 2 s; retries reuse the idempotency key |
The story in eleven beats: a split that isn't an outage (Part 2); the gate (Part 3); the fence and the promotion (Part 4); the flip (Part 5); where the load can go, and the ladder (Part 6); the orders caught in the middle (Part 7); Frankfurt comes back (Part 8); the bad deploy that stayed in one cell (Part 9); the drill and the budget (Part 10); one checkout through every layer (Part 11); and on AWS (Part 12). Each Part shows only its own slice of events; the full table is in Part 13.
Snapshot R1: normal Friday, 18:00
| Frankfurt | Ireland | Stockholm | |
|---|---|---|---|
| State | Home for EU events, serving 200 orders/s | Warm standby, serving canaries only | Control-table replica, canaries |
| Lease as seen there | home Frankfurt, epoch 7, ack_until renewed every 2 s | the same (strongly consistent) | the same |
writer_epoch in Aurora | 7 (writer) | 7 (read-only secondary, 0.8 s behind) | no database |
| Routing controls | eu-central-1-on On; eu-central-1-writable On (it must stay On through the 20:05 flip, or the mirror rule refuses turning eu-central-1-on Off) | eu-west-1-on Off; eu-west-1-writable Off | |
| Running vCPUs | 480 (60 × 8) | 120 of 640 quota (15 × 8); 45 stopped in the warm pool | |
| Stranded work | none | none | none |
What to remember from Part 1
- Failover is five decisions: is it the Region, do we move, is the standby safe, move the users, clean up.
- Each key has one home; failover moves homes, with a new epoch.
- RPO is set by replication, RTO by how long each decision takes.
Part 2. Is it the Region? Detecting from outside
A Region can't judge itself: when it is gone, its own alarms are gone with it. And a Region that looks dead from one place may be perfectly healthy. The first decision is therefore made by others, from more than one place, and it must tell a broken link from a dead Region. At 18:30, two hours before the real outage, Seatly gets exactly that test.
Trace: the link breaks at 18:30
| # | Time | Event | What the rule sees |
|---|---|---|---|
| 2 | 18:30:00 | The network link between Ireland and Frankfurt breaks. Replication to Ireland stops, and Ireland's lag grows by a second every second | The 18:30:00 canary run succeeded just before the break |
| 2 | 18:30:10, :20, :30 | Ireland's canaries miss three runs in a row: Ireland's vantage point fails at 18:30:30 | One of two vantage points failed |
| 3 | 18:30:30 | The rule checks the other signals: Stockholm's canaries still reach Frankfurt; Frankfurt's CPS report (read from the control table) is in band and fresh; Frankfurt still renews its lease (the last renewal, sent 18:30:29, set ack_until 18:30:39) through Stockholm, 2 of 3 replicas. Route 53's checkers still see Frankfurt healthy, since far more than 18% of them reach it, so DNS doesn't move either | No evacuation. A P2 alert for EU replication lag goes out |
| 3 | 18:31:00 | (A test in the simulator) An Ireland takeover attempted anyway would fail its condition: ack_until is 18:31:09, still in the future | The lease refuses too |
| 4 | 18:50:00 | Link restored; the lag is back to 0.8 s by 18:52 (our number) | Everything green |
Snapshot R2: the split
Synthesizing vector architecture diagram...
What to notice: only the Ireland–Frankfurt link is cut. Frankfurt still reaches Stockholm, so its lease stays fresh and its CPS keeps arriving. Ireland still reaches Stockholm, so it reads both. One failed vantage point plus a healthy business metric is a split, not an outage.
Ireland's canaries say Frankfurt is dead. What else must be true before we move, and what would a split look like?
Vantage points and a business metric
- Vantage points outside the Region. Canaries run in Ireland and in Stockholm, each attached to a VPC and probing Frankfurt's public checkout path every 10 s. Optionally, a third probe runs outside AWS; it is the only one that doesn't share the AWS backbone with replication.
- A business metric. Each Region writes its checkout starts per second (CPS) to the control table every 10 s. Frankfurt's CPS is "down" if it is below its normal band, or missing, for 60 s. The metric catches what probes miss: a front door that answers
200while checkouts fail. - Two independent deciders. The rule runs separately in Ireland and in Stockholm, each reading the control table. Both fire; every action they start is safe to repeat (setting a desired capacity of 60 twice is the same as once).
- Nothing on the EU path reads
us-east-1. Not the probes, not the metric, not the rule.
The rule has a second clause for a home that has fenced itself: Frankfurt lost the control table (its renewals fail), so it stopped writing at its lease stop time, yet its front door still answers and the canaries stay green. The self-fenced clause catches that case: Frankfurt's lease has lapsed and its CPS is down. The simulator's run of that case: Frankfurt's control-table replica fails at 19:30:00, its last renewal (sent 19:29:59) makes it stop at 19:30:09, its CPS report goes missing, and the rule fires at 19:31:00 with the canary clause false and the self-fenced clause true.
textEVACUATION RULE (runs every 10 s, in Ireland and in Stockholm, independently) canaries_failed = vantage(Ireland).misses >= 3 AND vantage(Stockholm).misses >= 3 cps_down = Frankfurt CPS below band, or no report, for >= 60 s self_fenced = lease.home = Frankfurt AND lease.ack_until < my wall clock if (canaries_failed AND cps_down) OR (self_fenced AND cps_down): page on-call; start safe-to-undo work (capacity); open the gate else if one vantage failed: alarm "possible split", never move
What a split does to the data at risk
A split doesn't move anyone, but it quietly removes the RPO promise. Suppose Frankfurt had died at 18:49:59, one second before the link came back. Ireland would have been 0.8 + 1,199 = 1,199.8 s behind: 200 × 1,199.8 = 239,960 acknowledged orders (about 240,000) that exist only in Frankfurt.
Aurora has a knob for this. rds.global_db_rpo (minimum 20 s) makes the primary block commits once every secondary is further behind than that. In the simulator, the lag reaches 20 s at 18:30:19.2, so Frankfurt stops taking orders for the remaining 19.7 minutes of the split: a write outage in a healthy Region, to protect a copy nobody is using. AWS also warns that with only two Regions, the parameter's default should stay in the secondary's parameter group; otherwise, after a failover, the new primary can find itself with no caught-up secondary and block its own commits. Seatly keeps the default and alarms on the lag instead: an RPO/availability trade made on purpose.
Route 53 health checks, precisely
Route 53 health checks still matter: they are how routing controls reach DNS (Part 5). But as a detector they need care.
| Fact | Value | What it means for a Region decision |
|---|---|---|
| Check interval | 30 s by default, or 10 s (a "fast" check, extra cost) | |
| Failure threshold | 3 by default | About 90 s with the defaults, "about" because many checkers report independently |
| Healthy rule | Healthy if more than 18% of the checkers say so | A split that leaves most checkers seeing Frankfurt keeps it healthy, which is right |
| All unhealthy | When every record in the group is unhealthy, Route 53 treats all of them as healthy: it fails open | A standby that fails its check by design makes DNS answer with the dead primary (Part 0) |
| Metrics | Health-check metrics in CloudWatch are available only in US East (N. Virginia) | Dashboards and alarms built on them depend on us-east-1 |
On the night itself (event 10x), Route 53's check for Frankfurt went unhealthy at about 20:01:30, and its metrics sat in us-east-1. Seatly shows it on a dashboard as advice only: the rule never waits for it and never reads it.
Hosts and Regions. Inside a cluster, failure detectors such as SWIM's "suspect, then let others refute" (which the Uber case study describes) or Cassandra's phi-accrual detector (default threshold 8) apply the same idea to machines: don't convict on one observer. This page applies it one level up, to a whole Region.
Trace: detection at 20:00
| # | Time | Event |
|---|---|---|
| 7 | 20:00:00.0 | Frankfurt goes dark. The 20:00:00 canary run and CPS report had just succeeded |
| 9 | 20:00:10, :20, :30 | Canaries in Ireland and Stockholm miss three runs in a row: both vantage points failed at 20:00:30 |
| 10 | 20:01:00 | Frankfurt's CPS report has been missing for 60 s (the last one arrived at 20:00:00). The rule fires in Ireland and in Stockholm at the same tick. Automation pages on-call and starts capacity at once (Part 6) |
| 10x | about 20:01:30 | Route 53's health check for Frankfurt goes unhealthy; advice only |
What to remember from Part 2
- Decide from outside, from at least two places, with what customers experience.
- A split breaks one path; a dead Region fails every path and its own metric.
- Anything read through the failing Region, or through
us-east-1, is advice.
Part 3. Do we move? The decision gate
Knowing that Frankfurt is gone doesn't mean moving is safe. Moving to a Region without capacity, without quota, or into a country the data may not enter makes things worse, and moving while the old writer may still write corrupts data. The second decision is a gate: automation gathers the evidence, and one approval opens it.
Trace: 20:01, the checklist
At 20:01:00 the rule has fired, and automation computes six checks.
Snapshot R3: the gate at 20:01
| Check | Question | Value at 20:01 | Result |
|---|---|---|---|
| (a) Region evidence | Is it the Region? | Both vantage points failed (20:00:30); CPS missing 60 s | Pass |
| (b) Target ready | Is capacity starting, is there quota, and does Ireland have what it needs to run? | Warm pool starting, InService by 20:02 (480 of 640 vCPUs); 600 needed with the surge, and the last 15 are new launches, healthy only at 20:24 (Part 6); keys, secrets, certificates and images present | Pass, with a known 5-minute headroom gap |
| (c) Residency | Is the target map legal? | The default "spread to spare capacity" map proposes Ireland 70% / us-east-1 30%. The us-east-1 share is refused (EU data stays in the EU), so the target is Ireland 100% | Pass, after the refusal |
| (d) The old writer's lease | Has Frankfurt's lease run out, plus the margin, on Ireland's clock? | ack_until 20:00:09.0 + 5 s = 20:00:14.0, long past | Pass |
| (e) Someone stays on | Will at least one EU routing control be On after the move? | The flip turns eu-west-1-on On as it turns eu-central-1-on Off | Pass |
| (f) Data at risk | How many acknowledged orders would the move give up? | Ireland's last AuroraGlobalDBRPOLag reading before the failure, 0.8 s, × 200/s = 160, below the 1,000 threshold, so the business owner doesn't need to sign | Pass |
Check (d) is also the gate's lower bound: no approval can take effect before ack_until + 5 s, whatever the human or the rules say. Here the bound (20:00:14) is hidden under detection (20:01:00); it matters when detection is fast.
| # | Time | Event |
|---|---|---|
| 11 | 20:01:00 → 20:02:00 | Capacity starts before approval, because it is safe to undo: the warm pool comes InService by 20:02 (Part 6) |
| 12 | 20:01:00 | The six checks above, all green after the residency refusal |
| 13 | 20:03:00 | On-call approves, 2 minutes after the page. Automation also flips the S3 access point at once (event 17b, Part 5) |
| 13p | (side row) 20:01:10 | With pre-approved rules (all six checks green, so the rules approve on their own in 10 s), approval comes at 20:01:10 |
Automation is 99% sure Frankfurt is dead. Why does a person still press the button, and when shouldn't they have to?
textGATE (automation; the approval is the only human step) on rule fired: start capacity (warm pool) -- safe to undo, start now checks = [region_evidence, target_ready, residency_map, lease_expired_plus_margin, one_control_stays_on, data_at_risk] if data_at_risk > 1000 orders: also require business owner approve when all checks pass AND (on-call approves OR pre-approved rules match) never before lease.ack_until + margin (on the taker's clock) then: take the lease, promote, bump the epoch (Part 4); then flip (Part 5)
Safety rules in the switch itself
ARC's routing controls come with safety rules that the cluster enforces on every change, so a wrong button press is refused by the switch, not by a runbook:
| Rule type | What it enforces | Seatly's rule |
|---|---|---|
| Assertion rule | A condition on the controls' states that must hold after any change | At least one of eu-central-1-on and eu-west-1-on is On, so no change can leave EU customers with no Region |
| Gating rule | Target controls may change only while gating controls meet a condition | eu-west-1-on may change only while at least one of {eu-west-1-writable, eu-west-1-read-only-approved} is On. The mirror rule: eu-central-1-on may change only while eu-central-1-writable is On |
The two gating controls are Seatly's design choice. Automation turns eu-west-1-writable On right after Ireland's database takes its first write at the new epoch (Part 4). The approver (on-call, or a pre-approved rule written for this case) turns eu-west-1-read-only-approved On only when choosing to move users into an explicitly read-only Ireland (Part 5). Until one of them is On, nobody can send EU users to Ireland, not a script, not a tired engineer. A gating rule blocks every change to its target, so a gating control must stay On until its target has moved, in either direction: eu-central-1-writable stays On until the 20:05 flip has turned eu-central-1-on Off, and only then goes Off.
When a rule blocks a move that must happen (a gating control stuck because its automation is down), ARC lets an update call override named safety rules ("Overriding safety rules to reroute traffic": the update call lists the rules to bypass). That is the break-glass for a stuck gate, used only by a written runbook, because it also bypasses the protection the rule exists for.
Why a move to a read-only Region is gated
The naive night sent users to Ireland 6.5 minutes in, where every write failed. A DNS move is only useful if the target can do what users came for. So there are exactly two legal orders:
- Promote and fence first, then flip. The main line.
- Flip into an explicitly approved read-only mode, with writes answering
503 Retry-After, and promote later. Part 5 prices it.
What is never legal is a health check moving users on its own to a Region that still refuses writes. That is why Seatly's DNS answers follow routing controls rather than Frankfurt's own health.
What to remember from Part 3
- Automate the evidence and the steps; gate the decision.
- The gate checks capacity, quota, residency, the old writer's lease, someone staying on, and the data at risk.
- Start what's safe to undo before approval.
Part 4. Fencing the old writer, promoting the new one
"Frankfurt is dark, so it can't write" is an assumption, not a fact. A dark Region is often a Region cut off from everyone else, with some app servers still running inside it. If Ireland starts writing while one of them still writes, two homes sell the same seats. The third decision makes the standby safe: the old home stops on its own, the new home waits until that stop is certain, and a number in the new database refuses anything older.
Recap: page 03's handover, at Region scale
Leases, Fencing Tokens & Distributed Locks (Part 7, "Single writers across a failover") gives the recipe: the writer holds a lease; the old writer stops when it can't renew, by its own monotonic clock; the new writer waits out the lease plus a margin; the new writer bumps the epoch in the store it will write before its first write; consumers compare (epoch, version). Page 03 also stops the holder 2 s before its lease ends (a holder margin). This page uses no holder margin: Frankfurt stops exactly at its send time + 10 s, and all the safety margin sits on the taker's side (5 s), which must be larger than the worst gap between Frankfurt's and Ireland's clocks (page 03, Part 5). What's new at Region scale is where the lease lives, whose clock reads it, and the fact that the old home has its own copy of the database.
A lease a split can't move
The lease item eu-writer = {home, epoch, ack_until, rvn} lives in the MRSC control table, with full replicas in Frankfurt, Ireland and Stockholm. A write to an MRSC table succeeds only while its Region can reach at least one other replica. AWS: an isolated Region "can only service eventually consistent reads". So:
- A split between Frankfurt and Ireland leaves Frankfurt with Stockholm: 2 of 3. Frankfurt keeps renewing, and any takeover from Ireland fails its condition (Part 2, event 3).
- A Frankfurt cut off from both can't renew. Its renewals fail, and it stops at its own stop time.
- Stockholm is there only to be the third copy and the second vantage point. It holds no customer data, which also keeps the residency map simple.
Page 03 (Part 8, "Across Regions") explains why an MREC global table can't hold a lock: its conditions are checked on the local replica, and both Regions can "win". The control table must be MRSC.
Whose clock?
| Value | Written by | Read by | Clock |
|---|---|---|---|
ack_until | Frankfurt, every 2 s | Ireland, Stockholm, and the gate | Frankfurt's wall clock at the send + 10 s |
| Frankfurt's stop time | Frankfurt, for itself | Nobody else | Frankfurt's monotonic clock at the send + 10 s |
| The taker's condition | Ireland | DynamoDB compares | Ireland's wall clock, minus the 5 s margin |
| The clock-free variant | Ireland | Ireland | Ireland's monotonic clock: rvn unchanged for 15 s |
ack_until crosses machines, so it is a wall-clock time, and the reader must allow for the two clocks disagreeing. DynamoDB conditions have no function that reads the current time, so :now in the condition is always the taker's clock.
textFrankfurt (home), every 2 s: UpdateItem eu-writer Condition: home = :me AND epoch = :e Update: SET ack_until = :my_wall_now_plus_10s, rvn = :new_random on success: stop_at = monotonic send time + 10 s before every write: if monotonic now >= stop_at, refuse the write Ireland (taker), after the gate approves: UpdateItem eu-writer Condition: epoch = :seen AND ack_until < :my_wall_now_minus_5s Update: SET home = :me, epoch = :seen + 1, rvn = :new_random clock-free alternative: Condition: rvn = :seen -- after rvn stayed unchanged for 15 s on my monotonic clock
What the simulator printed for the night: Frankfurt's last successful renewal was sent at 19:59:59.0, so ack_until = 20:00:09.0, and any Frankfurt server still alive stops writing at 20:00:09.0 on its own clock. With the two clocks agreeing, Ireland's condition could first pass at 20:00:09.0 + 5 = 20:00:14.0. Ireland polls the item every second (at x.5 s); it first saw the final rvn at 19:59:59.5, so the clock-free variant could claim from 20:00:14.5. The gate closed much later (20:03:00), so neither bound delayed anything tonight.
The margin is only as good as the clock assumption behind it. ack_until is written on Frankfurt's clock and compared with Ireland's, so what counts is how far Ireland's clock is ahead of Frankfurt's, not either clock's offset from true time. The simulator's replay with Frankfurt's clock exact and Ireland's wall clock ahead of it (only the gap between the two clocks counts: Frankfurt 3 s slow and Ireland 3 s fast behaves like the 6 s row):
| Ireland's clock ahead of Frankfurt's by | Condition first passes (true time) | Against Frankfurt's stop at 20:00:09.0 |
|---|---|---|
| 0 s | 20:00:14.0 | 5 s gap |
| 3 s | 20:00:11.0 | 2 s gap |
| 6 s (more than the margin) | 20:00:08.0 | 1 s overlap: two writers could act |
Offsets from true time add up. Frankfurt 3 s slow and Ireland 4.9 s fast are each "below the margin", yet their clocks are 7.9 s apart: the simulator shows Ireland's condition passing at 20:00:06.1, 2.9 s before Frankfurt stops. So the margin must exceed the worst gap between the two clocks. If each host may be up to x from true time, two hosts can differ by 2x. Alarm on any host whose offset from true time passes half the margin (2.5 s here), or use the clock-free variant, which depends only on Ireland's clock rate over 15 s.
The epoch lives in the database
The lease decides who should write. The epoch decides what a database accepts. Seatly's Aurora database has one row, writer_epoch, and every write transaction checks it:
sql-- Ireland's first transaction after promotion (epoch 8 taken in the control table): UPDATE writer_epoch SET epoch = 8 WHERE id = 'eu' AND epoch = 7; -- 1 row, or stop -- every later write transaction, from any app server: BEGIN; SELECT epoch FROM writer_epoch WHERE id = 'eu' FOR SHARE; -- if it isn't my epoch: ROLLBACK and stop UPDATE seats SET order_id = :order WHERE seat_id = :seat AND order_id IS NULL; INSERT INTO orders (order_id, epoch, ...) VALUES (:order, :my_epoch, ...); COMMIT; -- FOR SHARE keeps the epoch from changing until here
An app server that still believes it writes for epoch 7 (a Frankfurt server following Aurora's global writer endpoint to Ireland, a delayed retry, a worker that wasn't told) has its transaction refused in Ireland. The (epoch, version) pair also travels on every event the outbox emits, so consumers can park anything from an older epoch (the Change Streams & the Transactional Outbox page: the relay resumes under the new epoch).
Trace: the fence and the promotion
Synthesizing vector architecture diagram...
What to notice: nothing on this diagram needs Frankfurt to answer. Frankfurt stops itself at 20:00:09 on its own clock; Ireland's takeover is a condition on a table Frankfurt can't reach; Aurora's attempt to fence Frankfurt simply times out. The epoch is bumped before any customer write reaches Ireland.
| # | Time | Event | What protects the data |
|---|---|---|---|
| 7 | 20:00:00.0 | Frankfurt goes dark. Last renewal sent 19:59:59.0, so ack_until = 20:00:09.0 | |
| 8 | 20:00:09.0 | Any Frankfurt app server still alive but cut off stops writing, at its monotonic send reading + 10 s | This stop is the only thing protecting Frankfurt's own Aurora volume |
| 14 | 20:03:00.0 | Ireland takes the lease: home = eu-west-1, epoch = 8, only if epoch = 7 and ack_until < Ireland's clock − 5 s: 20:00:09.0 < 20:02:55.0, true | One home at a time |
| 15 | 20:03:00 → 20:05:00 | failover-global-cluster --allow-data-loss to Ireland; Ireland's writer is up at 20:05:00 (120 s, our number). Aurora's attempt to stop writes in Frankfurt times out, and Aurora records the event on the new primary | Aurora's fence is a best-effort extra |
| 16 | 20:05:00.2 | Ireland's first transaction sets writer_epoch = 8; automation turns eu-west-1-writable On | writer_epoch protects Ireland's database |
Aurora says it tried to fence Frankfurt. Why do we still need Frankfurt's own stop and our own epoch, and why must Frankfurt refuse writes even a week later?
Snapshot R4: after the flip
Synthesizing vector architecture diagram...
What to notice: Ireland holds the lease at epoch 8 and its database says 8. Frankfurt's volume still says 7 and holds the 160 orders Ireland never received; nothing in Frankfurt can write, because every server there stopped at 20:00:09 and will refuse until the lease names Frankfurt again.
Aurora: switchover or failover
Aurora Global Database has two ways to move the writer to another Region, and they differ in exactly one thing: whether the old primary is healthy.
| Switchover (planned) | Failover (--allow-data-loss) | |
|---|---|---|
| When | Both Regions healthy: a drill, a failback, maintenance | The primary Region is unavailable |
| What Aurora does | "Waits for the target secondary Region clusters to be fully synchronized", then moves the writer | Promotes the secondary as it is, then attempts to stop writes in the old primary Region |
| RPO | 0 | The lag at the moment of failure: 0.8 s here, 160 orders |
| Needs the old primary | Yes | No |
| Fence of the old primary | Not needed: it is demoted cleanly | Best-effort; if it times out, Aurora emits an event on the new primary |
| The old primary afterwards | Becomes a secondary | Aurora attempts a snapshot of its volume, named rds:unplanned-global-failover-<old cluster>-<timestamp>, then rebuilds it as a secondary on a new volume (Part 8) |
| On Seatly's night | Event 29 (failback) | Event 15 |
Three details the loops rely on:
- The global writer endpoint. Aurora gives a Global Database one writer endpoint whose DNS name follows the primary. AWS advises setting client DNS caches to a low value "such as 5 seconds", because a long cache keeps clients on the old writer (Part 5, the client-cache table).
- Headless secondaries. A secondary cluster may have no DB instance at all, which is cheap (the pilot-light rung in Part 6). AWS: before a switchover or failover to it, "you must add a DB instance to it". That time belongs in the recovery time.
- The target's control plane.
failover-global-clusteris an API call to RDS in Ireland. It can't be a pure data-plane action, so it must be rehearsed (Part 10).
Side row 16x, the switchover at failback: Saturday at 06:00 (Part 8), both Regions healthy, Aurora waits until Frankfurt's secondary has every commit, then moves the writer. Nothing is lost, and nothing needs fencing by timeout.
What to remember from Part 4
- The home stops by its own clock; the new home waits the lease plus a margin and bumps the epoch in the store it writes.
- Aurora's fence is best-effort. The old home's own stop and refusal keep it off its orphaned copy; the epoch in the new database refuses anything that reaches the new writer; the refusal never expires.
- Switchover loses nothing; failover loses the lag.
Part 5. Moving the users
Ireland can take writes at 20:05:00.2. EU customers are still knocking on Frankfurt's door. The fourth decision moves them, and it has two hard rules: the move must not need the dead Region, and it must not need a control plane that may be having its own bad night (for Route 53's record edits, that is us-east-1). Then comes the part everyone forgets: DNS answers are cached, so users arrive over minutes, not at the moment of the switch.
Trace: the flip at 20:05
| # | Time | Event |
|---|---|---|
| 17b | 20:03:00 | At approval, automation also flips the S3 Multi-Region Access Point's routes (Frankfurt's traffic dial to 0, Ireland's to 100). It needs no promotion, because the objects were already replicated, and AWS says traffic moves in about 2 minutes, so seat maps and ticket PDFs come from Ireland's bucket by about 20:05. Flipped with the routing controls at 20:05, they would stay on dark Frankfurt until about 20:07, after most users arrive. Existing connections are not cut. Replication Time Control had copied the objects: AWS's commitment is 99.9% of objects within 15 minutes |
| 17 | 20:05:00.5 | One UpdateRoutingControlStates call changes both controls at once: eu-central-1-on Off, eu-west-1-on On. The gating rule allows it because eu-west-1-writable is On; the mirror rule allows turning eu-central-1-on Off because eu-central-1-writable is still On; the assertion rule allows it because one EU control stays On. Right after the flip, automation turns eu-central-1-writable Off. The call goes to the ARC cluster's five Regional endpoints in turn: the first two are down for maintenance, the third answers. AWS keeps at least three of the five available for state changes |
| 18 | 20:05:15 | Route 53 answers api.seatly.eu with Ireland's address. ARC propagates the new state to all five cluster endpoints "within 5 seconds on average, and after no more than 15 seconds maximum"; the routing controls' health checks, and Route 53's answers, follow. Our example allows the full 15 s |
| 19 | 20:05:15 → 20:06:45 | The last clients whose cached answer was fresh at 20:05:15 re-resolve by 20:06:15 (TTL 60 s), and their retries, backing off up to 30 s, land by 20:06:45. EU impact from 20:00:00: 405 s = 6.75 minutes. With pre-approved rules (13p): 295 s ≈ 4.9 minutes |
| 20 | 20:05 → 20:11 | The long tail. Clients holding open connections to Frankfurt's front door get silence, not resets: nothing in Frankfurt is there to send a reset. They notice only when a request times out (2 s) or a heartbeat is missed, then reconnect and re-resolve. Sticky and WebSocket clients sit silent until their heartbeat times out. One ISP's resolver that ignores TTLs keeps 2% of EU clients (our number) on Frankfurt's addresses until 20:11 |
Three ways to move traffic
AWS offers three data-plane ways to send users to another Region. All three avoid control-plane calls during the move; they differ in who decides and what they can refuse.
| ARC routing controls (Seatly) | Route 53 failover records on real health checks | Global Accelerator | |
|---|---|---|---|
| Who decides | You: a person or pre-approved automation calls UpdateRoutingControlStates | The health checks, automatically | Global Accelerator's health checks, automatically |
| Can a gate be enforced? | Yes: safety rules (assertion and gating) in the cluster | No: a health check can't know whether Ireland's database is promoted | No for automatic failover. Traffic dials are a control-plane change through its API in US West (Oregon) |
| Control planes needed during the move | None: the cluster's data plane in five Regions, three needed | None: health checks and DNS answers are data plane | None for its automatic failover |
| What moves | New DNS lookups | New DNS lookups | New connections: the static IP addresses never change, so there is no DNS tail. Established TCP connections stay on the old endpoint until the 340 s idle timeout |
| Health checks | Routing controls drive the health checks | Route 53 checks: 30 s (or 10 s) interval, threshold 3; healthy if more than 18% agree; fails open when all records are unhealthy | EC2 and Elastic IP endpoints: Global Accelerator's own checks. ALB and NLB endpoints: their target groups' health |
| Residency | Goes only where you point it | Goes only to the configured secondary | Its failover ignores traffic dials and tries the three closest endpoint groups, then fails open: one accelerator for EU and US can send EU users to the US (Part 6) |
| Cost | 1,825 a month**, at most 2 clusters per account, shared by every application in it | Health-check charges only | An hourly charge per accelerator plus data-transfer charges |
The weakness of plain Route 53 failover records is not us-east-1 (their health checks and answers are data plane too); it is the missing gate and the fail-open rule. That is why Seatly's records follow routing controls instead of Frankfurt's own health.
The DNS tail and client caches
A DNS change is a promise about future lookups. Every cache between Route 53 and the code that opens a connection holds the old answer for a while:
| Cache | How long it holds Frankfurt's address | Where the number comes from |
|---|---|---|
Route 53's TTL on api.seatly.eu | 60 s | Seatly's setting; ARC's best-practice page calls 60 or 120 s a common choice |
| Recursive resolvers | Normally the TTL; some ignore it and keep answers longer (in our example, one ISP's resolver, 2% of EU clients, until 20:11) | Our number |
| The JVM's own cache | 30 s for successful lookups and 10 s for failed ones, when networkaddress.cache.ttl isn't set. It doesn't read the record's TTL, so it can stack on top of a resolver's cache | OpenJDK's java.security defaults (an implementation default, not a Java rule) |
| Database drivers and Aurora's global writer endpoint | AWS advises a DNS cache "such as 5 seconds" | AWS |
| Open connections (HTTP keep-alive, WebSocket, sticky sessions) | Until the connection is closed: a connection never re-resolves | |
| Old Region's servers, once back | Until the client re-resolves, but harmless: in redirect mode they answer every request with a redirect to the current home instead of serving it (Part 8, event 24) | |
| Retry backoff | Up to 30 s after the new answer before the next attempt | Seatly's setting |
Logins survive the move: sessions sit in the MREC global table, which replicates within about a second, so only the last second or so of session writes (a login at 19:59:59.5) is missing in Ireland until Frankfurt returns (Part 8).
The routing control flipped at 20:05:00.5. Why are some customers still in Frankfurt at 20:11?
The timing chain
Each decision takes time. Some steps run side by side (capacity starts at detection, while the gate and the promotion run), and some must follow one another (nobody can flip before the promotion, and nobody arrives before the flip). Side-by-side steps cost their maximum; dependent steps add up. That is our second formula:
The gate can't close before ack_until + margin (Part 3), so on a very fast detection the gate term grows to reach that bound.
- Human gate: 60 + max(60, 120 + 120) + 15 + 60 + 30 = 60 + 240 + 105 = 405 s = 6.75 minutes.
- Pre-approved rules: 60 + max(60, 10 + 120) + 15 + 60 + 30 = 60 + 130 + 105 = 295 s ≈ 4.9 minutes. Capacity (ready at 20:02:00) is still hidden under the gate and the promotion (done at 20:03:10).
Synthesizing vector architecture diagram...
What to notice: capacity runs beside the gate and promotion, so its 60 s cost nothing: the longest path is detect, gate, promotion, flip, TTL and backoff, which ends at 405 s (06:45 on the axis, 20:06:45 on the clock).
The night us-east-1goes dark (side replay U)
Seatly's US tier is backup and restore into us-west-2: the US recovery time is hours, by design (Part 6). The more useful question is what else breaks when us-east-1 is the Region that goes dark, and whether the EU plan notices:
| Call or feature | Plane | During a us-east-1 outage |
|---|---|---|
| Route 53 DNS answers | Data plane, served from 200+ edge locations | Keep working |
| Route 53 health checks | Data plane | Keep checking, and keep driving answers |
| Route 53 health-check metrics | Available only in US East (N. Virginia) | Dashboards and alarms built on them go blank |
| Route 53 record edits | Control plane | Wait. With accelerated recovery, turned on beforehand for a public hosted zone, AWS targets 60 minutes to resume record changes |
| ARC routing-control state changes | Data plane: the cluster's endpoints in five Regions | Keep working, as long as three of five answer |
| ARC configuration (creating controls and rules) | Control plane in US West (Oregon) | Unaffected by us-east-1, but not highly available: create everything in advance |
| Global Accelerator's API (endpoint groups, dials) | Control plane in US West (Oregon) (side row U2) | Unaffected by us-east-1; its automatic health-check failover is data plane |
The EU plan never needs us-east-1: its probes run in Ireland and Stockholm, its lease lives in three EU Regions, its routing switch is an ARC data-plane call, and its DNS answers were set up in advance.
Read-only mode (side row 19r)
If Aurora's promotion is slow or risky, the approver can move users into an explicitly read-only Ireland first: turn eu-west-1-read-only-approved On at 20:03:00 and flip then. Ireland serves browsing, event pages and seat maps, and answers every write with 503 Retry-After. Reads come back at 60 + max(60, 120) + 15 + 60 + 30 = 285 s (20:04:45). With pre-approved rules the gate takes 10 s, but capacity still has to be running before users arrive: 60 + max(60, 10) + 105 = 225 s (20:03:45). Writes still wait for the promotion and a second flip.
What to remember from Part 5
- Move traffic with a data-plane switch created in advance, never by editing records during the outage.
- Count the DNS tail: TTL, caches, backoff, and connections that wait for a heartbeat.
- Parallel steps cost their maximum; dependent steps add up.
Part 6. Where the load can go: capacity, quota, residency, and the ladder
A move is only as good as the Region it moves to. Ireland must be able to run Frankfurt's peak, plus the surge of retries that follows every outage, inside its own account limits, without moving any EU customer's data out of the EU, and with every key, secret and image it needs already in place. Part 3's check (b) and check (c) are where these are tested; this Part is how they're built.
Trace: 20:01, capacity starts
| # | Time | Event |
|---|---|---|
| 10 | 20:01:00 | The rule fires. Automation raises Ireland's Auto Scaling desired capacity from 15 to 60, drawing the 45 pre-initialized instances from the warm pool |
| 11 | 20:02:00 | The warm pool is InService (60 s, our number): 60 × 8 = 480 vCPUs running, within the quota of 640. This step used Ireland's Auto Scaling control plane, and it could have failed: AWS warns of "cold starts if an Availability Zone is out of capacity" |
| 11 | 20:01:00 | Sizing for the surge needs 75 instances (below), so automation also launches 15 new instances. They are new launches, not warm, so they take Seatly's measured 23 minutes: healthy at 20:24 |
Capacity options on equal terms
| Option | Time to serve at failover | Control plane needed at failover | Capacity risk at failover | Cost while idle | Vendor quota (running On-Demand vCPUs) | Scope |
|---|---|---|---|---|---|---|
| New launches | Seatly: 23 minutes | Yes: launch | High: the Region may be short of your instance types while everyone evacuates | None | Counts once running | Any AZ with capacity |
| Warm pool (stopped, pre-initialized; ours) | Seatly: 60 s | Yes: start | Medium: a start can fail when an AZ is out of capacity | EBS volumes and Elastic IPs only | Counts only once started | The group's AZs |
| Dark capacity (running, held out of service) | Seconds | None, if the app itself holds them out and lets them in | None | Full running cost | Counts (running) | Where they run |
| Capacity Reservations | A launch, into reserved capacity | Yes: launch | Low: the capacity is reserved | Billed whether used or not | Counts whether used or not | One AZ each |
| Hot standby (running, in service) | Zero | None | None | Full running cost | Counts (running) | Where they run |
Ireland's quota is 640 and 480 vCPUs are running after the warm pool starts. Is that enough, and what else must the target allow?
Why 512 looked fine (event K0, last month's game day). Ireland's quota was 512 vCPUs while Frankfurt's was 640. Stopped warm-pool instances don't count toward the running On-Demand quota, so every check that looked at Ireland's usage (120 vCPUs) against 512 passed. The game day started the warm pool: 480 running, 32 left, and the surge's 15 instances couldn't launch. Seatly raised the quota to 640. The rule it learned: quota parity with the home, checked against what will be running at failover. ARC's Region switch has built-in service-quota checks that compare the Regions' quotas every 24 hours and can request increases; Seatly runs the same comparison itself.
Sizing the target (event K1)
The simulator's utilization in Ireland, with load arriving at 20:06:45 and the surge lasting until about 20:11:45:
| Time | Load (instances' worth) | Running | Utilization |
|---|---|---|---|
| 20:07 | 40 × 1.5 = 60 | 60 | 100% |
| 20:10 | 60 | 60 | 100% |
| 20:12 | 40 (surge over) | 60 | 67% |
| 20:30 | 40 | 75 (the 15 new launches healthy at 20:24) | 53% |
The 80% target was not met during the surge. The warm pool covered the load itself, but the headroom depended on new launches, and new launches take 23 minutes. For five minutes Ireland ran with no headroom at all: slower responses and some timeouts are likely (the simulator doesn't charge them to the budget). The cheap fix is a warm pool of 60 instead of 45, so that all 75 are stopped-and-ready: stopped instances cost only their volumes. Sizing also works per AZ: Frankfurt runs 20 instances in each of three AZs so that losing one AZ still leaves 40, the peak; Ireland's 75 should be spread the same way.
Residency and prerequisites
The default target map was written by someone thinking only about capacity: "spread to spare capacity". Check (c) of the gate is where it meets the residency rule:
| Target map | Proposed | Allowed for EU data? | Result |
|---|---|---|---|
| Ireland | 70% | Yes | Kept |
us-east-1 | 30% | No | Refused; its share goes to Ireland |
| Final | Ireland 100%, which is why Ireland is sized for all of it |
Residency also limits how often you can move. With two EU Regions, the EU zone can fail over once: while Frankfurt is out, Ireland has no legal partner, so a second failure means waiting, not moving. Every copy counts, not only the one users are routed to: standbys, backups, witnesses and keys.
- An MRSC witness holds data (event K3). Had Seatly used an MRSC table for seats with Frankfurt, Ireland and a witness in
us-east-1, the witness "contains data written to global table replicas", so EU seat data would sit in the US. All three MRSC Regions must be inside the allowed zone. - Global Accelerator's failover ignores dials (side row K2). Suppose one accelerator served both EU and US, with endpoint groups Frankfurt (dial 100), Ireland (dial 0, the standby) and
us-east-1. If Ireland's health check passes, the accelerator moves EU users to Ireland as soon as it marks Frankfurt unhealthy, before promotion: exactly the ungated move Part 3 forbids. If Ireland's check fails by design until promotion, then from Frankfurt's failure until the promotion, the accelerator's failover ignores the dials, tries the closest endpoint groups, and sends EU users' new connections tous-east-1. The simulator's replay with Frankfurt and Ireland unhealthy routes an EU client tous-east-1. The fix is one accelerator (or one set of DNS records) per residency zone, so the only groups it can fail over to are legal ones. - Prerequisites (side row K4). Ireland must already have everything its fleet needs to start and decrypt: KMS multi-Region replica keys (related multi-Region keys share key material and key ID, so data encrypted in Frankfurt decrypts in Ireland "without re-encrypting or making a cross-Region call"; a replica works even when the primary is unreachable, and an existing single-Region key can't be converted, so this is decided when the data is first encrypted), secrets, TLS certificates and container images. A key or image that exists only in Frankfurt makes Ireland's servers unable to start or decrypt. Gate check (b) verifies them.
The ladder
"How ready is the standby?" is a ladder of choices, each rung costing more while idle and needing less at failover time. AWS's disaster-recovery whitepaper labels the rungs' recovery times as hours, tens of minutes, minutes and real time; the numbers below are Seatly's, from Formula 2 with the same inputs as the main line.
Synthesizing vector architecture diagram...
What to notice: each rung to the right has less to start at failover time. Every rung except hot standby and active-active still depends on something starting during the disaster; that dependency, not the RPO, is what the extra cost buys away.
Side replay L1: the same EU service on five rungs
The same EU service, the same 200 orders a second, the same 0.8 s lag, the same 2-minute human gate and 120 s promotion. Instances are counted across both EU Regions (Frankfurt's 60 included):
| Rung | Orders at risk (RPO) | RTO (Formula 2) | Instances running before | Starts during the failover | Control-plane calls at failover | Rehearsal |
|---|---|---|---|---|---|---|
| Backup and restore (hourly copies) | Up to an hour of orders, plus copy time | Hours | 60 | The database from backup, the whole fleet | Restore, launch, deploy | A full restore |
| Pilot light (headless Aurora secondary, no app servers) | 160 | 60 + max(1,380, 120 + 600 + 120) + 105 = 1,545 s ≈ 25.75 min (10 minutes to add a DB instance, our number; 23 minutes to launch the fleet) | 60 | A DB instance, then the fleet from zero | Create DB instance, Aurora failover, launch | Hard: everything starts on the day |
| Warm standby (Seatly) | 160 | 405 s = 6.75 min | 75 (and 45 stopped) | The warm pool | Warm-pool start, Aurora failover | Medium |
| Hot standby (all 60 running and in service in Ireland; like the warm row at arrival, it runs at 100% during the surge unless it keeps 75) | 160 | 405 s: the same, because capacity was never the longest path | 120 | Nothing | Aurora failover only | Easier: the fleet is always there |
| Active-active (each Region homes half the events, keeps dark capacity for the other half) | 100/s × 0.8 = 80 (Frankfurt wrote only half) | 405 s for the half that moves; 0 for the other half | 126 (63 per Region: (20 + 20 × 1.5) ÷ 0.8 = 62.5, rounded up) | Nothing | Aurora failover for Frankfurt's half | Capacity tested daily by real traffic |
Hot standby's RTO equals warm standby's here, because capacity (60 s) runs beside the gate and promotion (240 s). Both rows arrive with 60 instances and run at 100% through the surge; keeping 75 running (5× the warm row's 15) would hold 80%. What 4× the running cost in Ireland (60 instances instead of 15) buys is no capacity risk and no control-plane call for capacity during the disaster: AWS calls a recovery Region that can take the full load as deployed "statically stable". Active-active's gain is that only half the customers are hit (405 s × 0.2 = 81 s = 1.35 customer-weighted minutes instead of 2.7) and that its capacity is proven every day; its cost is dark capacity in both Regions and a conflict rule for anything homed in both. AWS Elastic Disaster Recovery, which replicates servers into a staging area, is a managed version of the pilot-light rung.
Active-active has near-zero RPO. Why is Seatly's warm standby's RPO the same 0.8 s?
Active-passive vs active-active
| Active-passive (warm or hot standby) | Active-active (homes split, dark capacity) | |
|---|---|---|
| What a Region failure moves | All of its homes | Only the failed Region's homes (half here) |
| Customer-weighted outage on Seatly's night | 405 s × 0.4 = 162 s = 2.7 min | 405 s × 0.2 = 81 s = 1.35 min |
| Data lost | The lag × its commit rate (160) | The lag × its commit rate (80) |
| Idle cost | The standby fleet (15 running + a warm pool, or 60 running) | Dark capacity in every Region |
| Is the standby tested? | Only by drills | Every day, by real traffic |
| Conflicts | None: one Region writes | None if every key has one home; LWW or worse if not |
| A split between the Regions | The home keeps writing; the standby's RPO grows | Each Region keeps writing its own homes; cross-home writes wait for their home |
Seatly's US tier sits on the bottom rung on purpose (side row L2): hourly backups copied to us-west-2, so its RPO is up to an hour plus the copy time and its RTO is hours. That is a stated business decision about the US catalogue, not an accident. Every workload gets its own rung.
Shared limits
Several resources are shared by everything that happens during a failover, and each must be counted with all its consumers:
| Shared resource | Who uses it at the same time |
|---|---|
| Ireland's capacity | Ireland's own traffic, the moved traffic, and the surge. And every other AWS customer evacuating the same Region is asking Ireland for capacity too |
| Ireland's vCPU quota | The warm pool once started, the surge launches, and every other workload in the account and Region. Capacity Reservations count whether used or not |
| The network path | Canaries from other Regions ride the same AWS backbone as replication; only a probe outside AWS is independent. The Ireland–Frankfurt path carries replication, Ireland's canaries and the lease traffic at once, which is why the second vantage point and the third copy sit in Stockholm |
| Target-side control planes | Aurora's failover-global-cluster and the warm-pool start both run in Ireland during the disaster; they can't be data-plane-only, so they are rehearsed (Part 10) |
| ARC | At most 2 clusters per account, shared by every application in it |
| The payment provider's rate limit | Live payments and the replayed unknowns (Part 7) |
| Ireland's database | Live writes, replayed requests and reconciliation reads. Reconcile from a side cluster (Part 8) |
| The error budget | Every incident in the month, including Seatly's own deploys (Part 10) |
What to remember from Part 6
- Capacity is running before the move, or starts from a warm pool within the gate's time with quota to spare; size for moved load × surge at most 80%.
- The target map is a residency decision made in daylight, and every copy counts.
- Every rung below hot standby depends on something starting during the disaster.
Part 7. Writes caught in the middle
Some orders were in the air when Frankfurt went dark. Some had been committed and acknowledged but never reached Ireland; some had been sent with no reply at all. A failover doesn't make them disappear: it makes them unknown, and unknown orders become lost money, double charges or double-sold seats unless something resumes them on purpose. This Part counts them and follows one seat through the gap.
Trace: 160 orders and 60 unknowns
| # | Time | Event |
|---|---|---|
| 5 | 19:59:59.2 → 20:00:00.0 | Orders committed in Frankfurt but not yet in Ireland: 200/s × 0.8 s = 160, nearly all acknowledged (the customer saw "confirmed"; a commit in the last milliseconds may have lost its reply and is also an unknown). Among them, o-4417, committed at 19:59:59.6, holds seat 14C |
| 6 | 20:00:00 | About 60 requests are in flight with no reply (200/s × 0.3 s mean latency, our number) |
| 21 | 20:06 → 20:30 | Clients retry their 60 unknowns as soon as they reach Ireland, reusing the same idempotency keys. For the 160 lost orders, the app's order journal sees 404 unknown order when it fetches the ticket, and re-sends the order under the same key: 120 of the 160 clients open the app before 20:30 (our number). The payment provider returns the saved charge for keys it has already seen, so nobody pays twice, and the orders are re-created with the same derived IDs. 40 orders remain only in Frankfurt's volume |
The rules behind this (derive every ID from the idempotency key, resume unknowns with the same key, and dedup in an inbox at the home) belong to Idempotency & Effectively-Once Processing; this page only counts what a Region move leaves for them. One detail is failover-specific: Ireland's dedup windows start empty. Anything Ireland keeps to recognize repeats (an inbox, a cache of recent keys) has only what was replicated. A request Frankfurt finished but Ireland never heard of looks brand new, which is why the IDs must be derived, so that a second copy collides with the first by ID in reconciliation.
Seat 14C
Synthesizing vector architecture diagram...
What to notice: nobody did anything wrong at 20:07:10. Ireland checked 14C against the only data it had, and the seat was free there. The double sale was created at 20:00:00, when 0.8 s of Frankfurt's commits became unreachable; reconciliation at 21:10 is the first moment anyone can see it.
Seat 14C was sold in Frankfurt at 19:59:59.6 and again in Ireland at 20:07:10. Which choice made that possible, and which one change prevents it?
Drill 05, both questions
The Booking That Existed in Frankfurt but Not in Virginia runs three Regions active-active on the seats, with asynchronous replication of 200 to 400 ms. Side row 22a replays it on Seatly: without homes, an EU phone routed to Ireland sees 14C free 0.4 s after Frankfurt sold it, and a second buyer can take it there.
- Q1, what went wrong: two Regions could both accept a claim on the same seat, each checking only its own copy, inside the replication window. The fix is exactly Seatly's home per event: a claim is a conditional write ("set 14C sold only if it is free") at the single owner of that event, wherever the request landed. Browsing stays local and may be a few hundred milliseconds stale; claiming does not. With homes, the window reopens only during a failover, which is event 22.
- Q2, why not make every write synchronous to every Region: read the last row of the table below. Every write, including the millions that don't need it, pays a round trip to the slowest Region, and a slow or cut-off Region stalls writes everywhere. Scope synchronous replication to contested writes.
Stores at failover
What each store promises across Regions is covered in Replication, Quorums & Read-Your-Writes. This table adds only what matters on the night:
| Store | Data lost when the home dies | Write latency | A split between Regions | Moving the writer | Transactions | TTL | Regions | Residency | Fencing an old writer |
|---|---|---|---|---|---|---|---|---|---|
| Single-Region home (no cross-Region copy) | Nothing is lost, but nothing is readable until the Region returns (or you restore a backup) | Local to the home | The home keeps writing; others can't | Not possible; restore instead | Yes | Yes | 1 | Simple | Not needed |
| Aurora Global Database (Seatly's orders) | The lag at failure: "typically measured in seconds"; 160 orders here | Local at the home | The home keeps writing; the secondary's lag grows (240,000 orders at risk after 20 minutes) | Switchover (RPO 0, needs a healthy primary) or failover (loses the lag) | Yes | Not applicable | A primary and its secondaries | Choose the Regions | Best-effort; add a lease and an epoch (Part 4) |
| DynamoDB global table, MREC (carts, sessions) | The replication delay ("typically within a second or less"); the tail arrives when the Region returns | Local in any Region | Both sides write; last writer wins when the link heals | Nothing to move: every replica takes writes | ACID only in the Region where the call was made | Yes | Any number | Choose the replica Regions | Can't hold a lock (page 03, Part 8) |
| DynamoDB global table, MRSC (the control table) | 0 | A round trip to at least one other Region | A Region that reaches another keeps writing; an isolated Region serves only eventually consistent reads | Nothing to move | No | No | Exactly 3 (3 replicas, or 2 and a witness) | All three in the allowed zone: a witness contains data | Its conditional writes can hold a lease |
| Synchronous to every Region (drill 05 Q2) | 0 | A round trip to the slowest Region, on every write | Every write stalls until the link heals | Nothing to move | Depends on the engine | All | All Regions must be legal | Needed only if two can write |
What a secondary accepts, it must forward
The payment provider (PSP) sends a webhook when a payment settles. Frankfurt's payments settle all evening, and their webhooks now arrive at Ireland.
| # | Time | Event |
|---|---|---|
| 23 | 20:00 → 20:30 | Webhooks for Frankfurt's payments fail until the flip (Frankfurt is dark; the PSP retries), then reach Ireland's hooks.seatly.eu. For orders Ireland doesn't have (the lost tail), the endpoint stores the webhook durably and forwards it to reconciliation, and answers success only after the store |
| 23x | (side row) | The trap: the handler answers 200 for unknown orders and discards them. The PSP stops retrying, and those payment confirmations are gone for good |
The general rule, for every secondary endpoint (a standby Region, a backup mail server, a regional webhook receiver): anything a secondary accepts is stored durably and forwarded to the key's current home, or to reconciliation. It never answers success and drops the data. The email loop (step 3.4, the backup MX at priority 20) goes one step further: its standby relays everything it accepts to the home continuously, in normal times too, so the forward path is exercised every day and not discovered broken during an outage.
What to remember from Part 7
- A failover loses the unshipped tail and leaves unknowns: resume them with the same keys and derived IDs.
- Contested writes need RPO 0, a single owner, or a compensation plan.
- Whatever a secondary accepts, it must store and forward.
Part 8. Failback: bringing the old Region back
At 20:40 Frankfurt's power returns. It has servers that were writers two hours ago, a database volume with 160 orders Ireland never saw, a queue of messages nobody processed, and caches full of old answers. The fifth decision is the cleanup: Frankfurt must come back as a replica, never as a writer, its leftovers must be sorted by ID, and the home moves back only when Seatly chooses.
Trace: Frankfurt returns at 20:40
| # | Time | Event |
|---|---|---|
| 24 | 20:40:00 | Frankfurt's power returns. Its app servers start, read the lease (home Ireland, epoch 8), and refuse every write, for good, redirecting callers to Ireland. The 2% of clients still holding a Frankfurt address (a pinned connection, a resolver that ignores TTLs) are redirected too |
| 25 | 20:45 | Aurora attempts a snapshot of Frankfurt's old volume at the point of failure (rds:unplanned-global-failover-<cluster>-<timestamp>), creates a new volume for Frankfurt, and adds it back as a secondary of Ireland. AWS says the rebuild takes "a few minutes to several hours"; 2 hours here (our number) |
A counterfactual the simulator keeps in the full table (24x): Frankfurt's servers trust Aurora's fence and their own configuration instead of the lease. One of them writes to a local queue that no longer drains anywhere, or to the old volume before Aurora detaches it. The write is acknowledged and never reaches Ireland. The refusal in event 24 is the old home's own protection; nothing else covers its orphaned copy.
Snapshot R5: Frankfurt back, refusing
| Frankfurt | Ireland | Stockholm | |
|---|---|---|---|
| State | Power back; refusing all writes, redirecting to Ireland | Home for EU events | Control-table replica |
| Lease as seen there | home Ireland, epoch 8 | home Ireland, epoch 8 | home Ireland, epoch 8 |
writer_epoch | Old volume: 7 (snapshot attempted); new volume: rebuilding as a secondary | 8 (writer) | |
| Routing controls | eu-central-1-on Off | eu-west-1-on On, eu-west-1-writable On | |
| Running vCPUs | App servers up but serving only redirects | 600 of 640 (75 × 8, since 20:24) | |
| Stranded work | 36 messages in Frankfurt's queue; 40 tail orders not yet re-created | 1 seat with two owners, not yet known |
Late arrivals and old queues
The MREC tail arrives late (event 26, 20:40). Carts and sessions written in Frankfurt's last second never reached Ireland. AWS: for an isolated MREC replica, "any data not yet replicated to other Regions will be replicated when the replica becomes healthy". They arrive now, and last writer wins: a cart a customer changed in Ireland since 20:06 keeps the newer version; a cart untouched since 19:59:59.5 gets its lost item back. That is fine for a cart; it would not be fine for a seat, which is why seats aren't in MREC.
The old queue (event 27, 20:50). Frankfurt's SQS queue holds 36 messages. Frankfurt's workers stay stopped: if they started, they would process against Frankfurt's stale database and its stale inbox. Instead, a drain job in Ireland reads the messages and checks each against Ireland's inbox (the home's record of what has been processed). Message IDs are derived from the order key and the event type, so a message about an order its customer re-created in Ireland carries the same ID as Ireland's own message and is recognised:
textDRAIN THE OLD QUEUE (a job at the home; the old Region's workers never start) for each message m in Frankfurt's queue: if m.message_id is in Ireland's inbox: drop it -- already processed elif m.order_id exists in Ireland: in one transaction: insert m.message_id into the inbox and apply the effect; then send any email or PSP call with m.message_id as its idempotency key else: send it to reconciliation, with the message
| Group | Count | Why | Result |
|---|---|---|---|
| Already in Ireland's inbox | 20 | Already processed: by Frankfurt's workers before 20:00 (the inbox rows replicated, but the messages weren't deleted yet; all 20 here), or, for an order its customer re-created, by Ireland's workers under the same derived message ID | Dropped |
| Order in Ireland, not in the inbox | 4 | The orders replicated, but Frankfurt's relay had already marked them published, so Ireland's relay never published them again | Processed now |
| Order not in Ireland | 12 | Orders from the lost tail whose customers hadn't come back (12 of the 40) | To reconciliation |
| Total | 36 = 20 + 4 + 12 |
The relay itself resumes in Ireland under epoch 8 (Change Streams & the Transactional Outbox); the old queue is drained, never replayed.
Reconciling the tail
| # | Time | Event |
|---|---|---|
| 28 | 21:10 | Reconciliation restores Frankfurt's snapshot into a side cluster (so the diff doesn't load Ireland's writer) and compares orders by ID after the last position Ireland had received. Of the 160 tail orders, 120 were re-created by their customers with the same derived IDs (nothing to do), and 40 exist only in Frankfurt. For each of the 40: honour it (insert it in Ireland at epoch 8, with its original ID) or refund it, by a written policy. On o-4417 vs o-5120 (seat 14C), one buyer is re-seated. If the snapshot attempt had failed, the tail would be rebuilt from the payment provider's settlement reports and the app order journals |
Frankfurt is healthy at 20:40. Why not flip traffic back now?
Failing back on purpose
| # | Time | Event |
|---|---|---|
| 29 | Sat 06:00:00 → 06:02:00.6 | Planned failback at a quiet hour: 1. (06:00:00) Ireland stops EU writes and releases the lease (it clears home and keeps epoch 8). eu-west-1-writable stays On for now, because a gating rule refuses every change to its target while its gating controls are Off. 2. Aurora switches over from Ireland to Frankfurt: it waits until Frankfurt has every commit, so the RPO is 0 (06:02:00; 120 s, our number). 3. (06:02:00.2) Frankfurt takes the lease at epoch 9, conditional on epoch 8, and sets writer_epoch = 9 in its database before the flip. 4. (06:02:00.3) Automation turns eu-central-1-writable On (the gating control of the mirror rule on eu-central-1-on). Then (06:02:00.5) the routing controls flip back in one UpdateRoutingControlStates batch, and only then (06:02:00.6) does eu-west-1-writable go Off. The team watches Frankfurt's CPS rise |
The failback's write pause (from the simulator; the switchover time is our number): EU writes stop at 06:00:00 and resume at 06:02:00.2, when Frankfurt holds epoch 9. That is about 2 minutes, the switchover itself. Clients whose DNS still points at Ireland get a redirect to Frankfurt from then on, so the DNS tail costs an extra hop, not more refusals. At 40% of customers, that is 2.0 × 0.4 = 0.8 customer-weighted minutes (Part 10 counts it).
The epoch goes up to 9, never back to 7: every writer, message and cached token from before the failback is then older than the current one, including anything left over from Frankfurt's first tenure.
This failback moves all EU events at once. A gradual return (5% of events a minute, or city by city as the ride-sharing loop does) needs a lease and an epoch per cell or per group of keys, because each group then has its own home during the move. Part 9 says why that is its own design.
Synthesizing vector architecture diagram...
What to notice: Frankfurt writes nothing until the last two boxes, and it gets the home back with a new epoch (9), not its old one. Every box before the switchover can take hours; none of them is urgent once Ireland is serving.
Snapshot R6: failback done, Saturday 06:00
| Frankfurt | Ireland | Stockholm | |
|---|---|---|---|
| State | Home for EU events | Warm standby again: 15 running, warm pool refilled | Control-table replica |
| Lease as seen there | home Frankfurt, epoch 9 | the same | the same |
writer_epoch | 9 (writer) | 9 (read-only secondary) | |
| Routing controls | eu-central-1-on On | eu-west-1-on Off, eu-west-1-writable Off | |
| Stranded work | none | none | none |
What to remember from Part 8
- The old Region comes back as a replica, never as a writer.
- Its tail is reconciled by ID; its queues drain against the home inbox, and its workers stay stopped.
- Fail back with a switchover and a new epoch, when it suits you.
Part 9. Shrinking the blast radius: cells and shuffle sharding
Most outages are not a Region going dark. They are a bad deploy, a bad config, or one tenant's traffic that breaks something. If such a failure takes down a whole Region, someone will propose evacuating it, and the evacuation carries the bad change to the next Region. The cheapest failover is the one you never need, because the failure stayed small. At 19:00, an hour before the real outage, Seatly's cells get tested.
Trace: the 19:00 deploy
| # | Time | Event |
|---|---|---|
| C1 | 19:00:00 | A config change reaches cell 2 of Frankfurt's four cells. Cell 2's errors jump; the deploy wave stops at the bake alarm, before cell 3; cell 2 is rolled back by 19:05. Impact: 5 minutes × ¼ of Frankfurt's customers × 0.4 of Seatly = 5 × 0.25 × 0.4 = 0.5 customer-weighted minutes |
| C1x | (side row) | Without cells, the change reaches all of Frankfurt at once: 5 × 0.4 = 2.0 minutes, four times as much. And with every Frankfurt checkout failing, someone proposes evacuating to Ireland, which would spend a Region failover (minutes and 160 lost orders) to escape a change that one cell's rollback undoes; and if the config came from a store shared by both Regions, Ireland would have it too |
Synthesizing vector architecture diagram...
What to notice: the router maps each event to exactly one cell, so the bad config hit only the quarter of events in cell 2. The wave stopped at the bake alarm; cells 3 and 4 never got the change.
Cells
A cell is a complete copy of the stack (front door, app servers, workers, its own database partition) that serves a fixed slice of keys. A thin cell router maps each key to its cell. Three rules make cells work:
- Deploy one cell at a time, bake, then continue; never two Regions at once. The cell is the unit you deploy and roll back.
- The router is critical and simple. It holds a mapping and nothing else, so it rarely needs to change.
- A cell is not automatically the unit you fail over. Moving one cell's keys to another Region is a smaller Region failover, and it needs its own lease and epoch per cell (the gradual failback of Part 8 needs the same).
A bad config broke cell 2. Why didn't anyone evacuate Frankfurt, and what would have happened if they had?
Shuffle sharding
Cells split customers into a few large groups; shuffle sharding gives each tenant its own small, random-looking combination of workers, so that almost no two tenants share the same set. Amazon's Builders' Library explains it with eight workers and each tenant on two of them:
- There are C(8, 2) = 8 × 7 ÷ 2 = 28 possible pairs, so 28 shards.
- A poisoned tenant (one whose requests crash workers) takes down its own two workers. Every other pair is one of three kinds:
| Other tenants' pairs | How many of the 28 | What happens to them |
|---|---|---|
| Share no worker with the poisoned pair | C(6, 2) = 15 | Untouched |
| Share one worker | 2 × 6 = 12 | Lose half their capacity, keep working if clients retry on the other worker |
| Are the same pair | 1 | Down with it: 1/28 of tenants |
With the same eight workers in four fixed pairs, a poisoned tenant takes down a quarter of all tenants; with shuffle sharding, 1/28, seven times fewer. The same idea scales: Route 53 gives each hosted zone 4 of 2,048 virtual name servers, about 730 billion combinations (C(2,048, 4) ≈ 7.3 × 10¹¹), arranged so that no two domains share more than 2. The metrics loop (step 3.1) puts each tenant's queries on 4 of 50 queriers: C(50, 4) = 230,300 shards. Shuffle sharding needs clients that retry on another worker; without retries, sharing one worker still hurts.
What to remember from Part 9
- Most outages are ours: keep them small before they're regional.
- Cells bound a failure to a slice of keys; shuffle shards make one tenant's problem almost nobody else's.
- A cell is the unit you deploy and roll back; making it the unit you fail over needs its own lease and epoch.
Part 10. Proving it: game days and the error budget
A failover plan that has never run is a guess. Quotas drift, keys get created in one Region only, runbooks name endpoints that were renamed, and the warm pool that looked full has no capacity behind it. And "four nines" means nothing until it is turned into minutes that one night can spend. This Part does both: it spends the month's budget, then rehearses what could spend it next time.
The budget
- 99.99% of a 30-day month: 30 × 1,440 = 43,200 minutes × 0.0001 = 4.32 minutes.
- 99.99% of an average month (30.44 days): 30.44 × 1,440 = 43,833.6 minutes × 0.0001 ≈ 4.38 minutes. State which month you use; the loops use both.
- Weight each outage by the share it hits. An outage that hits 40% of customers for 10 minutes spends 4 customer-weighted minutes.
This month's spending (event 30)
| Incident | Minutes × share hit | Spent | Share of 4.32 |
|---|---|---|---|
| The 18:30 split | No customer impact | 0 | 0% |
| The 19:00 cell-2 deploy | 5 × ¼ × 0.4 | 0.5 min | 11.6% |
| Frankfurt's night, human gate | 405 s = 6.75 min × 0.4 | 2.7 min | 62.5% |
| Month, human gate | 3.2 min | 74.1% | |
| (Frankfurt's night, pre-approved rules) | 295 s ≈ 4.92 min × 0.4 | 1.97 min | 45.5% |
| (Month, pre-approved rules) | 2.47 min | 57.1% | |
| Saturday's planned failback (writes paused from release to Frankfurt's epoch 9, about 2 min, Part 8) | 2.0 × 0.4 | 0.8 min, counted unless the SLO excludes announced maintenance | 18.5% (month with it: 4.0 min, 92.6%; pre-approved 3.27 min, 75.6%) |
Synthesizing vector architecture diagram...
What to notice: the flat line is the month's 4.32-minute budget. One Region failover alone spent 2.7 of it; the month's total bar (3.2) leaves 1.12 minutes for everything else. Without cells the deploy would have spent 2.0, and the month would have been over budget (4.7).
One Region failover can spend most of a 99.99% month even when it goes well. The naive night in Part 0 would have spent 10 minutes, 2.3 months' worth.
Rehearsing the split, the move and the return (event G1)
Next month's game day, one rehearsal per risk:
| What to rehearse | How | What it can't show |
|---|---|---|
| The split (Part 2) | AWS Fault Injection Service's Cross-Region: Connectivity scenario, run from Ireland toward Frankfurt. It blocks the experiment Region's VPC traffic to the other Region (transit gateways, route tables, VPC endpoints) and pauses S3, DynamoDB and MemoryDB cross-Region replication. It runs for 3 hours by default, up to 12 | It does not pause Aurora Global Database replication, so the drill can't rehearse the growing-lag case (the 240,000 orders of Part 2). It doesn't block canaries that aren't attached to a VPC, which is one reason Seatly's canaries are. The scenario includes no stop conditions: add CloudWatch alarms as stop conditions by hand. Its DynamoDB action cuts the experiment Region's replicas off from every other Region, so leave the control table untagged; otherwise Ireland also loses Stockholm, which is a different test from the 18:30 split |
| The flip (Part 5) | Change the routing controls through each of the five cluster endpoints in turn | That the endpoints you didn't reach are fine on the day |
| Promotion without loss (Part 4) | An Aurora switchover to Ireland and back, at a quiet hour. (FIS's aws:rds:failover-db-cluster action runs FailoverDBCluster, a failover inside one Aurora cluster, not a Global Database switchover, so it isn't used) | The real failover's lost tail and fence timeout |
| A warm pool that can't start (Part 6) | FIS's aws:ec2:asg-insufficient-instance-capacity-error action on Ireland's Auto Scaling group, which injects insufficient-capacity errors into the group's own launch requests. (The aws:ec2:api-… variant fails RunInstances, CreateCapacityReservation, StartInstances and CreateFleet only for the IAM roles you target; Auto Scaling starts warm-pool instances under its own service-linked role.) Whether it covers warm-pool starts is not documented: verify in a test run. It is meant to prove the gate notices that the warm pool didn't start and falls back (another AZ, a smaller target, or a person) | A real capacity shortage across a whole Region |
| Quota parity | Compare Ireland's quotas with Frankfurt's against what will be running (K0) |
Target-side control planes (Aurora's failover call and the warm-pool start both run in Ireland during the disaster) can't meet the "data planes only" rule. They are exactly the steps that must be rehearsed, because they are the ones that can fail on the day.
Scoring a drill. Score by the business metric, not by "the runbook finished": how long until EU CPS was back in band, and how long each phase took (detection, gate, promotion, flip, tail). The Netflix case study (step 3.5, "Did the evacuation work?") scores its evacuations the same way. ARC's Region switch uses application health alarms to report the actual recovery time of each run.
Your drill passed. Name three ways the real night could still differ.
What to remember from Part 10
- One Region failover can spend most of a 99.99% month: count it before it happens.
- Rehearse the split, the move and the return.
- A drill without stop conditions is an outage you scheduled.
Part 11. End to end through the layers
"Failover" touches every layer between a customer's phone and the disk, and each layer learns about the move in its own way and at its own speed. Following one checkout through all of them explains why the move takes minutes, and where the two protections of Part 4 sit.
One checkout after the flip (event 31, 20:07:30)
Synthesizing vector architecture diagram...
What to notice: the highlighted boxes are the guards: the worker's own lease check (cheap, but only as good as the worker's code); Protection 2, the epoch check inside Ireland's writer, where every write lands; and Protection 1 in Frankfurt's own servers, because nothing in Ireland can reach Frankfurt's old volume.
| Layer | What it knows | How it learns of the move | How it fails | Part |
|---|---|---|---|---|
| Phone app | Its order journal; the last address it used | A timeout or 404 unknown order; then re-resolves | Holds an open connection to a dark Region until a heartbeat times out | 5, 7 |
| Resolver and caches | A cached answer | Its TTL (or longer, if it ignores TTLs); the JVM's own 30 s | Serves Frankfurt's address minutes after the flip | 5 |
| Route 53 | Which records are healthy | The routing control's health check | Fails open if every record is unhealthy | 2, 5 |
| ARC routing controls | On/off per Region | An UpdateRoutingControlStates call | Needs 3 of 5 endpoints; refuses changes the safety rules forbid | 3, 5 |
| Front door and cell router | Which cell serves this event | Nothing: they serve whatever reaches them | A bad config in one cell (Part 9) | 9 |
| Worker | Its epoch and the lease | Reads the lease; told by the database when its epoch is old | A stale worker's write is refused | 4 |
| Aurora writer | writer_epoch | The promotion and the first transaction at epoch 8 | Promotion is a control-plane call in Ireland | 4 |
| Outbox and queue | (epoch, version) on each event | The relay resumes under the new epoch | The old Region's queue must be drained, not replayed | 8 |
| Stream jobs (per Region) | Their checkpoints | They move with the home | Replayed from a checkpoint in the wrong Region | Event Time, Watermarks & Checkpoints |
| The disk | Bytes | Nothing | A lost copy | Write-Ahead Log, fsync & Group Commit |
What to remember from Part 11
- Every layer learns of a move differently: DNS by TTL, connections by heartbeat, the database by epoch.
- Safety sits where writes land: the epoch check at the new writer, and the old home's own stop and refusal for its orphaned copy.
- Speed sits in what you prepared in advance.
Part 12. On AWS
The five decisions are the same on any cloud. On AWS, each has a service that carries it, a set of documented limits, and a few services that look like failover and aren't. The first rule of the whole page applies here too: a failover plan should use data-plane operations during the disaster, and call control planes only for steps that have been rehearsed.
Managed services that use it
| Service | Its role in a Region failover | What AWS documents |
|---|---|---|
| Amazon Application Recovery Controller (ARC) | Routing controls with safety rules (Seatly's switch); Region switch plans | Routing controls: each cluster is a data plane of endpoints in five Regions; "at least three out of the five" stay available for state changes, and you try them in turn; a state change reaches all five endpoints within 5 s on average and 15 s at most. UpdateRoutingControlStates changes several controls in one call. Safety rules are assertion rules and gating rules. The configuration API is in US West (Oregon) and is not highly available: create controls in advance. 70 per plan per month. Its execution blocks include ARC routing controls, Aurora Global Database, EC2 Auto Scaling groups, Route 53 health checks, manual approval and custom Lambda actions |
| Amazon Route 53 | Failover records whose answers follow the routing controls; the DNS every move rides on | Health checks every 30 s by default (10 s optional), failure threshold 3; healthy if more than 18% of checkers agree; when all records are unhealthy, all are treated as healthy; health-check metrics only in US East (N. Virginia). DNS answers are data plane, from 200+ edge locations; record edits are control plane. Accelerated recovery: opt-in, public hosted zones only, a 60-minute recovery-time target for resuming record changes |
| AWS Global Accelerator | Static anycast IPs with health-checked failover between endpoint groups | Own health checks for EC2 and Elastic IP endpoints; target-group health for ALB and NLB endpoints. Failover moves new connections; established TCP connections stay until the 340 s idle timeout (30 s for UDP). Failover ignores traffic dials, tries the three closest endpoint groups, then fails open. Traffic returns "in about 30 seconds or so" after recovery. Its API is in US West (Oregon) |
| Amazon Aurora (Global Database) | Orders and seats; switchover and failover | RPO "typically measured in seconds"; switchover RPO 0 and needs a healthy primary; failover with --allow-data-loss; write fencing is "a best-effort attempt", with an event if it times out; a snapshot rds:unplanned-global-failover-… is attempted; the old primary is rebuilt as a secondary in "a few minutes to several hours"; client DNS cache for the global writer endpoint "such as 5 seconds"; rds.global_db_rpo (minimum 20 s) blocks commits when every secondary lags more, and with two Regions its default should stay in the secondary's parameter group; a headless secondary needs a DB instance before a switchover or failover |
| Amazon DynamoDB (global tables) | MREC for carts and sessions; MRSC for the control table | MREC: asynchronous, "typically within a second or less", last writer wins; an isolated replica's unreplicated data is replicated "when the replica becomes healthy". MRSC: exactly three Regions (three replicas, or two and a witness, which "contains data"), RPO 0, no TTL, no transactions, ReplicatedWriteConflictException on concurrent writes to the same item; an isolated Region serves only eventually consistent reads. Transactions in global tables are ACID only in the Region where they were called. 99.999% availability SLA for global tables |
| Amazon S3 (Cross-Region Replication with RTC, Multi-Region Access Points) | Seat maps and ticket PDFs replicated to Ireland and served through an active-passive access point | RTC: "99.9 percent of those objects within 15 minutes". MRAP failover controls: set a Region's traffic dial to 100 (active) or 0 (passive); traffic moves in about 2 minutes, and existing connections aren't cut. The failover-control API is served in five Regions: US East (N. Virginia), US West (Oregon), Asia Pacific (Sydney), Asia Pacific (Tokyo) and Europe (Ireland). Seatly calls it in Ireland |
| AWS Fault Injection Service (FIS) | Game days | Cross-Region: Connectivity blocks the experiment Region's cross-Region traffic (transit gateways, route tables, VPC endpoints) and pauses S3, DynamoDB and MemoryDB replication; not Aurora Global Database; 3 hours by default, up to 12; no stop conditions included; its DynamoDB action cuts the experiment Region's replicas off from every other Region, so leave a lease table untagged unless that is the test. aws:ec2:asg-insufficient-instance-capacity-error injects insufficient-capacity errors into a target Auto Scaling group's requests (whether that includes warm-pool starts is not documented: verify in a test run); aws:ec2:api-insufficient-instance-capacity-error fails RunInstances, CreateCapacityReservation, StartInstances and CreateFleet only for targeted IAM roles. aws:rds:failover-db-cluster runs FailoverDBCluster, a failover within one cluster |
Running it yourself
| Option | What it is | Facts and sizing |
|---|---|---|
| Amazon EC2 | The standby fleet | The running On-Demand Standard instances quota (L-1216C47A) is counted in vCPUs, per Region, and counts running instances. Stopped warm-pool instances bill only their EBS volumes and Elastic IPs and don't count until they run. Capacity Reservations reserve capacity in one AZ, are billed whether used or not, and count toward the quota whether used or not |
| Amazon EC2 Auto Scaling | Warm pools; scaling out after the move | Warm pools keep pre-initialized instances stopped (or running, or hibernated). AWS: "You could also experience cold starts if an Availability Zone is out of capacity", and a warm-pool instance that fails to launch with an insufficient-capacity error is treated as a failed launch. Starting it is a control-plane call in the target Region |
| Amazon CloudWatch | Canaries attached to a VPC in Ireland and Stockholm; the CPS metric; lag alarms (AuroraGlobalDBRPOLag) | Put the alarms that decide the move in the Regions that decide it, never only in the home or in us-east-1 |
| Sizing in words | Target Region | Its own peak + the moved peak × the surge, per AZ, at most 80% busy; quota at least what will be running plus the surge, and at parity with the home; ARC cluster $1,825 a month |
Look-alikes that are not this mechanism
| Look-alike | Why it looks like Region failover | Why it isn't |
|---|---|---|
| Multi-AZ (RDS Multi-AZ, Aurora Replicas, ElastiCache replicas) | "Automatic failover" | Inside one Region: it survives an AZ, not a Region (Replication, Quorums & Read-Your-Writes) |
| ARC zonal shift and zonal autoshift | "ARC shifts traffic away from an impairment" | Between AZs in one Region; no additional charge |
| Route 53 latency routing, Global Accelerator's anycast | "Sends users to a healthy Region" | To the nearest healthy one, which is exactly why neither is the planned, residency-legal partner |
| Elastic Load Balancing | "Routes around failures" | Within one Region |
| AWS Backup cross-Region copies | "A copy in another Region" | The bottom rung of the ladder: hours to restore |
| CloudFront origin failover | "Fails over to another origin" | Per request, and only for GET, HEAD and OPTIONS; no promotion, no fence, no writes |
| RDS cross-Region read-replica promotion (PostgreSQL, MySQL) | "Promote the other Region" | A one-way operation: the replica "becomes a standalone DB instance" and can't be a replication target again. Replication is asynchronous, so RPO is the lag, and there is no managed switchover or rejoin (RDS for Oracle has a switchover; the others don't) |
What to remember from Part 12
- ARC routing controls and Region switch run on data planes spread across Regions; Route 53's record edits and health-check metrics depend on
us-east-1. - Aurora Global failover loses the lag and fences on a best-effort basis; switchover loses nothing.
- MREC's RPO is the replication delay; MRSC is three Regions, no transactions, no TTL.
Part 13. What you've learned
Back to Frankfurt at 20:00
The naive night cost about 25 minutes of broken checkout, 2.3 months of budget, and two writers. Seatly's Friday cost 6.75 minutes (4.9 with pre-approved rules). Here is what each piece did:
- Detection from outside (Part 2) ignored the 18:30 split, because only one vantage point failed and Frankfurt's CPS and lease stayed healthy (R2). At 20:00 both vantage points failed and CPS went missing: the rule fired at 20:01, from Ireland and Stockholm, without touching
us-east-1. - The gate (Part 3) refused the
us-east-1share of the default map, confirmed quota and prerequisites, and waited for a person, while the warm pool started anyway (R3). - The lease and the epoch (Part 4): Frankfurt stopped itself at 20:00:09; Ireland took the lease at epoch 8 only after
ack_until+ 5 s, promoted, and bumpedwriter_epochbefore any customer write (R4). - The flip (Part 5) went through the third ARC endpoint and was gated on
eu-west-1-writable; the DNS tail brought the bulk of clients over by 20:06:45. - The orders caught in the middle (Part 7): 120 of the 160 lost orders came back through the app journal with the same IDs; 40 and one double-sold seat went to reconciliation.
- Frankfurt came back refusing (R5), was rebuilt as a replica, had its queue drained against Ireland's inbox, and got its home back by switchover at epoch 9 (R6).
What it costs
- Data: 160 acknowledged orders missing until reconciliation, one seat sold twice; RPO 0 would cost a cross-Region round trip on every contested write.
- Time: 6.75 minutes (62.5% of a 99.99% month) with a human gate, or 4.9 (45.5%) with pre-approved rules.
- Money: a warm pool of 45 stopped instances (volumes only) plus 15 running, where a hot standby would keep all 60 running; ARC at 1,825 a month**.
- Discipline: an epoch on every write, a refusal that never expires, quota parity, and monthly rehearsals.
The whole story, event by event
Side rows and side replays are marked "(side)".
| # | Time | Event |
|---|---|---|
| K0 | last month (side) | Game day finds Ireland's quota at 512: the warm pool's 480 leave 32 vCPUs (4 instances) for a surge of 15. Raised to 640 |
| 1 | 18:00 | Baseline: Frankfurt home at epoch 7, 200 orders/s, lag 0.8 s; Ireland 15 running + 45 in the warm pool |
| 2 | 18:30:00 | The Ireland–Frankfurt link breaks; Ireland's vantage point fails at 18:30:30; lag grows 1 s per s |
| 3 | 18:30:30 | Stockholm's canaries fine, CPS in band, lease renewed via Stockholm: no evacuation; P2 lag alert |
| 3x | (side) | An Ireland-only rule would evacuate a healthy Frankfurt: 200 × 30.8 = 6,160 orders stranded; the lease condition would fail anyway |
| 3y | (side) | Had Frankfurt died at 18:49:59: 200 × 1,199.8 = 239,960 orders at risk; rds.global_db_rpo = 20 would stall Frankfurt's commits from 18:30:19.2 for 19.7 minutes |
| 3z | (side) | A self-fenced Frankfurt (control-table replica lost at 19:30:00, stop at 19:30:09) trips the rule's second clause at 19:31:00 |
| 4 | 18:50:00 | Link restored; lag 0.8 s by 18:52 |
| C1 | 19:00 | Bad config in cell 2; wave stopped; rolled back by 19:05: 0.5 customer-weighted min |
| C1x | (side) | Without cells: 2.0 min, and a proposal to spend a Region failover on a change one rollback undoes |
| 5 | 19:59:59.2 → 20:00:00 | 160 acknowledged orders not yet in Ireland; o-4417 holds seat 14C |
| 6 | 20:00:00 | About 60 requests in flight with no reply |
| 7 | 20:00:00.0 | Frankfurt goes dark; last renewal sent 19:59:59.0, ack_until 20:00:09.0 |
| 8 | 20:00:09.0 | Any surviving Frankfurt writer stops on its own monotonic clock |
| 9 | 20:00:10 / 20 / 30 | Both vantage points fail by 20:00:30 |
| 10 | 20:01:00 | Frankfurt's CPS missing for 60 s: the rule fires in Ireland and Stockholm; capacity starts |
| 10x | about 20:01:30 | Route 53's health check for Frankfurt goes unhealthy; advice only |
| 11 | 20:01:00 → 20:02:00 | Warm pool InService: 480 vCPUs ≤ 640; 15 new launches for the surge, healthy at 20:24 |
| 12 | 20:01:00 | The gate's six checks; us-east-1 share refused; 160 orders at risk < 1,000 |
| K2 | (side) | One accelerator for EU and US would send EU users' new connections to us-east-1 before promotion |
| K3 | (side) | An MRSC witness in us-east-1 would hold EU seat data |
| K4 | (side) | Keys, secrets, certificates and images must already be in Ireland |
| 13 | 20:03:00 | On-call approves |
| 17b | 20:03:00 | S3 Multi-Region Access Point routes flipped to Ireland at approval (about 2 minutes to take effect) |
| 13p | (side) 20:01:10 | Pre-approved rules approve |
| 14 | 20:03:00.0 | Ireland takes the lease: home Ireland, epoch 8, only if epoch 7 and ack_until < Ireland's clock − 5 s |
| 15 | 20:03:00 → 20:05:00 | Aurora failover with --allow-data-loss; fence of Frankfurt times out |
| 16 | 20:05:00.2 | writer_epoch = 8; eu-west-1-writable On |
| 16x | (side) | Switchover: waits for full sync, RPO 0, needs a healthy primary |
| 17 | 20:05:00.5 | One UpdateRoutingControlStates call through the third endpoint: Frankfurt Off, Ireland On |
| 18 | 20:05:15 | Route 53 answers with Ireland |
| 19 | 20:06:45 | Bulk of clients in Ireland: 405 s (295 s pre-approved) |
| 19r | (side) | Read-only mode: reads back at 285 s (225 s pre-approved) |
| 20 | 20:05 → 20:11 | Silent connections wait for heartbeats; 2% behind one resolver until 20:11 |
| K1 | 20:07 → 20:30 | Ireland at 100% through the surge, 67% after it, 53% once the 15 new instances join at 20:24 |
| 21 | 20:06 → 20:30 | 60 unknowns retried; 120 of the 160 lost orders re-sent under the same keys |
| 22 | 20:07:10 | Ireland sells 14C to o-5120: two owners |
| 22a | (side) | Without homes: 14C looks free in Ireland 0.4 s after Frankfurt sold it (drill 05) |
| 23 | 20:00 → 20:30 | Webhooks for unknown orders stored and forwarded to reconciliation |
| 23x | (side) | A handler that answers 200 and drops them loses the confirmations |
| U1 | (side) | us-east-1 dark: Route 53 metrics blank and record edits wait (60 minutes with accelerated recovery); DNS answers and ARC keep working |
| U2 | (side) | Global Accelerator's API is in US West (Oregon) |
| 24 | 20:40:00 | Frankfurt returns and refuses every write for good |
| 24x | (side) | Frankfurt trusting Aurora's fence instead: a write lands where it never reaches Ireland |
| 26 | 20:40 | MREC carts and sessions from Frankfurt's last second arrive, last writer wins |
| 25 | 20:45 | Snapshot of the old volume attempted; Frankfurt rebuilt as a secondary (2 h) |
| 27 | 20:50 | Frankfurt's 36 queued messages: 20 dropped, 4 processed, 12 to reconciliation |
| 28 | 21:10 | Reconciliation: 120 re-created, 40 honoured or refunded, 14C re-seated |
| 29 | Sat 06:00 → 06:02 | Failback: release, switchover, epoch 9, writer_epoch 9, eu-central-1-writable On, flip back, eu-west-1-writable Off; writes paused about 2 minutes |
| L1 | (side) | The ladder: pilot light 1,545 s, warm and hot standby 405 s, active-active 405 s for half |
| L2 | (side) | The US tier: backup and restore to us-west-2, RTO hours, by choice |
| 30 | end of month | Budget 4.32 min: 2.7 (night) + 0.5 (cell 2) = 3.2 min, 74.1%; 4.0 min (92.6%) if the failback's 0.8 counts |
| G1 | next month | Game day: FIS connectivity, each ARC endpoint, a switchover, the Auto Scaling insufficient-capacity action |
| 31 | 20:07:30 | One checkout through every layer (Part 11) |
The cheat card
| Topic | Remember |
|---|---|
| The five decisions | Is it the Region → do we move → is the standby safe → move the users → clean up |
| Home per key | One Region writes each key; failover moves homes with a new epoch |
| Orders at risk | Commit rate × lag: 200 × 0.8 = 160 |
| RTO | Detect + max(capacity, gate + promotion) + flip + TTL + backoff: 405 s (295 pre-approved) |
| Detection | Probes from two other Regions, plus a business metric where missing counts as down; never us-east-1, never Route 53 metrics |
| Split | One vantage point failed, metric fine: alarm, don't move |
| Gate | Evidence, capacity and quota, residency, lease + margin, one control stays on, data at risk; approval by a person or pre-approved rules |
| Lease | Home stops at its monotonic send + lease; taker waits ack_until + margin > worst gap between the two clocks, or rvn unchanged for the wait |
| Epoch | Bumped in the new home's database before its first write; checked in every write transaction |
| Two protections | Old home's own stop and permanent refusal (its orphaned copy); the epoch in the new database (everything reaching the new writer) |
| Aurora | Switchover: RPO 0, healthy primary. Failover: loses the lag, best-effort fence, snapshot attempted |
| Routing | ARC routing controls, pre-created, with assertion and gating rules; never edit records during the outage |
| DNS tail | TTL + resolvers + JVM 30 s + backoff + connections waiting for a heartbeat |
| Capacity | Moved load × surge ÷ 0.8, per AZ; warm pool within the gate's time; quota counts what will run |
| Residency | Target map decided in advance; every copy counts, including witnesses; one accelerator per zone |
| Failback | Replica first, reconcile by ID, drain old queues against the home inbox, switchover at a new epoch |
| Budget | 99.99% of 30 days = 4.32 min; weight by share hit |
Failure checklist
- Does any detection signal, dashboard or runbook step depend on the failed Region or on
us-east-1? - Does the rule require two vantage points and the business metric, and does a missing report count as down?
- Can DNS, a health check or an accelerator move users on its own to a Region that can't take writes?
- Does the old home stop by its own monotonic clock, and refuse writes on restart unless the lease names it, forever?
- Is the taker's margin larger than the worst gap between the holder's and the taker's clocks, and is each host's offset alarmed below half the margin?
- Is the epoch checked inside the new home's database on every write transaction, and bumped before the first one?
- Is the standby's quota at parity with the home, counted against what will be running after the warm pool and the surge start?
- Is every copy in the target map (standbys, witnesses, backups, keys) inside the allowed zone?
- Are keys, secrets, certificates and images already present in the standby Region?
- Do secondary endpoints store and forward everything they accept?
- After failover, do the old Region's workers stay stopped and its queues drain against the home's inbox?
- Has every step that calls a control plane in the target Region been rehearsed this quarter?
Think-first drills
Drill 1. A Region writes 500 orders a second with a replication lag of 1.5 s. A split between it and its standby lasts 4 minutes. How many acknowledged orders are at risk if the home dies (a) normally, and (b) at the end of the split? What would rds.global_db_rpo = 20 change?
Drill 2. Detection takes 60 s. Capacity is ready 60 s after detection. The gate takes 2 minutes, then promotion takes 90 s. The flip takes 15 s, the TTL is 30 s and the backoff 20 s. What is the RTO, and what share of a 99.99% (30-day) month does it spend if 30% of customers are affected?
Drill 3. Eight workers; each tenant is placed on 3 of them by shuffle sharding. How many shuffle shards are there, and what fraction of tenants share all three workers with a poisoned tenant? What if you instead split the eight workers into 2 cells of 4, and put each tenant on one whole cell?
Interview questions
| Question | Model answer |
|---|---|
| How do you decide a whole Region is down, and not just your view of it? | Probe it from outside, from at least two other Regions (VPC-attached canaries), and watch a business metric the Region reports into a store replicated in three Regions, counting a missing report as down. Move only when both vantage points fail and the metric is down or missing, or when the home has fenced itself and its metric is down. One failed vantage point with a healthy metric is a split: alarm, don't move. Never depend on the failed Region or on us-east-1 for the decision; Route 53 health-check metrics live in us-east-1, so they are advice. |
| Walk me through failing a database-backed service over to another Region: in what order, and why? | Detect from outside; start capacity at once (safe to undo); run the gate (capacity and quota, residency, the old lease plus margin, one Region stays on, data at risk) and get approval; take the lease with a condition on the old epoch; promote the database; bump the epoch in the new database before any write; only then flip pre-created routing controls, and count the DNS tail. Then resume unknowns with the same keys, reconcile the lost tail by ID, and bring the old Region back as a replica. The order exists so that users never land on a Region that can't write and two writers never overlap. |
| How do you stop the old primary from writing after you promote the new one, and for how long? | Two protections. The old home holds a lease in a three-Region strongly consistent table and stops at its own monotonic send time plus the lease; on restart it writes only if the lease names it, forever. That protects its own orphaned copy, which nothing in the new Region can reach. The new home bumps an epoch inside its database and every write transaction checks it, which refuses anything old that reaches the new writer. Aurora's own fence is best-effort, so it is a backup, not the plan. |
| Your standby runs at 25% capacity. What has to be true for the failover to work, and what does it cost? | The rest must start within the gate's time and fit the quota: here a warm pool of stopped, pre-initialized instances started at detection, with the quota counting what will run (480 vCPUs, plus the surge to 600, under 640). Size for moved load × surge at 80% per AZ, and rehearse a warm pool that can't start. It costs a quota at parity with the home, EBS volumes for the stopped instances, and a dependency on a control plane and on spare capacity during the disaster, which a hot standby (4× the running cost) removes. |
| Active-passive or active-active: what does each lose, and how long does each take? | Both lose the same lag if they replicate asynchronously: RPO is set by replication, not by the shape. Active-active moves only the failed Region's homes, so fewer customers are hit (half here: 1.35 instead of 2.7 customer-weighted minutes), and its capacity is tested daily; the moved half still waits for detection, gate, promotion and the DNS tail. Active-passive is simpler: no conflict rule, one writer. Near-zero RPO needs MRSC, a switchover (planned only) or synchronous replication for the writes that need it. |
| The old Region is healthy again. How do you bring it back? | It refuses writes because the lease names the new home. Rebuild it as a replica (Aurora attempts a snapshot of the old volume and adds a new one as a secondary), let late MREC data arrive, drain its queues against the new home's inbox with its workers stopped, and reconcile its lost tail by ID from the snapshot (or from the payment provider and client journals). Then fail back on purpose, at a quiet hour: release the lease, switch over with RPO 0, take the lease at a new epoch, bump the epoch in the database, flip in one batch and watch the business metric. |
Where to go next
- Leases, Fencing Tokens & Distributed Locks: the lease and margin arithmetic behind Part 4 (Part 7 there), and why an MREC table can't hold a lock (Part 8 there).
- Idempotency & Effectively-Once Processing: derived IDs, resuming unknowns and inboxes, which Parts 7 and 8 only count.
- Change Streams & the Transactional Outbox: the relay that resumes under the new epoch (Part 8).
- Replication, Quorums & Read-Your-Writes: what MREC, MRSC and Aurora Global Database promise (Parts 1 and 7).
- Event Time, Watermarks & Checkpoints: moving stream jobs with their home (Part 11).
- Write-Ahead Log, fsync & Group Commit: durability of one copy, the layer below all of this.
- Sharding, Hot Keys & Rebalancing: partitioning and the permanent "moved" refusal this page applies to a Region (Parts 1, 4 and 9).
- Drill: The Booking That Existed in Frankfurt but Not in Virginia, both questions answered in Part 7.
- Loops that move a Region: the key-value store (steps 3.1, 3.3, 3.4, 3.5), the message queue (steps 3.1, 3.2, 3.3), chat (step 3.5), payment (step 3.5), the digital wallet (step 3.4), hotel reservations (step 3.6), the transactional outbox (step 3.4), ride-sharing (step 3.6), Google Drive (step 3.4), email (step 3.4), the mobile stock-trading app (step 3.3), the Netflix case study (steps 3.2 to 3.5) and the Uber case study (step 3.4).