Shopify: Pods, Flash Sales and Payment Resilience
This page is one interview loop in three rounds, built on a real company's public engineering history. All three rounds design the same system: the part of Shopify that serves storefronts, takes buyers through checkout and charges their cards, for millions of shops that share one platform. Each round is an era. In each era, a new kind of load forced Shopify to change how that shared platform keeps one shop's big moment from hurting everyone else.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Era | About 2006–2015: one Rails app, one MySQL database | About 2015–2018: shards, pods and flash-sale defenses | About 2018–today: Google Cloud, Black Friday at planet scale, resilient payments |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Shops (published) | 100,000 (2014) | 600,000+ (March 2018) | "Millions of merchants" (2025) |
| Traffic we plan for | 3,500 storefront requests/s and 10 orders/s at the daily peak; one flash sale of 20,000 buyers (assumptions) | 100 pods; one sale of 100,000 buyers for 5,000 units, admitted at 250 checkouts/s (assumptions) | BFCM 2025 peak: 489M requests/min at the edge, 117M at the app servers, $5.1M of sales a minute (published) |
| Target | Stay up as merchants grow; never oversell, never charge twice | One merchant's sale never degrades the others | Checkout keeps working while payment providers, pods or a region fail |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
How to read a case study. Every claim about what Shopify actually did comes from Shopify's engineering blog, a talk by a Shopify engineer, Shopify's open-source code or Shopify's own announcements, and each round ends with a Sources list. We mark those claims Shopify published or (published). Where Shopify hasn't published the details, we say so and show a design that fits, labeled as ours. Numbers marked assumption are ours, chosen to make the arithmetic concrete. They are not Shopify's internal figures.
Infrastructure note. Shopify ran its own data centers for its first twelve years. In March 2018 it announced it was building its cloud with Google, and by then it had moved more than half of its data-center workloads to Google Cloud; a 2020 post says Shopify has partnered with Google Cloud Platform as its main cloud vendor since 2017 and runs on Google Kubernetes Engine. Where this page shows AWS services (Aurora, ElastiCache, DynamoDB and so on), that is a translation for this course, not Shopify's setup.
Loop Opener: Why Shopify?
A Mall Where One Shop's Sale Can Jam Every Entrance
Picture a shopping mall with a million small shops and a few famous ones, all sharing the same doors, corridors and cash registers. Most days, traffic spreads out. Then a celebrity announces a new lipstick on social media, and in ten seconds a hundred thousand people run for the same shop. If the mall is built naively, the crowd jams every entrance, and the bakery at the other end loses its customers too.
That is Shopify's problem. In Shopify's own words (2017), its merchants put it to the test daily: "on any given day, they'll hurl Super Bowl-sized traffic, often without notice." Its first big flash sale came in 2007, when fans rushed to buy T-shirts after the Super Bowl. In February 2016, a Kylie Cosmetics sale "took down not just her store, but all others on the database shard where her store was allocated". Shopify's stated policy (2017): "Shopify decided to never fire a customer for the amount of traffic they bring."
Synthesizing vector architecture diagram...
Read it as three eras. Round 1 is one application and one database; Round 2 splits the data into isolated pods and defends checkout during flash sales; Round 3 prepares the whole platform, including payments, for the biggest weekend of the year.
A few words we'll use all page:
| Word | What it means on this page |
|---|---|
| Tenant, shop, merchant | One business selling on Shopify. Shopify's data model hangs almost everything off a shop_id. |
| Storefront | The shop's public pages: home, collections, product pages. Mostly reads. |
| Checkout | The steps from "I want to pay" to an order: contact details, shipping, payment. Mostly writes, plus calls to outside payment providers. |
| Flash sale | A sale for a short time, often with limited stock, announced in advance. Shopify (2022 talk): the product "can sell out in seconds, even if there are thousands of items in inventory". |
| BFCM | Black Friday Cyber Monday: Shopify's name for the weekend from Black Friday to Cyber Monday, its busiest of the year. |
| GMV | Gross merchandise volume: the total value of orders placed through the platform. Shopify's BFCM "sales" figures are GMV. |
| Shard | A slice of the data on its own database server. |
| Pod | Shopify's word (not a Kubernetes pod) for a set of shops that live on a fully isolated set of datastores: its own MySQL, Redis and Memcached. |
| Payment provider | The outside company (a payment gateway or processor) that talks to card networks and banks for us. |
What Makes It Hard
- Traffic is spiky per tenant, not just per season. One shop can bring orders of magnitude more traffic than its baseline for a few minutes (Shopify, 2017). You can't plan capacity per shop.
- Stock is scarce and must not oversell. Thousands of buyers race for a few thousand units. Every "yes" must be backed by a real unit.
- Checkout depends on outside systems. A payment provider that slows down can hold our workers hostage, and a timeout doesn't tell us whether the card was charged.
- Everyone shares the platform. A shop's sale, a slow query or a failing provider must stay inside its own blast radius.
The Question the Whole Loop Answers
How do we let one merchant's biggest moment happen without hurting anyone else, and without overselling or charging anyone twice?
The answer grows every round:
- Round 1: cache the reads, make every stock decrement and every payment safe under concurrency and retries, and split background work by priority, all on one database.
- Round 2: isolate shops into pods with a router in front, move shops between pods without downtime, and put a fair queue and a stock gate in front of checkout during flash sales.
- Round 3: fail fast around slow dependencies with circuit breakers and bulkheads, prove it with failure injection and full-scale load tests, resolve unknown payment outcomes, and run pods across cloud regions that can take over for each other.
Round 1 · Mid-level · "Era 1: One Rails App, One Database"
~35 min · SDE II (L5) · about 2006–2015 · 100,000 shops (published, 2014) · 3,500 storefront requests/s and 10 orders/s at the daily peak (planning assumptions) · stay up as merchants grow, never oversell, never charge twice
R1.1 Establish Design Scope
The interviewer sets the scene: "It's 2014. We host 100,000 online stores. There's one Ruby on Rails application and one big MySQL database for all of them. Merchants run their own sales, and some run flash sales that bring a crowd in a minute. Last month one of them made the whole platform slow. Tell us how you'd keep it running, and what you'd fix first."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Who are the tenants? | Shops. "100,000 businesses now use Shopify to power their stores" (Shopify, 2014). They share one app and one database; every row carries a shop_id. | Every query is scoped by shop_id, and every index starts with it. |
| What data? | Products and variants, inventory counts, carts, checkouts, orders, customers. | Reads dominate on storefronts; writes dominate in checkout (R1.2). |
| How do payments work? | Through outside payment providers. We send a charge request and wait for the answer. | Every charge must be safe to retry (step 1.3). |
| How does checkout write? | Shopify (2017): "every checkout session created a new record in MySQL" and "every step in the flow modified that same record". | A checkout is several writes, and a crowd of checkouts is a burst of writes (R1.7). |
| What runs in the background? | Emails, webhooks to merchants' apps, payment follow-ups. Shopify's first job system, Delayed Job, stored jobs in the database (published, 2022 talk). | Jobs compete with checkout for the same database (step 1.4). |
| What's breaking? | Database contention. One shop's sale adds thousands of writes a second, and every shop on the database feels it. | The shared database is everyone's single point of contention. We fix what we can this round; Round 2 splits it. |
Out of scope for this round: sharding and pods (Round 2), waiting rooms (Round 2), multiple regions (Round 3).
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Operation |
|---|---|
| "Buyers browse a shop" | GET storefront pages: home, collection, product; each shows prices and "in stock" |
| "Add to cart" | The cart lives in the buyer's browser until checkout starts (Shopify, 2022 talk: the cart "is purely client side until you actually start storing some buyer information") |
| "Start checkout" | createCheckout(shop, line items) stores the buyer's contact and shipping details |
| "Pay" | completeCheckout(checkout, payment) charges the card through a provider |
| "Get an order" | A successful payment turns the checkout into an order and decrements stock |
| "Tell everyone" | Confirmation email, webhooks to the merchant's apps, analytics, all in the background |
Not yet: pods, a checkout queue, a stock gate, resilience to a whole region failing.
R1.3 Non-Functional Requirements: the Questions
We name each quality first; the numbers come in R1.7.
- Inventory correctness. Never sell a unit we don't have, unless the merchant chose to (Shopify lets a merchant keep selling when out of stock, per product variant:
inventoryPolicyDENYorCONTINUE). - Payment correctness. A buyer is charged at most once per order, even when requests are retried.
- Availability. One shop's traffic must not take down the others. (This round can only soften that; Round 2 enforces it.)
- Storefront latency. Product pages must load fast: most traffic is window shopping, and slow pages cost sales.
- Checkout latency. The buyer is staring at a spinner. Shopify's 2022 payments post suggests a one-second connect timeout and a five-second read timeout as "a decent starting point", from the view of a person who won't wait longer.
R1.4 The API
This API is illustrative: Shopify's internal checkout endpoints aren't public in this form, so the paths and fields below are ours.
Start a checkout
httpPOST /v1/shops/tshirthero/checkouts HTTP/1.1 Content-Type: application/json Idempotency-Key: 01J9ZQ3C6X8R2M4T7V1B5N0K9D { "line_items": [ { "variant_id": 4471, "quantity": 1 } ], "email": "sam@example.com", "shipping_address": { "country": "CA", "zip": "H2X 1Y4" } }
httpHTTP/1.1 201 Created Content-Type: application/json { "checkout_id": "chk_5f1c2a9e", "status": "OPEN", "total": { "amount": "32.00", "currency": "CAD" } }
Complete it (pay)
httpPOST /v1/shops/tshirthero/checkouts/chk_5f1c2a9e/complete HTTP/1.1 Content-Type: application/json Idempotency-Key: 01J9ZQ4F2H7W3K8P1Q6S0D5M2A { "payment": { "card_token": "ct_91ab", "amount": "32.00", "currency": "CAD" } }
httpHTTP/1.1 200 OK Content-Type: application/json { "checkout_id": "chk_5f1c2a9e", "status": "COMPLETED", "order_id": "ord_8842" }
Idempotency-Keynames one attempt. A retry with the same key returns the same result instead of charging again (step 1.3). We use a ULID (a sortable ID: 48 bits of time, then 80 random bits), the choice Shopify published in 2022, because keys that increase over time insert at the end of a B-tree index.card_token, never a card number. Shopify (2022 talk) keeps its main app out of card data entirely: the card fields are iframes on a separate domain that post to a service called CardSink, which returns a token. We come back to this in Round 3.
Status codes
| Code | Meaning |
|---|---|
201 Created / 200 OK | Done; a repeat of the same key returns the stored response |
202 Accepted | Payment outcome not known yet; poll GET .../checkouts/{id} |
409 Conflict | The same key is still being processed; retry shortly |
409 Conflict with "error": "SOLD_OUT" | A line item's stock ran out |
422 Unprocessable Entity | Card declined, or the same key was sent with a different body |
A checkout's states
Synthesizing vector architecture diagram...
A checkout only moves along these arrows. UNKNOWN is a real state, not an error: a timeout means we don't know yet, and the buyer must not be allowed to start a second charge while we find out.
Recap
- Storefronts are read-heavy and cacheable; checkout is write-heavy and talks to outside providers.
- Every write that matters is idempotent, and every checkout has a small set of explicit states.
R1.5 Design Evolution: Making One Database Survive
Each step is a problem, your turn to think, the answer, and what it costs us.
Step 1.0: The Baseline
Synthesizing vector architecture diagram...
One application, one database. Storefront reads, checkout writes and the background job queue all land on the same MySQL server.
Shopify published this shape: "Our main tool of choice for building backend system is Ruby on Rails with MySQL, Redis, and memcached as our datastores" (2022 talk), and its main monolith is "under continuous development since at least 2006" (2020). Until 2015, Shopify kept "buying a larger database server" (2018).
Step 1.1: Storefront Reads Crush the Database
The problem: at the daily peak, storefronts send 3,500 requests a second. An uncached product page runs about 15 queries, so the database answers 52,000 queries a second just for window shoppers, and checkout writes wait behind them. What would you do?
Primitive: Distributed Cache Patterns and Eviction · Drill: The Product Page That Melted Redis (answered here: the cold-cache stampede and why "cache forever" is the wrong trade)
Step 1.2: Two Buyers Got the Last Unit
The problem: a shop has one hoodie left. Two buyers click Pay at the same moment. Both requests read available = 1, both see enough stock, both create an order. The merchant now owes a hoodie it doesn't have.
What would you do?
Synthesizing vector architecture diagram...
The second update waits for the first to commit, then re-reads the row and finds nothing left. The row lock is the queue.
Primitive: Database Isolation Levels, ACID and Concurrency Anomalies · Loop primitive: Write-Ahead Log, fsync & Group Commit (why a commit's flush is part of every lock hold) · Loop: Design a Hotel Reservation System (Round 1, step 1.2: the same race for the last room)
Step 1.3: A Buyer Was Charged Twice
The problem: a buyer clicked Pay, the provider charged the card, and the response was lost on the way back. The browser retried. The retry looked like a new payment, and the buyer was charged twice. What would you do?
The key table (our schema; same database and shard as the checkout, so the claim and the order can share a transaction):
sqlCREATE TABLE payment_requests ( shop_id BIGINT UNSIGNED NOT NULL, idempotency_key CHAR(26) NOT NULL, -- ULID checkout_id BIGINT UNSIGNED NOT NULL, request_hash BINARY(32) NOT NULL, -- SHA-256 of the body; a different body under the same key gets 422 status ENUM('IN_FLIGHT','COMPLETED','UNKNOWN') NOT NULL, locked_until TIMESTAMP(3) NOT NULL, -- lease on IN_FLIGHT, from the database clock: NOW(3) + 30 s provider_charge VARCHAR(64) NULL, response_code SMALLINT NULL, response_body JSON NULL, created_at TIMESTAMP(3) NOT NULL DEFAULT CURRENT_TIMESTAMP(3), PRIMARY KEY (shop_id, idempotency_key) ) ENGINE=InnoDB;
Loop primitive: Idempotency & Effectively-Once Processing (Part 3: the durable guard, claimed before the side effect) · Loop: Design a Payment Processing System (Round 1, steps 1.1 and 1.2, in depth)
Step 1.4: Background Jobs Pile Up
The problem: every order enqueues four jobs: a confirmation email, webhooks to the merchant's apps, analytics and a fulfillment notice. During a flash sale the queue grows to hundreds of thousands of jobs. Payment follow-ups wait behind marketing emails, and the jobs table, which lives in the same MySQL, adds its own inserts, updates and deletes to the load. What would you do?
Loop primitive: Change Streams & the Transactional Outbox (Part 1: why writing to two systems fails; Part 3: the relay) · Primitive: Message Queues vs Event Streams
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | One Rails app, one MySQL | One database for everyone |
| 1.1 | Storefront reads crush the database | Page cache with per-shop versions, a CDN, request coalescing | Invalidation; stale pages |
| 1.2 | Two buyers got the last unit | Conditional decrement, hot-row update last | A row that serializes sales |
| 1.3 | A buyer was charged twice | Idempotency key claimed before the provider call, passed to the provider | An extra insert; UNKNOWN outcomes |
| 1.4 | Jobs pile up | Resque on Redis, queues by priority, an outbox | More pools; idempotent jobs |
R1.6 Architecture v1
Synthesizing vector architecture diagram...
Reads go to the cache first, then to replicas for browsing; anything that decides stock or money goes to the primary. Jobs leave the database through the outbox and wait in Redis, one queue per priority.
Trace: one checkout completes (our design around the published shape)
Synthesizing vector architecture diagram...
The key is claimed and committed before the provider is called, and the hot inventory row is touched last, just before the commit. If the stock update finds 0 rows, the order rolls back and the charge is voided (we build that path properly in Round 3).
Our AWS translation (not Shopify's setup). The monolith on EKS behind a Network Load Balancer; CloudFront for assets; one Aurora MySQL cluster with readers; ElastiCache for the page cache and for the job queues.
Sources for this round
- Shopify, 100,000 Online Stores Now Use Shopify, Shopify blog, 2014: the 100,000-store milestone.
- Eskildsen, Building and Testing Resilient Ruby on Rails Applications, Shopify Engineering, January 2015: BFCM 2014 preparation; the session store example.
- Stolarsky, Surviving Flashes of High-Write Traffic Using Scriptable Load Balancers (Part I), Shopify Engineering, February 2017: the 2007 flash sale; one checkout record updated at every step; the dynamic caching layer that "helped but not enough".
- Denis, A Pods Architecture To Allow Shopify To Scale, Shopify Engineering, March 2018: buying ever-larger database servers until 2015.
- Müller, Under Deconstruction: The State of Shopify's Monolith, Shopify Engineering, September 2020: the monolith in development since at least 2006.
- de Water, Shopify's Architecture to Handle the World's Biggest Flash Sales, QCon Plus talk, published by InfoQ in October 2022: Rails, MySQL, Redis and Memcached; the client-side cart; CardSink; Delayed Job and Resque.
- de Water, 10 Tips for Building Resilient Payment Systems, Shopify Engineering, July 2022: timeouts; idempotency keys and ULIDs; "once in a million".
Everything else in this round (the cache design, the API shapes, the tables, the job tiers, the outbox and every number in R1.7 not marked published) is our design or our assumption.
R1.7 Numbers
Figures marked published come from the sources above; the rest are assumptions (2014 planning figures) or derived from them.
Storefront reads at the daily peak
| Quantity | Arithmetic | Result |
|---|---|---|
| Shops | Published (2014) | 100,000 |
| Storefront requests a day | 100,000 × 1,000 per shop (assumption) | 100M |
| Average | 100M ÷ 86,400 s | 1,157/s |
| Daily peak | × 3 (assumption) | ≈ 3,470/s |
| Queries per uncached page | Assumption | 15 |
| Queries with no cache | 3,470 × 15 | ≈ 52,000/s |
| Uncached pages at an 85% hit rate | 3,470 × 0.15 | ≈ 521/s |
| Queries with the cache | 521 × 15 | ≈ 7,800/s |
Checkout writes at the daily peak
| Quantity | Arithmetic | Result |
|---|---|---|
| Orders a day | 100,000 shops × 3 (assumption) | 300,000 |
| Peak orders | 300,000 ÷ 86,400 × 3 | 10.4/s |
| Checkouts started | 3 per order (assumption) | 31.2/s |
| Writes to checkout rows | 31.2 × 6 steps (assumption) | 187/s |
| Order writes | 10.4 × 10 rows (assumption) | 104/s |
| Job-table writes (Delayed Job style) | 10.4 × 4 jobs × 3 writes (insert, lock, delete) | 125/s |
| All peak writes | 187 + 104 + 125 | ≈ 416/s |
One flash sale on the shared database
| Quantity | Arithmetic | Result |
|---|---|---|
| Buyers reaching checkout | 20,000 in 60 s (assumption) | 333/s |
| Extra checkout-row writes | 333 × 6 | 2,000/s: 4.8× the whole platform's normal peak; the total becomes 2,416/s, 5.8× |
| Buyers pressing Pay | 60% of them (assumption) | 200/s |
| Hot-row ceiling, update first | 1 ÷ 8 ms | 125/s |
| Backlog on the hot row | 200 − 125 | +75 waiting transactions each second |
| Free connections | 1,000 max_connections − 300 in normal use (assumptions) | 700 |
| Time to exhaust connections | 700 ÷ 75 | ≈ 9.3 s |
After about nine seconds, every shop on the database starts getting "too many connections" errors, not just the one running the sale. With the hot-row update moved last (about 1,000/s ceiling), this backlog doesn't form, but the 2,000 extra writes a second remain. That is the wall this era hits: Shopify's February 2016 outage took down every shop on the sale's database shard.
Synthesizing vector architecture diagram...
One shop's sale is almost five times the whole platform's normal write peak. Caching can't help: these are writes.
R1.8 Trade-Offs
A monolith vs services (Shopify's published stance)
Shopify chose a modular monolith: "we would keep all of the code in one codebase, but ensure that boundaries were defined and respected between different components" (2019). By 2020 the core monolith had "over 2.8 million lines of Ruby code", and in the 2022 talk it was deployed "around 40 times a day". Boundaries are checked automatically with Packwerk, an open-source static analysis tool that flags one component reaching into another's internals (2022 talk).
| Modular monolith (Shopify's choice) | Microservices | |
|---|---|---|
| Calls between parts | In-process: no network, no timeouts | Remote: timeouts, retries, partial failure |
| Transactions | One database transaction can cover order, payment record and stock | Needs sagas or outboxes between services |
| Deploys | One app, deployed often | Each service on its own schedule |
| Boundaries | Enforced by tooling (Packwerk) | Enforced by the network |
| Scaling the data | Shard the data under the one app (Round 2) | Each service owns and scales its own store |
The honest point: a monolith doesn't stop you from scaling the data. Round 2 shards the data under the same app.
Where the stock check lives
| Check in the app, then write | Conditional UPDATE (chosen) | SELECT ... FOR UPDATE, then write | |
|---|---|---|---|
| Correct under concurrency | No | Yes | Yes |
| Round trips | 2 | 1 | 2, with the lock held between them |
| Lock hold | None (and wrong) | Shortest | Longest |
R1.9 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Database saturated by one shop's sale | Slow queries, then "too many connections" for every shop | This round: hot-row update last, caches for reads, low-priority jobs paused. The real fix is Round 2: isolate shops and queue the crowd. |
| Cache stampede after a deploy or a version bump | A spike of identical queries | Request coalescing and serving the old page while one request rebuilds it; jittered expiry. |
| Payment provider slow | Checkout requests and payment jobs wait for seconds; workers run out | Low timeouts (one second to connect, five to read), a separate worker pool for payments. Round 3 adds circuit breakers. |
| Provider times out after charging | The buyer retries | The same idempotency key; the provider recognizes the retry. The checkout stays UNKNOWN until we know. |
| Session store (Redis) down | Buyers can't sign in | Shopify's own example (2015): the session store only matters for sign-in, so "customers can still purchase products as guests". Degrade, don't fail. |
| Relay stops | Emails and webhooks stop; orders still commit | Outbox rows wait safely; alarm on the oldest unsent row's age. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Idempotent payment requests and jobs; checkout keeps working when the session store fails; queues split by priority so emails can't starve payments REL 4 · REL 5 |
| Performance Efficiency | A versioned page cache and replicas for browsing; the hot-row update placed last PERF 3 |
| Security | Card numbers never reach the app: card fields post to a separate tokenizing service, and the app only sees a token SEC 7 |
| Operational Excellence | Low timeouts everywhere, chosen from what a waiting buyer will accept; alarms on queue age OPS 8 |
| Cost Optimization | Skipped this round: one database, and the constraint is contention, not money. |
| Sustainability | Skipped this round: the fleet is small. |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Separates cacheable storefront reads from checkout writes, and knows caches don't fix writes.
- Makes the stock decrement one atomic conditional statement, and knows where the lock is held.
- Claims an idempotency key durably before calling the provider, and passes a key to the provider too.
- Splits background work by priority, and spots the dual write when jobs leave the database.
- Computes why one shop's sale can saturate a shared database.
Follow-up questions
-
"Why not
SELECT ... FOR UPDATEon the inventory row, then decide in Ruby?" Answer: it's correct, but it holds the row lock across a round trip to the app and back, which lowers the hot row's ceiling. The conditionalUPDATEdecides inside the database in one step. -
"The idempotency key is in the same database as the order. What if that database is down?" Answer: then checkout is down anyway: we can't create orders. Putting the key next to the order is what lets the claim, the order and the stock change commit together. A key store in a separate system would add a dual write.
-
"A buyer pays, the stock update finds 0 rows. Now what?" Answer: roll back the order, keep the payment request's record, and void the authorization with the provider (a void releases the hold on the card; nothing was captured). Tell the buyer the item sold out. Round 2 makes this rare by holding stock before payment.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Joins instead of a cache" | Same rows read for every page view. |
| "Check stock, then write" | Two requests both pass the check. |
| "Disable the button to stop double charges" | Retries come from browsers, proxies and jobs too. |
| "One job queue, more workers" | Payments still wait behind emails. |
"MySQL CHECK protects stock" | Only enforced from MySQL 8.0.16; the conditional WHERE is the guard. |
Round 2 · Senior · "Era 2: Pods and Flash-Sale Defenses"
~40 min · Senior SDE (L6) · about 2015–2018 · 600,000+ merchants (published, March 2018) · 100 pods and one sale of 100,000 buyers for 5,000 units (planning figures) · one merchant's sale never degrades the others
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "Shopify is one Rails monolith, and until 2015 it ran on one big MySQL database for every shop. Storefronts are reads, so we cached rendered pages in Memcached with a per-shop version for instant invalidation, coalesced rebuilds to avoid stampedes, and kept stock and money decisions on the primary. Stock is a conditional
UPDATE ... WHERE available >= qty, placed last in the transaction, because the row lock lasts until commit and one hot row can only take about 1 ÷ (lock hold time) orders a second. Payments claim an idempotency key with an insert-if-absent before calling the provider, and pass a key to the provider too; a timeout isUNKNOWN, not failed. Jobs moved from the database to Redis queues split by priority, fed through an outbox. Open costs: one database for everyone, so one flash sale adds several times the platform's normal write load, and every shop on the database goes down with it."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: reads cached, writes safe, but every shop's writes still land on one database.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Reads crush the database | Versioned page cache, coalescing | Stale pages |
| 1.2 | Two buyers, one unit | Conditional decrement, placed last | A hot row |
| 1.3 | Double charge | Idempotency key claimed before the call | UNKNOWN outcomes |
| 1.4 | Jobs pile up | Redis queues by priority, an outbox | Idempotent jobs |
Open costs: a single shared database, no way to hold back a crowd, and no way to keep one shop's sale away from its neighbors.
R2.1 The Scope Raise
Interviewer: "It's 2016. We've sharded the database, and it bought us capacity, but in February a celebrity's lipstick sale took down every shop on her shard. She sells again next week, and the week after. We're heading for hundreds of thousands of merchants. We need one merchant's sale to never hurt the others, a crowd at checkout that doesn't crash anything and doesn't feel like a lottery, no overselling, and bots not buying everything."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How did sharding go? | Shopify (2018): "what we gained in performance and scalability we lost in resilience". Code like Sharding.with_each_shard meant that "If any of our shards went down, that entire action would be unavailable across the platform." | Isolation, not just sharding (step 2.1). |
| How big is a sale? | Kylie Cosmetics sales came "every seven or eight days", and "Items quickly sold out, sometimes in minutes" (Shopify, 2017); by 2018 Shopify described a Kylie flash sale that sells out "in 20 seconds". | We plan for 100,000 buyers in the first minute for 5,000 units (assumption). |
| What exactly fails? | "Checkout with its massive number of writes was still taking down segments of the platform" (2017). | Hold the crowd back before checkout writes (step 2.3). |
| Can we rewrite checkout to write less? | Shopify (2017): "The ideal solution would be a refactor of the entire checkout flow to condense the number of writes into one", but "the next crippling sale was the following week". | Protect first, refactor later. |
| How many merchants? | "600,000+ merchants on our platform today" (March 2018). | Many shops per database unit; they must be movable. |
| What about bots? | Limited products "could fetch double or triple the original price on the secondary market", and "Bots also hammer our systems much more than real buyers do" (2022 talk). | Bot filtering at the edge and per-buyer limits (step 2.5). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Shops | 100,000 | 600,000+ (published, 2018) |
| Data layout | One MySQL | ~100 pods, each with its own MySQL, Redis and Memcached (count is our assumption) |
| Biggest burst | 20,000 buyers in a minute | 100,000 buyers in a minute for 5,000 units (assumption) |
| Crowd at checkout | Everyone let through | A throttle with a fair queue, sized to the pod |
| Stock | Checked at payment | Held for a few minutes at checkout, with a cheap gate in front |
| Target | Stay up | A sale never degrades other shops |
The "Not yet" list from R1.2 comes back: pods, a checkout queue and a stock gate are in scope now.
R2.2 What Breaks in the Round 1 Design
| Round 1 piece | What breaks at the new scale |
|---|---|
| One database for every shop | One sale adds several times the platform's write peak; a failure anywhere is a failure everywhere. |
| Shards without isolation | Cross-shard code paths make every shard a dependency of every request that loops over shards. |
| Letting every buyer into checkout | 100,000 checkouts start in a minute, each writing its checkout row at every step. The shard's writes and connections run out in seconds. |
| Stock checked only at the database | Every hopeless request after the sellout still takes a row lock to learn "no". |
| Shared app workers | Workers waiting on a slow shard aren't serving anyone else. |
R2.3 New Requirements and API Additions
1. Every request belongs to one pod. Shopify (2021) keeps a routing table that maps each shop's domain to a pod ID, "in a separate database that isn't sharded". The shape below is ours:
sqlCREATE TABLE shop_routes ( domain VARCHAR(255) NOT NULL, -- 'tshirthero.com' or 'tshirthero.myshopify.com' shop_id BIGINT UNSIGNED NOT NULL, pod_id SMALLINT UNSIGNED NOT NULL, route_version BIGINT UNSIGNED NOT NULL, -- bumped on every move PRIMARY KEY (domain), KEY idx_shop (shop_id) ) ENGINE=InnoDB;
A request then carries the answer to the app as a header (Shopify's load balancers add one; the name here is ours):
httpGET /products/hoodie HTTP/1.1 Host: tshirthero.com X-Pod-Id: 17 X-Route-Version: 42
2. A queue pass for checkout during a sale. When the checkout throttle is engaged for a shop, a buyer who gets through receives a signed cookie. Shopify published that the cookie is "securely signed", carries the buyer's first-attempt timestamp while they wait, and lets an admitted buyer "skip the throttle for the rest of their session". This payload is our design:
json{ "v": 1, "kid": "q-2016-11", "shop_id": 812, "first_seen_ms": 1478530800123, "admitted_ms": 1478530811456, "exp_ms": 1478531711456, "session_hash": "b3f1c0e2", "pass_id": "01BX5ZZKBKACTAV9WEVGEMMVRZ" }
The cookie is base64url(payload) + "." + base64url(HMAC-SHA256(key[kid], payload)). Every load balancer verifies it without a data store:
- Reject a cookie over 1 KB or one that doesn't parse.
- Look up
kidin the load balancer's local set of signing keys (we rotate keys; the old key stays valid until every pass it signed has expired). - Recompute the HMAC and compare in constant time. A mismatch means forged or edited.
- Check
shop_idagainst the shop Sorting Hat resolved for this request. A pass for one shop is useless at another. - Check
exp_msagainst the load balancer's clock, allowing 2 seconds of clock skew. The edge wrote both timestamps, so the buyer can't move them. - Check
session_hashagainst a hash of the buyer's session cookie, so a pass copied to another browser fails. - Let the request through. The pass is not single-use at the edge (that would need a shared store); instead, the pod's database allows one stock hold per
pass_id(next item), which is what stops a shared pass from buying twice.
3. Stock is held during checkout. Shopify hasn't published how its checkout holds stock during a sale, so this is a design that fits: when a buyer with a pass starts checkout, we hold the units for 10 minutes; payment turns the hold into a sale; an expired hold gives the units back.
sqlCREATE TABLE inventory_holds ( shop_id BIGINT UNSIGNED NOT NULL, hold_id CHAR(26) NOT NULL, -- ULID variant_id BIGINT UNSIGNED NOT NULL, bucket TINYINT UNSIGNED NOT NULL, -- which stock bucket it came from (step 2.4) quantity INT UNSIGNED NOT NULL, checkout_id BIGINT UNSIGNED NOT NULL, pass_id CHAR(26) NULL, -- one hold per queue pass status ENUM('ACTIVE','PAYING','CONVERTED','RELEASED') NOT NULL, expires_at TIMESTAMP(3) NOT NULL, -- set from the database clock PRIMARY KEY (shop_id, hold_id), UNIQUE KEY uq_pass (shop_id, pass_id), -- NULLs don't collide in MySQL, so shops without a queue are unaffected KEY idx_expiry (status, expires_at) ) ENGINE=InnoDB;
Synthesizing vector architecture diagram...
A hold is released exactly once, by the transition that wins its conditional update. Only an ACTIVE hold has a timer. A hold in PAYING is never released by a timer: it leaves PAYING only when its payment attempt is final (DECLINED sends it back to ACTIVE with a new timer, VOIDED releases it, success converts it), so a slow card challenge or an UNKNOWN payment can't lose the buyer's unit.
Recap
- Every request and job names one pod; the routing table is a lookup, so any shop can live on any pod.
- A queue pass is verified with a key and a clock, no shared store.
- Stock is held with an expiry set by the database's clock, and released exactly once.
R2.4 Design Evolution: Isolate the Tenants, Then Defend Checkout
Step 2.1: One Database Unit for Hundreds of Thousands of Shops
The problem: we've already sharded MySQL, but code still loops over every shard, so when one shard is down, features break for every shop. And a sale on one shard still hurts the whole platform through shared caches, queues and workers. What would you do?
Synthesizing vector architecture diagram...
The load balancer decides the pod before the app sees the request. Workers are shared, datastores are not, and nothing in the request path reads two pods. Cross-shop questions go to the warehouse.
Primitive: Database Sharding and Partition Keys · Drill: One Customer, One Shard, One Outage (answered here and in step 2.2: the biggest tenant gets its own pod through a pinned route before its load arrives, and platform-wide reports read the warehouse, not the pods)
Step 2.2: One Merchant's Sale Slows Its Pod-Mates
The problem: pod 1 holds two fast-growing merchants who, it turns out, run their flash sales on the same days. Their pod's database is at 90% while others idle at 25%. Next week both have sales again. What would you do?
Loop primitive: Leases, Fencing Tokens & Distributed Locks (Part 4: the pause and the zombie writer) · Loop primitive: Change Streams & the Transactional Outbox (Part 5: reading the database's own log)
Step 2.3: A Hundred Thousand Buyers Hit Checkout at Once
The problem: a sale opens. 100,000 buyers press Checkout within a minute. The pod can safely complete about 250 checkouts a second. Everyone else will hit the database anyway unless something holds them back, and whatever holds them back must not look like the site crashed. What would you do?
Synthesizing vector architecture diagram...
Everything above the app runs in the load balancers with no shared store: the cookie carries the buyer's place in line, and each load balancer computes its own threshold.
Primitive: Distributed Rate Limiting · Drill: The Partner Whose Retry Loop Took Down Everyone Else (answered here: per-node limits must be divided across nodes, and a reset-every-period bucket allows a double burst at the boundary where a token bucket doesn't)
Step 2.4: The Hot Stock Row Locks Up
The problem: 250 buyers a second get into checkout for one hoodie with 5,000 units. Each needs a unit held, and each payment turns a hold into a sale. After the sellout, admitted buyers still ask for stock that isn't there. Every one of those asks takes the same row lock. What would you do?
Synthesizing vector architecture diagram...
The gate turns away the crowd after the sellout; the database decides every real claim. Dotted lines are the return path that keeps the gate from underselling.
Primitive: Distributed Locks and Leases · Loop primitive: Leases, Fencing Tokens & Distributed Locks (Part 3: whose clock decides; Part 5: conditional writes are the check) · Loop: Design a Hotel Reservation System (Round 2, steps 2.2 and 2.3)
Step 2.5: Bots Buy Everything
The problem: the sneaker drop sold out in 20 seconds, and the merchant's customers say they never had a chance. Resale sites list the pairs at triple the price an hour later. Bots opened thousands of sessions, joined the queue early and checked out faster than any person. What would you do?
Primitive: Bot Defense, Sybil Resistance and Registration Abuse · Primitive: Cloudflare Turnstile Managed Challenges
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | One database unit for everyone | Pods with isolated datastores, shared workers, Sorting Hat and a routing table | No cross-shop queries in the app |
| 2.2 | A sale slows its pod-mates | Shop moves with Ghostferry, a writer lock and a routing update; per-shop caps | Seconds of blocked writes per move |
| 2.3 | 100,000 buyers at checkout | Edge leaky bucket, a cached queue page, a fair threshold from a P controller | Buyers wait; the rate is set by load tests |
| 2.4 | The hot stock row | Redis gate, database holds with expiry, bucketed stock rows | A counter to keep honest |
| 2.5 | Bots buy everything | Edge scoring, session-bound passes, per-buyer limits | False positives |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
A request is filtered, routed and, during a sale, throttled before any Rails code runs. The shared workers talk to exactly one pod per unit of work. Moving a shop is a copy plus a routing-table update.
Trace 1: a flash-sale checkout (our design around the published shape)
Synthesizing vector architecture diagram...
The buyer's place in line lives in their cookie; their unit lives in a hold row. The gate only decides who gets to ask the database.
Trace 2: moving a shop from pod 1 to pod 2 (published phases, our timings)
Synthesizing vector architecture diagram...
Only the part between taking and releasing the writer lock blocks the merchant: 2.5 seconds in our chain.
Our AWS translation (not Shopify's setup). The edge as OpenResty on EKS behind a Network Load Balancer, with CloudFront and AWS WAF in front; each pod as its own Aurora MySQL cluster, ElastiCache (Valkey) for the gate and job queues, and ElastiCache for Memcached; the routing table in a small unsharded Aurora MySQL cluster (a DynamoDB table would also fit); shared Rails workers on EKS.
Sources for this round
- Stolarsky, Surviving Flashes of High-Write Traffic Using Scriptable Load Balancers, Part I and Part II, Shopify Engineering, February 2017: Kylie's weekly sales and the February 2016 shard outage; Nginx and Lua; the leaky bucket with 5-second periods; the cached queue page and signed cookie; the 40-minute lottery; timestamps, the threshold and the P controller; why not a data store.
- Denis, A Pods Architecture To Allow Shopify To Scale, Shopify Engineering, March 2018: sharding in 2015;
with_each_shard; pods; shared workers; one pod per unit of work; Sorting Hat; pairs of data centers and the Pod Mover. - Neufeld, Shopify's Infrastructure Collaboration with Google, Shopify Engineering, March 2018: 600,000+ merchants; Kylie selling out in 20 seconds; the Shop Mover's 2.5-second average downtime.
- Madan, Shard Balancing: Moving Shops Confidently with Zero-Downtime at Terabyte-scale, Shopify Engineering, September 2021: pods with MySQL, Redis and Memcached; the routing table; four-times and two-times utilization spreads; Ghostferry's phases; the multi-reader, single-writer lock; verification and pruning; TLA+.
- de Water, Shopify's Architecture to Handle the World's Biggest Flash Sales, QCon Plus talk, published by InfoQ in October 2022: shared stateless workers; OpenResty, bot blocking and Sorting Hat; round-robin assignment; extra-large merchants on their own pods; move downtime; the warehouse fed from replicas.
- Shopify open source: Ghostferry (checked September 2026).
Everything else in this round (the pod count, the pass payload and its checks, the hold table, the gate, the buckets, the bot layers, the translation and every number in R2.6 not marked published) is our design or our assumption.
R2.6 Numbers and Cost
All figures are assumptions unless marked published or derived.
Pods
| Quantity | Arithmetic | Result |
|---|---|---|
| Merchants | Published (March 2018) | 600,000+ |
| Pods | Assumption | 100 |
| Shops per pod | 600,000 ÷ 100 | 6,000 |
| Shops hurt by one pod's outage | 1 pod of 100 | 1%, down from 100% |
The sale and the admission rate
| Quantity | Arithmetic | Result |
|---|---|---|
| Buyers in the first minute | Assumption | 100,000 |
| Units | Assumption; 1 per buyer | 5,000 |
| Pod primary, sustainable writes | Assumption, from load tests | 8,000/s |
| Other shops on the pod at their peak | Assumption | 1,500/s |
| Budget for the sale, keeping 25% headroom | 8,000 × 0.75 − 1,500 | 4,500 writes/s |
| Writes per purchase | 6 checkout steps + 1 hold (with its bucket decrement) + 10 order rows + 1 hold convert | 18 |
| Admission rate | 4,500 ÷ 18 | 250 checkouts/s |
| Per 5-second period (published period) | 250 × 5 | 1,250 |
| Per load balancer, 10 of them | 1,250 ÷ 10 | 125 per period |
| Time to sell out | 5,000 ÷ 250 | 20 s |
| Queue polling at the edge | 100,000 waiting × one poll every 5 s | 20,000 requests/s, none reaching Rails |
The stock rows
| Quantity | Arithmetic | Result |
|---|---|---|
| Stock-row updates per unit sold | The hold's decrement; a convert touches only the hold row; a release adds one only for abandoned holds | 1 |
| Updates on one stock row | 250 × 1 | 250/s |
| Busy time at 1 ms per lock hold | 250 × 0.001 | 25% |
| Busy time at 8 ms per lock hold | 250 × 0.008 | 200%: impossible, a queue forms |
| With 10 buckets | 250 ÷ 10 = 25/s × 0.001 | 2.5% each |
| Near the sellout, 2 buckets left | 250 ÷ 2 = 125/s × 0.001 | 12.5% each |
Holds
| Quantity | Arithmetic | Result |
|---|---|---|
| Units held after 20 s | All 5,000 | Gate reads 0; "sold out" |
| Buyers who pay | 70% (assumption), average 90 s after the hold | 3,500 units sold by about t = 110 s |
| Buyers who abandon | 30% (assumption) | 1,500 units come back at their holds' expiry, t ≈ 600–620 s |
| Second wave | 1,500 ÷ 250/s | 6 s to hand them to buyers still in line |
| With a 5-minute hold instead | Same 30% | The 1,500 units come back at t ≈ 300–320 s, but more real buyers time out mid-payment |
Keeping the line open after "sold out" is what makes the second wave fair: returned units go to the next buyers in line, not to whoever refreshes fastest.
The gate after a Redis failover
| Quantity | Arithmetic | Result |
|---|---|---|
| Replication delay at the moment of failure | Assumption | 20 ms |
| Gate decrements lost | 250/s × 0.02 s | ≈ 5: the gate shows 5 units that are already held |
| Effect | The database refuses those 5 holds and the gate is corrected down | No oversell |
A shop move
| Quantity | Arithmetic | Result |
|---|---|---|
| Shop data | Assumption | 50 GB |
| Copy rate | Assumption, throttled to protect both pods | 50 MB/s |
| Batch copy | 50,000 MB ÷ 50 MB/s | 1,000 s ≈ 17 min, shop online throughout |
| Blocked writes at cutover | Chain in step 2.2 | 2.5 s (published average, 2018: 2.5 s) |
Cost: spreading app workers across three AZs (AWS translation, us-east-1; a 30-day month of 720 hours on every cost on this page). Each pod's writer lives in one Availability Zone, so about two-thirds of app-to-database traffic crosses zones, and AWS charges $0.01/GB in each direction: $0.02 for every GB that crosses.
| Quantity | Arithmetic | Result |
|---|---|---|
| Average queries per pod | Assumption | 7,000/s |
| Bytes per query, both directions | Assumption | 1.5 KB |
| Traffic per pod | 7,000 × 1.5 KB | 10.5 MB/s |
| Crossing zones | × 2/3 | 7 MB/s = 18,144 GB a month |
| Cost per pod | 18,144 × $0.02 | $363 a month |
| 100 pods | × 100 | $36,288 a month |
It's the price of surviving a zone loss with workers in every zone. Keeping workers in the writer's zone would save it, but a zone failure would then take the workers and the writer at once.
R2.7 Trade-Offs
Pod size
| Small pods (few shops each) | Large pods (many shops each) | |
|---|---|---|
| Blast radius | Small | Large |
| Datastores to run | Many | Few |
| Headroom for a surprise sale | Less per pod: one big sale fills it | More spare capacity to absorb a spike |
| Moves needed | More (shops outgrow pods sooner) | Fewer |
Shopify's answer, in public: limits found by scale tests ("let's just keep throwing more checkouts at it until we find the limit"), extra-large merchants on their own pods, and continuous rebalancing (2022 talk).
Queue fairness
| Random polling (Shopify v1) | Arrival order with a threshold (Shopify v2) | Central queue with positions | |
|---|---|---|---|
| Fair | No: a lottery | Yes, close to arrival order | Exactly |
| State | None | In the buyer's signed cookie | A shared store |
| New failure point | None | None | The store, at the worst moment |
| Across data centers | Works | Works: each load balancer computes its own threshold | Needs cross-site replication |
Gate vs database checks
| Database only | Redis gate + database (chosen) | Redis only | |
|---|---|---|---|
| Oversell | Never | Never: the database decides | Possible after a failover |
| Cost of a "no" after sellout | A row lock | About 1 ms in memory | About 1 ms in memory |
| Moving parts | One | Two, kept in sync | One |
A watch-list note for translations. If you move holds to DynamoDB, don't rely on TTL to release them: TTL deletes items only eventually (typically within a few days), so an expired hold must be ignored by a condition on expires_at. DynamoDB conditions have no server time, so the writer supplies now, and with global tables a condition checks only the local Region's copy.
R2.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Redis gate fails over mid-sale | A few seconds of gate errors, then a counter a few units high | The app's circuit for the gate opens and checkout falls back to database holds at a lower admission rate; the database refuses holds the gate over-promised and the gate is corrected down; re-seed from the primary, never a replica. No oversell. |
| Forged or edited queue pass | Passes that fail HMAC checks | Rejected at the edge; alarm on the failure rate. Keys rotate with a kid; a leaked key is removed from every load balancer's key set, which invalidates its passes. |
| A real pass shared with a bot farm | Many sessions presenting one pass | The session binding fails for all but one; the database's one-hold-per-pass key stops even a stolen session from buying twice. |
| Connection exhaustion in a pod | Pod 17's MySQL at its connection limit; its other shops fail; shared workers pile up waiting | A bulkhead per pod datastore on every worker host: Shopify's Semian caps how many workers may use one resource at once. Shopify's 2015 example: if workers talk to mysql_shard_0 10% of the time, "The probability that five workers are talking to it at the same time is 0.001%", so capping at five fails the sixth instantly. With 625 hosts of 32 workers and 6 tickets per host for pod 17, at most 3,750 workers can wait on pod 17, under its 4,000-connection limit (our numbers); the other 16,250 keep serving the other pods. |
| Shop move can't get the writer lock | Long-running units of work on the shop | "If the shop mover cannot acquire the writer lock in time, it fails the move" (published). Retry later; nothing changed. |
| A paused request wakes after a move | A write aimed at the old pod | Its first statement reads the old pod's never-pruned shop_copies row under a shared lock, finds MOVED, and aborts (our design); verification alerts if any query reaches the old copy (published). |
| Routing table database down | No new moves or signups | Load balancers keep serving from their local copy; shops keep working. |
R2.9 Production Gotchas
| Gotcha | Why it hurts | What we do |
|---|---|---|
| Direct database stock checks for the crowd | Every "is it in stock?" takes the hot row's lock | Gate in memory; database only for real claims |
| A global shared database | One tenant's sale is everyone's outage | Pods: isolated datastores, one pod per unit of work |
| Code that loops over every shard | Any shard down breaks the feature everywhere | Cross-shop work goes to the warehouse or a separate app |
| Unbounded job backlogs | A sale's emails and webhooks delay payment jobs and hold releases | Tiered queues with dedicated workers; shed low first; alarm on the oldest job's age |
A plain 429 for buyers | Looks like a crash; buyers retry harder | A cached queue page that polls |
| Per-node limits set to the global limit | N nodes admit N times the limit | Divide by the node count |
| Re-seeding the gate from a replica | Shows units already held | Re-seed from the primary |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | Pods as fault-isolation cells, one pod per unit of work, bulkheads per pod datastore REL 10; an edge throttle and queue that shed load before the database REL 5; online shop moves with a writer lock, verification and a permanent fence, so load can follow shops as they grow REL 7 |
| Performance Efficiency | A memory gate for "no" answers, bucketed stock rows, the queue page served from the edge cache PERF 3 · PERF 4 |
| Security | Bot filtering and challenges at the edge SEC 5; signed, expiring, session-bound queue passes with rotating keys, and one hold per pass |
| Cost Optimization | A queue that shapes demand to the capacity we have instead of provisioning for the crowd COST 9; the cross-AZ price of workers in every zone, known and chosen COST 8 |
| Operational Excellence | Moves verified by tooling and a formal specification; signals for queue length, admission rate, gate corrections and hold expiries OPS 8 |
| Sustainability | Skipped this round: the gains here are isolation and correctness; efficiency comes in Round 3. |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Explains why sharding without isolation lowers availability, and designs pods with one pod per unit of work.
- Chooses a routing table over a hash so tenants can be pinned and moved.
- Walks through an online move: copy, binlog tail, cutover under a lock, routing update, fence, prune, with a timing chain.
- Designs a stateless, fair queue at the edge and sizes the admission rate from the database's write budget.
- Separates a fast gate from the database's final decision, and fixes a hot row with buckets.
- Knows asynchronous replication loses acknowledged writes, and where that's harmless.
Follow-up questions
-
"Why did Shopify put the throttle in the load balancers instead of Rails?" Answer: turning a buyer away in Rails already costs a worker and often a database connection, which is the capacity we're protecting. The edge is event-driven Nginx that handles huge connection counts cheaply, and Lua let a small team change it without touching the monolith (Shopify, 2017).
-
"Your threshold controller overshoots. What happens?" Answer: the bucket still caps what reaches the pod, so overshoot means some admitted buyers are turned back to the queue page for a moment, not an overload. The controller only decides who is next; the bucket decides how many.
-
"Why not hold stock at add-to-cart?" Answer: carts are client-side and most never convert; holding at cart would lock stock for window shoppers and bots. Holding at checkout, with a short expiry, reserves stock only for buyers who've shown intent and passed the queue.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "A bigger database" | There's always a largest server; Shopify reached it in 2015. |
| "Sharding is isolation" | Code that loops over shards makes every shard a dependency. |
"429 for the crowd" | Buyers read it as a crash and retry. |
| "Queue positions in Redis" | A new point of failure, under the heaviest load, across data centers. |
| "Stock in Redis only" | A failover can forget decrements and sell units twice. |
| "The shop lock fences the old pod" | A paused writer outlives its lock; the old copy must refuse writes. |
Round 3 · Architect · "Era 3: Black Friday at Planet Scale, Resilient Payments"
~45 min · Principal (L7) · about 2018–today · BFCM 2025: $14.6B of sales, 489M requests/min at the edge and 117M at the app servers (published) · checkout keeps working while payment providers, pods or a region fail
R3.0 Where We Left Off
Round 2 in 60 seconds. "Sharding the database added capacity but not isolation, so Shopify built pods: each pod is a set of shops on a fully isolated set of datastores (MySQL, Redis, Memcached), while app servers, job workers and load balancers are shared. Every request or job touches exactly one pod. Sorting Hat, a Lua module in the OpenResty load balancers, looks up the shop in a routing table and adds the pod as a header. Because routing is a table, shops are pinned and moved: Ghostferry copies a shop's rows, tails the binlog, and cuts over under a per-shop writer lock in a few seconds. During flash sales, a leaky bucket at the edge admits checkouts at the rate the pod can take and parks everyone else on a cached queue page; a first-attempt timestamp in a signed cookie and a proportional controller make the line fair without any shared store. Stock sits behind a Redis gate that says 'no' cheaply, while the database's conditional updates and expiring holds decide every real unit. Open costs: checkout still calls outside payment providers that can be slow or down, nobody has proven which failures we survive, capacity for the biggest weekend is a guess, and a whole site can fail."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: isolation by pod, admission at the edge, stock decided by the database. The payment providers on the right are still a shared, fragile dependency.
Round 2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | One database unit for everyone | Pods, Sorting Hat, a routing table | No cross-shop queries |
| 2.2 | A sale slows its pod-mates | Online shop moves (Ghostferry) | Seconds of blocked writes |
| 2.3 | 100,000 buyers at checkout | Edge leaky bucket, fair queue | Waiting buyers |
| 2.4 | The hot stock row | Redis gate, database holds, buckets | A counter to keep honest |
| 2.5 | Bots | Edge scoring, bound passes, per-buyer limits | False positives |
Open costs: slow providers, unknown failure modes, capacity by guesswork, and a single site that can still fail.
R3.1 The Scope Raise
Interviewer: "It's today. We run on Google Cloud. Black Friday Cyber Monday breaks our records every year, and our merchants plan their year around it. Checkout talks to our own payments stack in many countries and to many third-party providers. When one of them slows down, checkout must not slow down with it. We want proof, before the weekend, of which failures we survive. And a region can fail on Black Friday."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How big is the weekend? | BFCM 2025: "$14.6 billion in sales", peaking at "$5.1 million per minute" at 12:01 p.m. EST on Black Friday; "Performance peaked at 489 million requests per minute on edge and more than 117 million requests per minute on app servers" (Shopify, December 2025). | ≈ 1.95M app-server requests a second at the peak (R3.6). |
| How does a card payment flow? | Card fields are iframes that post to CardSink, which encrypts the card and returns a token; a background job in the monolith passes the token to CardServer, which decrypts it and calls the right processor "using an adapter library" (2022 talk). | A provider abstraction already exists; the job workers that call providers are what a slow provider ties up (step 3.1). |
| How many providers? | Shopify Payments ran in 17 countries at the time of the 2022 talk, "We also support many payment integrations from third parties". | Failures come per provider and per country (step 3.1). |
| What happens when a provider slows? | Shopify developed Semian "to protect Net::HTTP, MySQL, Redis, and gRPC services with a circuit breaker in Ruby" (2022). | Fail fast, with fallbacks (step 3.1). |
| How do we know what we survive? | Toxiproxy for failure injection, resiliency matrices, game days, and load tests "in production" with a tool called Genghis (2015, 2020, 2022, 2025). | Test the failure modes, not just the load (steps 3.2 and 3.3). |
| How much capacity? | "We model traffic patterns from historical data and merchant growth, submitting our estimates to our cloud providers so they don't run out of cloud" (2025). 2025's program ran "Bimonthly fire drills all year, simulating 150% of last year's BFCM load". | Plan, reserve and prove capacity (step 3.3). |
| And regions? | A pod "is active in a single region at a time but exists in two with replication set up from the active to the non-active region" (2022 talk). In 2025, "We executed regional failovers, evacuating traffic from core US and EU regions". | Plan for losing a region (step 3.5). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Peak load | One flash sale on one pod | BFCM: 117M app-server requests/min (published, 2025) |
| Payment providers | Called and waited for | Circuit breakers and bulkheads per provider and country |
| Proof | Load test the sale path | Failure injection, game days, full-platform scale tests |
| Capacity | Pod write budgets | Forecast, reserve and test for the weekend, plus failover |
| Payment outcomes | Idempotent attempts | Unknown-state resolution, webhooks and reconciliation |
| Sites | Pods in pairs of data centers | Pods across cloud regions; a region can be evacuated |
R3.2 What Breaks in the Round 2 Design
| Round 2 piece | What breaks at the new scale |
|---|---|
| Payment calls with only timeouts | A provider that answers in 5 seconds instead of 1.5 holds workers more than three times as long; the shared worker pool fills and every payment waits (step 3.1). |
| Failure behavior nobody has tested | Each dependency's failure has a guessed effect; guesses drift as the code changes (step 3.2). |
| Capacity from last year plus a margin | The first time the whole platform runs at peak is the peak itself (step 3.3). |
| Idempotent attempts, but outcomes by one path | A lost reply, a late webhook and a buyer's second click can each move the same payment (step 3.4). |
| A pod's standby in the same place | A regional failure takes out both copies, or leaves a standby with no capacity to run (step 3.5). |
R3.3 New Requirements and API Additions
A payment attempt as a resource. Shopify's payment service (2019) takes the idempotency key as an input of every mutation, "a 'first class citizen' of the API", and stores each incoming request "uniquely identified by the client and idempotency key", with a progress column of named steps. This DTO is ours:
json{ "client": "checkout", "idempotency_key": "01JBQ7D2M9X4K6T8V3P1R5S0NA", "shop_id": 812, "checkout_id": 55120931, "amount": { "value": "96.00", "currency": "USD" }, "card_token": "cs_tok_7Hq2", "merchant_country": "US" }
Synthesizing vector architecture diagram...
Every arrow is a conditional update that names the state it leaves, so two paths that learn the same outcome (the job's reply and the provider's webhook) can't both apply it.
A circuit and bulkhead per provider and country. Shopify (2022) adds "the merchant's country code to the endpoint host and port to create a more fine-grained Semian identifier", so a local outage in one country doesn't block payments in others. This YAML is ours, modeled on Semian's parameters: the timeouts belong to the HTTP client, and the _s suffixes and the timeouts block are our notation, not Semian's names. The values are ours:
yamlresource: "gateway.example-psp.com:443:CA" # host:port:country timeouts: open_s: 1 # Shopify's suggested starting point (2022) read_s: 5 circuit_breaker: error_threshold: 3 # errors within error_threshold_timeout (defaults to error_timeout) that open the circuit error_timeout_s: 30 # how long it stays open; then the next real request is the half-open call half_open_resource_timeout_s: 3 # the half-open call's own timeout, above healthy p99 success_threshold: 2 # consecutive successes to close again bulkhead: quota: 0.33 # at most ceil(0.33 x workers on the host) in flight to this resource
A resiliency matrix per team. Shopify asks teams to write "a user-centric resiliency matrix, documenting the expected user experience under various scenarios" (2020). An example in that spirit (our rows):
| Dependency down or slow | Browse | Add to cart | Check out | Sign in |
|---|---|---|---|---|
| Session Redis | Yes | Yes | Yes, as a guest | No |
| Pod MySQL primary | Cached pages only | Yes (client-side) | No | No |
| One payment provider in one country | Yes | Yes | Yes, with other methods | Yes |
| Stock gate Redis | Yes | Yes | Yes, slower admission | Yes |
| Email provider | Yes | Yes | Yes, email later | Yes |
A capacity plan per component (our shape; Shopify partitions "each component into a separate GCP project, which makes it a lot easier to think of quotas per every project", 2020):
yamlcomponent: storefront-renderer forecast_peak_rpm: 150000000 safety_margin: 0.5 # plan for 1.5x the forecast failover: survive_one_region # capacity in the survivors, not just in total regions: [us-central, us-east, europe-west4] quota_checked: 2026-09-15 scale_test_dates: [2026-05-12, 2026-07-14, 2026-09-15]
Recap
- A payment attempt is a resource with explicit states, changed only by conditional updates.
- Circuits and bulkheads are keyed by provider, endpoint and country.
- Failure behavior is written down per team, then tested.
R3.4 Design Evolution: Fail Fast, Prove It, and Survive the Weekend
Step 3.1: A Slow Payment Provider Ties Up All the Workers
The problem: at the Black Friday peak, 850 payments a second run through the payment job workers. Half go to one provider, whose answers slow from 1.5 seconds to the 5-second read timeout. Within a minute every payment is slow, including those going to other providers, and checkout pages spin. What would you do?
Synthesizing vector architecture diagram...
A Semian circuit, per process and per resource. While Open, calls raise instantly; the bulkhead limits how many workers can be stuck while the circuit is still Closed.
Primitive: Circuit Breaker, Bulkhead and Fault-Tolerance Patterns · Drill: The Slow Recommendation Service That Took Down Checkout (answered here: why a slow dependency exhausts the shared workers, by Little's law, and why a tiny timeout on every call fails healthy calls where a breaker reacts to a sustained pattern) · Loop: Design a Payment Processing System (Round 2, step 2.4: a breaker per PSP)
Step 3.2: We Don't Know What Breaks Until It Breaks
The problem: every team says its service degrades gracefully. Nobody has seen checkout run with the session store down, with a provider at 5 seconds, or with a pod's primary gone. The first real test will be Black Friday. What would you do?
Synthesizing vector architecture diagram...
The matrix is the model; tests and game days check the model against reality, and every disagreement changes either the system or the matrix.
Step 3.3: Capacity for the Biggest Weekend
The problem: last year's BFCM peak was 117 million requests a minute at the app servers. Sales grew 27% last year, some merchants plan record sales, and a region may have to be evacuated mid-weekend. How much do we provision, and how do we know it works? What would you do?
Primitive: Cloud Disaster Recovery and Multi-Region Active-Active
Step 3.4: Payments Must Never Double-Charge, Whatever Fails
The problem: three things happen in the same minute. A charge's reply is lost, so our job marks it UNKNOWN. The provider's webhook for that same charge arrives saying "approved", while the resolver is also asking. And the buyer, staring at a spinner, presses Pay again. Separately, a payment succeeds but the order transaction fails, and the charge must be undone.
What would you do?
Synthesizing vector architecture diagram...
Three paths learn the same fact; the conditional update lets exactly one of them act on it. The buyer's second click is refused because the first attempt isn't final.
Loop primitive: Idempotency & Effectively-Once Processing (Part 3: claim, record, run, answer; Part 4: the late original) · Primitive: Two-Phase Commit and Saga Orchestration · Drill: The Flight Booking That Charged Without a Seat (answered here: a failed compensation is retried idempotently from an outbox, then becomes an anomaly for people; and a provider can't join a two-phase commit, so compensation replaces it) · Loop: Design a Payment Processing System (Round 1, steps 1.2 and 1.4: unknowns and webhooks)
Step 3.5: Where Does It Run, and What If a Region Fails?
The problem: Shopify has moved from its own data centers to Google Cloud. Pods must now live in cloud regions. On Black Friday, one region becomes unreachable. Its pods' shops must come back elsewhere, with their data, without two copies of a pod both taking writes. What would you do?
Synthesizing vector architecture diagram...
Detection comes from outside, region A's lease runs out before anything is promoted, the move is gated, and the epoch makes the new primary the only one that accepts the new routes.
Loop primitive: Replication, Quorums & Read-Your-Writes (Part 2: sync, async, semi-sync; Part 6: the leader crashes; Part 7: a partition and a split-brain attempt) · Primitive: Cloud Disaster Recovery and Multi-Region Active-Active
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | A slow provider ties up workers | Semian circuit breakers and bulkheads, keyed by country | Quick "try another method"; tuning |
| 3.2 | We don't know what breaks | Resiliency matrix, Toxiproxy tests, game days, a benchmark gateway | Engineering time |
| 3.3 | Capacity for the weekend | Forecasts with margins, scale tests on production, slower change | Idle capacity and test cost |
| 3.4 | Never double-charge | Attempts with conditional transitions, one live attempt, voids from an outbox, reconciliation | Delayed confirmations |
| 3.5 | Where it runs; a region fails | Google Cloud, pods in region pairs, gated epoch failover | Minutes of outage for one region's pods |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Each region runs the edge, the shared workers and the pods active there, and holds replicas of the other region's pods. Payment calls leave through circuits and bulkheads keyed by provider and country.
Synthesizing vector architecture diagram...
The readiness cycle, as Shopify describes it: plan, reserve, test the load and the failures, write the runbooks, freeze nothing but slow everything, and feed the weekend back into next year's forecast.
Trace: a payment when provider X's Canadian path is down (our design around the published pieces)
Synthesizing vector architecture diagram...
The open circuit is scoped to one country's path. Because nothing was sent to the provider, offering another method can't cause a double charge.
Our AWS translation (not Shopify's setup). Two or three Regions, each with CloudFront and AWS WAF at the edge, OpenResty on EKS behind a Network Load Balancer, and per-pod Aurora MySQL clusters; each pod as an Aurora Global Database whose secondary sits in the pod's paired Region (Aurora's managed failover to a secondary is itself an asynchronous-replication failover, so the RPO is its lag); ElastiCache in each Region; the routing table as a DynamoDB global table, written only from a single control-plane home Region because global tables resolve concurrent writes by last writer wins and replicate in a few seconds; Route 53 Application Recovery Controller routing controls as the gated switch; CloudWatch alarms fed by probes running in other Regions. Before a failover, check the survivor's On-Demand vCPU quota.
Sources for this round
- Eskildsen, Building and Testing Resilient Ruby on Rails Applications, Shopify Engineering, January 2015: the dependency matrix; Toxiproxy; circuit breakers; Semian's bulkheads and the
0.001%reasoning. - Neufeld, Shopify's Infrastructure Collaboration with Google, Shopify Engineering, March 2018: building its cloud with Google; GKE; over 50% of workloads migrated.
- Denis, A Pods Architecture To Allow Shopify To Scale, Shopify Engineering, March 2018: pod pairs of data centers; the Pod Mover; evacuating a data center pod by pod.
- Jefferson, Building Resilient GraphQL APIs Using Idempotency, Shopify Engineering, August 2019: the payment service; keys as first-class inputs; the lock and
409; incoming requests and recovery points. - Polan, Your Circuit Breaker is Misconfigured, Shopify Engineering, February 2020: Semian's parameters; per-process state; 263% to 4%.
- Tang and Shatrov, Capacity Planning at Scale, Shopify Engineering, December 2020: forecasts, safety margins and failover buffers; a GCP project per component; scale-up tests since 2018; GKE capacity for test windows.
- McIlmoyl, Resiliency Planning for High-Traffic Events, Shopify Engineering, December 2020: the user-centric resiliency matrix; game days; slowing the rate of change; over $5.1 billion in BFCM 2020 sales.
- Vaillancourt, How Shopify Reduced Storefront Response Times with a Rewrite, Shopify Engineering, August 2020: the storefront renderer, separate from the monolith, reading from dedicated read replicas.
- Inch, Pummelling the Platform–Performance Testing Shopify, Shopify Engineering, December 2020: load and stress tests; Go and Lua flows; HAR-based tests.
- Leinwand, Shopify's cloud, load and modular code in 2022, Shopify Engineering, January 2022: BFCM 2021 peaks of 32M app-server and 34M load-balancer requests a minute; a 50.7M load test.
- de Water, 10 Tips for Building Resilient Payment Systems, Shopify Engineering, July 2022: timeouts; Semian and country identifiers; Little's law; idempotency; reconciliation and anomalies; the benchmark gateway.
- de Water, Shopify's Architecture to Handle the World's Biggest Flash Sales, QCon Plus talk, published by InfoQ in October 2022: CardSink and CardServer; 17 countries; pods in two regions; Genghis; Resumption.
- Petroski and Frail, How we prepare Shopify for BFCM, Shopify Engineering, November 2025: 2024 records; capacity planning with cloud providers; game days; the Resiliency Matrix; five scale tests; regional failovers.
- Shopify BFCM announcements: 2021 (November 2021), 2022 (November 2022), 2023 (November 2023), 2024 (December 2024) and 2025 (December 2025); Shopify, Performance up, complexity down: killer updates from Shopify engineering, January 2024 (the 2023 core application server peak); @ShopifyEng on X, BFCM 2022 stats, November 2022 (75.98M requests a minute).
- Shopify open source: Semian and Toxiproxy (checked September 2026).
Everything else in this round (the attempt states, the resolver and void design, the worker and bulkhead numbers, the region sizing, the failover chain, the translation, and every number in R3.6 not marked published) is our design or our assumption.
R3.6 Numbers and Cost
BFCM, as published
| Year | Sales (GMV) | Peak sales | Peak requests | Source |
|---|---|---|---|---|
| 2020 | Over $5.1B | – | – | Shopify Engineering, December 2020 |
| 2021 | $6.3B | More than $3.1M/min | More than 32M/min at app servers; more than 34M/min at load balancers | Shopify news, November 2021; Shopify Engineering, January 2022 |
| 2022 | $7.5B | More than $3.5M/min | 75.98M/min "to our commerce platform" | Shopify news and @ShopifyEng on X, November 2022 |
| 2023 | $9.3B | $4.2M/min | 967K/s = 58M/min at the "core application server" | Shopify news, November 2023 and January 2024 |
| 2024 | $11.5B | $4.6M/min | 284M/min at the edge; more than 80M/min at app servers | Shopify news, December 2024 |
| 2025 | $14.6B | $5.1M/min | 489M/min at the edge; more than 117M/min at app servers; 31.8M API requests/min | Shopify press release, December 2025 |
The request figures don't all measure the same layer: the 2022 figure is "our commerce platform"; the 2023 figure is the "core application server", which we treat as the app-server layer, so a year-to-year line across them would mislead. The app-server series (2021, 2023, 2024, 2025) is consistent:
Synthesizing vector architecture diagram...
From 2021 to 2025 the peak app-server rate grew about 3.7 times.
What 2025's peak means per second
| Quantity | Arithmetic | Result |
|---|---|---|
| Edge | 489M ÷ 60 | 8.15M requests/s |
| App servers | 117M ÷ 60 | 1.95M requests/s |
| Share of edge requests that never reached the core app servers | (489 − 117) ÷ 489 | about 76% (edge cache, or other services such as the storefront renderer), if every app request came through the edge |
| Growth 2024 → 2025 | 117 ÷ 80 (app); 489 ÷ 284 (edge); sales +27% (published) | 1.46×; 1.72× |
| Average database writes over the weekend | 1.75T (published) ÷ 421,200 s (the 117-hour window Shopify used for 2023, assumed for 2025) | ≈ 4.2M/s |
| Average database queries | 14.8T (published) ÷ 421,200 s | ≈ 35M/s |
Payments at the peak (assumptions unless marked)
| Quantity | Arithmetic | Result |
|---|---|---|
| Orders a minute | $5.1M/min (published) ÷ $100 average order (assumption) | 51,000 |
| Payments a second | 51,000 ÷ 60 | 850 |
| 2025's scale test 4 | 80,000+ checkouts/min (published) ÷ 51,000 | about 1.6× our derived peak (if "checkouts" means completed orders) |
| "Once in a million" at this rate | 850 × 86,400 ÷ 1,000,000 | ≈ 73 times a day |
| Busy payment workers, normal | 850 × 1.5 s | 1,275 of 3,000: 43% |
| Provider X (half) at 5 s | 425 × 5 + 425 × 1.5 | 2,763: 92% |
| With a 1,000-ticket bulkhead on X | 1,000 + 638 | 1,638: 55%; X served at 200/s, 225/s told "try another method" |
| Breaker half-open waste, e = 30 s, h = 3 s | 3 ÷ 33 | ≈ 9% of a worker per failing resource |
| Time for one process to open its circuit | 3 timeouts × 5 s | 15 s |
Capacity for next year (assumptions)
| Quantity | Arithmetic | Result |
|---|---|---|
| CPU at the 2025 app-server peak | 1.95M/s × 20 ms | 39,000 vCPUs busy |
| At 60% target utilization | 39,000 ÷ 0.6 | 65,000 vCPUs |
| Two regions, each able to carry all | 65,000 × 2 | 130,000 (2×) |
| Three regions, pods paired evenly: each survivor takes half a failed region | each region 1/3 + 1/6 = 1/2 of 65,000 = 32,500; × 3 | 97,500 (1.5×) |
| Headroom for 10 days at $0.04 per vCPU-hour (assumed blended price) | 32,500 or 65,000 extra × 240 h × $0.04 | $312,000 (three regions) or $624,000 (two) |
| One 3-hour scale test on the full three-region fleet | 97,500 × 3 × $0.04 | $11,700 |
| Load generators for 146M requests/min | 2.43M/s ÷ 1,000 per vCPU = 2,433 vCPUs × 3 h × $0.04 | ≈ $292 |
| Merchant sales at risk in a 10-minute checkout outage at the peak | 10 × $5.1M | $51M of GMV (merchants' sales, not Shopify's revenue) |
The load test itself is cheap; the fleet it proves is not. Shopify's 2020 note on why the cloud helps: running on GKE "allowed us to grab extra compute capacity just for the window of the exercise, and only pay for those hours when we needed it".
A test at 146M requests a minute sits well above 2025's 117M app-server peak, but Shopify doesn't say which layer its test figures measure, so compare them with care. In 2021 the comparison is clean: a load-balancer test at 50.7M against a real load-balancer peak of 34M, about 1.5×.
Region failover (assumptions)
| Quantity | Arithmetic | Result |
|---|---|---|
| Time to serve a failed region's pods elsewhere | 35 + 60 + 5 + 60 + 5 + 60 (chain in step 3.5) | ≈ 225 s worst case |
| Payment records missing from promoted replicas | Two regions: 425/s × 2 s of lag; three regions: 283/s × 2 s | ≈ 850 or ≈ 567, rebuilt from provider records |
| Lease vs promotion | 30 s lease against 85–95 s of detection and approval | The old region stops serving at least 55 s before promotion |
R3.7 Trade-Offs
A modular monolith with pods vs microservices (closing the loop)
| Modular monolith on pods (Shopify) | Microservices per domain | |
|---|---|---|
| Scaling data | Pods: add a pod, move shops | Each service shards its own store |
| Isolation | By tenant (pod) and by dependency (Semian) | By service; a tenant's spike hits every service it touches |
| Checkout's many writes | One transaction in one pod | A saga across services |
| Card data | Kept out of the monolith by CardSink and CardServer | The same idea: isolate the sensitive service |
| Change | Hundreds of developers on one app, deployed about 40 times a day, boundaries enforced by Packwerk | Independent deploys, network contracts |
Shopify's stated view (2022 talk): the monolith won't be broken into microservices "anytime soon", because things "that pertain to a single shop, like an order or a chargeback or a refund, it's just easier to keep it in a monolith". It does split where the traffic profile differs: storefront rendering moved to a separate application, reading from "dedicated read replicas" (2020).
Failing fast vs retrying
| Retry (with backoff) | Fail fast (breaker open) | |
|---|---|---|
| Good for | Transient blips on a healthy dependency | A dependency that is down or saturated |
| Risk | Extra load on a struggling dependency | Refusing a request that might have succeeded |
| Payments | Only with the same key, only for the same attempt | Only when nothing was sent, so no double charge |
Headroom
| Two regions | Three regions | |
|---|---|---|
| Fleet to survive one loss | 2× | 1.5× |
| Datastores per pod | 2 copies | 2 copies (each pod paired once) |
| Pods moved in a failure | Half of all pods | A third of all pods |
| Operations | Simpler | More routing combinations to test |
R3.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Webhook timeout and a double-charge race | The job's reply lost, the webhook and the resolver both learn "approved", the buyer clicks again | One conditional transition wins; the order's unique key blocks a second order; a new attempt is refused while one is UNKNOWN (step 3.4). |
| A pod's primary fails | That pod's shops can't check out; others are unaffected | Promote a replica in the region; only this pod's shops notice. Payment records lost to asynchronous replication are rebuilt from provider records. |
| A provider-wide outage | A provider's circuit opens in every country | Show other payment methods; never re-send an UNKNOWN attempt to another provider; route to an alternate processor only attempts we know were never sent (as in the payment loop's step 2.4). Alert the provider's account team. |
| A breaker opens on a blip | A healthy provider refused for error_timeout | error_threshold set from measured error rates; alarm on circuit opens; a lower error_timeout for resources whose false opens cost more than their half-open calls. |
| A region is unreachable | Outside probes fail; that region's pods stop serving | Gated failover with epochs (step 3.5), about 4 minutes; survivors pre-sized; old primaries rejoin as replicas. |
| A load test leaks into real traffic | Benchmark orders or payments where they shouldn't be | Benchmark stores and the benchmark gateway only; test traffic identified and dropped by any real provider path. |
| Surprise bottleneck at scale | A limit no dashboard showed: VMs per network, packets per Memcached server, logging throughput (Shopify's 2020 examples) | Found in scale tests months ahead; fixed and re-tested. |
R3.9 Runbook and Incident Response
Shopify's 2022 payments post lists the four golden signals and one nuance for payments: "we distinguish between payment failures and errors". A decline isn't an incident; a provider's 500 is.
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Checkout completion rate, per pod and platform-wide | Drops 20% below the same minute last week for 5 min | P1 | One pod, one region, one provider, or everything? |
| Queue admission rate vs the pod's budget | Admission below 50% of budget while the queue grows | P2 | Is the pod slow, or is the bucket set too low? |
| Pod saturation: MySQL connections, lock waits, write latency | > 80% of connections, or p99 commit latency 3× baseline | P1 | A hot shop? Tighten that shop's throttle; plan a move. |
| Breaker state per provider and country | Any circuit open > 2 min | P2 | Provider incident or our misconfiguration? |
| Payment error rate (not declines) per provider | > 2% of attempts | P1 | Check the provider's status; confirm fallbacks are shown. |
UNKNOWN attempts older than 15 minutes | Any growth | P2 | Resolver health; provider lookups; reconciliation queue. |
| Payment latency p99 per provider | > 2× baseline for 5 min | P2 | Watch worker saturation before the breaker opens. |
| Outside probes per region | 3 consecutive failures | P1 | Start the region runbook. |
Procedure: pod overload during a sale REL 5
- Confirm it's one pod. Saturation on one pod's MySQL, normal elsewhere.
- Find the shop. Top shops by writes and lock waits on that pod.
- Throttle, don't block. Engage or tighten that shop's checkout throttle; the queue page absorbs the crowd.
- Protect neighbors. Check the pod's bulkhead tickets are holding; shed low-priority jobs for the pod.
- Afterwards, move. Schedule the shop onto a quieter or dedicated pod with the Shop Mover, outside the sale.
Procedure: tuning the waiting room COST 9
- Read the budget. Pod write budget from the last load test, minus other shops' current load.
- Set the rate. Admission = budget ÷ writes per purchase; divide by the number of load balancers.
- Watch the controller. Admitted rate should track the bucket limit; big oscillations mean the controller is overcorrecting.
- Watch fairness. Median and p95 queue times should stay close; a widening gap means buyers are skipping the line.
- Keep the line after sellout so returned stock goes to the next buyers.
Procedure: provider degradation REL 11
- Scope it. Which provider, which countries, errors or latency?
- Check our side. Worker saturation, bulkhead tickets, open circuits: are we protecting ourselves?
- Degrade honestly. Other payment methods shown; no re-routing of unknown attempts.
- After recovery, drain
UNKNOWNattempts through the resolver and check reconciliation anomalies.
Procedure: region evacuation REL 13 · OPS 10
- Confirm from outside that the region, not a service, is down.
- Check the survivors have capacity and quota for the extra pods, and hold those pods' data where it's allowed to live.
- Flip the routing control for the region's pods and confirm each pod's epoch moved.
- Watch checkout completion for the moved shops and the payment reconciliation queue.
- Fence and rejoin the old region's primaries as replicas; reconcile their late writes before moving pods back one at a time.
Shopify's incident process (2022) names three roles: the Incident Manager on Call, a Support Response Manager for public communication, and the service owners, with a retrospective "within a week after the incident occurred".
Go deeper: CLI and SQL checks (AWS translation; replace names with real ones)
text# 1. Alarms currently firing for checkout and payments aws cloudwatch describe-alarms --state-value ALARM --alarm-name-prefix checkout- # 2. The state of a pod's stock-gate cache aws elasticache describe-replication-groups --replication-group-id pod-017-gate # 3. A pod's database cluster and its members aws rds describe-db-clusters --db-cluster-identifier pod-017 # 4. Unplanned failover of a pod's global database to its paired Region (accepts the replication lag as data loss) aws rds failover-global-cluster --global-cluster-identifier pod-017-global --target-db-cluster-identifier arn:aws:rds:us-east-2:111122223333:cluster:pod-017-use2 --allow-data-loss # 5. Read, then flip, a Region's routing control (run against a cluster endpoint) aws route53-recovery-cluster get-routing-control-state --routing-control-arn arn:aws:route53-recovery-control::111122223333:controlpanel/abc/routingcontrol/def --region us-west-2 --endpoint-url https://host-aaa.us-west-2.example.com/v1 aws route53-recovery-cluster update-routing-control-state --routing-control-arn arn:aws:route53-recovery-control::111122223333:controlpanel/abc/routingcontrol/def --routing-control-state Off --region us-west-2 --endpoint-url https://host-aaa.us-west-2.example.com/v1 # 6. The vCPU quota for Running On-Demand Standard instances in the survivor Region aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A --region us-east-2 # 7. Reserve capacity for the weekend in one AZ aws ec2 create-capacity-reservation --instance-type c7g.4xlarge --instance-platform Linux/UNIX --availability-zone us-east-1a --instance-count 50 --end-date-type limited --end-date 2026-12-03T00:00:00Z # 8. Which pod serves a shop (strongly consistent read in the table's home Region) aws dynamodb get-item --table-name shop-routes --key '{"domain":{"S":"tshirthero.com"}}' --consistent-read # 9. On a pod's MySQL: who is waiting for which lock SELECT * FROM performance_schema.data_lock_waits; # 10. On a MySQL replica: how far behind is it? SHOW REPLICA STATUS\G
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | Circuit breakers and bulkheads per provider and country REL 5; resiliency matrices, Toxiproxy tests and game days REL 12; pods paired across regions with gated, epoch-fenced failover and a stated RPO REL 13; survivors sized and quota-checked for a region loss REL 1 |
| Performance Efficiency | Scale tests on production at the forecast peak and above, with benchmark stores on every pod PERF 5; storefront rendering split out and served from read replicas PERF 3 |
| Security | Card data classified and confined to CardSink and CardServer, so the monolith never sees it SEC 7; card data encrypted by CardSink and decrypted only by CardServer SEC 8; security incidents run through the same IMOC-led process and retrospectives (our design) SEC 10 |
| Cost Optimization | Three regions instead of two cut failover headroom from 2× to 1.5× COST 6; capacity reserved for the weekend and test windows only COST 7; building Semian, Toxiproxy and Ghostferry where no product fit, and open-sourcing them COST 11 |
| Operational Excellence | Golden signals with failures separated from errors OPS 8; IMOC-led incidents OPS 10; retrospectives within a week and a readiness cycle fed by each BFCM OPS 11 |
| Sustainability | Extra capacity only for test windows and the weekend, released afterwards SUS 2; about three of every four edge requests at the 2025 peak never reached the core app servers (edge cache, or other services such as the storefront renderer; derived, R3.6) SUS 3 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Computes, with Little's law, how one slow dependency saturates shared workers, and sizes bulkheads and breaker parameters from the numbers.
- Treats failure behavior as something to write down and test: matrices, failure injection, game days, a fake provider for load tests.
- Plans capacity for the failover case, not just for growth, and compares region layouts by headroom.
- Designs payments so every path that learns an outcome uses the same conditional transition, with one live attempt per checkout, voids driven from an outbox, and reconciliation against the provider.
- Designs a region failover end to end: outside detection, a gate, epochs, capacity and residency, a timing chain and a stated RPO with a repair source.
- Separates what Shopify published from what it designs, and reads published numbers for which layer they measure.
Follow-up questions
-
"Why not make the breaker's
error_timeout1 second so recovery is instant?" Answer: every open circuit then lets a real request through as the half-open call every second, and each can wait its full half-open timeout while the provider is still down: witherror_timeout1 s and h = 3 s, 3 ÷ 4 = 75% of a worker's time goes to half-open calls, each of them a real buyer's payment that fails. Shopify's own tuning went the other way, to 30 seconds, trading slower recovery for a 4% cost instead of 263%. -
"The resolver and the webhook disagree: one says approved, one says declined." Answer: the first to apply its conditional transition wins, and the attempt is final; the disagreement becomes a reconciliation anomaly. Card outcomes don't flip in reality, so a disagreement means a bug or a lookup of the wrong attempt, which is exactly what anomalies are for.
-
"Could every pod be active in two regions to avoid the 4-minute failover?" Answer: not for stock and orders: two writable copies can each sell the last unit, and last-writer-wins would silently drop an order. Active-active fits data that's read-mostly or mergeable (storefront rendering from replicas, the routing table written from one home region). For checkout, one writer per pod with a fast, gated failover is the honest design, with a short lease so a cut-off region stops writing before the other one is promoted.
-
"Is Semian still how Shopify does this?" Answer: Semian has been in production at Shopify since October 2014 (README), and Shopify's 2022 payments post describes it protecting
Net::HTTP, MySQL, Redis and gRPC. The repository is active, with adapters for MySQL, Redis,Net::HTTPand Active Record.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Longer timeouts for slow providers" | Each call holds a worker longer; the pool fills faster. |
| "50 ms timeouts everywhere" | Fails healthy calls and manufactures unknown payments. |
| "Unit tests prove resilience" | They test the author's assumptions, not the failure modes. |
| "Autoscale on Black Friday" | Databases don't autoscale, and the cloud may not have the machines. |
| "Retry the payment until it works" | Without the same key, each retry is a new charge. |
| "Health checks in the failed region trigger failover" | They fail with it. |
| "Active-active checkout" | Two writers sell the last unit twice. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scope: shops, storefront, checkout, payments; what's breaking | Restate Round 1 in 60 seconds | Restate Round 2 in 60 seconds |
| 5–15 min | Requirements; checkout API with idempotency keys; checkout states | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.4: cache → conditional stock → idempotent payments → job tiers and outbox | Steps 2.1–2.5: pods → shop moves → fair queue → stock gate and holds → bots | Steps 3.1–3.5: breakers and bulkheads → failure testing → capacity → unknown payments → regions |
| 40–50 min | Reads with and without cache; one sale on a shared database; the hot row | Write budget → admission rate; hot-row math; holds; move timing; cross-AZ cost | Published BFCM numbers; Little's law for workers; region sizing; failover chain and RPO |
| 50–60 min | Failures and pillar check | Failures, gotchas, pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint. For payments in depth, see the payment processing loop; for holds, flash sales and gate counters, the hotel reservation loop; for another company that outgrew its first database, the Uber case study.
The Two Sentences That Matter Most
- Opening any round: "This is a multi-tenant platform where one merchant's flash sale looks like an attack on everyone else, so I'll isolate tenants, admit the crowd at the rate the database can take, and let the database, never a cache, decide every unit of stock and every charge."
- When scale arrives: "I'll put every shop in exactly one pod, every outside call behind a breaker and a bulkhead, every payment outcome behind one conditional transition, and then prove each of those with failure injection and full-scale tests before the weekend that matters."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "How do you stop one merchant from taking down the others?" (REL 10) | Pods: each shop's data lives on one isolated set of datastores, every request touches one pod, and hot shops are moved or given their own pod. | 2 | Steps 2.1, 2.2 |
| "A million people hit checkout at once. What happens?" (REL 5) | The edge admits them at the pod's measured rate and parks the rest on a cached queue page, in arrival order. | 2 | Step 2.3 | |
| "A payment provider slows down. Does checkout?" (REL 5) | No: a bulkhead caps the workers it can hold, and a circuit keyed by provider and country fails fast with a fallback. | 3 | Step 3.1 | |
| "How do you know you'll survive Black Friday?" (REL 12) | A resiliency matrix per team, Toxiproxy tests in CI, game days, and scale tests on production at above the forecast, including failovers. | 3 | Steps 3.2, 3.3 | |
| "What if a region fails?" (REL 13) | Each pod has a replica in a paired region; outside probes, a gate, a lease that stops the cut-off region, and an epoch bump move it in about four minutes, and the seconds of lag are rebuilt from providers' records. | 3 | Step 3.5 | |
| "How do you avoid charging twice?" (REL 4) | An idempotency key claimed before the call and passed to the provider, one live attempt per checkout, and one conditional transition for every path that learns the outcome. | 1, 3 | Steps 1.3, 3.4 | |
| Performance | "How do you serve a sold-out product to a hundred thousand people?" (PERF 3) | From memory and the edge: a Redis gate says "no" in a millisecond, and the queue page is cached; only real claims reach the database. | 2 | Step 2.4 |
| "How do you know your capacity?" (PERF 5) | By measuring it: weekly load tests on benchmark stores on every pod, with a fake gateway that behaves like real providers. | 3 | Steps 3.2, 3.3 | |
| Security | "How do you keep card data safe in a big monolith?" (SEC 7) | The monolith never sees it: card fields post to a separate service that returns a token, and only another separate service can decrypt it. | 1, 3 | R1.4, R3.1 |
| "How do you stop bots?" (SEC 5) | Score and challenge at the edge, bind queue passes to sessions, and limit units per buyer in the database. | 2 | Step 2.5 | |
| Cost | "Why a queue instead of more servers?" (COST 9) | The database can't autoscale; shaping demand to capacity costs almost nothing, while provisioning for the crowd would cost a fleet. | 2 | Step 2.3 |
| "What does surviving a region loss cost?" (COST 6) | Spare capacity: 2× the fleet with two regions, 1.5× with three, about $312,000 for a 10-day window in our three-region example. | 3 | R3.6 | |
| Operations | "What do you watch on Black Friday?" (OPS 8) | Checkout completion, admission rate, pod saturation, breaker states, payment errors (not declines), unknown attempts, and probes from outside each region. | 3 | R3.9 |
| Sustainability | "How do you avoid paying for peak all year?" (SUS 2) | Scale up only for test windows and the weekend, and remember that about three of every four edge requests at the 2025 peak never reached the core app servers (edge cache, or other services such as the storefront renderer). | 3 | R3.6, R3.10 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Tenancy | Scopes every query by shop | Pods, a routing table and online shop moves | Pods across regions, sized for failover |
| Traffic spikes | Caches reads | A fair, stateless queue sized from the write budget | Scale tests above the forecast, with failovers |
| Stock | One conditional decrement | A memory gate, database holds with expiry, bucketed rows | Holds that survive payment delays and database failovers |
| Payments | Idempotency keys claimed before the call | Holds taken before payment | Unknown-state resolution, one live attempt, voids from an outbox, reconciliation |
| Dependencies | Low timeouts, degrade when a store is down | Bulkheads per pod datastore | Breakers and bulkheads per provider and country, tuned with math |
| Honesty | Labels assumptions | Separates Shopify's published mechanisms from our additions | Reads published numbers for which layer they measure |
Sources
All sources used on this page, oldest first.
- Shopify, 100,000 Online Stores Now Use Shopify, Shopify blog, 2014.
- Eskildsen, Building and Testing Resilient Ruby on Rails Applications, Shopify Engineering, January 2015.
- Stolarsky, Surviving Flashes of High-Write Traffic Using Scriptable Load Balancers, Part I and Part II, Shopify Engineering, February 2017.
- Denis, A Pods Architecture To Allow Shopify To Scale, Shopify Engineering, March 2018.
- Neufeld, Shopify's Infrastructure Collaboration with Google, Shopify Engineering, March 2018.
- Westeinde, Deconstructing the Monolith, Shopify Engineering, February 2019.
- Jefferson, Building Resilient GraphQL APIs Using Idempotency, Shopify Engineering, August 2019.
- Polan, Your Circuit Breaker is Misconfigured, Shopify Engineering, February 2020.
- Müller, Under Deconstruction: The State of Shopify's Monolith, Shopify Engineering, September 2020.
- Vaillancourt, How Shopify Reduced Storefront Response Times with a Rewrite, August 2020.
- Inch, Pummelling the Platform–Performance Testing Shopify, Shopify Engineering, December 2020.
- Tang and Shatrov, Capacity Planning at Scale, Shopify Engineering, December 2020.
- McIlmoyl, Resiliency Planning for High-Traffic Events, Shopify Engineering, December 2020.
- Madan, Shard Balancing: Moving Shops Confidently with Zero-Downtime at Terabyte-scale, Shopify Engineering, September 2021.
- Shopify, BFCM 2021 results, November 2021.
- Leinwand, Shopify's cloud, load and modular code in 2022, Shopify Engineering, January 2022.
- de Water, 10 Tips for Building Resilient Payment Systems, Shopify Engineering, July 2022.
- de Water, Shopify's Architecture to Handle the World's Biggest Flash Sales, QCon Plus talk, published by InfoQ in October 2022.
- Shopify, BFCM 2022 results, November 2022, and @ShopifyEng on X, BFCM 2022 stats, November 2022.
- Shopify, BFCM 2023 results, November 2023.
- Shopify, Performance up, complexity down: killer updates from Shopify engineering, January 2024.
- Shopify, BFCM 2024 results, December 2024.
- Petroski and Frail, How we prepare Shopify for BFCM, Shopify Engineering, November 2025.
- Shopify, BFCM 2025 press release, December 2025.
- Shopify open source: Semian, Toxiproxy and Ghostferry (checked September 2026).
- Shopify API reference, ProductVariantInventoryPolicy (
DENYandCONTINUE). - MySQL 8.4 Reference Manual, CHECK Constraints and Semisynchronous Replication.
- AWS, EC2 data transfer pricing (
us-east-1), Aurora Global Database documentation and the AWS CLI reference (checked September 2026), for the translation's figures and commands.