Design a Mobile Stock Trading App
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
In this loop the phone is part of the system. It has a flaky network, a screen that redraws 120 times a second, a secure chip that can sign things, and a user whose thumb is on the Buy button. Every round designs both halves: the app on the device and the backend it talks to. We are the broker's app, not the exchange. How an exchange matches orders is the subject of the matching engine loop, and how money moves through a ledger is the subject of the digital wallet and payment processing loops. We link to them instead of teaching them again.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | A small brokerage app: watchlist, quote screen, buy and sell | A popular retail trading app on meme-stock days | A regulated broker-dealer at national scale |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Traffic | 50K users online at the open, 10 symbols each; 100K orders/day, ~100/s at peak | 3M online at the open, 15 symbols each (45M subscriptions); 2M orders/day, 5K/s at the open | 6M online at the open (90M subscriptions); 8M orders/day with options, 20K/s at the open |
| On the device | A live watchlist; order status that is never ambiguous | + smooth 120 Hz scrolling and charts; biometric-signed orders; no trade on a stale price | + a release that can't break trading; device and session risk signals |
| Targets | Quotes about a second fresh; no duplicate orders; 99.9% in market hours | Order in, to broker, P99 < 150 ms; quotes marked stale after 2.5 s; 99.99% in market hours | Complete, tamper-proof records; a region lost at 10 a.m. loses no accepted order; order entry back within 15 min |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is a Trading App?
You Already Know One: a Live Price Board With a Buy Button
Picture the departures board at a train station, except the numbers are prices and they change many times a second. Under the board there's a button. You watch a price, decide it's right, and press the button. Someone at the other end must do exactly what you pressed, once, and tell you honestly what happened.
A mobile trading app is that board and that button in your pocket. The board is fed by the stock market; the button sends an order to a broker, who sends it on to a market to be executed.
A few words carry this whole loop. Each gets one line now:
| Term | One-line definition |
|---|---|
| Quote | The current prices for one stock: the best bid (highest price someone will pay) and the best ask (lowest price someone will sell for), with sizes, plus the last trade price |
| Spread | Ask minus bid: what you lose if you buy and immediately sell |
| NBBO | National best bid and offer: the best bid and ask across all US exchanges together |
| Tick | In this loop, one change to a quote. (Exchanges also use "tick" for the smallest price step; we won't.) |
| Conflation | Keeping only the newest quote per symbol over a short window and dropping the ones in between |
| Market order | "Buy 10 now at whatever the price is": it fills fast, at an uncertain price |
| Limit order | "Buy 10 at 227.50 or better": a price guarantee, but it may not fill |
| Ack | The broker or market saying "I have your order". It is not a trade. |
| Fill | A trade: some or all of the order executed, at a known price |
| Buying power | How much the account can spend on new orders right now: cash, minus money already promised to open orders, plus margin if the account has it |
| Halt | Trading in a stock stopped by the market, for news or a volatility pause |
Synthesizing vector architecture diagram...
Two flows cross in our backend. Market data flows down and is huge; orders flow up and are few, but each one is money.
What Makes It Hard
- Prices move faster than a phone can show them. A busy stock can change a hundred times a second. A person can't read that, a cellular radio shouldn't carry it, and a phone shouldn't redraw on every change.
- The network drops just as you tap Buy. The app often can't know whether the order arrived.
- A duplicated order is real money. Two orders instead of one means twice the shares, twice the cost, and an angry customer.
- A stale price is a trap. A user who taps Buy while looking at a price from 30 seconds ago is trading blind.
- Everyone arrives at 9:30 a.m. Eastern. The market opens, millions of people open the app in the same minute, and the busiest seconds of the day come first.
The Question the Whole Loop Answers
How do we show prices that are fresh enough to trust, and make every order happen exactly once, only if the user meant it?
The answer gets sharper every round:
- Round 1: stream quotes instead of polling, give every order a key born on the phone so retries can't duplicate it, and let the server, never the phone, decide buying power and order state.
- Round 2: shrink the market-data firehose to what a phone can use (conflation), keep the screen smooth, refuse to trade on stale prices, and sign every order with a key only the user's biometric can unlock.
- Round 3: run it as a regulated financial service: records that can be proven years later, routing that can be defended, a second region that loses no accepted order, and releases that can't break trading.
Round 1 · Mid-level · "Quotes and Orders for a Small Brokerage"
~35 min · SDE II (L5) · 1 region, 3 AZs · 50K users online at the open · 500K quote deliveries/s at peak · 100K orders/day, ~100/s at peak · 99.9% in market hours
R1.1 Establish Design Scope
The interviewer says: "We're a small brokerage. Our customers want a phone app with a watchlist, a quote screen, and a way to buy and sell. Design it." We ask before we draw.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Which instruments? | US stocks and ETFs only. | One kind of quote and one kind of order. Options come in Round 3. |
| Which order types? | Market and limit, good for the day. | Two order types; a day order expires at the close (step 1.3). |
| How fresh must quotes be? | About a second. | We stream, but we don't need every change: at most one update per symbol per second (step 1.1). |
| Real-time or delayed prices? | Real-time on the quote and order screens for signed-in customers. | Real-time data is licensed per user by the exchanges, through our vendor; delayed data (commonly 15 minutes old) is cheaper and fine for public pages (R1.8). |
| Do we match orders ourselves? | No. We send them to a partner broker, which executes and clears them. | We build order entry, risk checks and status; the partner does execution. The exchange side is the matching engine loop. |
| Charts? | Simple: today's price line. | A small history API; no heavy charting yet. |
| Authentication? | A normal login. | Signed tokens on every call. Stronger order signing comes in Round 2. |
| How many users? | About 500K accounts; 50K online at the open, each watching about 10 symbols. | We derive the traffic in R1.7. |
Out of scope for this round: options, margin accounts, smooth 120 Hz charts, biometric order signing, and the scale of a meme-stock morning.
The interviewer will widen this scope later. Write your out-of-scope list where you can see it: in a multi-round loop, some of it comes back.
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Requirement |
|---|---|
| "A watchlist" | The app shows live quotes (bid, ask, last) for the user's symbols |
| "A quote screen" | Quote detail with today's price line |
| "Buy and sell" | Place a market or limit order; cancel an open order |
| "Did it work?" | Every order has one clear status, updated live: received, working, partly filled, filled, cancelled, rejected, expired |
| "What do I own?" | Holdings and cash, from the server's books |
Not yet: 120 Hz rendering, biometric signing, the opening-bell scale. The interviewer will bring them back.
R1.3 Non-Functional Requirements: the Questions
Numbers come in R1.7. For now, the questions and why each matters:
- Freshness. How old may a price on screen be before it's misleading? And how does the app know its prices are old, instead of silently showing the last one?
- No duplicate orders. A tap must produce at most one order, whatever the network does. And the user must never have to guess whether an order exists.
- Clear status. "Accepted by us", "acknowledged by the market" and "filled" are three different facts. The app must never show one as another.
- Availability during market hours. The regular US session runs 9:30 a.m. to 4:00 p.m. Eastern, 6.5 hours a day, about 252 days a year: hours. So 99.9% in market hours means at most minutes a year, about 8 minutes a month, and it's measured when customers need us, not at 3 a.m. on a Sunday.
R1.4 The API
Three pieces: an order API, one streaming connection for quotes and order updates, and a small history API for the chart. Every call carries the user's access token.
1. Place an order.
httpPOST /v1/orders HTTP/1.1 Host: api.example-broker.com Authorization: Bearer <token> Content-Type: application/json { "client_order_id": "0192f3c4-5b1e-7a42-9c3d-2f8e1a6b7c90", "account_id": "acc_5521", "symbol": "AAPL", "side": "BUY", "type": "LIMIT", "quantity": 10, "limit_price": "227.50", "time_in_force": "DAY" }
httpHTTP/1.1 201 Created Content-Type: application/json { "order_id": "ord_8KZ2M4Q1", "client_order_id": "0192f3c4-5b1e-7a42-9c3d-2f8e1a6b7c90", "status": "RECEIVED", "reserved_cash": "2275.00", "account_event_seq": 1042 }
client_order_idis the idempotency key (a unique ID the phone creates once per order intent, so a repeat of the same request is recognised). It's a UUIDv7, whose first bits are a timestamp, so IDs sort roughly by creation. It lives in the body, not a header, because in Round 2 it must sit inside the bytes the user signs.- Prices are decimal strings, never floating point, so 227.50 stays 227.50.
201means "we recorded it and reserved the money". Not "it traded", not even "the market has it". Those come later as status updates.- The same request again returns
200 OKwith this same body.
2. Look up and cancel.
httpGET /v1/orders/ord_8KZ2M4Q1 HTTP/1.1 Authorization: Bearer <token>
httpDELETE /v1/orders/ord_8KZ2M4Q1 HTTP/1.1 Authorization: Bearer <token>
A cancel returns 202 Accepted with "status": "PENDING_CANCEL": cancelling is a request to the market, and a fill can still win the race (step 1.3).
3. Resolve an order in doubt. When the app never got an answer to POST /v1/orders, it asks by its own key:
httpPOST /v1/orders:resolve HTTP/1.1 Authorization: Bearer <token> Content-Type: application/json { "account_id": "acc_5521", "client_order_id": "0192f3c4-5b1e-7a42-9c3d-2f8e1a6b7c90" }
It returns the order if it exists, or "status": "VOIDED", which means "it doesn't exist, and now it never will". Step 1.2 explains why this call writes.
4. The stream. One WebSocket per app session (a WebSocket is a long-lived, two-way connection that starts as an HTTP request and then carries messages in both directions). The app says what it wants:
json{ "op": "subscribe", "symbols": ["AAPL", "MSFT", "NVDA"] }
The server answers with a snapshot of each symbol's current quote, then updates, then heartbeats (small "I'm alive" frames) when nothing else is sent:
json{ "type": "quote", "symbol": "AAPL", "bid": "227.48", "ask": "227.52", "bid_size": 300, "ask_size": 200, "last": "227.50", "volume": 18422310, "server_ts": "2026-09-28T13:30:01.004Z" }
json{ "type": "order", "account_event_seq": 1043, "order_id": "ord_8KZ2M4Q1", "status": "NEW" }
json{ "type": "heartbeat", "server_ts": "2026-09-28T13:30:02.004Z" }
5. Today's chart. GET /v1/charts/AAPL?interval=1m&day=today returns one-minute bars (open, high, low, close, volume).
| Status | When | What the app does |
|---|---|---|
201 Created / 200 OK | Order recorded (first time / a repeat) | Show "Received", then follow updates on the stream |
400 Bad Request | Malformed order (a negative quantity, an unknown symbol) | Show the reason; nothing was created |
401 Unauthorized | Token expired | Refresh the token and resend the same request, same key |
403 Forbidden | The account isn't the caller's, or can't trade | Show the reason |
409 Conflict ORDER_VOIDED | The key was voided by a resolve call | The order doesn't exist; the user may place a new one |
422 INSUFFICIENT_BUYING_POWER | Not enough buying power | Show available buying power |
422 IDEMPOTENCY_KEY_REUSED | Same key, different order details | A client bug; change nothing |
503 BROKER_UNAVAILABLE | Our partner broker is down | Say plainly the order was not placed (R1.9) |
Recap
- One order API whose key is born on the phone; one resolve call for orders in doubt.
- One WebSocket for quotes and order updates, with heartbeats.
201means recorded and reserved, nothing more.
Let's build it, starting with the simplest thing that works.
R1.5 Design Evolution: From Polling and Posting to Streams and Keys
Every step below follows the same pattern: a problem, your turn to think, the answer, and what the answer costs us. The cost is always the next problem.
Step 1.0: The Baseline
The app polls GET /v1/quotes?symbols=... every second. The Buy button posts an order; the order service checks the balance, calls the partner broker, and returns what the broker said.
Synthesizing vector architecture diagram...
Everything is request and response. The app asks for prices; the order waits on the broker.
What's good about it: simple, and easy to reason about.
What it costs us: 50K phones × 1 poll a second is 50K requests a second, most returning the same prices; the prices are up to a second old before they even leave; and a retried order post is a second order.
Step 1.1: Polling Is Slow and Wasteful
The problem: 50,000 phones poll once a second. The quote API serves 50K requests a second, each with a new TLS round trip on a bad network, and users complain prices "jump" rather than move. What would you do?
Primitive: WebSocket, SSE & Long Polling
Step 1.2: The Network Dropped After I Tapped Buy, So I Tapped Again
The problem: a user on a train taps Buy. The request reaches us, we record the order and send it to the broker, but our response is lost in a tunnel. The app shows "Something went wrong". The user taps Buy again, and now owns 20 shares instead of 10. What would you do?
Primitive: Distributed Unique ID Generators
Step 1.3: "Is My Order Done?"
The problem: the app shows a green "Order placed!" toast as soon as the 201 arrives. Users read it as "I bought the shares". Some of those orders are later rejected by the broker, some sit unfilled all day, and one user sold shares he didn't yet own because the app said "bought".
What would you do?
Synthesizing vector architecture diagram...
Only RECEIVED, ROUTED, PENDING_CANCEL and a risk REJECTED are ours to set; every other state is a fact reported by the broker. A cancel can fail (the broker rejects it), which returns the order to where it was. Four states end the order.
Primitive: Change Data Capture & the Outbox Pattern
Step 1.4: The User Doesn't Have the Money
The problem: Maya has $3,000 of cash. On her phone she places a buy for $2,500; at the same moment, on her tablet, she places another $2,500 buy. The app checked her balance before each and both passed. Both orders reach the market. She now owes $2,000 she doesn't have. What would you do?
Primitive: Database Isolation Levels, ACID & Concurrency Anomalies · Drill: ACID isolation write skew anomaly (both of its questions are answered just above)
Step 1.5: Holdings Disagree With Orders
The problem: the app kept its own list of holdings: add 10 shares when the user taps Buy, remove them on Sell. After a partial fill, a rejected order and a reinstall, the app says Maya owns 14 shares of AAPL. The partner broker says 4. What would you do?
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | Polling for quotes; the order waits on the broker | Poll load; duplicates on retry |
| 1.1 | Polling is slow and wasteful | One WebSocket per session; snapshot, then at most one update per symbol per second; heartbeats; closed in the background | Long-lived connections |
| 1.2 | Retried taps duplicate orders | client_order_id born on the phone; unique key claimed in the order's transaction; 5 s server deadline < 10 s app timeout; a resolve call that voids; a local journal, never an offline queue | A key per order; a "checking" state |
| 1.3 | "Is my order done?" | Server-owned state machine; execution reports applied once by execution ID; a numbered event per account; never an optimistic fill | In-between states to design |
| 1.4 | Not enough money | Reserve buying power under the account row lock; collars for market buys; reserved shares for sells | Orders per account serialized |
| 1.5 | Holdings disagree | Server ledger; settled vs unsettled (T+1); nightly reconciliation; cached and labelled on the phone | A ledger and a reconciliation job |
Two costs stay open for Round 2: every quote goes to every watcher (which won't survive millions of users), and anyone holding a stolen session token can trade.
R1.6 Architecture v1
Now the concepts get AWS names.
Synthesizing vector architecture diagram...
Quotes flow in from the vendor, through the quote service, out through the gateways. Orders flow the other way: through the order API into one Aurora transaction, then out to the broker by the router. Everything the broker reports comes back through the router, into the database, and out to the phone as a numbered event.
The pieces:
- Quote service (2 tasks, active and standby) holds the vendor connection and the latest quote per symbol in memory. At the open, our users' ~5,000 watched symbols send it up to about 200K changes a second (R1.7).
- Stream gateways on ECS (Fargate) behind a Network Load Balancer (a layer-4 load balancer that passes TCP connections straight through). They hold the WebSockets, the subscriptions, and send each client one frame a second.
- Order API on ECS behind an Application Load Balancer: authenticates, validates, and runs the order transaction.
- Aurora PostgreSQL (a writer and a reader in another AZ) holds accounts, orders, keys, events, the ledger and an outbox.
- Outbox relay and router. The order transaction also writes an
outboxrow (the transactional outbox: a table of messages to send, written in the same transaction as the change, so "saved" and "will be sent" can't disagree). The router reads it, sends the order to the partner broker over FIX (the standard messaging protocol between brokers and markets) through a site-to-site VPN, and writes the broker's execution reports back. The outbox loop teaches the relay.
Schema (Aurora PostgreSQL; the ledger's own tables follow the digital wallet loop)
sqlCREATE TABLE accounts ( account_id TEXT PRIMARY KEY, user_id TEXT NOT NULL, -- the owner; checked against the token on every write cash_settled NUMERIC(18,2) NOT NULL, cash_unsettled NUMERIC(18,2) NOT NULL, cash_reserved NUMERIC(18,2) NOT NULL DEFAULT 0, -- promised to open buy orders event_seq BIGINT NOT NULL DEFAULT 0 -- last order event number ); CREATE TABLE positions ( account_id TEXT NOT NULL REFERENCES accounts, symbol TEXT NOT NULL, quantity NUMERIC(18,6) NOT NULL, qty_reserved NUMERIC(18,6) NOT NULL DEFAULT 0, -- promised to open sell orders PRIMARY KEY (account_id, symbol) ); CREATE TABLE order_keys ( account_id TEXT NOT NULL, client_order_id UUID NOT NULL, request_hash BYTEA, -- NULL for a VOID row state TEXT NOT NULL CHECK (state IN ('ORDER', 'VOID')), order_id TEXT, PRIMARY KEY (account_id, client_order_id) ); CREATE TABLE orders ( order_id TEXT PRIMARY KEY, account_id TEXT NOT NULL, client_order_id UUID NOT NULL, device_id TEXT NOT NULL, symbol TEXT NOT NULL, side TEXT NOT NULL CHECK (side IN ('BUY', 'SELL')), type TEXT NOT NULL CHECK (type IN ('MARKET', 'LIMIT')), quantity NUMERIC(18,6) NOT NULL CHECK (quantity > 0), limit_price NUMERIC(18,4), -- the collar price for a market buy time_in_force TEXT NOT NULL, status TEXT NOT NULL, filled_qty NUMERIC(18,6) NOT NULL DEFAULT 0, avg_fill_price NUMERIC(18,6), reserved_amount NUMERIC(18,2) NOT NULL, created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE INDEX orders_by_account ON orders (account_id, created_at DESC); CREATE TABLE executions ( broker TEXT NOT NULL, exec_id TEXT NOT NULL, -- the broker's execution ID order_id TEXT NOT NULL REFERENCES orders, quantity NUMERIC(18,6) NOT NULL, price NUMERIC(18,6) NOT NULL, executed_at TIMESTAMPTZ NOT NULL, PRIMARY KEY (broker, exec_id) -- a resent report is applied once ); CREATE TABLE order_events ( account_id TEXT NOT NULL, event_seq BIGINT NOT NULL, -- dense per account, set under the row lock order_id TEXT NOT NULL, payload JSONB NOT NULL, PRIMARY KEY (account_id, event_seq) ); CREATE TABLE device_last_order ( device_id TEXT PRIMARY KEY, account_id TEXT NOT NULL, client_order_id UUID NOT NULL, -- the last one accepted from this device accepted_at TIMESTAMPTZ NOT NULL -- server time; "newer" is judged by this, not the UUIDv7 );
Trace 1: a quote update.
Synthesizing vector architecture diagram...
Two vendor changes, one frame to the phone: the second replaced the first before the gateway's next second came round.
Trace 2: an order with a lost response and a retry.
Synthesizing vector architecture diagram...
The retry creates nothing. It reads the order the first request made, and the user sees one order moving through its states.
R1.7 Numbers
Quotes (assumptions: 50K users online at the open, 10 symbols each, every watched symbol changes at least once a second at the open, which is the worst case; JSON frames)
| Quantity | Math | Value |
|---|---|---|
| Subscriptions | 50K × 10 | 500K |
| Quote deliveries/s, peak | 500K × 1 update/s | 500K |
| Frame per client per second | 10 symbols × ~120 B of JSON + ~80 B of WebSocket, TLS and TCP/IP headers | ~1.28 KB |
| Egress, peak | 50K × 1.28 KB = 64 MB/s | ≈ 512 Mbps |
| Per user | 1.28 KB × 3,600 s | ≈ 4.6 MB per hour watching |
| Vendor input at the open | ~5,000 watched symbols × ~40 changes/s (a guess for the open) | ≈ 200K changes/s into one service |
Orders (assumptions: 100K orders a day; 30% of them in the first half hour; the busiest seconds are 5× that half hour's average)
| Quantity | Math | Value |
|---|---|---|
| Orders in the first 30 min | 100K × 30% | 30K |
| Average in that window | 30,000 ÷ 1,800 s | ≈ 16.7/s |
| Busiest seconds | 16.7 × 5 | ≈ 83/s; plan for 100/s |
| Execution reports | ~1.2 fills + 1 ack per order | ≈ 220K/day |
Monthly egress (assumption: averaged over the 6.5-hour session, stream traffic is about a quarter of the open's peak: fewer users after the first hour, and quiet symbols send nothing): ; TB a month.
Fleet (per-task rates are assumptions to confirm by load test; every fleet must carry the peak with one AZ lost, rounded per AZ)
| Fleet | Per task (2 vCPU, 4 GB) | Peak need | With one AZ lost | Tasks |
|---|---|---|---|---|
| Stream gateways | 20K connections | 50K ÷ 20K = 2.5 → 3 | 2 AZs carry 3 → 2 per AZ | 6 |
| Order API | 250 orders/s | 1 | 1 per AZ | 3 |
| Quote service, router | active + standby each | 4 |
Monthly cost (us-east-1 on-demand list prices; rounded)
| Item | Math | Monthly |
|---|---|---|
| ECS on Fargate (Graviton) | 13 tasks × (2 × $0.03238 + 4 × $0.00356)/h ≈ $0.079/h × 730 h | ≈ $750 |
| Aurora PostgreSQL | writer + reader, db.r7g.large at $0.276/h × 730, plus storage and I/O | ≈ $500 |
| Data transfer out | 7.9 TB × $0.09/GB (first 10 TB tier) | ≈ $710 |
| NLB and ALB | 2 × $0.0225/h × 730; NLB bytes 7,900 GB × $0.006 per NLCU-hour (1 GB = 1 NLCU-hour), ALB LCUs small | ≈ $90 |
| Site-to-site VPN to the broker | 2 VPN connections × $0.05/h × 730 | ≈ $75 |
| CloudWatch, logs, misc. | ≈ $300 | |
| Total AWS | ≈ $2.4K/month |
Say the headline: the AWS bill is small; the market-data licence probably isn't. Exchanges charge for real-time data per user (with different fees for professional and non-professional users) or through enterprise arrangements, and the vendor passes that on. Get the vendor's fee schedule before choosing between real-time and delayed data for each screen.
R1.8 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| WebSocket vs SSE | WebSocket: one connection carries subscribe and unsubscribe messages up, and quotes, order events and heartbeats down, in order | SSE is simpler: plain HTTP, built-in reconnect with Last-Event-ID, and friendlier to some proxies. But it's one-way: every subscription change becomes a separate request that must find the server holding this client's stream. For a chat app, where clients send constantly, the case for WebSocket is even stronger. |
| Our own fan-out vs the vendor's SDK in the app | Our own gateways | Vendors offer SDKs that stream quotes straight to each phone. That's less to build, but the vendor then sees every user, our licence counts get harder to control, and we can't add our own heartbeat, staleness rules or order events on the same connection. |
| Real-time vs delayed | Real-time on quote and order screens for signed-in customers; delayed on public pages | Real-time costs licence fees per user. A US rule (the "vendor display rule" in Regulation NMS) also expects a consolidated display (the NBBO and last sale) wherever a customer can make a trading decision, so the order ticket shows consolidated real-time data, not a cheaper single-exchange feed alone. Confirm the exact obligations with compliance. |
| One update per second vs every change | One per second | Users don't see every tick. Round 2 makes the window smaller and smarter. |
| Synchronous broker call vs outbox and router | Outbox and router: our 201 doesn't wait for the broker | The user sees "Received" and then "Working" a moment later, instead of one answer. In return, a slow broker never holds our API threads, and a crash between commit and send can't lose an order. |
R1.9 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| The quote stream drops | No frames, no heartbeats | After 3 seconds without any frame, every price on screen turns grey with "Prices paused". Market orders are blocked; the app reconnects with jittered backoff and resubscribes, and the snapshot refreshes every row. |
| The partner broker is down | The router can't connect; sends fail | New orders fail fast with 503 BROKER_UNAVAILABLE, and the app says plainly "Your order was not placed". We never accept an order we can't send and hold it for later: an endpoint that accepts an order must be able to forward it. Orders already sent keep their state; the router asks the broker for their status on reconnect. |
| Duplicate submit | Two requests with one key | The unique key and the resolve call (step 1.2). |
| The app is offline | No connection at all | Quotes grey out; the Buy button is disabled. No order is queued. |
| The API crashes after commit, before the router sends | An order stuck in RECEIVED | The outbox row is still there; the router sends it on its next pass. |
| The broker resends an execution report | The same fill twice | The unique execution ID applies it once. |
| The Aurora writer fails | Order writes fail for a short time | Aurora promotes the reader in another AZ, typically within about a minute; the order API retries its connection. Quotes don't depend on Aurora at all. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Idempotent orders with keys born on the phone and a resolve call that fences them; an outbox so accepted orders reach the broker; fleets in three AZs sized to lose one REL 4 · REL 10 |
| Performance Efficiency | Streams instead of polling; at most one update per symbol per second; quotes never touch the database PERF 4 |
| Security | Light this round: the acting user comes from the token; every order and cancel checks that the caller owns the account and the order SEC 3 |
| Cost Optimization | About $2.4K a month on AWS; the market-data licence is the line to check first COST 5 |
| Operational Excellence | Skipped this round. |
| Sustainability | Streams close when the app is in the background; one frame a second per client, not one per change SUS 3 |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Streams quotes and explains the WebSocket vs SSE choice by what the client sends, not by fashion.
- Makes orders idempotent with a key born on the phone and claimed atomically on the server, and handles the order whose outcome is unknown.
- Refuses to queue orders offline, and refuses optimistic fills.
- Designs the order state machine, and knows ack from fill.
- Puts buying power on the server under a lock, and can name the anomaly the lock prevents.
- Derives traffic and cost, and knows that data licensing matters more than servers at this size.
Follow-up questions
-
"Why not let the app decide the order is a duplicate?" Answer: the app can't see what the server did with a request whose response was lost. It can only keep the same key and ask. The decision has to be one atomic write on the server.
-
"What if the user edits the price and resubmits while the first order is unknown?" Answer: the edited order is a new intent with a new key, so the app must first resolve the unknown one. While an order is
UNKNOWN, the ticket for that symbol shows "Checking your last order…" and won't send a new one until the resolve call returns. -
"The phone's clock is wrong by an hour. What breaks?" Answer: nothing that matters. Keys only need to be unique (a UUIDv7 has 74 random bits besides its timestamp). The stale timer uses the phone's monotonic clock (time since boot), not wall time. Times on screen come from
server_ts.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Retry the order until it succeeds" | Without a key, each retry is a new order. |
| "Queue orders offline and send them later" | The user decided on a price that no longer exists. |
| "Show filled when the API returns" | The 201 means recorded, not traded. |
| "Check buying power in the app" | Two devices, one balance, and a modified app. |
| "Check whether a similar order exists first" | Check-then-act races, and a user may really want two. |
Round 2 · Senior · "3M Traders at the Opening Bell"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 10M accounts, 3M online at 9:30 · 45M subscriptions, 90M quote deliveries/s at peak · 2M orders/day, 5K/s at the open · order in, to broker, P99 < 150 ms · 99.99% in market hours
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built a trading app for a small brokerage: 50K users online at the open, 10 symbols each, 100K orders a day. Quotes stream over one WebSocket per app session: a snapshot on subscribe, then at most one frame a second with every changed symbol, and heartbeats when nothing changes; after 3 seconds of silence, prices grey out and market orders are blocked. Orders carry a
client_order_idborn on the phone. The server claims it with a unique key in the same Aurora transaction that locks the account row, reserves buying power and inserts the order, so a retry returns the original. The server abandons requests after 5 seconds; the app waits 10 before asking, and its resolve call voids a key that doesn't exist yet, so a late original can't create a second order. We never queue orders offline and never show a fill we weren't told about. The server owns the state machine: received, routed, working, partly filled, filled, cancelled, rejected, expired, with a numbered event per account. An outbox and a router send orders to our partner broker over FIX; execution reports come back and are applied once by execution ID. Holdings come from a ledger, settled on T+1, reconciled nightly with the partner. About $2.4K a month on AWS. Two costs are open: every quote goes to every watcher, and a stolen session token can trade."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: quotes down one path, orders up another, meeting only in the app.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Polling | One WebSocket; one frame a second; heartbeats | Long-lived connections |
| 1.2 | Duplicate orders | Key born on the phone; unique claim in the order's transaction; deadline; resolve call | A "checking" state |
| 1.3 | Unclear status | Server state machine; execution IDs; numbered events | In-between states |
| 1.4 | Not enough money | Reservation under the account row lock | Serialized orders per account |
| 1.5 | Holdings drift | Ledger; T+1; nightly reconciliation | A reconciliation job |
Open costs: fan-out that grows with every watcher; session tokens alone authorize trades.
R2.1 The Scope Raise
Interviewer: "We're a popular retail app now: 10 million accounts, and on a meme-stock morning 3 million people are online at 9:30, each watching about 15 symbols. The hot symbols change a hundred times a second. People scroll and pinch charts on 120 Hz phones and expect it to feel like glass. Last month someone stole a session token and traded in a customer's account, so orders must be signed with the user's biometric. Nobody may ever trade on a stale price. And the first minutes bring 5,000 orders a second."
A scope raise is not the end of scoping. We ask back, and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How many symbols, and how fast do they change? | About 10,000 US stocks and ETFs. The busiest change about 100 times a second at the open; most far less. | Input is per symbol and bounded; the problem is the fan-out (step 2.1). |
| How fresh must a displayed price be? | Half a second is fine. A price that stopped updating must look stopped. | A 500 ms conflation window, and a staleness watchdog (steps 2.1, 2.4). |
| What does "smooth" mean exactly? | No dropped frames while scrolling or panning a chart during a volatile open, on 120 Hz screens. | A frame-synced render pipeline off the main thread (step 2.2). |
| What must a stolen session not be able to do? | Place or change orders. Viewing is less critical. | Orders signed by a hardware key unlocked by the user's biometric (step 2.5). |
| What latency do you promise for orders? | From the order reaching us to the broker receiving it: P99 under 150 ms. | A budget we can own; the phone's network isn't in it (R2.6). |
| How many orders? | 2M a day; about 5,000 a second in the first minutes after 9:30. | Admission control and pre-scaling (step 2.6). |
| How available? | 99.99% during market hours. | About 9.8 minutes a year: min. |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Users online at the open | 50K | 3M (10M accounts) |
| Subscriptions | 500K | 45M (3M × 15) |
| Quote freshness | ~1 s | 500 ms; stale prices marked and not tradable |
| Rendering | Basic lists | 120 Hz scrolling and charts, no dropped frames |
| Order authorization | Session token | Token + a per-order signature from a biometric-unlocked hardware key |
| Orders | 100K/day; ~100/s peak | 2M/day; 5K/s at the open; P99 < 150 ms in, to broker |
| Availability | 99.9% in market hours | 99.99% in market hours (≈ 9.8 min a year) |
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | What breaks at the new scope |
|---|---|
| One quote service, one frame per client per second | Input at the open is up to 1M changes a second across 10,000 symbols; one process can't parse it, and a crash takes every quote down. |
| JSON frames, ~120 B per symbol | Even conflated to 2 updates a second, 45M subscriptions × 2 × 120 B = 10.8 GB/s, about 86 Gbps. |
| "Send every change" as the fallback idea | The common estimate of "45M × 100 = 4.5B events/s" is not the market's tick rate; it's what we'd deliver if every watched symbol forwarded 100 ticks a second to every watcher: 4.5B × 40 B = 180 GB/s, 1.44 Tbps, and 1,500 updates a second per phone. Step 2.1 does this math properly. |
| Redraw a row on every message | During a spike, parsing and layout on the main thread blow the 8.33 ms frame budget: stutters and frozen touches. |
| A 3-second timer on receipt of frames | After an elevator ride, the connection delivers a burst of buffered, old frames; the timer resets on old news, and the user trades on it. |
| Session tokens authorize trades | The incident we just had. |
| Order events broadcast to every gateway | 5K orders/s × ~3 events × 225 gateways ≈ 3.4M internal messages a second, almost all thrown away. |
| Everything scales on demand | At 9:30 the demand arrives within a minute; auto scaling reacts in minutes. |
R2.3 New Requirements and API Additions
1. A binary quote record, with its own sequence. Frames are now binary (a compact encoding such as Protocol Buffers or SBE; the choice doesn't matter here). One WebSocket message carries a frame header and one or more records.
| Field | Typical size | What it is | Why |
|---|---|---|---|
symbol_id | 3 B | A number from today's symbol table, which the app downloads once a day | Tickers as text cost more bytes |
epoch | 4 B | The conflator shard's epoch (a number that goes up each time a new process takes over the shard) | A new owner's numbers never mix with an old owner's (step 2.3); 4 bytes, because a 1-byte epoch would wrap after 255 takeovers |
cseq | 4 B | The conflated sequence number: +1 each time the conflator emits this symbol | Ordering and gap detection on our stream, not the exchange's |
kind + field_mask | 2 B | UPDATE (only the changed fields follow) or SNAPSHOT (all fields), and which fields are present | Deltas are smaller; snapshots reset |
bid, ask, last | ~12 B | Prices as integers in hundredths of a cent, variable-length encoded | Exact, compact |
bid_size, ask_size, last_size | ~8 B | Sizes | |
volume | ~5 B | Today's volume | |
status | 1 B | Flags: HALTED, LIMIT_STATE, STALE | Halts and staleness are part of the quote |
| Total, full record | ~39–43 B | We plan with 40 B; a typical UPDATE is 15–25 B |
The frame header carries the protocol version, the frame type (DATA, HEARTBEAT, ORDER_EVENT, TIME) and server_ts in milliseconds.
2. Snapshots. In-band, on the same WebSocket, ordered with the updates:
json{ "op": "snapshot", "symbols": ["NVDA"] }
And over HTTP, one URL per symbol, cacheable for 1 second at CloudFront, for the moments before the WebSocket is up (a cold launch, a widget, a notification tap):
httpGET /v1/quotes/NVDA HTTP/1.1 Host: api.example-broker.com
httpHTTP/1.1 200 OK Cache-Control: public, max-age=1 Content-Type: application/json { "symbol": "NVDA", "epoch": 7, "cseq": 481203, "bid": "131.42", "ask": "131.45", "last": "131.44", "status": [], "server_ts": "2026-09-28T13:30:00.512Z" }
3. A signed order. The app signs the exact bytes it sends; the server parses what was signed, never a re-serialized copy (two JSON encoders rarely agree byte for byte).
httpPOST /v1/orders HTTP/1.1 Host: api.example-broker.com Authorization: Bearer <token> Content-Type: application/json { "key_id": "dk_7f3a91c2", "payload": "eyJjbGllbnRfb3JkZXJfaWQiOiIwMTkyZjNjNC01YjFlLTdhNDIt...", "signature": "MEUCIQDs2x9...AiB4kQ" }
payload is the base64url of these bytes, and signature is ECDSA P-256 over their SHA-256:
json{ "client_order_id": "0192f3c4-5b1e-7a42-9c3d-2f8e1a6b7c90", "account_id": "acc_5521", "device_id": "dev_91c0", "symbol": "NVDA", "side": "BUY", "type": "MARKET", "quantity": 20, "time_in_force": "DAY", "quote_ref": { "symbol": "NVDA", "epoch": 7, "cseq": 481203 }, "shown_price": "131.45", "signed_at_client": "2026-09-28T13:30:01.330Z" }
client_order_idis inside the signature. If it were only a header, a captured signed order could be replayed under a fresh key and accepted as a new order.quote_refnames the quote the user saw, by the server's own identifiers. The server checks its age with its own clock (step 2.4);signed_at_clientis for support, never for decisions.
4. New error codes.
| Status and code | Meaning | What the app does |
|---|---|---|
422 QUOTE_STALE | The quote the user saw was too old when the order arrived | Show the fresh quote (in the body); ask again |
422 PRICE_MOVED | A market order's price moved more than the allowed band since the quote | Show the new price; ask again |
422 SYMBOL_HALTED | Trading is halted; market orders are refused | Show the halt badge |
401 SIGNATURE_INVALID | The signature doesn't verify | Don't retry; report; offer re-pairing |
403 DEVICE_KEY_REVOKED | The key was revoked (a new biometric, sign-out elsewhere) | Start re-pairing (step 2.5) |
429 RATE_LIMITED | Account or fleet admission | Honour Retry-After; cancels use their own limit |
R2.4 Design Evolution: Volume, Frames, Freshness and Intent
Step 2.1: 100 Ticks a Second, Times Millions of Watchers
The problem: the hot symbols change 100 times a second at the open, 3M people are online, and they hold 45M subscriptions between them. The first proposal is "send every change to everyone who watches it". What would you do? And first: how many messages a second is that, really?
Synthesizing vector architecture diagram...
Volume shrinks 50× in the conflators (1M to 20K a second), then grows 4,500× in the gateways (20K to 90M), and only there.
Step 2.2: The UI Stutters During Volatility
The problem: the network callback updates the watchlist's view model for every record. During a volatile open, rows relayout constantly, a pinch on the chart hitches, and the Android profiler shows the main thread busy 30 ms at a time. What would you do?
Synthesizing vector architecture diagram...
Records arrive whenever the network delivers them; the screen reads them only at its own rhythm. Nothing the network does can make a frame late.
Step 2.3: A Record Arrived Out of Order, and a Gap Appeared
The problem: after a reconnect to a different gateway, a phone briefly shows NVDA's price go 131.45 → 131.38 → 131.45: an older record arrived after a newer one. Elsewhere, a phone missed one delta record and has shown the wrong bid size for a minute. A teammate proposes: "Put the exchange's sequence number in each record; if it jumps, fetch a snapshot." What would you do?
Primitive: Distributed Locks & Leases
Step 2.4: The User Traded on a 30-Second-Old Price in an Elevator
The problem: Priya watches a price of 120.00 and steps into an elevator. The connection stalls without closing. The market jumps to 135.00. When the doors open, she taps Buy on a market order for 100 shares while the app still shows 120.00. Her fill is 1,500 dollars worse than she expected. What would you do?
Synthesizing vector architecture diagram...
Market orders are allowed only in LIVE. RESYNC and STALE dim the price; HALTED shows a badge.
Step 2.5: A Stolen Session Token Placed Trades
The problem: malware on one customer's laptop stole their access and refresh tokens from a browser session, and the attacker placed orders through our API. Tokens say who is calling; they say nothing about whether that person meant this order. What would you do?
Synthesizing vector architecture diagram...
The private key never leaves the secure hardware, and it signs nothing without the biometric. The server trusts the math, not a flag.
Primitive: OAuth2, OIDC & Distributed Token Authentication
Step 2.6: The Opening Bell Floods Us
The problem: at 9:29 three million people open the app. Each opens a WebSocket, refreshes a token, and loads 15 snapshots; at 9:30:00 the orders start at 5,000 a second. Auto scaling notices a few minutes later, which is after the busiest minute of the day. Last Tuesday a trading bot on one account sent 300 orders a second, and legitimate cancels waited behind them. What would you do?
Primitive: Distributed Rate Limiting · Drill: Rate limiting a partner API (both of its questions, shared counters and token bucket vs fixed window, are answered just above)
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Ticks × watchers | Feed handlers → 8 conflator shards (500 ms) → every gateway → per-client flush every 500 ms; per-client conflation for slow phones | Users don't see every tick |
| 2.2 | UI stutters | Decode off the main thread into a latest-value table; draw once per frame, visible rows only; 120 Hz opt-in on iPhone | Rendering discipline |
| 2.3 | Out of order, gaps | cseq on the conflated stream; UPDATE vs SNAPSHOT rules; in-band snapshots; gateways convert their own skips; leased single writer with an epoch | A protocol every client implements |
| 2.4 | Stale price in an elevator | 1 s heartbeats with server time; offset measured on the connection; stale after 2.5 s by either clock; shard heartbeats; server re-checks quote_ref at 5 s and 1% | Blocked market orders on bad networks |
| 2.5 | Stolen tokens trade | P-256 key in Secure Enclave or StrongBox, biometric per use; attestation; server verifies exact bytes; re-pairing when biometrics change | A prompt per order; recovery flows |
| 2.6 | The opening flood | Scheduled pre-scaling; in-memory and per-symbol edge snapshots; shed before TLS; pre-warmed sessions; per-account token buckets in Valkey; cancels first | Idle capacity each morning; some 429s |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
Three streams. Market data: vendor → feed handlers → conflators → every gateway, order API task and alert engine. Orders: phone → ALB → order API → Aurora → routers → broker. Order events: routers → Valkey channel for that account → the one gateway holding that account's connection → phone.
What changed from Round 1:
- The market-data tier is split into feed handlers, 8 conflator shards with hot standbys, and a gateway fleet; the order API and alert engines subscribe to the same conflated stream.
- Order events go through per-account channels. When a connection authenticates, its gateway subscribes to the account's channel in Valkey (sharded pub/sub, so channels spread over the cluster's shards). Routers publish each numbered event to the account's channel. Pub/sub doesn't store messages, so an event published while a phone is reconnecting is simply missed; the per-account
event_seq(step 1.3) makes the gap visible, and the app fetches the missing events from the order API. The numbered log in Aurora is the source of truth; the channel is only the fast path. - Routers partition accounts over 4 FIX sessions to the broker, each an active router with a standby, over AWS Direct Connect. A router holds its session with a lease and an epoch, like the conflators, and the broker accepts only one logon per session, which fences a second router at the other end. After an order commits, the order API hands it to its router directly (the fast path); the outbox relay only sends orders the fast path missed. Before sending, the router claims the order with
UPDATE orders SET status = 'ROUTED' WHERE order_id = $1 AND status = 'RECEIVED', so the fast path and the relay can't both send it. - DynamoDB holds device keys, sessions, price alerts and the conflator leases.
Server tables added in Round 2
| Table | Key | Attributes |
|---|---|---|
DeviceKeys (DynamoDB) | PK = USER#<user_id>, SK = KEY#<key_id> | device_id, public_key (SPKI), security_level (STRONGBOX, TEE, SECURE_ENCLAVE), status (ACTIVE, REVOKED), created_at, revoked_at |
Sessions (DynamoDB) | PK = SESSION#<session_id> | user_id, device_id, refresh_hash, previous_refresh_hash, rotated_at (for the 60 s grace window), status |
ConflatorLease (DynamoDB) | PK = SHARD#<n> | owner, epoch, version (a counter, never a timestamp) |
PriceAlerts (DynamoDB) | PK = USER#<user_id>, SK = ALERT#<alert_id> | symbol, direction (ABOVE, BELOW), threshold, status (ARMED, FIRED), fired_at, version (incremented on every change) |
Price alerts ("tell me when NVDA goes above 140")
- An index fed from the change stream. Alert engines own symbols by hash, like conflators. Each keeps, per symbol, two sorted lists of armed thresholds (above and below). The table's DynamoDB Stream is consumed by a Lambda function that forwards each record to the engine that owns its symbol. At startup an engine notes its stream position first: it starts accepting forwarded records into a buffer, then loads its symbols' alerts from the table, then applies the buffered and later records, keeping a record only if its alert
versionis newer than the one it holds. Loading first and subscribing second would lose any change written in between. Streams order records per item only, and an alert's create, edit and delete are one item, so they arrive in order; different alerts don't need a shared order. The stream consumer is a Lambda function with bisect on error (a failing batch is split in half until the bad record is isolated) and an on-failure destination, and a nightly job compares each engine's index with the table and repairs differences. - Evaluation. On each conflated record, the engine finds crossed thresholds by binary search in the sorted lists: cost grows with alerts that fire, not with alerts that exist.
- Firing once. The engine writes
status = FIRED"only if stillARMED", then sends the push. A retry after a crash between the two re-sends the push, so each push carries a collapse ID (apns-collapse-idon iOS, the notification tag on Android) derived from the alert ID: a second copy replaces the first on the lock screen instead of adding another. - Fill notifications work the same way, with the order ID and fill number as the collapse ID and tag.
Trace 1: the open.
Synthesizing vector architecture diagram...
Trace 2: the elevator.
Synthesizing vector architecture diagram...
The burst of old frames is received "now" but dated 30 seconds ago; the server-time rule keeps the price dimmed until a fresh frame arrives, and the user confirms at 135.00, not 120.00.
Trace 3: a signed order after a network drop.
Synthesizing vector architecture diagram...
No second signature and no second order. The resolve call either finds the order or makes sure it can never appear.
R2.6 Numbers and Cost
Market data (the interviewer's figures, derived; 40 B per record; the peak assumes every watched symbol changes in every 500 ms window, which is the worst case)
| Quantity | Math | Value |
|---|---|---|
| Subscriptions | 3M × 15 | 45M |
| Ticks in, upper bound | 10,000 symbols × 100/s | 1M/s ≈ 40 MB/s ≈ 320 Mbps |
| Conflated records | 10,000 × 2/s | 20K/s ≈ 800 KB/s per full copy |
| Deliveries out | 45M × 2/s | 90M/s |
| Payload out | 90M × 40 B = 3.6 GB/s | 28.8 Gbps |
| Frames out | 3M clients × 2/s | 6M/s |
| Bytes per frame on the wire | 15 × 40 B + ~80 B of WebSocket, TLS and TCP/IP headers | ~680 B |
| Egress on the wire, peak | 6M × 680 B = 4.08 GB/s | ≈ 32.6 Gbps |
| Per phone | 2 × 680 B = 1.36 KB/s | ≈ 4.9 MB per hour watching |
| Heartbeats when quiet | 3M × 1/s × ~110 B | ≈ 2.6 Gbps, sent only when there's no data, so the peak is the larger of the two, not the sum |
Synthesizing vector architecture diagram...
Conflation is the big cut (1.44 Tbps of payload to 28.8 Gbps); the binary format then cuts JSON's 86.4 Gbps to about 32.6 Gbps on the wire, headers included.
Orders
| Quantity | Math | Value |
|---|---|---|
| Orders per day | given | 2M |
| Peak | given | 5K/s |
| Order ingress at peak | 5,000 × 1.2 KB (payload, signature, headers) = 6 MB/s | 48 Mbps |
| Aurora transactions at peak | orders 5K/s + execution reports ~6K/s (acks and fills) | ~11K/s; sized for 7.5K orders/s with fills |
Order latency budget (P99, from the order reaching our load balancer to the broker receiving it; the steps depend on each other, so they add)
| Step | P99 |
|---|---|
| ALB, on an already open TLS connection | 2 ms |
| Token check, account token bucket in Valkey | 2 ms |
| Device key (cached) and ECDSA verification | 2 ms |
quote_ref check and risk rules, in memory | 3 ms |
| Aurora transaction: key, lock, reserve, order, event, outbox, commit (busy at the open) | 25 ms |
| Hand-off to the router | 3 ms |
Router claim (UPDATE … WHERE status = 'RECEIVED') | 10 ms |
| FIX over Direct Connect to the broker | 5 ms |
| Total | ≈ 52 ms, leaving headroom to 150 ms for pauses and a reconnect |
An order that misses the fast path waits for the outbox relay, which polls every second: rare, and outside the fast path's budget, so we alarm on how often it happens.
Fleet sizing (per-task rates are assumptions to confirm by load test; every fleet carries the peak with one AZ lost, rounded per AZ)
| Fleet (2 vCPU, 4 GB) | Per task | Peak need | With one AZ lost | Tasks |
|---|---|---|---|---|
| Stream gateways | 20K connections (~218 Mbps out at peak) | 3M ÷ 20K = 150 | 75 per AZ | 225 |
| Order API | 250 orders/s | 5,000 ÷ 250 = 20 | 10 per AZ | 30 |
| Conflators | 125K ticks/s | 8 shards | active + standby in different AZs | 16 |
| Feed handlers | one vendor line each | 2 lines | active + standby | 4 |
| Routers | one FIX session | 4 sessions | active + standby | 8 |
| Alert engines, push senders, auth | 2 per AZ each | 18 |
Quotas to raise before launch: Fargate vCPUs (225 + 76 tasks × 2 vCPU ≈ 600 vCPUs at the open); DynamoDB warm throughput on the session table; FCM's default of 600,000 messages a minute (10K a second) per project, which a mass alert firing at the open can reach, so push senders pace themselves under it.
Monthly cost (us-east-1 list prices; 21 trading days; stream traffic averaged over the 23,400-second session at a quarter of the open's peak, as in Round 1; rounded)
| Item | Math | Monthly |
|---|---|---|
| Data transfer out, streams | 4.08 GB/s × 25% × 23,400 s × 21 ≈ 501 TB: 10 TB × $90 + 40 × $85 + 100 × $70 + 351 × $50 per TB | ≈ $28.9K |
| NLB | processed bytes dominate: 501,000 GB ≈ 501,000 NLCU-hours × $0.006, + $0.0225/h (3M active connections are only 30 NLCU) | ≈ $3.0K |
| Cross-AZ, conflated stream to gateways | 225 × 800 KB/s = 180 MB/s, two thirds cross-AZ, averaged at half: 60 MB/s × 23,400 s × 21 ≈ 29.5 TB × $0.02/GB | ≈ $0.6K |
| Gateways, Fargate (Graviton, $0.079/h) | 225 tasks × 210 h (9 a.m.–7 p.m. on trading days) + 15 tasks × 520 h | ≈ $4.4K |
| Other Fargate tasks | 76 tasks × $0.079 × 730 | ≈ $4.4K |
| Aurora PostgreSQL, I/O-Optimized | writer + reader db.r7g.4xlarge ≈ 2 × $2.87 × 730, + storage | ≈ $4.5K |
| ElastiCache for Valkey | 3 shards × 2 × cache.r7g.large ≈ 6 × $0.175 × 730 | ≈ $0.8K |
| DynamoDB | sessions (~4M daily users × 4 refreshes of 15-minute tokens ≈ 16M writes a day × 30.4 ≈ 486M a month × $0.625 per million ≈ $0.3K), keys, alerts, leases, on demand | ≈ $0.5K |
| CloudFront, snapshots and charts | 3M users × 60 requests × 21 ≈ 3.78B × $0.01 per 10K + 18.9 TB × $0.080–0.085 | ≈ $5.3K |
| Direct Connect to the broker | two 1 Gbps connections, port hours and cross-connects | ≈ $0.6K |
| CloudWatch, WAF, logs, misc. | ≈ $6K | |
| Total AWS | ≈ $59K/month |
Say the headline: half the bill is bytes leaving AWS, and every byte saved per record is multiplied by 90M a second. The market-data licences for 3M real-time users are a separate bill, and likely the larger one.
The pricing trap: a managed WebSocket service. API Gateway WebSocket APIs charge per message, tiered ($1.00 per million for the first 1B messages a month, then $0.80 per million; each message up to 32 KB), plus $0.25 per million connection minutes. Our average of 1.5M frames a second over the session is messages a month: 1\text{B} \times \1.00 + 736\text{B} \times $0.80 per million ≈ **\590K, about $0.6M**, before connection minutes. Per-message pricing is great for chat; it's the wrong model for a firehose.
R2.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| Conflation window | 500 ms at the conflator, 500 ms flush per client | Fresher (100 ms) means 5× the deliveries and bytes; staler (1 s) feels laggy on a hot stock. The window is per product decision; a "pro" screen could get 100 ms for one symbol. |
| SSE vs WebSocket | WebSocket | In-band snapshots, subscriptions and time sync need the client to talk; SSE would need a second channel for all three. We give up SSE's plain-HTTP simplicity and built-in resume. |
| Blocking vs warning on stale prices | Block market orders; confirm limit orders | Some users on bad networks can't place market orders. A limit order caps the damage by itself; a market order on a stale price has no cap. |
| Biometric friction vs security | A prompt per order | A second per order, and users without strong biometrics can't trade from that device. A time window ("no prompt for 5 minutes") would make the key usable without a fresh biometric, which is exactly what a stolen, unlocked phone needs. |
| Deltas vs full records | Deltas with gap detection; snapshots on doubt | A stricter client protocol, for about half the bytes of full records. |
| Our gateways vs API Gateway WebSockets | Our gateways on Fargate behind an NLB | We run and patch a fleet; we save about $0.6M a month (R2.6). |
R2.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| Biometric key invalidated | Signing fails right after the user added a fingerprint | The app catches it, blocks trading on this device, and re-pairs with a second factor; the server revokes the old key (step 2.5). We alarm if invalidations spike for one app version: a bug there locks people out of trading. |
| Out-of-order or duplicate records | A price jumps back after a gateway switch | Dropped by the (epoch, cseq) rules; gaps trigger in-band snapshots (step 2.3). |
| A conflator dies | A shard's symbols stop updating | Timing, every delay counted: the gateway marks them STALE after 1.5 s of shard silence, and phones see it at their next flush (≤ 0.5 s), about 2 s after the death. The standby takes the lease after 5 s without renewal, emits snapshots under a new epoch, and prices are live again after about 5.5 s. |
| The opening stampede beats the plan | Gateway accept limits hit; 429s | Shed before TLS; clients back off with full jitter; the per-account and per-task buckets protect Aurora; cancels keep flowing. |
| Network drop mid-order | No response | Resolve by key after 10 s (step 1.2); no second signature needed. |
| A router crashes after claiming, before sending | An order stuck in ROUTED with no ack | The standby router takes the session and, for every ROUTED order without an ack, asks the broker for its status by our order ID (FIX order status request). Unknown to the broker: sent again with the same client order ID and the "possible resend" flag, so a copy that did arrive can be recognised. |
| Rooted or jailbroken device | Hooked biometric APIs | Hooks can fake a true, but not a signature from a hardware key that needs a real biometric. Attestation and integrity verdicts (Play Integrity, App Attest) feed risk scores; a device that fails them can view, not trade, a policy choice. |
| A halt or a circuit breaker | A symbol, or the whole market, stops trading | The vendor feed carries halt status; the record's HALTED flag shows a badge and market orders get SYMBOL_HALTED. Individual stocks pause under the Limit Up-Limit Down plan (a 15-second limit state, then a 5-minute pause); a market-wide circuit breaker halts everything for 15 minutes when the S&P 500 falls 7% or 13% before 3:25 p.m., and for the day at 20%. The matching engine loop, steps 2.6 and 3.6, explains them from the exchange's side. What happens to open orders in a halt is up to the venue and broker; we show what they report. |
| Holidays and early closes | The app says "market open" on the day after Thanksgiving at 2 p.m. | The market calendar has a base schedule plus a separate override layer (holidays, 1 p.m. early closes, emergency closures), checked first. An override never edits the base schedule. |
| Reconnect storm after a gateway deploy | 20K clients per task drop at once | Deploys drain one task at a time outside market hours (a release freeze arrives in Round 3); clients reconnect with full jitter; gateways shed before TLS and check tokens locally, so no auth service sees the storm. Each reconnect also re-subscribes its account's order-event channel in Valkey, and the jitter spreads those subscribe commands too. |
| Clock skew | A phone hours off | Staleness uses the monotonic clock and the measured offset; the server judges quotes by quote_ref and its own clock. The phone's wall clock decides nothing. |
| App version skew | An old app can't decode a new record field | The frame header carries a protocol version negotiated at connect; new fields are optional and old parsers skip them. An app too old for the binary protocol gets JSON frames from a compatibility path until its share is negligible. |
Drill: WebSocket reconnect thundering herd (the storm rows above and step 2.6 answer its first question; R1.8 and R2.7 answer why WebSocket over SSE)
R2.9 Production Gotchas
| Gotcha | Symptom | Cause | Fix |
|---|---|---|---|
| Private keys in app preferences | Keys copied from backups or rooted phones | A software key in SharedPreferences or UserDefaults | Keys created in the Secure Enclave or Keystore, never exportable |
| UI updates on every packet | Jank during volatility | Network callbacks touch views | Latest-value table; draw once per frame |
| Client-only buying power | Two devices overspend | The phone checks, the server trusts | Reservation under the account row lock (step 1.4) |
| Gap detection on exchange sequence numbers | Every record looks like a gap after conflation | Using the input's numbers on the output stream | cseq on the conflated stream |
| Staleness by receipt time | Old prices look live after a stall | A burst of buffered frames resets the timer | Age by server time plus the measured offset |
| Biometric result as a boolean | Hooked APIs trade | Trusting biometric_ok: true | Verify a hardware signature on the server |
| Signing a re-serialized object | Valid orders fail verification after an app update | Client and server serialize JSON differently | Sign and send the exact bytes |
| One snapshot URL per watchlist | Origin melts at 9:30 | Nothing is shared in the cache | One URL per symbol |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | Leased conflators with epochs and hot standbys; in-band snapshots after gaps; pre-scaling for the open; admission that sheds before expensive work; quotas raised ahead REL 7 · REL 5 · REL 1 |
| Performance Efficiency | Conflation (1.44 Tbps → 28.8 Gbps of payload); binary deltas; a 52 ms order budget; drawing once per frame, visible rows only PERF 4 · PERF 1 |
| Security | Per-order signatures from biometric-unlocked hardware keys; attestation at enrollment; exact-bytes verification; ownership checks on key, account and device; refresh rotation with a grace window SEC 2 · SEC 3 |
| Cost Optimization | ≈ $59K a month, derived; egress is half, so bytes per record matter; our gateways instead of ~$0.6M a month of per-message fees COST 5 · COST 8 |
| Operational Excellence | Light this round: the alarms below, counted by distinct installs OPS 8 |
| Sustainability | Streams only while visible; fewer, smaller records per phone; gateways scaled down outside market hours SUS 2 · SUS 3 |
Alarms and first actions (every user-facing rate counts distinct installs, not events: one phone in a reconnect loop can send thousands of stale reports)
| Signal | Alarm | First action |
|---|---|---|
| Order in → broker P99 | > 150 ms for 3 min | Per-stage timings: Aurora transaction, router claim, FIX send |
Installs with a STALE price | > 1% of connected installs for 1 min | Is one shard stale (a conflator), one AZ (gateways), or everyone (the vendor)? |
| Tick-to-screen latency (conflator emit → drawn, from sampled clients) | p95 > 1.5 s | Gateway flush lag, then client frame times |
| Client gap rate | > 0.1% of installs per 5 min | A gateway bug converting skips, or a conflator epoch storm |
| Fast-path misses (orders sent by the relay) | > 0.1% of orders | Hand-off failures between order API and routers |
| Biometric invalidations | 3× baseline for one app version | A key-handling bug in that version |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Separates ticks in (per symbol, bounded) from deliveries out (per subscriber) and puts conflation where it cuts the most.
- Handles slow clients by conflating per client instead of queueing.
- Puts the sequence on the conflated stream, with a fenced single writer, and knows why exchange sequence numbers don't work after conflation.
- Designs the render loop around the frame budget, not around message arrival.
- Catches the elevator burst: staleness by server time with a measured offset, and a server-side check by quote reference.
- Signs orders with hardware keys and verifies exact bytes; plans for key invalidation and recovery.
- Prepares for the open on a schedule, sheds before expensive work, and prices the managed alternative.
Follow-up questions
-
"A user wants every tick for one symbol, for day trading." Answer: a separate, opt-in subscription class: one or two symbols at a 100 ms window from the conflator (a second emit schedule for those symbols), licensed and priced for it. It multiplies deliveries only for the few who ask; the default stays at 500 ms.
-
"Why not skip the server check if the app already blocks stale prices?" Answer: the app is outside our control: an old version, a modified app, or a bug. The server's check uses only its own clock and its own records, so it holds even when the client is wrong.
-
"What if the Secure Enclave key is fine but the user's session was stolen on another device?" Answer: the thief has a token but no registered key, so they can't sign orders. They'd have to enroll a new device, which needs a login with a second factor, and Round 3 adds new-device risk checks and alerts to the user's other devices.
Interview gotchas from this round
| Gotcha | Why it's wrong |
|---|---|
| "4.5B ticks a second" | The market produces at most ~1M; 4.5B is deliveries without conflation. |
| "Check sequence gaps with exchange numbers" | Conflation skips them by design. |
| "Queue updates for slow clients" | Memory grows without bound and prices arrive late; replace instead. |
| "Reset the stale timer on every frame" | A burst of old frames looks fresh. |
"Send biometric_ok: true" | A boolean is trivially forged. |
| "Auto scaling handles the open" | It reacts minutes after the busiest minute. |
Round 3 · Architect · "A Regulated Broker: Audit, DR and Fairness"
~45 min · Principal (L7) · us-east-1 primary for orders, us-east-2 warm for orders, market data active in both · 25M accounts, 6M online at the open · 90M subscriptions · 8M orders/day with options, 20K/s at the open · RPO 0 for accepted orders · order entry back within 15 min of a region loss
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Rounds 1 and 2.
Round 2 in 60 seconds. "We run a retail trading app with 10 million accounts and 3 million people online at 9:30, watching 45 million symbol subscriptions. Feed handlers take up to a million ticks a second from the vendor; 8 conflator shards keep the latest state per symbol and emit each changed symbol every 500 ms, 20,000 records a second for the whole market. Every gateway gets that stream and flushes each client every 500 ms: 90 million deliveries a second, 28.8 Gbps of payload, about 32.6 on the wire. Slow clients get replaced records, never queues. Records carry
(epoch, cseq)on our conflated stream, so the app drops old ones and asks for in-band snapshots on gaps; conflators hold a lease and a new owner bumps the epoch. The app draws once per frame from a latest-value table. Prices go stale after 2.5 s without a frame or when the newest frame is 2.5 s old by server time, using an offset measured on the connection; stale blocks market orders, and the server re-checks the quote reference at 5 s. Orders are signed per use by a P-256 key in the Secure Enclave or StrongBox, unlocked by a current biometric; the server verifies the exact bytes. We pre-scale for the open, shed before TLS, and admit orders with per-account token buckets, cancels first. Order in to broker is about 52 ms at P99, and AWS costs about $59K a month, half of it egress. Open costs: logs aren't records, one broker route, one region, a new device is trusted too easily, releases can break trading, and the risk checks only understand stocks."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: a conflated market-data tier and a signed, idempotent order path, in one region.
Steps so far
| Step | Problem | Component |
|---|---|---|
| 1.1 | Polling | One WebSocket per session |
| 1.2 | Duplicate orders | Key born on the phone; unique claim; deadline; resolve call |
| 1.3–1.5 | Status, money, holdings | Server state machine; reservation under the account lock; ledger with T+1 |
| 2.1 | Ticks × watchers | Conflation per symbol, fan-out, per-client conflation |
| 2.2 | Jank | Latest-value table; draw once per frame |
| 2.3 | Order and gaps | (epoch, cseq) on the conflated stream; in-band snapshots |
| 2.4 | Stale prices | Two-clock watchdog; server-side quote_ref check |
| 2.5 | Stolen tokens | Biometric-unlocked hardware keys; exact-bytes signatures |
| 2.6 | The open | Pre-scaling; edge and in-memory snapshots; admission by priority |
Open costs: no provable records; one route; one region; weak new-device trust; risky releases; stock-only risk checks.
R3.1 The Scope Raise
Interviewer: "We're a registered broker-dealer now: 25 million accounts, 6 million online at the open, and options trading. We route to several market centers ourselves. Regulators can ask us about any order from years ago, and we must be able to defend where we sent it. If our main region fails at 10 a.m., no accepted order may vanish and customers must not be left guessing about their positions. Account takeovers drained a few hundred accounts last quarter. And last spring an app release broke the order ticket on a trading morning."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How big is the day now? | 8M orders a day, about a quarter of them options; 20K a second at the open. | The order path goes into account cells (R3.5); options need new risk checks (step 3.6). |
| Which records, for how long? | Every order event, routing decision and customer communication, retrievable for years and provably unaltered. | A write-once audit trail fed from the order log (step 3.1). |
| Do we still send everything to one broker? | No: several exchanges and wholesalers, chosen per order. | A routing engine that records why (step 3.2). |
| What exactly must survive a region loss? | Every order we told a customer we'd received. Order entry back within 15 minutes. | RPO 0 for accepted orders, a warm second region, reconciliation (step 3.3). |
| Where may data live? | In the US. | us-east-1 and us-east-2, with a witness in us-west-2 (step 3.3). |
| How do takeovers happen? | Phished passwords, a new phone enrolled, positions sold, cash sent to a new bank account. | Device binding, risk-based step-up, cooling-off on money leaving (step 3.4). |
| How do we ship? | Weekly app releases; backend deploys any time. | Release freezes, staged rollouts, kill switches, a fallback ticket, a trading-version policy (step 3.5). |
About "five nines during market hours". 99.999% of 1,638 market hours a year is seconds of downtime a year (the "5.26 minutes a year" often quoted is five nines of a 24/7 year). A single regional failover takes minutes, so no design meets five nines through a region loss. We say so and set two targets: 99.99% in market hours for normal operation (about 9.8 minutes a year), and a region loss handled as a disaster with RPO 0 for accepted orders and order entry back within 15 minutes.
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Accounts / online at the open | 10M / 3M | 25M / 6M |
| Subscriptions | 45M | 90M |
| Orders | 2M/day; 5K/s | 8M/day incl. options; 20K/s |
| Execution | One partner broker | Several venues, routed by us |
| Records | Database rows and logs | Write-once audit trail, kept 6 years |
| Footprint | 1 region | Orders: primary + warm DR; market data: active in both |
| Account security | Hardware-signed orders | + new-device limits, risk-based step-up, cooling-off on withdrawals |
| Releases | Any time | Freezes in market hours; kill switches; a fallback ticket |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| Application logs as the history | Logs are incomplete, change format between versions, and can be edited or deleted. They're not records. |
| One broker route | We now owe customers best execution across venues, and we must show why each order went where it went. |
| One region for everything | A regional outage at 10 a.m. stops trading for everyone, and orders accepted in its last second exist only there. |
| A registered device key is enough | An attacker with a phished password and SMS code enrolls their own phone and signs perfectly valid orders. |
| Deploy whenever | A gateway deploy at 9:31 drops 20K connections per task; a store release can't be recalled from phones that have it. |
| Stock-only risk checks | A multi-leg options order can lose far more than its premium; a naked short call's loss has no cap. |
| One Aurora cluster | 20K orders a second plus fills, on one writer, leaves no headroom. |
R3.3 New Requirements and API Additions
1. The audit record. Every order event becomes one of these, written once:
json{ "record_id": "ae_01J9ZQ4M2T8V", "account_id": "acc_5521", "account_event_seq": 88213, "epoch": 3, "order_id": "ord_9QH1", "client_order_id": "0192f3c4-5b1e-7a42-9c3d-2f8e1a6b7c90", "type": "ORDER_RECEIVED", "occurred_at": "2026-09-28T13:30:01.412318Z", "clock_source": "amazon-time-sync", "actor": { "user_id": "u_318", "device_id": "dev_91c0", "key_id": "dk_7f3a91c2", "app_version": "8.4.1", "ip": "203.0.113.24" }, "evidence": { "payload_sha256": "9f2c…", "signature": "MEUCIQDs2x9…", "quote_ref": { "symbol": "NVDA", "epoch": 7, "cseq": 481203 }, "shown_price": "131.45", "disclosures_version": "opt-2026-07" }, "body": { "symbol": "NVDA", "side": "BUY", "type": "MARKET", "quantity": 20, "time_in_force": "DAY" } }
Other types: RISK_PASSED or RISK_REJECTED (with the rule versions used), ROUTE_DECIDED, SENT, ACKED, FILL, CANCEL_REQUESTED, CANCELLED, REJECTED, EXPIRED, RECONCILED.
2. The routing decision record.
json{ "type": "ROUTE_DECIDED", "order_id": "ord_9QH1", "decided_at": "2026-09-28T13:30:01.431907Z", "policy_version": "route-2026-09-15", "nbbo": { "bid": "131.42", "ask": "131.45", "source": "consolidated", "as_of": "2026-09-28T13:30:01.430Z" }, "candidates": [ { "venue": "WHOLESALER_A", "score": 0.91, "expected_price_improvement_bps": 1.8, "fill_rate": 0.998, "status": "UP" }, { "venue": "EXCHANGE_X", "score": 0.84, "expected_price_improvement_bps": 0.0, "fill_rate": 0.97, "fee_per_share": "0.0030", "status": "UP" }, { "venue": "EXCHANGE_Y", "score": null, "status": "CIRCUIT_OPEN" } ], "chosen": "WHOLESALER_A", "reason": "highest score for NVDA, market, 1-100 shares, this quarter's stats" }
3. Order status "unknown", then reconciled. After a region failover, an order the new region can't yet confirm is shown honestly:
json{ "order_id": "ord_9QH1", "status": "UNKNOWN", "reason": "REGION_FAILOVER", "message": "We're confirming this order with the market. Its funds stay reserved until we do." }
It later moves to a real state, with a RECONCILED audit record saying how we learned it.
4. Device and session risk signals. Every sensitive call carries an integrity token (a Play Integrity verdict or an App Attest assertion), the device ID and the app version; the server adds network, location and behaviour signals and returns a decision:
json{ "decision": "STEP_UP", "methods": ["passkey", "trusted_device_approval"], "reason": "NEW_DEVICE_24H" }
R3.4 Design Evolution: Records, Routes, Regions and Releases
Step 3.1: Every Order Must Be Provable Years Later
The problem: a regulator asks about one customer's options order from 2027: who placed it, from which device, what price they saw, which risk rules it passed, where we sent it and why, and every fill. Our answer today would be stitched together from application logs, some of which were rotated away. What would you do?
Primitive: Event Sourcing & CQRS
Step 3.2: "Best Execution"
The problem: we route every order to one wholesaler because it's simple and it pays us rebates. A regulator asks how we know customers got the best reasonably available terms, and a customer complains that her limit order sat unfilled while the stock traded through her price on an exchange. What would you do?
Primitive: Circuit Breaker, Bulkhead & Fault Tolerance
Step 3.3: A Region Fails at 10 a.m.
The problem: us-east-1 becomes unreachable at 10:00. Customers have orders we acknowledged with 201 a moment earlier; some were already at the market, some weren't; fills are arriving at a region we can't reach. The first proposal: "Run order entry active-active in both regions with DynamoDB global tables."
What would you do?
Recovery objectives
| What | RPO | RTO | How |
|---|---|---|---|
Accepted orders (we sent a 201) | 0 | — | MRSC journal written before the 201 |
| Fills | 0 | — | Drop copies to both regions |
| Ledger and order details | ~1 s, then rebuilt | — | Aurora Global Database + journal + drop copies |
| Quotes | — | ~1–2 min | Already running in us-east-2; clients reconnect with jitter |
| Order entry | — | ≤ 15 min target | The gated sequence below |
Failover timing (every delay counted; the parallel steps take the max)
| Step | Time | Elapsed |
|---|---|---|
| Health checks from outside the region fail (10 s interval × 3) and page | 30 s | 0:30 |
| On-call confirms and approves (the gate) | ≤ 3 min | 3:30 |
| In parallel: fence the old FIX sessions (~2 min with the brokers); promote Aurora secondaries (~1–2 min); scale the order API 30 → 150 tasks (~2 min) | max ≈ 2 min | 5:30 |
| Bump the epoch; log on the us-east-2 FIX sessions | 30 s | 6:00 |
| Flip the routing control; DNS TTL 60 s; phones reconnect with jitter | 1 min | 7:00: order entry open |
| Reconcile orders without a terminal state, in the background | ~3 min | ~10:00: all UNKNOWNs resolved |
We plan capacity and residency before the day: the us-east-2 quotas (Fargate vCPUs, Aurora instances, NLB) already match production, and every Region in the design is in the US, so failover never moves data somewhere it isn't allowed to be.
Primitive: Cloud Disaster Recovery & Multi-Region Active-Active · Drill: Multi-region replication consistency (its first question is the async-replication wrong answer above; its second is why synchronous replication stays narrow)
Step 3.4: Account Takeover Drains Accounts
The problem: an attacker phishes a customer's password and SMS code, signs in on their own phone, enrolls a device key (step 2.5 works perfectly for them), sells the customer's positions, links a new bank account and withdraws the cash. Separately, a credential-stuffing botnet tries leaked passwords from 400,000 residential IP addresses. What would you do?
Primitive: Bot Defense, Sybil Resistance & Registration Abuse · Drill: Credential stuffing surge (per-IP limits and SMS-on-every-login, both answered above) · Drill: JWT revocation token blacklist (point 5 answers both)
Step 3.5: A Release Broke Order Entry on a Trading Day
The problem: a weekly app release with a redesigned order ticket crashes on submit for some Android phones. It shipped Tuesday night; by Wednesday's open, 30% of Android users have it. Separately, a backend deploy at 9:31 restarted gateways and dropped 2 million connections in the busiest minute. What would you do?
Step 3.6: Options and Margin Need Complex Risk Checks
The problem: options are live. One customer sold calls without owning the stock (a loss with no ceiling); another placed a four-leg spread that the stock-only check priced as its premium, not its real risk. The app's options screen already shows "Max loss", but that's computed on the phone. What would you do?
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Provable records | Audit records from order_events → Kinesis → Firehose → S3 Object Lock (compliance mode); dense-sequence completeness checks; hash chain; repair from the source; bisect on error | A pipeline to prove complete daily |
| 3.2 | Best execution | Routing engine on measured venue quality; versioned policy; routing records; quarterly review; venue circuit breakers | Routing complexity; a review process |
| 3.3 | Region loss at 10 a.m. | Market data active in both regions; one order primary with an epoch; MRSC journal before 201; Aurora Global Database; drop copies to both regions; gated failover; reconcile, never resend late | A second region; ~20 ms per order |
| 3.4 | Account takeover | Probation for new devices; risk-based step-up; cooling-off on withdrawals; per-account and per-credential limits; revocation with an in-memory denylist | Friction at risky moments |
| 3.5 | Releases break trading | Backend freeze in market hours; staged store releases; flags with a kill layer on the stream; fallback ticket; min_trading_version that still allows cancels | Slower delivery; a fallback to maintain |
| 3.6 | Options and margin risk | Server-side risk service; approval levels; strategy-based buying power; rules as versioned data; ≤ 10 ms | Compute in the hot path |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Market data runs in both regions all the time. Orders run in one: the journal is the only synchronous cross-region write, and the venues' drop copies keep the standby informed of every fill. The routing control decides where orders go, and it moves only after the old region is fenced.
Scaling the order path to 20K a second: account cells. An order touches one account, so we split accounts over 4 cells, each a complete Round 2 order path (an Aurora cluster and its share of order API, risk and router tasks) handling up to 5K orders a second. A small cell directory (account → cell, a DynamoDB global table cached in every task) routes each request. A bad deploy or a hot account affects one cell. Sharding a ledger is covered in depth in the digital wallet loop, step 2.1.
Order latency budget, Round 3 (P99, dependent steps add): Round 2's 52 ms, plus the MRSC journal write (~20 ms: a round trip to us-east-2 of roughly 10–15 ms, plus the write), plus options risk (up to 10 ms instead of 3), plus the routing decision (2 ms): about 81 ms, still well under 150 ms.
Trace 1: a region failure during an order.
Synthesizing vector architecture diagram...
The resolve reads the journal, not us-east-2's lagging Aurora copy, so it can't wrongly void a key the lost region accepted. The journal is why the order isn't lost; the fence is why "unknown to the venue" is final; and the rule "never send it late" is why the customer isn't surprised by a trade at 10:07.
Trace 2: a routing decision, audited later.
Synthesizing vector architecture diagram...
Trace 3: a takeover attempt, blocked.
Synthesizing vector architecture diagram...
The attacker had everything the old design asked for. What they didn't have was a device the account already trusted.
R3.6 Numbers and Cost
Traffic (25M accounts; 6M online at the open; 15 symbols each; the same record and frame sizes as Round 2)
| Quantity | Math | Value |
|---|---|---|
| Subscriptions | 6M × 15 | 90M |
| Deliveries | 90M × 2/s | 180M/s |
| Payload out | 180M × 40 B = 7.2 GB/s | 57.6 Gbps |
| On the wire, peak | 6M × 2 × 680 B = 8.16 GB/s | ≈ 65.3 Gbps |
| Orders | 8M/day; 20K/s at the open; × 1.2 KB | ingress 24 MB/s ≈ 192 Mbps |
Fan-out capacity in both regions. Each region must carry all 6M connections if the other is lost: gateways per region, 100 per AZ. Normally each carries about 3M at 50% use, which also covers losing an AZ inside a region (). 600 gateways in market hours.
Order path (per task 200 orders/s, lower than Round 2's 250 because of options risk; AZ-loss sizing): , so 50 per AZ, 150 tasks in us-east-1; 30 warm in us-east-2, scaled to 150 during a failover.
Audit records (assumptions: ~10 events per order, ~1 KB each, compressing about 4×)
| Quantity | Math | Value |
|---|---|---|
| Records per trading day | 8M × 10 | 80M |
| Peak record rate | 20K orders/s × ~4 events at entry | ~80K/s |
| Raw bytes per day | 80M × 1 KB | 80 GB |
| Compressed per day | 80 GB ÷ 4 | 20 GB |
| Per year | 20 GB × 252 trading days | ≈ 5 TB |
| Kept at steady state (6 years) | 5 TB × 6 | ≈ 30 TB |
Firehose bills each record in 5 KB increments, so we pack about four 1 KB events per record: 20M records a day billed at 5 KB is 100 GB a day, instead of 400 GB if each event were its own record.
Journal (MRSC): 2 writes per order (received, terminal) of ~2 KB (2 write units) in 2 replica Regions: write units a trading day; stored 30 calendar days, about 21–22 trading days: GB per replica.
Monthly cost (us-east-1 and us-east-2 list prices; 21 trading days; stream traffic averaged at a quarter of the peak over the session; rounded)
| Item | Math | Monthly |
|---|---|---|
| Data transfer out, streams | 8.16 GB/s × 25% × 23,400 s × 21 ≈ 1,002 TB: 10 TB × $90 + 40 × $85 + 100 × $70 + 852 × $50 per TB (tiers applied to the total; check how your account aggregates them across Regions) | ≈ $53.9K |
| NLBs | ≈ 1,002,000 GB processed ≈ NLCU-hours × $0.006 | ≈ $6.0K |
| Cross-AZ, conflated stream | 600 gateways × 800 KB/s, two thirds cross-AZ, averaged at half: ≈ 78.6 TB × $0.02/GB | ≈ $1.6K |
| Gateways, Fargate | 600 tasks × 210 h + 30 × 520 h, × $0.079 | ≈ $11.2K |
| Order API | (150 + 30) × $0.079 × 730 | ≈ $10.4K |
| Market data tiers, both regions | 52 tasks (conflators, feed handlers, on-demand options quotes) × $0.079 × 730 | ≈ $3.0K |
| Risk, routing, routers, alerts, push, auth | 90 tasks × $0.079 × 730 | ≈ $5.2K |
| Aurora, 4 cells | (writer + reader + us-east-2 secondary) × 4 = 12 × db.r7g.4xlarge I/O-Optimized ≈ $2.87/h × 730, + storage and global replication ≈ $2K | ≈ $27.1K |
| Order journal, MRSC | 64M write units × 21 ≈ 1.34B × $0.625 per million (assuming replicated writes bill like on-demand writes), + ~0.7 TB stored (2 × 340 GB) × $0.25/GB, + deletes | ≈ $1.2K |
| Valkey, both regions | 12 × cache.r7g.large ≈ $0.175 × 730 | ≈ $1.5K |
| DynamoDB, other (sessions, keys, alerts, cell directory as global tables) | ≈ $1.5K | |
| Audit pipeline | Kinesis ~100 shards × $0.015 × 730 ≈ $1.1K; Firehose 2.1 TB × $0.029/GB ≈ $60; S3: year one ~5 TB Standard ≈ $120, older years in Glacier Deep Archive (Object Lock retention stays with the version) ≈ $25 | ≈ $1.3K |
| Route 53 ARC cluster | $2.50/h × 730 | ≈ $1.8K |
| Direct Connect, both regions | ports and cross-connects to brokers and venues | ≈ $1.2K |
| Cross-region traffic | Aurora replication, drop copies, events to us-east-2's Valkey | ≈ $0.5K |
| CloudFront, snapshots and charts | about twice Round 2 | ≈ $10.6K |
| CloudWatch, WAF, security services, logs | ≈ $15K | |
| Total AWS | ≈ $153K/month |
What DR costs: the us-east-2 Aurora secondaries (4 × $2.87 × 730 ≈ $8.4K), the warm order API and routers (≈ $2.3K), the gateways beyond what one region already needs (a single region sized to lose an AZ would run 450, so the extra is ~150: 150 × $0.079 × 210 ≈ $2.5K), us-east-2's own feed handlers and conflators (26 tasks × $0.079 × 730 ≈ $1.5K), ARC ($1.8K), and a second set of connections and a journal replica (≈ $1.2K): about $17.7K a month, roughly 12% of the bill (a rough split: some of it also serves nearby users every day), for RPO 0 on accepted orders and 15 minutes of RTO. Retention storage for six years of records is tiny next to it. Egress is still the largest line, and market-data licences for 6M real-time users are still outside this table.
R3.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| DR: sync vs async | Synchronous only for the small journal (~20 ms per order); async for everything else, rebuilt from the journal and drop copies | A slower order path, and a rebuild procedure to rehearse. Fully synchronous would slow every write; fully async would lose the orders of the last second. |
| Active-active orders vs one primary | One primary, gated failover | Minutes of no order entry on a region loss, in exchange for no double orders and no double spending across regions. |
| Resend vs cancel orders the market never saw | Cancel and tell the customer | Some customers must re-enter an order after an outage. A late market order is a trade they didn't choose. |
| Routing complexity vs execution quality | A scored, versioned engine with reviews | Harder to build and explain than one route; defensible to regulators and better for customers. |
| Friction vs security | Friction only on risky moments: new devices, new banks, big withdrawals, liquidation | Honest users with a new phone wait or approve from their old one. |
| Release speed vs trading safety | Freezes in market hours; kill switches; a fallback ticket | Fixes wait for the close unless they're emergencies; old code paths stay alive for years. |
| Keep everything six years vs per-record periods | Six years for all | A little more storage, for one rule nobody can misapply. |
Closing the loop. The opening question was: how do we show prices that are fresh enough to trust, and make every order happen exactly once, only if the user meant it? The answer is now:
- Fresh enough to trust: the market's firehose is conflated to what a phone can use, sequenced on our own stream, drawn once per frame, and marked stale by two clocks the phone can trust, with the server checking again (Rounds 1 and 2).
- Exactly once: a key born on the phone, claimed atomically and never forgotten; a resolve call for doubts; a journal that survives a region; the market's own reports as the final word (Rounds 1 and 3).
- Only if the user meant it: a signature from a key only their current biometric unlocks, a fresh price, a device the account trusts, and never an order sent later than they meant it (Rounds 2 and 3).
- And provably so: every step recorded once, kept write-once, and reviewable years later (Round 3).
R3.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| A region outage mid-order | us-east-1 unreachable | The gated failover of step 3.3: fence, promote, bump the epoch, move DNS, reconcile. Quotes recover in 1–2 minutes, order entry in about 7, every unknown order in about 10. |
| The market-data vendor fails | A feed goes silent or wrong | Feed handlers take two sources (the vendor's two sites or a second vendor) and compare them; if the primary is silent for 1 second or disagrees persistently, they switch. Conflators are the sequencer, so cseq continues; clients see no gap. |
| A routing venue is down or slow | Errors or slow acks from one venue | Its circuit opens and the routing engine drops it from the candidates. Each venue's session has its own threads and queue (a bulkhead), so a slow venue can't exhaust the router's workers and block orders to the others. Orders already sent there are not retried elsewhere on a timeout: a timeout doesn't mean the order didn't execute, so we ask the venue for their status when it's back. For the same reason we don't set aggressive 50 ms timeouts that would treat normal jitter as failure. |
| A mass stale-quote event | Every price dims at once | Market orders are blocked everywhere, which is correct. The status page and an in-app banner (a server-sent message) say "Market data delayed"; we check the vendor, then the consolidated feed, and whether the market itself halted (a market-wide circuit breaker shows as HALTED, not STALE). |
| An MRSC Region is impaired | Journal writes slow or fail | With replicas in us-east-1 and us-east-2 and a witness in us-west-2, two of three keep writing. If journal writes fail, orders fail with 503: we never answer 201 without the journal. |
| A poison record in the audit consumers | One consumer stuck | Bisect on error isolates it; the rest flow; the record is repaired from order_events. |
| A takeover wave | Many new-device logins from one network | Risk scores rise, step-ups go up, bot defences engage; per-credential limits stop stuffing. |
| A bad release | Crash-free installs drop for one version | Kill the feature (about 37 s for phones in use); the fallback ticket keeps trading possible; halt the rollout. |
Drill: Circuit breaker cascading thread stall (the venue row answers both of its questions: thread exhaustion from a slow dependency, and why aggressive timeouts alone are wrong)
R3.9 Runbook and Incident Response
Golden signals OPS 8 · REL 6 (user-facing rates count distinct installs or accounts)
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Tick-to-screen latency (sampled clients) | p95 > 1.5 s for 3 min | P2 | Conflator emit lag, gateway flush lag, client frame times |
| Installs with stale prices | > 1% of connected installs for 1 min | P1 | One shard, one AZ, one region, or the vendor? |
| Order in → venue, P99 | > 150 ms for 3 min | P1 | Per stage: journal, Aurora, risk, router |
| Reject rate by reason | QUOTE_STALE, RATE_LIMITED, risk rejects: 3× baseline | P2 | A stale shard, a bot, or a rule change |
| Duplicate order attempts caught | Key conflicts and resolve calls: 3× baseline | P3 | Lost responses at the edge, or a client retry bug |
| Fan-out lag | Gateway flush time p99 > 250 ms | P2 | Hot gateways; slow-client share |
| Journal write latency | p99 > 50 ms | P2 | MRSC Region health |
| Aurora global replication lag | > 5 s | P2 | Cross-region link; write load |
| Drop copy gap | Fills on a venue session missing from its drop copy for 1 min | P1 | Venue connectivity; our DR readiness is degraded |
| Crash-free installs, trading screens | < 99.5% for a new version with ≥ 10,000 installs | P1 | Kill the feature; halt the rollout |
Opening-bell procedure OPS 10 · REL 7
Synthesizing vector architecture diagram...
- 8:30: both feed sources healthy in both regions; drop copies flowing to both.
- 8:50: FIX sessions logged on; each venue's circuit closed.
- 9:00: scheduled scaling has run; check desired and running counts (CLI 2) and that Fargate quotas leave headroom.
- 9:20–9:30: watch connections ramp (CLI 1) and snapshot load.
- 9:30–10:00: watch order P99,
429rates, stale installs; the freeze is in force. - 10:30: let scale-in bring gateways down as connections fall.
Vendor feed failover procedure REL 11
- Confirm it's the source: one source silent or diverging while the other is healthy, in both regions.
- Switch the feed handlers to the healthy source (automatic after 1 s of silence; manual for "wrong but not silent" data).
- Watch stale installs fall and conflators keep
cseqcontinuous (no client gap spike). - Tell compliance if displayed data was wrong for any period: it may matter for customer complaints and trade reviews.
Go deeper: CLI playbook
Plain commands an on-call engineer runs, one at a time. Replace the names with real ones.
text# 1. Active WebSocket connections through a region's NLB aws cloudwatch get-metric-statistics --namespace AWS/NetworkELB --metric-name ActiveFlowCount --dimensions Name=LoadBalancer,Value=net/stream-nlb/0123456789abcdef --statistics Maximum --period 60 --start-time 2026-09-28T13:00:00Z --end-time 2026-09-28T14:00:00Z # 2. Is the gateway service at open size? aws ecs describe-services --cluster market-data --services stream-gateway --query "services[0].[desiredCount,runningCount]" # 3. The scheduled pre-scale, in Eastern time aws application-autoscaling put-scheduled-action --service-namespace ecs --scalable-dimension ecs:service:DesiredCount --resource-id service/market-data/stream-gateway --scheduled-action-name open-prescale --schedule "cron(0 9 ? * MON-FRI *)" --timezone "America/New_York" --scalable-target-action MinCapacity=300,MaxCapacity=360 # 4. Aurora global replication lag for one cell's secondary aws cloudwatch get-metric-statistics --namespace AWS/RDS --metric-name AuroraGlobalDBReplicationLag --dimensions Name=DBClusterIdentifier,Value=orders-cell-1-use2 --statistics Maximum --period 60 --start-time 2026-09-28T13:00:00Z --end-time 2026-09-28T14:00:00Z --region us-east-2 # 5. Failover: promote one cell's secondary (unplanned; accepts replication lag as loss) aws rds failover-global-cluster --global-cluster-identifier orders-cell-1 --target-db-cluster-identifier arn:aws:rds:us-east-2:111122223333:cluster:orders-cell-1-use2 --allow-data-loss # 6. Failover: send order traffic to us-east-2 (run against one of the cluster's endpoints) aws route53-recovery-cluster update-routing-control-state --routing-control-arn arn:aws:route53-recovery-control::111122223333:controlpanel/abc/routingcontrol/use2-orders --routing-control-state On --region us-west-2 --endpoint-url https://host-example.us-west-2.example.aws
The failover commands run only in the gated order of step 3.3: fence first, then promote (5), then the epoch bump and FIX logons, then traffic (6).
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | Market data active in two regions; RPO 0 for accepted orders through an MRSC journal; gated failover with an outside-the-region detector, a fence and an epoch; account cells; venue circuit breakers REL 13 · REL 10 · REL 11 |
| Performance Efficiency | An 81 ms order path with a cross-region journal and options risk; in-memory risk per account shard; fan-out sized per region PERF 1 · PERF 4 |
| Security | New-device probation, risk-based step-up and cooling-off; revocation with an in-memory denylist; audit records written once with Object Lock and a signed hash chain; records classified and retained six years SEC 2 · SEC 4 · SEC 7 · SEC 8 |
| Cost Optimization | ≈ $153K a month, derived; DR is about 12%; egress is still a third; Firehose records packed to avoid 5 KB rounding; old records in Deep Archive COST 8 · COST 5 |
| Operational Excellence | Market-hours freezes; staged releases; kill switches on the live stream (~37 s); a fallback ticket; an opening-bell procedure; rehearsed failover OPS 6 · OPS 10 · OPS 8 |
| Sustainability | Gateways scaled to the trading day; streams only while visible; audit data moved to Deep Archive after a year SUS 2 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Defines availability in market hours honestly, and separates normal operation from disaster with explicit RPO and RTO.
- Knows why asynchronous replication breaks dedup guards, and keeps synchronous replication narrow: a journal of what we promised.
- Fences the old region at the market (the FIX sessions), not only in our own code, and refuses to send stale orders late.
- Builds records from the same events that drive the system, write-once, provably complete, and keeps legal claims general unless verified.
- Treats routing as a measured, versioned, reviewable decision.
- Puts friction where attackers profit, and makes revocation fast without a lookup per request.
- Makes releases boring on trading days, and keeps a fallback path for the one screen that must never break.
- Prices DR and says what it buys.
Follow-up questions
-
"Why not make us-east-2 a full active-active order region for users closer to it?" Answer: the gain is a few milliseconds for some users; the cost is two writers for the same accounts, which needs either cross-region locking on every order or accepting double spending during lag. An account could have a home region, with cells assigned per region, but then each home region needs its own standby, and we've built two of this design. We'd do it for residency, not for latency.
-
"The venue says it never received an order, but its message is still in flight on the old session." Answer: that's why we fence first. Once the venue has ended the old session, nothing more from it can arrive, and "never received" becomes final. Before the fence, the answer is "not yet", and the order stays
UNKNOWN. -
"A regulator asks for everything about one account over three years." Answer: a query over the archived records by
account_id(Athena over the Parquet files for recent years; a restore from Deep Archive for older ones, which takes hours, fine for a request with notice), checked for completeness by the dense sequence, and verified against the signed chain heads.
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: instruments, order types, freshness, who executes, real-time vs delayed | Restate the Round 1 design in 60 seconds | Restate the Round 2 design in 60 seconds |
| 5–15 min | Requirements; ack vs fill; API with the key in the body; the stream | Scope raise → what breaks; redo the tick math | Scope raise → what breaks; define "five nines in market hours" |
| 15–40 min | Steps 1.0–1.5: stream, idempotent orders and the resolve call, state machine, reservation, ledger | Steps 2.1–2.6: conflation, render loop, cseq and epochs, staleness, signing, the open | Steps 3.1–3.6: audit trail, routing, DR, takeover, releases, options risk |
| 40–50 min | Numbers: deliveries, orders, cost; licences | Numbers: 90M deliveries/s, 32.6 Gbps, fleets, the latency budget, cost | Numbers: two-region fan-out, records, DR cost |
| 50–60 min | Failures + pillar check | Failures + pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening a round: "Before I design: how fresh must prices be, who executes the orders, and what exactly does the user see between tapping Buy and owning the shares?"
- When the scope is raised: "Here's what breaks, and I'll fix it in this order: anything that can create an order the user didn't mean or lose one we accepted, then anything that shows a price that isn't true, then smoothness, then cost."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What if the order is sent twice?" (REL 4) | A key born on the phone is claimed in the order's own transaction and never expires; a resolve call settles doubts for good. | 1 | Step 1.2 |
| "What happens at 9:30?" (REL 7) | We pre-scale on a schedule, serve snapshots from memory and the edge, and admit orders by priority, cancels first. | 2 | Step 2.6 | |
| "What if a region fails?" (REL 13) | Market data already runs in both; accepted orders are in a synchronous journal; we fence, promote and reconcile, and never send a stale order late. | 3 | Step 3.3 | |
| Performance | "How do you stream to millions?" (PERF 4) | Conflate per symbol, fan out the conflated stream, conflate again per client: 1M ticks in, 90M small deliveries out. | 2 | Step 2.1 |
| "Why doesn't the UI stutter?" (PERF 1) | The network writes a latest-value table off the main thread; the screen reads it once per frame. | 2 | Step 2.2 | |
| Security | "What stops a stolen session?" (SEC 2) | Each order is signed by a hardware key only the user's current biometric unlocks, and new devices start limited. | 2–3 | Steps 2.5, 3.4 |
| "Can a user touch someone else's account?" (SEC 3) | The user comes from the token, and every order, cancel, key and bank link is checked against it. | 1–3 | Steps 1.4, 2.5, 3.4 | |
| "How do you prove what happened?" (SEC 4) | Audit records from the order log, stored write-once and hash-chained, checked complete daily. | 3 | Step 3.1 | |
| Cost | "Where does the money go?" (COST 8) | About $2.4K, $59K and $153K a month; streaming egress is the biggest line, so bytes per record matter, and licences sit outside AWS. | 1–3 | R1.7, R2.6, R3.6 |
| "Why not a managed WebSocket service?" (COST 5) | We priced it: per-message fees would be about $0.6M a month for our firehose. | 2 | R2.6 | |
| Operations | "How do you ship safely?" (OPS 6) | No deploys in market hours, staged store releases, kill switches that reach phones in about 37 seconds, and a fallback ticket. | 3 | Step 3.5 |
| "How do you know users see fresh prices?" (OPS 8) | Tick-to-screen latency and stale installs, counted by distinct installs. | 2–3 | R2.10, R3.9 | |
| Sustainability | "How do you save battery and data?" (SUS 2) | Streams only while visible, conflated binary records, and fleets scaled to the trading day. | 1–2 | Steps 1.1, 2.1 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Market data | Streams with snapshots and heartbeats; one update per second | Conflation per symbol and per client; (epoch, cseq); in-band snapshots; the corrected tick math | Active in two regions, sized for failover; dual sources |
| Orders | Keys born on the phone; atomic claim; resolve call; never queued offline | Signed exact bytes; quote_ref checks; admission by priority | A cross-region journal; fencing at the market; reconcile, never resend late |
| Money and risk | Reservation under the account lock; ledger; T+1 | Per-account limits; server-side staleness checks | Strategy-based options risk; rules as versioned data |
| Device | Clear states; no optimistic fills | Frame-synced rendering; two-clock staleness; hardware keys | Probation for new devices; a fallback ticket; trading-version policy |
| Numbers | Deliveries, orders, cost; licences | 90M deliveries/s, 32.6 Gbps, fleets, a latency budget, the managed-service trap | Two-region capacity, records, DR cost, honest availability |
| Well-Architected trade-offs | Streaming vs polling | Freshness vs bytes; friction vs security | DR sync vs latency; release speed vs safety |
| Evolving under new scope | Builds from the baseline, one problem at a time | Opens with what could create an unintended order | Designs for the day the region, the vendor or the release fails |