Design a Digital Wallet System
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | Students top up and pay each other or the canteen | A national peer-to-peer wallet with merchants, payroll and bank withdrawals | A cross-border wallet in 30 countries, with regulators watching |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Volume | ~1M wallets; 200K transfers/day; ~30/s in the busiest minute | 50M wallets; 20M transfers/day; 5,000/s peak, 20,000/s ceiling; 50K balance reads/s | ~300M wallets; 100M transfers/day across 3 home regions |
| Data | One database, ~55 GB a year | 16 ledger shards; ~5.48 TB of ledger a year | 68 shards in 3 regions; 7 years of tamper-evident archives |
| Footprint | 1 region, 3 AZs | 1 region, 3 AZs; survives a database failover mid-transfer | 3 home regions, each with a DR region |
| Targets | No negative balances, no lost updates; 99.9% | Transfer P99 < 20 ms same shard, < 60 ms cross-shard; 99.99% | Regulator-grade audit; a region can fail without double-spend |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is a Digital Wallet?
You Already Know One: a Prepaid Card With a Notebook
Think of a prepaid canteen card and a small notebook next to the till. Every time money goes onto the card or comes off it, the cashier writes a line: the date, what happened, and how much. The number on the card is your balance, but the truth is the notebook. If the card says 12.40 and nobody can explain how it got there, the card is wrong, not the notebook.
A digital wallet works the same way:
| Thing | In the notebook picture | In our system |
|---|---|---|
| Balance | The number on the card | A cached running total of the account's entries |
| A movement of money | One line in the notebook | A ledger entry: which account, how much, which transfer it belongs to |
| A transfer from Alice to Bob | Two lines: minus on Alice's page, plus on Bob's | Two entries that sum to zero |
| "Where did my money go?" | Read the notebook | Read the account's entries, newest first |
What Makes It Hard
Three things turn a notebook into a hard design problem:
- Concurrency. Thousands of transfers touch the same balances at once. Two of them must never both spend the same 40.00.
- Size. The notebook outgrows one database, and then Alice's page and Bob's page live on different machines. A transfer between them can no longer be one database transaction.
- Accounting. Every cent must always be somewhere: in a wallet, in flight between machines, on its way to a bank, or held for review. "It was lost during a retry" is never an acceptable answer.
Throughout the loop we protect four invariants. An invariant is a rule that must be true at every moment, not just at the end of the day:
- No user balance goes negative.
- Debits equal credits: every transfer's entries sum to zero, so money is never created or destroyed.
- No lost or double updates: each transfer changes each balance exactly once.
- Money in flight is always accounted for: anything that has left one balance and not yet reached another sits in a named account we can add up.
The Question the Whole Loop Answers
How do we move money between balances concurrently, at scale, so that every balance is right and every cent is accounted for?
The answer gets sharper every round:
- Round 1: one database, a double-entry ledger, row locks in a fixed order, and idempotency keys.
- Round 2: the ledger split across many databases, money "in flight" between them, hot merchant accounts, and fast cached balances that never go backwards.
- Round 3: many currencies, compliance inside the transfer path, proof that the ledger matches the real bank accounts, and regions that can fail without anyone spending the same money twice.
This loop is about our own ledger under concurrency. Talking to card processors, banks and their timeouts is the payment system's problem; we link to the payment processing interview loop where we need it instead of repeating it.
Round 1 · Mid-level · "A Wallet for One Campus App"
~35 min · SDE II (L5) · 1 region, 3 AZs · ~30 transfers/s in the busiest minute · 99.9%
R1.1 Establish Design Scope
The interviewer says: "Design the wallet for a campus app. Students top up with a card and pay each other or the canteen." Before we draw anything, we ask questions, and we say out loud what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How does money get in? | By card, through the payment processor (PSP) we already use. | Card top-ups are asynchronous: we only credit the wallet once the PSP confirms (step 1.5). The PSP details belong to the payment system. |
| Is a transfer between students instant? | Yes. Bob should see the money as soon as Alice taps "send". | A transfer is one database transaction that debits and credits together. |
| Can a balance go below zero? | Never. No overdrafts, no credit. | "No negative balance" is an invariant we must enforce under concurrency (step 1.2). |
| Can students withdraw to a bank account? | Not yet. | No outbound bank rails this round. |
| How many users? | About 1M wallets across our partner campuses. About 100K are active on a normal day, and each makes about 2 transfers. | About 200K transfers a day, peaking at lunch. Small for one database (R1.7). |
| Do users need history? | Yes, every movement, and support must be able to explain any balance. | The balance must be derived from a record of movements, not stored on its own (step 1.1). |
| One currency? | Yes, US dollars. | We store every amount as an integer number of cents (the currency's "minor unit"). Never floating point: 0.1 + 0.2 is not exactly 0.3 in binary floating point. |
Out of scope for this round:
- Withdrawals to bank accounts.
- Merchants beyond the canteen, and anything like payroll.
- More than one currency.
The interviewer will widen this scope later. Write your out-of-scope list where you can see it: in a multi-round loop, all three of these come back.
R1.2 Functional Requirements, Derived Step by Step
We read the problem one phrase at a time and turn each phrase into a requirement:
| Phrase from the problem | Requirement |
|---|---|
| "Students top up with a card" | top_up(wallet, amount, card) creates a pending top-up; the wallet is credited when the PSP confirms |
| "Pay each other or the canteen" | transfer(from, to, amount) moves money between two wallets in one step; the canteen is just another wallet |
| "See their balance" | balance(wallet) returns the current balance |
| "Support can explain any balance" | entries(wallet) lists every movement, newest first, with the balance after each one |
Not yet: withdrawals, merchants with special needs, and currencies. The interviewer may bring these back.
R1.3 Non-Functional Requirements: the Questions
Numbers come in R1.7. For now, the questions, in the order a wallet cares about them:
- Correctness first. No negative balances, no lost updates, and debits equal credits. A slow wallet is annoying; a wrong one is a financial incident. When correctness and speed conflict, correctness wins.
- Durability. A transfer the app showed as "sent" must survive any single failure, including a database failover.
- Latency. A transfer should feel instant at the canteen till: we aim for P99 under 100 ms on the server side.
- Availability. 99.9%: about 43.8 minutes of downtime a month. During an outage it is fine to refuse transfers; it is never fine to process them wrongly.
R1.4 The API
Three endpoints. All amounts are integers in cents.
Make a transfer
httpPOST /v1/transfers HTTP/1.1 Host: api.campuswallet.example Authorization: Bearer <access token for Alice> Idempotency-Key: 6f1c2a9e-3b7d-4c55-9a0e-2f8d1b7c4e11 Content-Type: application/json { "from_wallet_id": "w_1001", "to_wallet_id": "w_2002", "amount_minor": 2500, "currency": "USD", "memo": "Pizza" }
httpHTTP/1.1 200 OK Content-Type: application/json { "transfer_id": "t_7301", "status": "COMPLETED", "amount_minor": 2500, "currency": "USD", "from_balance_after_minor": 7500, "from_version": 2 }
The Idempotency-Key is a random ID the app creates once per "send" tap and reuses on every retry of that tap. It lets the server recognize a retry and return the first answer instead of moving the money again (step 1.4).
Read a balance
httpGET /v1/wallets/w_1001/balance HTTP/1.1 Authorization: Bearer <access token for Alice>
httpHTTP/1.1 200 OK Content-Type: application/json { "wallet_id": "w_1001", "currency": "USD", "balance_minor": 7500, "version": 2 }
Read the history
httpGET /v1/wallets/w_1001/entries?limit=50 HTTP/1.1 Authorization: Bearer <access token for Alice>
httpHTTP/1.1 200 OK Content-Type: application/json { "entries": [ { "entry_id": "e_103", "transfer_id": "t_7301", "amount_minor": -2500, "balance_after_minor": 7500, "at": "2026-09-27T12:01:07Z" }, { "entry_id": "e_102", "transfer_id": "t_7300", "amount_minor": 10000, "balance_after_minor": 10000, "at": "2026-09-27T11:58:40Z" } ], "next_cursor": "e_102" }
This is the first page, so the request has no cursor. next_cursor is the last entry ID we returned; passing it back asks for "entries older than this one". Unlike page numbers, a cursor doesn't skip or repeat rows when new entries arrive while the user scrolls.
| Status | When |
|---|---|
200 OK | The transfer completed, or this is a retry of one that did (same body as the first time) |
400 Bad Request | Amount not a positive integer, sender equals recipient, unknown currency |
401 / 403 | Not signed in, or the wallet in from_wallet_id isn't yours |
404 Not Found | A wallet doesn't exist |
422 Unprocessable Entity | INSUFFICIENT_FUNDS, ACCOUNT_FROZEN, or IDEMPOTENCY_KEY_REUSED (the same key with a different request body) |
503 Service Unavailable | We can't safely process transfers right now (for example mid-failover); retry with the same key |
Amounts are JSON numbers here. That is safe because JavaScript reads integers exactly up to cents, which is 90 trillion dollars.
Recap
- Top up by card (asynchronous), transfer, balance, history.
- Every amount is an integer in cents.
- Every transfer carries an idempotency key.
- Correctness beats latency: no negative balances, no lost updates, debits equal credits.
- About 200K transfers a day.
Let's build it, starting with the simplest thing that works.
R1.5 Design Evolution: From a Balance Column to a Ledger
Every step below follows the same pattern: a problem, your turn to think, the answer, and what the answer costs us. The cost is always the next problem.
Step 1.0: The Baseline
One table with a balance column per wallet. A transfer runs two statements: subtract from Alice, add to Bob.
Synthesizing vector architecture diagram...
The simplest wallet: one row per user, and a transfer is two updates.
What's good about it: it's easy to read and fast.
What it costs us: nothing records why a balance is what it is, and nothing stops two requests from racing each other. We fix the first problem first, because every later fix builds on it.
Step 1.1: "We Can't Explain a Balance"
The problem: a student writes to support: "My balance says 12.40 but it should be 37.40." Support opens the database and sees one number: 1240. There is no record of how it got there. What would you do?
Alice's first two movements as entries (sign convention: a negative amount takes money out of an account, a positive amount adds it):
| entry_id | transfer_id | account | amount_minor | balance_after | version |
|---|---|---|---|---|---|
| 101 | t_7300 (top-up) | psp_clearing (system) | −10,000 | −10,000 | 1 |
| 102 | t_7300 (top-up) | Alice | +10,000 | 10,000 | 1 |
| 103 | t_7301 (P2P) | Alice | −2,500 | 7,500 | 2 |
| 104 | t_7301 (P2P) | Bob | +2,500 | 2,500 | 1 |
Each transfer's rows sum to zero (−10,000 + 10,000; −2,500 + 2,500), so the whole table sums to zero. System accounts like psp_clearing may be negative in this convention: −10,000 means "the PSP owes us 100.00". Only user accounts carry the "never negative" rule.
Accountants use debits and credits with a "normal balance" per account type instead of signs. The signed-amount form is equivalent and easier to check in SQL: the invariant is simply "the amounts of every transfer sum to zero".
Step 1.2: "Two Transfers Read $50 and Both Spent $40"
The problem: Alice has 50.00. Her phone and her laptop each send 40.00, a few milliseconds apart. Both requests read her balance (5000), both see "enough", and both go through. She has spent 80.00 out of 50.00. What would you do? And which isolation level are you relying on?
What the other isolation levels would do here, since interviewers ask:
| Level (PostgreSQL) | Same race without FOR UPDATE | With FOR UPDATE |
|---|---|---|
| Read Committed (default) | Lost update or negative balance | Second transaction waits, then re-reads the newest row and fails the check |
| Repeatable Read (snapshot isolation) | The second writer to the same row gets a serialization failure and must retry | Same: locking a row changed since our snapshot raises a serialization failure; the app retries the whole transaction |
| Serializable (SSI) | One of the two gets a serialization failure and must retry | Same, plus protection against write skew (below) |
The anomaly row locks alone don't catch: write skew. Suppose the campus adds a rule: "a student can send at most 200.00 per day." The transfer sums today's outgoing entries and checks the limit. Two concurrent 15.00 transfers from Alice each read "180.00 sent today", each pass, and each insert a new entry. Under snapshot isolation (Repeatable Read), neither transaction writes a row the other wrote, so nothing conflicts, and Alice ends at 210.00. That is write skew: two transactions read an overlapping set, check a rule over it, and write different rows. Snapshot isolation only detects two writers of the same row. Our design is safe anyway, because both transfers first lock Alice's account row, so the second one waits and re-reads the sum after the first commits. Locking one row that stands for the whole rule ("materializing the conflict") is the usual fix.
Why not just run everything at Serializable? PostgreSQL's Serializable (serializable snapshot isolation) does catch write skew, but it works by aborting transactions it can't prove safe, with serialization failures the application must retry. Under contention, on exactly the busy accounts we care about, the abort and retry rate climbs, and throughput drops. It also tracks predicate locks, which costs memory and CPU. And it only covers one database: from Round 2 on, a transfer spans two databases and no isolation level helps. So we choose Read Committed plus explicit locks on the rows that carry each rule, and we write down which rule each lock protects.
Primitive: Database Isolation Levels, ACID & Concurrency Anomalies · Drill: ACID isolation write skew anomaly (both of its questions are answered just above)
Step 1.3: "Alice→Bob and Bob→Alice at the Same Time Deadlock"
The problem: Alice sends Bob 10.00 at the same moment Bob sends Alice 5.00. Transaction 1 locks Alice, then asks for Bob. Transaction 2 locks Bob, then asks for Alice. Each waits for the other forever. What would you do?
Synthesizing vector architecture diagram...
Both transfers ask for account 1001 first, so the second one simply waits its turn instead of holding Bob and waiting for Alice.
Step 1.4: "The App Retried and the Transfer Ran Twice"
The problem: Alice taps "send 25.00" on a weak campus Wi-Fi. The server commits the transfer, but the response is lost. The app times out and retries. Bob receives 50.00. What would you do?
The transfer, as pseudocode:
texttransfer(caller, idem_key, from, to, amount): reject 400 unless amount > 0 and from != to BEGIN (Read Committed) insert idempotency_keys(caller, idem_key, hash(request)) on unique violation: ROLLBACK if stored hash != hash(request): return 422 IDEMPOTENCY_KEY_REUSED return stored response for id in sorted([from, to]): SELECT ... FROM accounts WHERE account_id = id FOR UPDATE if either account is not ACTIVE: record REJECTED(ACCOUNT_FROZEN), store response, COMMIT, return 422 if from.balance_minor < amount: record REJECTED(INSUFFICIENT_FUNDS), store response, COMMIT, return 422 UPDATE from: balance_minor -= amount, version += 1, entry_seq += 1 UPDATE to: balance_minor += amount, version += 1, entry_seq += 1 INSERT transfer (COMPLETED) INSERT entries: (from, -amount, new balance, new version, new entry_seq), (to, +amount, ...) store response on the idempotency row COMMIT return 200 with the stored response
Rejections are committed too, with their response. A retry of a rejected transfer must return the same rejection, not try again: Alice may have received money in between, and a retry that suddenly succeeds is a second, different payment she didn't confirm.
Step 1.5: "A Top-Up by Card: When Is the Money in the Wallet?"
The problem: Alice tops up 100.00 by card. The card processor (PSP) takes a few seconds to answer, sometimes times out, and occasionally declines after first saying "processing". What would you do? When do we add the 100.00 to her balance?
Synthesizing vector architecture diagram...
Entries appear only on the arrow into SUCCEEDED, so a declined card never touches the ledger.
One more case the interviewer may raise: a chargeback weeks later, after Alice spent the money. We don't force her balance negative (the constraint would refuse, correctly). We post the chargeback against a system account for "amounts customers owe us", freeze the wallet, and let collections handle it.
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | A balance column, two updates | No history; races |
| 1.1 | Can't explain a balance | Double-entry ledger; balance cached in the same transaction; system accounts | More rows; a rule for every engineer |
| 1.2 | Two spends of the same money | SELECT ... FOR UPDATE at Read Committed; check constraint; locks for each rule (write skew) | Lock waits on busy accounts |
| 1.3 | Opposite transfers deadlock | Lock accounts in sorted ID order, one per statement | A rule in every code path |
| 1.4 | Retries double-spend | Idempotency key inserted in the same transaction; stored response | A key table, 30-day retention |
| 1.5 | Top-up money not yet real | Pending top-up; credit only on the PSP's confirmation | Top-ups take seconds |
The cost that survives into Round 2 is the one we haven't hit yet: everything assumes one database.
R1.6 Architecture v1
Now the concepts get AWS names.
Synthesizing vector architecture diagram...
Follow the solid arrows: every request lands on a stateless task, and every money decision happens in one Aurora writer. The reader exists to take over if the writer fails; the PSP talks to us only through the top-up flow.
The pieces:
- Wallet service: stateless tasks on ECS Fargate behind an Application Load Balancer, one per AZ. Any task can serve any request because all state is in the database.
- Aurora PostgreSQL: one writer and one reader in a second AZ. Aurora keeps six copies of the data across three AZs in its storage layer, whatever instances you run; the reader is there so a failover promotes an existing instance (typically under 60 seconds, often under 30, per AWS) instead of creating a new one (typically under 10 minutes).
- The PSP: called only for top-ups, never inside a transfer transaction.
Schema
sqlCREATE TABLE accounts ( account_id BIGINT PRIMARY KEY, owner_id BIGINT, -- NULL for system accounts kind TEXT NOT NULL CHECK (kind IN ('USER', 'SYSTEM')), currency CHAR(3) NOT NULL, status TEXT NOT NULL DEFAULT 'ACTIVE' CHECK (status IN ('ACTIVE', 'FROZEN', 'CLOSED')), balance_minor BIGINT NOT NULL DEFAULT 0, -- cache of SUM(entries), same transaction version BIGINT NOT NULL DEFAULT 0, -- +1 on every change, including freezes (fences the cache) entry_seq BIGINT NOT NULL DEFAULT 0, -- +1 only when an entry is written (proves completeness) CONSTRAINT user_balance_not_negative CHECK (kind = 'SYSTEM' OR balance_minor >= 0) ); CREATE TABLE transfers ( transfer_id BIGINT PRIMARY KEY, kind TEXT NOT NULL, -- 'P2P', 'TOP_UP', 'CHARGEBACK' from_account BIGINT NOT NULL REFERENCES accounts, to_account BIGINT NOT NULL REFERENCES accounts, amount_minor BIGINT NOT NULL CHECK (amount_minor > 0), currency CHAR(3) NOT NULL, status TEXT NOT NULL, -- 'COMPLETED', 'REJECTED' reject_reason TEXT, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), CHECK (from_account <> to_account) ); CREATE TABLE ledger_entries ( entry_id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY, transfer_id BIGINT NOT NULL REFERENCES transfers, account_id BIGINT NOT NULL REFERENCES accounts, amount_minor BIGINT NOT NULL CHECK (amount_minor <> 0), -- signed currency CHAR(3) NOT NULL, balance_after BIGINT NOT NULL, account_version BIGINT NOT NULL, entry_seq BIGINT NOT NULL, -- 1, 2, 3 ... per account, no gaps created_at TIMESTAMPTZ NOT NULL DEFAULT now(), UNIQUE (account_id, account_version), -- two entries can never claim the same version UNIQUE (account_id, entry_seq) ); CREATE TABLE idempotency_keys ( caller_id BIGINT NOT NULL, idem_key TEXT NOT NULL, request_hash BYTEA NOT NULL, transfer_id BIGINT, response JSONB, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), PRIMARY KEY (caller_id, idem_key) );
The UNIQUE (account_id, account_version) constraint is cheap insurance: if a bug ever let two transactions update the same account without the lock, both would try to write the same version and one would fail. We keep two counters on purpose. version goes up on every change to the account row, including a freeze that writes no entry, because the cache (Round 2) must see status changes too; so entry versions can have gaps. entry_seq goes up only when an entry is written, so an account's entries are numbered 1, 2, 3 with no gaps, and a missing number means a missing entry. The service's database role may INSERT into ledger_entries but has no UPDATE or DELETE grant on it: the ledger is append-only by permission, not by convention. Mistakes are fixed with new, balancing entries.
The history endpoint reads one account's entries newest first, so we add one index:
sqlCREATE INDEX entries_by_account ON ledger_entries (account_id, entry_id DESC);
Trace: Alice pays Bob 25.00
Synthesizing vector architecture diagram...
Everything between BEGIN and COMMIT is one atomic step: either the key, the transfer, both balance changes and both entries exist, or none of them do.
R1.7 Numbers
Traffic
| Quantity | Math | Value |
|---|---|---|
| Transfers per day | 100K active students × 2 | 200,000 |
| Average rate | 200,000 ÷ 86,400 s | ≈ 2.3/s |
| Lunch hour | assume 20% of the day's transfers between 12:00 and 13:00: 40,000 ÷ 3,600 s | ≈ 11/s |
| Busiest minute | assume 2.7× the lunch-hour average (queues at the till) | ≈ 30/s |
| Balance and history reads | assume 10 per transfer | ≈ 300/s in the busiest minute |
Rows and storage per transfer. One transfer writes one transfers row, two entries and one idempotency key, and updates two account rows. We plan on about 750 bytes per transfer including index entries (the same figure Round 2 uses).
Seven years would be about 384 GB. One Aurora cluster holds up to 256 TiB on recent engine versions (128 TiB on older ones), so storage is no concern.
One database's headroom. We plan on about 1,000 short money transactions per second for a db.r7g.large writer (2 vCPUs). That is a planning assumption, not a measured fact: we confirm it with a load test that runs our real transfer transaction. At 30/s we use about 3% of it. Throughput is not this round's problem; correctness under concurrency is.
Availability. 99.9% of a 730-hour month allows minutes. One Aurora failover (under 60 s typically) uses about 2% of that.
Monthly cost (us-east-1 on-demand prices, 730 hours a month)
| Item | Math | Monthly |
|---|---|---|
Aurora writer + reader, db.r7g.large, Aurora Standard | 2 × $0.276/h × 730 h | ≈ $403 |
| Aurora storage | ~55 GB after a year × $0.10/GB-month | ≈ $6 |
| Aurora I/O | assume ~30 I/Os per transfer: 200K × 30 × 30.4 days ≈ 183M × $0.20 per million | ≈ $37 |
| Fargate, 3 tasks × 0.5 vCPU, 1 GB, Arm | 3 × (0.5 × $0.03238 + 1 × $0.00356)/h × 730 h | ≈ $43 |
| Application Load Balancer | $0.0225/h × 730 h + a few capacity units | ≈ $22 |
| CloudWatch metrics, alarms, logs | ≈ $20 | |
| Total | ≈ $530/month |
The PSP's per-transaction fees are the payment system's cost, not counted here. Backups up to the size of the cluster are included in Aurora's price.
R1.8 Trade-Offs
Pessimistic locks vs optimistic concurrency
| Approach | How it works | Good when | Bad when |
|---|---|---|---|
Pessimistic: SELECT ... FOR UPDATE (our choice) | Lock the rows, then decide and write | Conflicts are common, like a canteen account that many students pay at once | A lock is held while the app does slow work (never call anything external inside the transaction) |
| Optimistic: version check | Read the row and its version; write with WHERE version = :read_version; if 0 rows changed, someone else won, so re-read and retry | Conflicts are rare | Conflicts are common: every loser redoes its work, and retries pile up exactly when the account is busiest |
| One conditional update | UPDATE accounts SET balance_minor = balance_minor - 2500, version = version + 1 WHERE account_id = 1001 AND balance_minor >= 2500 | Only one row decides the outcome | We need to read both rows before deciding (both must be active, same currency) |
The conditional update is safe under Read Committed too: if another transaction changes the row first, PostgreSQL waits, then re-checks the WHERE condition against the newest version. It's a good answer for a single-row debit; we lock explicitly because a transfer's decision depends on two rows.
A cached balance column vs summing entries
Cached balance_minor (our choice) | Sum the entries on every read | |
|---|---|---|
| Read cost | One row | Grows with history: the canteen account gains thousands of entries a day |
| Risk | The cache and the entries could disagree if someone updates one without the other | None: the entries are the truth |
| How we keep it honest | Same transaction; balance_after and version on every entry; a nightly check (R1.9) | – |
A middle ground used by many ledgers: store periodic balance snapshots ("balance as of entry 9,000,000") and sum only the entries after the latest one.
R1.9 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| The Aurora writer fails mid-transfer | A burst of connection errors for under a minute (typically) | Uncommitted transactions roll back; nothing half-done survives. The reader is promoted. The app retries with the same idempotency key: if the first attempt committed just before the crash, the retry returns the stored response; if not, the retry runs it for the first time. |
| A slow transaction holds a lock | Other transfers on the same account wait; P99 latency rises | lock_timeout (say 2 s) and statement_timeout make waiters give up with an error instead of queueing forever; the app retries with the same key. The real fix is never doing slow work, such as a network call, inside the transaction. |
| A deadlock anyway | Error 40P01 in the logs | Should be impossible with sorted locking. Any deadlock means a code path broke the order; we alarm on the deadlock count and find it. |
| An entry imbalance | The nightly check finds a transfer whose entries don't sum to zero, or an account whose balance_minor isn't the sum of its entries | A P1 alarm. Freeze the affected accounts, find the code path, and repair with new balancing entries, never by editing rows. |
The nightly check, as two short queries:
sql-- 1. Every transfer's entries sum to zero SELECT transfer_id FROM ledger_entries GROUP BY transfer_id HAVING SUM(amount_minor) <> 0; -- 2. Every cached balance equals the sum of its entries SELECT a.account_id FROM accounts a JOIN (SELECT account_id, SUM(amount_minor) AS s FROM ledger_entries GROUP BY account_id) e USING (account_id) WHERE a.balance_minor <> e.s;
Both queries return zero rows on a healthy night. At this size they take minutes on the reader; in Round 2 they move to the analytics copy of the data.
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | One writer and a reader in another AZ; failover rolls back uncommitted work; every retry is safe because of idempotency keys stored in the same transaction REL 4 · REL 10 · REL 11 |
| Performance Efficiency | A relational database for a relational, transactional problem; the balance cached in the account row so reads are one row; history by an index on (account_id, entry_id) PERF 3 |
| Security | The service's database role can't update or delete ledger entries; users can only send from their own wallet; TLS everywhere, encryption at rest on Aurora SEC 3 · SEC 8 · SEC 9 |
| Cost Optimization | About $530/month; a small writer at ~3% busy, sized for availability, and we say why COST 6 |
| Operational Excellence | Light this round: the nightly balance check, a deadlock-count alarm and lock-wait metrics OPS 8 |
| Sustainability | Skipped this round: the whole system is two small database instances and three small tasks. |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Stores money as integers in minor units and says why.
- Makes the ledger the source of truth, with the balance as a cache updated in the same transaction.
- Finds the double-spend race and fixes it with row locks in a transaction, not with an application-side check.
- Names the isolation level and what it does and doesn't guarantee.
- Prevents deadlocks with a global lock order.
- Makes transfers idempotent with a key stored in the same transaction.
- Credits top-ups only on the processor's confirmation.
Follow-up questions
-
"Why store the rejected transfer? Nothing happened." Answer: so a retry returns the same answer. Without it, Alice's retry could succeed a minute later, after Bob paid her back, and she'd have made a payment she thought had failed. The idempotency key's promise is "same request, same result", including rejections.
-
"Could we use the
transferstable's primary key as the idempotency key instead of a separate table?" Answer: yes, if the client could choose the transfer ID. A separate table keeps the client's key (a random string, scoped by caller) apart from our internal IDs, stores the request hash to catch key reuse, and lets us delete keys after 30 days without touching the ledger. Either works if the key is written in the same transaction. -
"The nightly check found an account whose balance is 500 more than its entries. What do you do?" Answer: freeze the account so the difference can't be spent, then find the cause from the entries'
balance_afterandentry_seqchain: the first entry whosebalance_afterdoesn't follow from the one before shows where it went wrong. Repair with a new transfer between the account and a "corrections" system account, approved by a second person. Never edit or delete the existing rows: the audit trail must show the mistake and the fix.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Floats are fine for money" | Binary floating point can't represent most decimal fractions exactly; sums drift by fractions of a cent. |
| "Check the balance in the app, then update" | Two requests both pass the check; Read Committed doesn't stop it. Lock the row, then check. |
| "Serializable fixes everything" | It aborts under contention, costs throughput, and does nothing once a transfer spans two databases. |
| "Just retry on deadlock" | Each deadlock costs a second. Lock in a global order and deadlocks can't form. |
| "Credit the top-up, reverse it if the card fails" | The money can be spent before the decline. Credit only on confirmation. |
Round 2 · Senior · "50M Wallets, Sharded, With Hot Merchants"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 20M transfers/day: 5,000/s peak, 20,000/s ceiling · 50K balance reads/s · transfer P99 < 20 ms same shard, < 60 ms cross-shard · 99.99%
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built a wallet for a campus app: 1M wallets, 200K transfers a day, about 30 a second at the busiest minute, one region, 99.9%. Money is stored as integer cents. The truth is a double-entry ledger: every transfer writes entries that sum to zero, and system accounts like
psp_clearingstand for money outside the wallet. Each account row caches its balance and a version, updated in the same transaction as the entries, with a check constraint that user balances never go negative. A transfer runs at Read Committed, locks both account rowsFOR UPDATEin sorted ID order so opposite transfers can't deadlock, and starts by inserting an idempotency key in the same transaction, so retries return the stored answer. Top-ups are credited only when the card processor confirms. It all runs in one Aurora PostgreSQL writer with a reader for failover, about $530 a month. The open cost: everything assumes one database."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: every money decision is one local transaction in one database.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Can't explain a balance | Double-entry ledger; balance cached in the same transaction | More rows |
| 1.2 | Double spend | FOR UPDATE at Read Committed; non-negative check constraint | Lock waits |
| 1.3 | Deadlocks | Sorted lock order | Discipline in every code path |
| 1.4 | Retries double-spend | Idempotency key in the same transaction | A key table |
| 1.5 | Top-up not yet real | Credit on the PSP's confirmation | Top-ups take seconds |
Open costs: one database's write capacity; nothing publishes changes to other systems; every balance read hits the database.
R2.1 The Scope Raise
Interviewer: "The campus app became a national wallet. We have 50 million wallets and 10 million active a day, making 20 million transfers a day. Evenings and paydays run about 20 times the average, around 5,000 a second, and during big events like a championship final or a holiday sale we've seen bursts toward 20,000 a second. The home screen shows the balance, so that's about 50,000 balance reads a second at peak."
Interviewer: "Big merchants and payroll wallets now use us: one merchant can receive thousands of payments a second. Users can withdraw to their bank accounts. The analytics and fraud teams want every event. And a database failover in the middle of a transfer must not lose or duplicate a cent. 99.99%."
A scope raise is not the end of scoping. We ask back, and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Is the 20,000/s ceiling something we must absorb at any moment, or are the events known in advance? | Both happen, but the big ones are on the calendar. | We size the number of shards for the ceiling (splitting a ledger later is slow), and the instance size for normal peaks, scaling up ahead of known events (R2.6). |
| What share of transfers pay merchants? | About half. A handful of merchants and payroll wallets take thousands a second. | Hot accounts need their own design (step 2.3), and it can also make most merchant payments local to one shard. |
| Must the recipient see the money instantly? | The sender must see "sent" at once. The recipient within about a second is fine. | Transfers between two databases may finish asynchronously, with a status (step 2.2). |
| How fresh must the balance on the home screen be? | Right after a user's own transfer, it must show the new balance. Otherwise a second of delay is fine. | A cache is allowed, but it must never go backwards and must honour "at least the version I just saw" (step 2.4). |
| Which bank rails for withdrawals, and can they fail after we send? | Standard bank transfers, and instant payments where the user's bank supports them. Banks can send a payment back days later. | A withdrawal is a small state machine with ledger entries at each step (step 2.6). |
| What do analytics and fraud need exactly? | Every ledger change, complete, in order per account, within seconds. | Change data capture from each database's log, partitioned by account (step 2.5). |
| Anything about the data we must keep? | Seven years of ledger history. | Old ledger rows move from the databases to S3 (R2.6). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Wallets | ~1M | 50M; 10M active a day |
| Transfers | 200K/day, ~30/s busiest minute | 20M/day: ~231.5/s average, 5,000/s peak, 20,000/s ceiling |
| Balance reads | ~300/s | 50,000/s |
| Account types | Students and the canteen | + hot merchants and payroll wallets |
| Money out | None | Withdrawals to banks, which can bounce |
| Consumers of events | None | Analytics, fraud, notifications |
| Latency | P99 < 100 ms | P99 < 20 ms same shard, < 60 ms cross-shard |
| Availability | 99.9% (43.8 min/month) | 99.99% (4.4 min/month) |
| Retention | Unspecified | 7 years of ledger |
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | What breaks at the new scope |
|---|---|
| One Aurora writer | 20,000 transfers/s, each several row writes plus index updates and a commit, is far past what we'd plan for one writer. Aurora scales reads with replicas, but a cluster has exactly one writer. |
| A transfer is one local transaction | Once wallets live on different databases, Alice's row and Bob's row can't be locked in one transaction. |
| One row per account | A merchant receiving 5,000 payments/s means 5,000 transactions a second queueing on one row lock. |
| Every balance read hits the database | 50,000 reads/s on the writers competes with transfers for the same CPU. Readers lag slightly, which breaks "show me my new balance". |
| Nothing publishes changes | Analytics, fraud and notifications would have to poll the ledger, or the service would have to "publish after commit", which loses events when the process dies between the two. |
| No withdrawals | Money leaving for a bank takes time and can come back. Round 1 has no state for "on its way out". |
The order we fix it in: capacity (2.1), the transfers that capacity creates between databases (2.2), hot accounts (2.3), reads (2.4), events (2.5), then withdrawals (2.6).
R2.3 New Requirements and API Additions
A transfer can now finish asynchronously
httpHTTP/1.1 202 Accepted Content-Type: application/json { "transfer_id": "t_9f2c01", "status": "PROCESSING", "amount_minor": 2500, "from_balance_after_minor": 7500, "from_version": 42 }
202 PROCESSING means the money has left Alice's balance and is on its way to Bob. It finishes as COMPLETED, or as RETURNED if Bob's side refuses (a frozen account, say), in which case the money comes back to Alice. When the whole transfer completes within its latency budget, the same call returns 200 COMPLETED instead.
httpGET /v1/transfers/t_9f2c01 HTTP/1.1
httpHTTP/1.1 200 OK Content-Type: application/json { "transfer_id": "t_9f2c01", "status": "COMPLETED", "completed_at": "2026-09-27T18:04:11.083Z" }
Read-your-own-writes on balances. The balance endpoint takes the version the client last saw:
httpGET /v1/wallets/w_1001/balance?min_version=42 HTTP/1.1
The server never answers with a version below 42 (step 2.4).
Withdrawals
httpPOST /v1/withdrawals HTTP/1.1 Idempotency-Key: 0b8e5d7a-44f1-4f0e-8a51-7c2d9e6b3a10 Content-Type: application/json { "wallet_id": "w_1001", "bank_account_id": "ba_77", "amount_minor": 5000, "speed": "INSTANT_IF_AVAILABLE" }
httpHTTP/1.1 202 Accepted Content-Type: application/json { "withdrawal_id": "wd_5521", "status": "RESERVED", "amount_minor": 5000 }
Statuses: RESERVED → SENT → SETTLED, or RETURNED / FAILED with the money back in the wallet (step 2.6).
Merchant balances
httpGET /v1/merchants/m_300/balance HTTP/1.1
httpHTTP/1.1 200 OK Content-Type: application/json { "merchant_id": "m_300", "currency": "USD", "available_for_payout_minor": 18250000, "pending_sweep_minor": 412300, "total_minor": 18662300 }
The split between "available for payout" and "pending sweep" comes from how hot accounts work (step 2.3).
The event stream contract. Two topics, both at-least-once (a consumer may see an event twice and must ignore the repeat):
| Topic | Key (sets the partition) | Event fields | Order guarantee | How to deduplicate |
|---|---|---|---|---|
ledger-entries | account_id | entry_id, account_id, account_version, transfer_id, amount_minor, currency, balance_after, committed_at | Per account, in account_version order | (account_id, account_version) |
transfer-events | transfer_id | transfer_id, event_type (RESERVED, CONFIRMED, REJECTED, RETURNED, COMPLETED), from, to, amount_minor, shard | Per transfer | event_id |
R2.4 Design Evolution: One Ledger on Many Machines
Step 2.1: "One Database Can't Take 5,000–20,000 Transfers/s"
The problem: at 20,000 transfers/s, the single Aurora writer would need to run tens of thousands of short money transactions a second, each several row writes, index updates and a commit. Aurora gives us up to 15 readers, but only one writer per cluster. What would you do?
Synthesizing vector architecture diagram...
Two hops: a fixed hash to a bucket, then a lookup to a shard. The directory is what lets us move one bucket without rehashing everything.
Go deeper: when one customer outgrows its shard, and reports across all shards. Suppose a payroll provider with a million employees is about to join, and all its wallets hash into a few buckets on shard 5. On launch day, shard 5 would take far more than its share. Before the launch we carve out those buckets: copy them to a new shard, then switch their directory entries (a short write freeze on just those buckets while the switch happens). Its hottest single accounts get striped as in step 2.3. The opposite question, "one report over every wallet", is answered by not querying 16 shards at once: every shard's changes stream to S3 (step 2.5), and platform-wide reports run there with Athena, a few seconds to minutes behind the ledger.
Primitive: Database Sharding & Partition Keys · Drill: Sharding tenant hotspot
Step 2.2: "Alice Is on Shard 3, Bob on Shard 9"
The problem: Alice (shard 3) sends Bob (shard 9) 25.00. No single transaction can lock both rows. If we debit Alice and crash before crediting Bob, 25.00 has vanished; if we credit Bob first and Alice's debit fails, we've created 25.00. What would you do? And where is the money while it moves?
The four entries of one cross-shard transfer
| Shard | Step | Account | amount_minor | Shard's running sum |
|---|---|---|---|---|
| 3 | Try | Alice | −2,500 | |
| 3 | Try | in_transit (shard 3) | +2,500 | 0 |
| 9 | Confirm | in_transit (shard 9) | −2,500 | |
| 9 | Confirm | Bob | +2,500 | 0 |
Every local transaction balances on its own shard, so "debits equal credits" holds per shard at every moment. The in_transit accounts, added across all shards, hold exactly the money in flight: +2,500 between the Try and the Confirm, 0 after.
The in-transit check, covering every path. Each transfer touches in_transit through exactly these paths: Try (+amount on the sender's shard), Confirm (−amount on the receiver's shard) or Cancel (−amount on the sender's shard). So per transfer ID, the in_transit entries across all shards must sum to 0 (finished) or +amount (in flight). A continuous job over the event stream checks this, and a nightly job recounts every transfer over the full copy in S3. The two shards' changes reach the stream independently, so the streaming check can briefly see −amount (the receiver's Confirm arrived before the sender's Try); it therefore applies the 5-minute grace to any non-zero sum, and only the nightly full recount is authoritative for a −amount result:
Per-transfer in_transit sum | Meaning | Action |
|---|---|---|
| 0 | Confirmed or cancelled | none |
| +amount, younger than 5 min | In flight | none |
| +amount, older than 5 min | Stuck saga | P2 alarm; the worker re-drives it; a person looks after 30 min |
| any other non-zero sum, younger than 5 min | Probably events arriving out of order between shards | none yet |
| −amount, −2 × amount or anything else, after 5 min in the stream or in the nightly recount | A confirm and a cancel both happened, or a bug wrote a one-sided entry | P1: freeze both wallets, investigate |
Synthesizing vector architecture diagram...
The timeout doesn't trigger the cancel; only shard 9's recorded REJECTED does. The cancel is a new balanced pair, so the ledger shows both the attempt and the return.
Answering "what if the cancel itself fails halfway?" It can't fail halfway: the cancel is one local transaction on shard 3, so it commits completely or not at all. If it doesn't commit (shard 3 is failing over), the worker retries it from the outbox with backoff; it is idempotent because it only runs while the transfer's status is RESERVED. If it keeps failing, the transfer stays in flight, where the in-transit check above sees it and pages a person after 30 minutes. At no point is the money unaccounted for: it is in in_transit, which is exactly what that account is for.
Primitives: Two-Phase Commit & Saga Orchestration · Change Data Capture & Outbox · Drill: Saga orchestration compensating failure
Step 2.3: "A Merchant Row Gets 5,000 Credits/s"
The problem: a big merchant receives 5,000 payments a second during a sale. Every payment locks the merchant's account row. Each transaction takes the lock after its idempotency insert and holds it through the commit; we plan on about 8 ms (a planning assumption to measure under load), so one row can serve at most transactions a second. Payments queue, time out and retry. What would you do? And how does the merchant later pay money out without ever going negative?
Synthesizing vector architecture diagram...
Money flows one way: payers credit a stripe on their own shard, sweeps collect into the main account, and only the main account pays out. No path ever debits a stripe by more than it holds.
Step 2.4: "50K Balance Reads/s Hit the Primary"
The problem: 50,000 balance reads a second at peak, from the app's home screen, compete with transfers for the shard writers' CPU. Sending them to Aurora readers helps with load, but readers lag slightly: a user who just paid sees the old balance, thinks the payment failed, and pays again. What would you do?
A cold cache. After a cache node restart, the first wave of reads all miss. Each service task coalesces concurrent misses for the same account into one writer read (a "single-flight" per key), and the writer lookups are primary-key reads spread over 16 shards: at 50,000 misses/s that's about 3,100 point reads per shard for a few seconds, small next to the transfer load.
Why not cache forever and skip the TTL? We almost do: the stream keeps entries current, and the version fence makes them safe. We keep a long TTL only as a backstop for the cases where no event arrives to correct a key: a bug in the updater, or a key filled from a path we didn't think of.
Primitive: Distributed Cache Patterns & Eviction · Drill: Caching hot product page (its cold-cache and no-TTL questions are answered just above)
Step 2.5: "Events Are Lost or Out of Order"
The problem: the wallet service commits a transfer, then publishes "transfer completed" to the stream for the cache, fraud and notifications. Sometimes the process dies between the two, and the event is never sent. Sometimes two events for one account arrive in the wrong order. What would you do?
Why CDC instead of a job that polls the outbox table? A poller runs SELECT ... WHERE published = false ORDER BY id every few hundred milliseconds on every shard, then updates each row it published: that is extra queries, extra writes, index churn on the hot outbox table, and latency up to the polling interval. Reading the WAL costs the database almost nothing beyond keeping the log, and publishes as soon as the commit is in the log. The price is running Debezium and watching its slot.
If a slot is ever lost (for example, some engine versions need the slot recreated after a failover), we don't trust the stream to be complete. We recreate the slot, then republish from the stored source of truth: the outbox and ledger rows written since the last event the stream has for that shard. We keep outbox rows for 7 days for exactly this. Consumers ignore the repeats.
Primitive: Change Data Capture & Outbox · Drill: CDC outbox dual-write drift
Step 2.6: "Withdrawals Take Days and Can Bounce"
The problem: Alice withdraws 50.00 to her bank. A standard bank transfer settles hours or a day later, and the bank can send it back afterwards (the account was closed, the details were wrong). An instant payment settles in seconds but only if her bank supports it. What would you do? What does the ledger say at each moment?
Synthesizing vector architecture diagram...
Every arrow that moves money writes a balanced pair; the one arrow users don't see coming is SETTLED to RETURNED, which is why it exists.
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | One writer can't keep up | 16 Aurora shards; 1,024 buckets; a directory; system accounts per shard | 15 in 16 transfers cross shards |
| 2.2 | Transfers across shards | Try–confirm–cancel saga through in_transit; never cancel on a timeout; outbox as the second driver; in-transit check | Money visibly in flight; 3 local transactions |
| 2.3 | Hot merchant row | 32/64/128 credit stripes on every shard, paid from the payer's shard; sweeps into one main account; payouts only from main; payroll debit stripes | Payout lag ~10 s; 129-account totals |
| 2.4 | Reads hit the writers | Valkey cache, fenced by version; min_version; refills from the writer; money decisions on the writer | A cache cluster; stream lag |
| 2.5 | Events lost or reordered | Debezium CDC from each shard's WAL to MSK, keyed by account; outbox; idempotent consumers | Connectors and slot lag to watch |
| 2.6 | Withdrawals bounce | RESERVED → SENT → SETTLED / RETURNED / FAILED, entries at each step; bank-file reconciliation | Pending money; a daily process |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
Follow the money: every change commits in one shard, the WAL carries it to MSK, and everything else (cache, sagas, analytics, fraud) is a consumer of that one ordered log. The only things that write to shards are the API, the saga worker, the withdrawal engine and reconciliation, all through the same routing.
Trace 1: a same-shard transfer. Exactly Round 1's transaction on one shard, followed by CDC: the entries reach MSK, the cache updater writes Alice's new balance with version 43 and Bob's with his new version, and Bob's push notification goes out from the notifications consumer.
Trace 2: a cross-shard transfer, with the direct path.
Synthesizing vector architecture diagram...
The user waits for the two dependent transactions and the hops between them; marking the transfer complete on shard 3 happens after the reply. If the confirm doesn't answer within the budget, the user gets 202 PROCESSING and the saga worker finishes the job (the compensation trace is in step 2.2).
Trace 3: a hot-merchant payment. Alice (shard 3) pays merchant m_300. The API looks up m_300's stripes on shard 3 (about 8 of its 128), picks one at random, and runs one local transaction: lock Alice and the stripe in sorted ID order, −2500 / +2500. Within 10 seconds the sweeper moves that stripe's balance to the main account on shard 0 with a saga.
Trace 4: a withdrawal that comes back.
Synthesizing vector architecture diagram...
The return arrives after settlement, so it is its own balanced pair; nothing earlier is edited.
R2.6 Numbers and Cost
Traffic
| Quantity | Math | Value |
|---|---|---|
| Average | 20,000,000 ÷ 86,400 s | ≈ 231.5/s |
| Peak | 231.5 × 20 | ≈ 4,630/s, provisioned as 5,000/s |
| Ceiling | given (events) | 20,000/s |
| Balance reads at peak | 10 reads per transfer × 5,000 | 50,000/s |
Storage. About 750 bytes per transfer (the transfer row, two entries and their share of indexes and the idempotency key):
Cross-shard transfers write four entries instead of two, so treat 750 B as a planning average and check it against real table sizes. Aurora bills each gigabyte once: its six copies across three AZs and all of a cluster's readers share one storage volume, so we don't multiply by 3. We also don't keep 38 TB in Aurora. The ledger tables are partitioned by month; we keep 400 days online (≈ 15 GB × 400 = 6 TB across 16 shards, about 375 GB each), and older partitions are dropped once their export to S3 (step 2.5's stream, plus a monthly verified export) is confirmed. History older than 400 days is served from S3 with Athena.
How many local transactions the shards run. Assume, from the scope raise, half of transfers pay merchants and are local thanks to stripes on every shard; the other half are P2P between random wallets, of which 1 in 16 is local and 15 in 16 cross shards at 3 local transactions each:
| Transfers/s | Local transactions/s | Per shard (÷ 16) | |
|---|---|---|---|
| Peak | 5,000 | 9,688 | 605 |
| Ceiling | 20,000 | 38,750 | 2,422 |
Shard count and size. We plan on about 500 short money transactions per second per vCPU of an Aurora writer, an assumption confirmed by load tests with our real transactions: 2,000/s for a db.r7g.xlarge (4 vCPUs), 4,000/s for a db.r7g.2xlarge (8 vCPUs).
- Shard count is set by the ceiling, because adding shards means moving buckets, which is slow and risky during an event. At the ceiling on 2xlarge writers at no more than 70% busy: , so at least 14. We choose 16 for headroom and an even 64 buckets per shard.
- Instance size is set by the normal peak. On xlarge writers, the peak runs at . Before a known event we scale every shard to 2xlarge (add a bigger reader, then fail over to it: about 30–60 seconds of write errors per shard, done shard by shard at a quiet hour), which puts the ceiling at .
- An unplanned surge on xlarge writers reaches 70% at transfers/s. Above that, the API sheds load with
429andRetry-Afterbefore starting any database work, rather than letting every shard slow down.
Losing an AZ. Each shard's reader sits in a different AZ from its writer, so a lost AZ fails over the writers that were in it (under 60 s typically) to readers of the same size: shard capacity is unchanged. Fargate tasks are sized so that two AZs can carry the peak on their own, rounded per service (each service is a separate bulkhead):
| Service | Peak need (assumptions) | Tasks of 2 vCPU | Per AZ, so 2 AZs suffice | Total |
|---|---|---|---|---|
| Read API | 50,000 reads/s ÷ 1,000 per vCPU = 50 vCPU | 25 | ⌈25 ÷ 2⌉ = 13 | 39 |
| Write API | 5,000 transfers/s ÷ 250 per vCPU = 20 vCPU | 10 | ⌈10 ÷ 2⌉ = 5 | 15 |
| Total | 54 |
For a 20,000/s event we pre-scale ahead of time: the write API needs 20,000 ÷ 250 = 80 vCPU = 40 tasks, so 20 per AZ and 60 in total, and the read API grows in proportion to the event's read traffic.
Cache size. 10M active wallets × about 200 bytes of data (balance, version, status) ≈ 2 GB. With the key and Valkey's per-key overhead, assume about 350 bytes each: 3.5 GB. One cache.r7g.large (13.07 GiB) holds it with room; we run one primary and two replicas, one per AZ. 50,000 reads/s over three nodes is about 17,000 each. The cache updater writes about 2 keys per transfer: 10,000 fenced writes/s at peak, all on the primary.
Stream size. About 1.5 KB of events per transfer (two or four entries and one or two transfer events): 20M × 1.5 KB = 30 GB/day; 7.5 MB/s at peak, 30 MB/s at the ceiling, before Kafka's three-way replication.
Latency budgets (P99 targets)
| Path | Dependent steps | Budget |
|---|---|---|
| Same-shard transfer | ALB and service ~2 ms + the local transaction ~12 ms (8 statements and an Aurora commit) | < 20 ms |
| Cross-shard transfer | ~2 ms + Try ~12 ms + hop ~1 ms + Confirm ~12 ms + hop ~1 ms | ≈ 28 ms typical, < 60 ms P99; past that, reply 202 |
| Balance read, cache hit | ALB and service ~2 ms + Valkey ~1 ms | < 5 ms |
Availability. 99.99% of a month is about 4.4 minutes. One shard failover (under 60 s) affects 1/16 of wallets, about 3.75 "whole-system seconds". The monthly budget is about 263 seconds, so on paper it covers roughly 263 ÷ 3.75 ≈ 70 single-shard failovers a month, but it must also cover every other outage, so failovers are planned and done one shard at a time.
Monthly cost (us-east-1 on-demand prices, 730 hours)
| Item | Math | Monthly |
|---|---|---|
Aurora, 32 × db.r7g.xlarge, I/O-Optimized | 32 × ($0.553 × 1.3)/h × 730 h | ≈ $16,790 |
| Aurora storage, 6 TB at steady state | 6,000 GB × $0.225 (I/O-Optimized) | ≈ $1,350 |
Valkey, 3 × cache.r7g.large | 3 × $0.175/h × 730 h | ≈ $383 |
MSK, 3 × kafka.m7g.xlarge | 3 × $0.408/h × 730 h | ≈ $894 |
| MSK storage | 30 GB/day × 7 days × 3 copies ≈ 630 GB, provisioned 1 TB × $0.10 | ≈ $100 |
| MSK Connect, 16 Debezium connectors × 1 MCU | 16 × $0.11/h × 730 h | ≈ $1,285 |
| Fargate API, 54 tasks × 2 vCPU, 4 GB, Arm | 54 × (2 × $0.03238 + 4 × $0.00356)/h × 730 h | ≈ $3,114 |
| Fargate workers (saga, cache, withdrawal, sweeps, recon), 12 × 1 vCPU, 2 GB | 12 × $0.0395/h × 730 h | ≈ $346 |
| ALB | ≈ $45 | |
| S3 archive and lake | ≈ $100 | |
| CloudWatch | ≈ $300 | |
| Total | ≈ $24,700/month |
Why I/O-Optimized: with thousands of commits a second, per-I/O charges would be a large and unpredictable part of the bill; I/O-Optimized drops them for a 30% higher instance price and a higher storage price. AWS's own guidance is to consider it when I/O is more than about a quarter of Aurora spend; we check the bill after a month and can switch. The event weeks cost extra: 2xlarge writers double the Aurora compute line for those days.
Per transfer: 20M × 30.4 days ≈ 608M transfers a month, so about $0.04 per thousand transfers.
The kafka broker size (3 × m7g.xlarge for 30 MB/s at the ceiling) and the 250 transfers/s per vCPU for the write API are planning assumptions we confirm with load tests.
R2.7 Trade-Offs
Where does the ledger live? Verified facts only; "scales" means by adding nodes or shards, and no option removes the hot-row limit.
| Aurora PostgreSQL, our shards (chosen) | Aurora PostgreSQL Limitless Database | DynamoDB transactions | Event sourcing on Kafka | Distributed SQL (Spanner, CockroachDB, Aurora DSQL) | |
|---|---|---|---|---|---|
| Transactions | Full ACID inside a shard; our saga across shards | Distributed transactions across its shards, managed for you | TransactWriteItems: up to 100 actions, 4 MB, all or nothing; conflicting transactions are cancelled | None in the database: a single writer per account partition decides in order | Distributed ACID transactions |
| Isolation | Read Committed + our row locks (any level we choose) | Read Committed and Repeatable Read; Serializable is not supported | Serializable per transaction | Whatever the writer enforces | Spanner: external consistency (strict serializability); CockroachDB: Serializable by default; DSQL: strong snapshot isolation (like Repeatable Read), optimistic |
| "Never negative" | Check constraint + FOR UPDATE | PostgreSQL row locks; check which constraints its sharded tables support | A condition expression on every debit | Code in the writer | Constraints; under optimistic control, conflicting writers abort at commit |
| Hot account | Stripes | Stripes | One item takes at most 1,000 write units/s (a 1 KB transactional write costs 2): stripes | One partition's throughput: stripes | Contention or aborts: stripes |
| What we run | 16 clusters, a router, a saga | Fewer moving parts; newer, less control over placement | Nothing to patch; ledger queries by key only | Kafka as the source of truth, projections, replays | Spanner is Google Cloud; CockroachDB is self-run or its vendor's cloud; DSQL is AWS |
We chose our own shards because the transfer transaction, the constraints and the lock order are exactly Round 1's, just repeated 16 times, and because we control which wallets live together (the merchant stripes on every shard depend on that). Limitless would be the first thing to re-evaluate if running 16 clusters and the saga becomes the bigger burden.
Stripe count. More stripes mean less lock contention and more rows to add up for a total and more sweeps. We size from the formula in step 2.3 and raise S when the stripe lock-wait alarm fires.
Cache freshness. Users who pass min_version always see at least their own last write. Everyone else sees the stream's lag. Making every read strongly consistent would mean reading the writers, which is the load we built the cache to remove.
Primitive: Event Sourcing & CQRS
R2.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| A cross-shard confirm fails | 202 PROCESSING replies; saga retry metrics rise | The worker retries the confirm (idempotent by transfer ID). Only a recorded REJECTED on the receiving shard leads to a cancel, a new balanced pair. Anything in flight for 5 minutes alarms. |
| Aurora failover mid-transfer | Write errors on one shard for under a minute | On the sender's shard before the Try commits: it rolls back; the client retries with its key. After the Try commits: the outbox row survives and the worker finishes the saga. On the receiver's shard: the confirm is retried after the failover. |
| Cache and database disagree | A user sees an old balance | Every cache write is fenced by version, so the cache can't go backwards; min_version sends lagging reads to the writer; refills come from the writer; a 24-hour TTL is the backstop. The runbook can delete a key, since the next read refills it from the writer. |
| Replica lag | History pages missing the latest entry | History reads pass min_version too: if the reader's copy of the account is older, the read goes to the writer. |
| A hot-account lock storm | Lock waits on a merchant's stripes rise; payments to it slow down | The stripe lock-wait alarm fires; raise S (new stripes start empty, and payers pick them at once); lock_timeout stops waiters queueing forever. |
| A shard hot spot | One shard's CPU far above the others | Find the bucket(s): a big new customer or a merchant without stripes. Stripe the account, or carve its buckets out to their own shard through the directory. |
| The CDC connector stalls | Stream lag climbs; slot lag climbs on one shard | Cache reads fall back to the writer via min_version; restart the connector; if the slot is lost, republish from the outbox and ledger rows (step 2.5). |
R2.9 Production Gotchas
| Gotcha | Symptom | Cause | Fix |
|---|---|---|---|
| Locks in random order | Deadlocks (40P01), second-long stalls at peak | A new code path, often a batch job or refund, locked accounts in its own order | One shared "lock accounts" function that sorts; alarm on any deadlock |
| Balance column without entries | The nightly check finds accounts whose balance isn't the sum of their entries | Someone "fixed" a balance with an UPDATE | No UPDATE/DELETE grants on the ledger; balance changes only through the transfer function; corrections as new transfers |
| A network call inside the transaction | Lock waits and connection pool exhaustion when a downstream service is slow | Sending a push notification or calling fraud scoring between BEGIN and COMMIT | Nothing external inside a transaction: notifications come from the stream after commit; fraud scoring runs before BEGIN |
| Single-row hot accounts | One merchant's payments time out during a sale while the rest of the shard is fine | 5,000 transactions a second on one row lock | Credit stripes on every shard, sized by the formula |
| Balances from replicas without fencing | Users see their old balance after paying and pay again | Reads served from a lagging reader or a cache with no version check | Version fence on every cache write; min_version on reads |
| "Check the total" across stripes | A merchant overdrawn in total | The rule was only on the total and stripes were allowed to go negative individually; two payouts summed the stripes without locks and debited different stripes (write skew) | Payouts only from the main account; debit stripes pre-split for payroll |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | Every shard has a reader in another AZ; saga steps are idempotent and never compensate on a timeout; money in flight is in named accounts with an age alarm; load shedding above the planned capacity REL 4 · REL 5 · REL 10 · REL 11 |
| Performance Efficiency | Shards sized from the ceiling, instances from the peak; stripes that turn merchant payments into local transactions; a version-fenced cache for 50K reads/s PERF 1 · PERF 3 |
| Security | Only the transfer function writes ledger entries; per-service database roles (the cache updater can't touch shards; reconciliation can only post through the transfer function); bank files and payment instructions over authenticated, encrypted channels SEC 3 · SEC 8 · SEC 9 |
| Cost Optimization | About $24,700 a month (≈ $0.04 per thousand transfers); I/O-Optimized chosen from the I/O pattern; bigger instances only for event weeks; 400 days in Aurora and the rest in S3 COST 5 · COST 6 |
| Operational Excellence | Alarms with first actions (below); planned, one-shard-at-a-time scaling for events OPS 6 · OPS 8 |
| Sustainability | Light this round: Graviton instances throughout; compute scaled up only for known events; cold ledger history moved out of the databases to S3 SUS 2 · SUS 4 |
Alarms and first actions
| Signal | Alarm | First action |
|---|---|---|
| Ledger imbalance (any transfer or shard whose entries don't sum to zero) | any | P1: freeze the affected accounts; find the code path |
| In-transit age | any transfer in flight > 5 min | Check the receiving shard's health; re-drive the saga |
| Saga cancels | > 10/min | Look for a shard rejecting confirms (frozen accounts in bulk? a bad deploy?) |
| Stripe lock wait P99 | > 50 ms for 5 min | Raise that merchant's stripe count |
| Cache version lag (stream position vs latest committed version) | > 2 s for 5 min | Check the cache updater and the connector |
| CDC slot lag per shard | growing for 10 min | Restart the connector; check MSK health |
| Deadlocks | any | Find the code path that broke the lock order |
| Unmatched bank statement lines | any at the daily run | Reconciliation on-call investigates the break |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Shards by wallet with a directory, and says what that does to the share of cross-shard transfers.
- Uses a saga with an in-transit account instead of 2PC, and explains why: no locks held across machines.
- Knows the saga's one dangerous shortcut, cancelling on a timeout, and makes the receiver's recorded outcome the only trigger.
- Stripes hot accounts and gets the debit side right: never "check the total", spend only from one account or from pre-split stripes.
- Keeps a cache that can't go backwards, and never uses it for money decisions.
- Publishes events from the log, not after commit.
- Models withdrawals as states with entries at every step, including returns after settlement.
- Sizes shards from the ceiling and instances from the peak.
Follow-up questions
-
"Why not lock both shards' rows with 2PC just for transfers under 1 ms apart?" Answer: the cost of 2PC isn't average latency, it's the failure case: a coordinator that dies after prepare leaves locks on both accounts until it recovers. With a saga, a crash leaves money in
in_transit, which is visible, counted and finished by the worker, while both accounts keep working. -
"A user sends money to themselves across shards, their USD wallet on shard 3 to their savings pocket on shard 9. Anything special?" Answer: it shouldn't happen: all accounts of one wallet hash on the wallet ID into the same bucket, so their own pockets are on one shard and the transfer is local. That co-location is a deliberate part of the bucket design.
-
"How would you move bucket 717 from shard 11 to a new shard without downtime?" Answer: copy the bucket's accounts and recent entries to the new shard while it keeps serving (logical replication or a backfill plus the CDC stream), then take a short write freeze on just that bucket: stop new transactions for it, wait for in-flight ones and for the copy to catch up, bump the directory version to point the bucket at the new shard, and release. Sagas in flight keep working because steps are addressed by account, which the router resolves through the new directory. Transfers to that bucket during the freeze get a quick retryable error.
Interview gotchas from this round
| Gotcha | Why it's wrong |
|---|---|
| "2PC across shards" | Prepared transactions hold row locks until the coordinator decides; a coordinator failure blocks the accounts. |
| "Cancel the saga after a timeout" | The confirm may have succeeded; cancelling then creates money. Only a recorded rejection cancels. |
| "Check the merchant's total before a payout" | Without locking every stripe, and with the rule only on the total, two payouts both pass: write skew. With a rule per stripe, big payouts must lock many stripes: the hot lock is back. |
| "Invalidate the cache on change" | A delete followed by a slow refill can put an old value back; write new values with a version fence instead. |
| "Publish the event after commit" | A crash between the two loses the event; capture it from the log. |
Round 3 · Architect · "Cross-Border, Regulated, and Provably Right"
~45 min · Principal (L7) · 3 home regions, each with a DR region · ~300M wallets, 100M transfers/day · 99.99% per region for money movement · a region can fail without double-spend
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Rounds 1 and 2.
Round 2 in 60 seconds. "We run a national wallet: 50M wallets, 20M transfers a day, 5,000 a second at peak and 20,000 during events, 50,000 balance reads a second, 99.99% in one region. The ledger is double-entry, in integer cents, and sharded: a wallet hashes into one of 1,024 buckets, a directory maps buckets to 16 Aurora PostgreSQL shards, and each shard runs Round 1's transaction: an idempotency key, rows locked
FOR UPDATEin sorted order, entries that sum to zero, a non-negative check. Transfers across shards are a try–confirm–cancel saga through anin_transitsystem account; we never cancel on a timeout, only on the receiver's recorded rejection, and a check makes sure every transfer's in-transit entries net to zero or to the amount in flight. Hot merchants are striped 32 to 128 ways across every shard, payers credit a stripe on their own shard, sweeps collect into one main account, and payouts spend only from that account. Balances are served from Valkey with a version fence andmin_version. Every change leaves through Debezium CDC to MSK, keyed by account. Withdrawals are a state machine with entries at each step and bank-file reconciliation. About $24,700 a month. Open costs: one currency, no compliance states, nothing ties our ledger to the real bank balances, one region, and ordinary logs as the audit trail."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: many shards running the same local transaction, and one ordered log that feeds everything else.
Steps so far
| Step | Problem | Component |
|---|---|---|
| 1.1 | Explain a balance | Double-entry ledger, balance cached in the same transaction |
| 1.2–1.3 | Double spend, deadlocks | FOR UPDATE in sorted order at Read Committed; check constraint |
| 1.4–1.5 | Retries, top-ups | Idempotency key in the transaction; credit on the PSP's confirmation |
| 2.1 | One writer | 16 shards, buckets, a directory |
| 2.2 | Cross-shard transfers | Saga through in_transit; cancel only on a recorded rejection |
| 2.3 | Hot merchants | Stripes on every shard, sweeps, payouts from one account |
| 2.4–2.5 | Reads, events | Version-fenced cache; CDC keyed by account |
| 2.6 | Withdrawals | State machine with entries; bank-file reconciliation |
Open costs: one currency; no compliance states; no proof the ledger matches the bank; one region; logs as the audit trail.
R3.1 The Scope Raise
Interviewer: "We're now in 30 countries, each with its own currency, about 300 million wallets and 100 million transfers a day. Users send money across borders, so a US user pays a German user in euros. Regulators in each country require identity checks, limits by verification level, sanctions screening, and the ability to hold a suspicious transfer for review."
Interviewer: "Customer money has to sit in safeguarded bank accounts, and auditors will ask us to prove those accounts match our ledger every day. EU residents' data must stay in the EU. A whole AWS region can fail, and we must never let anyone spend the same money twice because of it. Auditors also want tamper evidence for 7 years. And one more: the board has asked for five nines."
We ask back, and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Can one wallet hold several currencies? | Yes: a German user may hold EUR and USD. | One account per (wallet, currency), all in the wallet's bucket so conversions between a user's own currencies are local (step 3.1). |
| Who carries the exchange-rate risk while the user decides? | We show a rate the user can accept for a short time; treasury hedges our positions. | A quote with an expiry, locked when the user confirms, and FX system accounts per currency (step 3.1). |
| Which checks must happen before money moves, and which can come after? | Sanctions and limits before. Suspicious patterns may either block or hold for a human. | Compliance moves into the transfer path, with a held state (step 3.2). |
| How is customer money held? | In safeguarded bank accounts, per legal entity and per currency, separate from company money. | A daily safeguarding reconciliation from ledger to bank statement (step 3.3). |
| Where must data live? | EU residents in the EU. Other countries have their own rules; compliance gives us a table per country. | A home region per wallet, chosen at sign-up from the country of residence (step 3.4). |
| What do we promise if a region fails? | Other regions keep working. The failed region's users are back within an hour. And no double-spend. | A DR region per home region, with a failover procedure that decides what happens to the last seconds of writes (step 3.4). |
| What's our biggest fraud loss today? | Account takeovers: someone logs in as the user and drains the wallet. | Risk checks and friction on money leaving the wallet (step 3.6). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Wallets | 50M | ~300M in 30 countries |
| Transfers | 20M/day, one currency | 100M/day, 30 currencies, some cross-border with FX |
| Regions | 1 region, 3 AZs | 3 home regions (Americas, EU, Asia-Pacific), each with a DR region |
| Checks before posting | Balance, status | + KYC tier limits, sanctions, risk scoring, holds |
| Proof | Internal checks | Daily safeguarding reconciliation; 7 years of tamper-evident audit |
| Data residency | None | Home region per wallet |
| Survives | An AZ; a database failover | A region, without double-spend |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| One currency | An account balance in "cents" means nothing without a currency, and sums across currencies are meaningless. A USD→EUR transfer has no way to post. |
| Transfers are "allowed" or "insufficient funds" | No place for "held for review", no limits per verification level, no sanctions check. |
| The ledger checks itself | Nothing proves that customer balances match money in real bank accounts. |
| One region | A region outage stops every wallet, and EU data would sit outside the EU. |
| Logs and database backups as history | Anyone with enough access could change a row and its backup; nothing would show it. |
| Password login is enough to send money | A stolen password empties the wallet in one transfer. |
| 16 shards in one region | The Americas region alone grows past what 16 shards were sized for; it needs about 27 (R3.6), which means moving buckets to new shards with the Round 2 bucket-move procedure. |
The order we fix it in: currencies (3.1), compliance (3.2), proof of money (3.3), regions and residency (3.4), proof of history (3.5), then account takeover (3.6).
R3.3 New Requirements and API Additions
FX quote, then confirm
httpPOST /v1/fx/quotes HTTP/1.1 Content-Type: application/json { "from_wallet_id": "w_us_1001", "to_wallet_id": "w_de_2002", "sell_currency": "USD", "sell_amount_minor": 10000, "buy_currency": "EUR" }
httpHTTP/1.1 201 Created Content-Type: application/json { "quote_id": "q_58c1", "sell": { "currency": "USD", "amount_minor": 10000 }, "fee": { "currency": "USD", "amount_minor": 50 }, "rate": "0.9208", "buy": { "currency": "EUR", "amount_minor": 9162 }, "expires_at": "2026-09-27T14:03:30Z" }
httpPOST /v1/transfers HTTP/1.1 Idempotency-Key: 3d0f6a2b-9b1e-4c8e-b7aa-51e2f0c9d4a7 Content-Type: application/json { "quote_id": "q_58c1" }
The rate is a decimal string, never a float. The amounts are final and shown to the user before confirming: rounding happens once, at quote time.
Holds. A transfer can now answer 202 with "status": "HELD". The money has left the sender's available balance but hasn't moved to the recipient. It ends as COMPLETED (released) or RETURNED (reversed). What we tell the user about why is a compliance decision, not an engineering one.
KYC tiers and limits. A KYC tier ("know your customer") is how much we have verified about a user's identity. The numbers below are illustrative; real thresholds come from each country's rules and our risk policy.
| Tier | Verified | Example limits |
|---|---|---|
| 0 | Phone and email | Receive only; balance cap 250 |
| 1 | Government ID | Send up to 1,000 per rolling 24 hours; domestic only |
| 2 | ID, address and source of funds | Up to 10,000 per rolling 24 hours; cross-border |
We use a rolling 24 hours rather than "per calendar day" as a design choice: it has no time-zone or daylight-saving edge where a user could send a day's limit at 23:59 and again at 00:01.
Safeguarding report (internal)
httpGET /internal/v1/safeguarding/entity/eu-01/currency/EUR?statement_date=2026-09-26 HTTP/1.1
It returns every line of the reconciliation in step 3.3 and whether the day balanced.
Residency attribute. Every wallet has home_region, set at sign-up from the country of residence using compliance's country table, and changed only by a supervised migration. The router uses it before the bucket: region, then bucket, then shard.
R3.4 Design Evolution: Currencies, Law and Proof
Step 3.1: "USD to EUR Across Wallets"
The problem: Alice in the US sends 100.00 USD to Bob in Germany, who holds euros. The rate moves every second. Alice looks at the confirmation screen for 40 seconds before tapping "send". What would you do? At what rate do we convert, and how do the entries balance when two currencies are involved?
Step 3.2: "A Transfer Looks Like Money Laundering"
The problem: a new account receives 40 small payments from different people in an hour and immediately tries to send the total abroad. Compliance wants that stopped before the money leaves, not discovered in next week's report. What would you do? Where do the checks run, and what happens to the money while a human decides?
Synthesizing vector architecture diagram...
Money only moves on the arrows out of Try, and every one of those moves is a balanced pair. A held transfer is still the sender's money, sitting in an account she can't spend from.
Step 3.3: "Does Our Ledger Match the Money in the Bank?"
The problem: the EU entity's ledger says customers own 50,000,000.00 EUR. The EU safeguarding bank account holds 51,100,000.00 EUR at yesterday's cut-off. An auditor asks: "Is that right? Prove it." What would you do?
The cut-off, computed. The EU bank cuts its statement at 23:59:59 local time. On 2026-09-26 Frankfurt is on summer time (CEST, UTC+2; it switches back to UTC+1 on 2026-10-25), so the ledger side uses entries timestamped up to 2026-09-26T21:59:59Z. On a winter date the same local cut-off is 22:59:59Z. Getting this wrong by an hour creates a break made only of an hour's traffic.
Step 3.4: "Residency, and a Region Can Fail"
The problem: EU wallets must live in the EU, and a whole region can fail. Someone proposes running every wallet's ledger in all three regions at once, "active-active", so any region can serve any user. What would you do? And how do we fail over without anyone spending the same money twice?
The Frankfurt-and-Virginia question, answered. "A user's balance in Frankfurt shows 100.00 while Virginia already spent it": with home regions, there is no Virginia copy that can spend a Frankfurt wallet's money, and no region shows another region's replica as an answer to a money question. And "why not make every write synchronous across regions?": because we'd pay the inter-region round trip on all 100M transfers a day and couple each region's availability to its partner, to protect against a failure that step 4 already contains for the only money that matters, the money that leaves.
Primitive: Cloud Disaster Recovery & Multi-Region Active-Active · Drill: Replication multi-region consistency
Step 3.5: "Prove Nothing Was Changed in 7 Years"
The problem: an auditor asks: "How do we know that no one, including your own database administrators, changed or deleted a ledger entry from 2027?" What would you do?
Synthesizing vector architecture diagram...
Each day's record includes the previous day's hash, so altering day 1 would change every record after it, including the signed head the auditor already has.
Step 3.6: "Account Takeover Drains Wallets"
The problem: an attacker phishes Alice's password, logs in from a new phone, adds a new payee, and sends her whole balance abroad in two minutes. What would you do?
Primitive: Bot Defense, Sybil Resistance & Registration Abuse
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | USD to EUR | 60 s quote stored on the sender's shard, checked with the database clock; legs balance per currency through fx_position; rounding at quote time | FX risk inside the window |
| 3.2 | Suspicious transfers | Screening, tier and risk before BEGIN; limits under the row lock; held accounts; release or reverse; fail closed for risky flows | Some transfers wait for a human |
| 3.3 | Ledger vs bank | Daily safeguarding reconciliation, one line per money path, cut-off converted to UTC correctly | A daily process and a buffer |
| 3.4 | Residency and region loss | Home region per wallet; cross-region sagas; Aurora Global Database DR; outbound money waits for DR; fenced promotion; lost-tail reconciliation | ~2 s cross-region; 1 h RTO |
| 3.5 | Tamper evidence | Append-only ledger; daily hash chain; S3 Object Lock compliance mode; KMS-signed head held by the auditor | Storage and a daily job |
| 3.6 | Account takeover | Device risk, step-up, payee cooling-off, velocity limits, holds after credential changes | Friction |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Each region is Round 2's system, sized for its own traffic. Regions talk to each other only through sagas for cross-border transfers, and each replicates only to its own DR region. The audit export stays in each region's own bucket, so residency holds for the archive too.
Around each regional ledger sit the pieces this round added: the FX service (rates and spreads), the compliance engine (screening status, tiers, limits policy, case management for holds), the safeguarding reconciliation job per legal entity and currency, and the audit export.
Trace 1: a cross-border transfer with FX.
Synthesizing vector architecture diagram...
The user is answered right after the Try. The confirm waits for the US DR replica, and the CONFIRMED reply waits for the EU DR replica, so neither side ever acts on a step that exists in only one region.
Trace 2: a held transfer. A tier-1 account, 3 days old, sends 900.00 to a payee added an hour ago. The risk score says hold. The Try moves 900.00 from Alice's available account to her held account (one local transaction) and the API answers 202 HELD. An analyst reviews the case the next morning and releases it; the release is an ordinary saga from the held account to the payee, with the analyst's ID on the transfer.
Trace 3: a region failover.
Synthesizing vector architecture diagram...
The order matters: the old writers are fenced before the new ones open. Nothing lost in the tail had left the region, because outbound money waited for DR durability.
R3.6 Numbers and Cost
Traffic per home region (a planning split of 100M transfers a day; peak = 20 × average and ceiling = 4 × peak, as in Round 2)
| Region | Transfers/day | Average | Peak | Ceiling |
|---|---|---|---|---|
| Americas (us-east-1) | 40M | 463/s | 9,259/s | 37,037/s |
| EU (eu-central-1) | 35M | 405/s | 8,102/s | 32,407/s |
| APAC (ap-southeast-1) | 25M | 289/s | 5,787/s | 23,148/s |
| Total | 100M | 1,157/s |
The regional peaks fall at different hours of the UTC day, so the global peak is lower than the sum of the regional ones; each region is sized for its own.
Shards per region. At most 2.0 local transactions per transfer (half the traffic is local merchant payments; the other half costs up to 3 local transactions per transfer, in-region or cross-region), and the Round 2 rule: at the ceiling, on 2xlarge writers (4,000/s each), at no more than 70%:
| Region | Local transactions/s at the ceiling | ÷ 2,800 | Shards | Normal peak per shard on xlarge |
|---|---|---|---|---|
| Americas | 37,037 × 2 = 74,074 | 26.5 | 27 | 9,259 × 2 ÷ 27 = 686/s, 34% of 2,000 |
| EU | 32,407 × 2 = 64,815 | 23.1 | 24 | 8,102 × 2 ÷ 24 = 675/s, 34% |
| APAC | 23,148 × 2 = 46,296 | 16.5 | 17 | 5,787 × 2 ÷ 17 = 681/s, 34% |
| Total | 68 |
With a directory, the shard count doesn't need to be a power of two.
Compliance-check latency budget (same region, cross-shard, P99)
| Step | Runs | Budget |
|---|---|---|
| ALB, authentication, routing | first | 3 ms |
| Sanctions status, KYC tier, risk score | in parallel, so the slowest counts: max(3, 2, 15) | 15 ms |
| Try on the sender's shard (quote check and limits inside) | after the checks | 15 ms |
| Hop to the receiver's shard, Confirm, hop back | after the Try | 1 + 15 + 1 ms |
| Service overhead | 2 ms | |
| Total | 3 + 15 + 15 + 17 + 2 | 52 ms, inside the 60 ms target |
Cross-region transfer, every delay in the chain (typical, not P99): Try 15 ms, wait for the outbox row to show on the DR reader (replication lag, typically under 1 s, so budget 1,000 ms), plus up to 100 ms for the check interval, then the inter-region hop (roughly 45 ms each way between us-east-1 and eu-central-1; measure your own), Confirm 15 ms, the same DR wait on the receiving side (1,000 + 100 ms), and the reply 45 ms: ms. So cross-region transfers reply 202 PROCESSING and usually complete in about two seconds.
Availability, honestly. Five nines is 99.999% of a month: seconds. One Aurora writer failover typically takes up to 60 seconds, so a single failover on any shard would spend more than the month's budget for the wallets on it. A ledger with one writer per shard can't promise five nines for money movement, and we say so. What we promise: 99.99% per region for money movement, higher for balance reads served from the cache, and for a declared regional disaster a separate objective: RTO 1 hour, with no loss of money that left the region and any lost in-region seconds recovered from the failover snapshot.
Audit storage over 7 years. Assume the compressed daily export averages 150 bytes per transfer (the transfer and its entries in Parquet):
The newest 90 days stay in S3 Standard (≈ 1.35 TB × $0.023 ≈ $31/month); older files move by lifecycle rule to S3 Glacier Deep Archive (≈ 37 TB × $0.00099 ≈ $37/month at year 7), where an auditor's request is restored within hours. Object Lock retention stays in force when objects change storage class.
Online ledger storage. 400 days in Aurora, at 750 bytes per transfer: Americas 30 GB/day × 400 = 12 TB, EU 26.25 GB/day × 400 = 10.5 TB, APAC 18.75 GB/day × 400 = 7.5 TB: 30 TB, billed once per cluster, plus the same again in the DR regions' secondary clusters.
Fleet sizes, rounded per service so each survives an AZ loss (Round 2's planning figures: 1,000 reads/s and 250 transfers/s per vCPU, 2-vCPU tasks, 10 reads per transfer)
| Region | Read API tasks (per AZ × 3) | Write API tasks (per AZ × 3) | Total |
|---|---|---|---|
| Americas | 92,590 ÷ 1,000 = 93 vCPU → 47 tasks → 24 × 3 = 72 | 9,259 ÷ 250 = 38 vCPU → 19 tasks → 10 × 3 = 30 | 102 |
| EU | 81,020 → 82 vCPU → 41 tasks → 21 × 3 = 63 | 8,102 → 33 vCPU → 17 tasks → 9 × 3 = 27 | 90 |
| APAC | 57,870 → 58 vCPU → 29 tasks → 15 × 3 = 45 | 5,787 → 24 vCPU → 12 tasks → 6 × 3 = 18 | 63 |
| Total | 255 |
Monthly cost (us-east-1 on-demand prices for every line; EU and APAC Regions list roughly 10–20% higher for these instances, so budget about 10% more overall)
| Item | Math | Monthly |
|---|---|---|
Aurora primaries, 68 shards × 2 × db.r7g.xlarge, I/O-Optimized | 136 × $0.719/h × 730 h | ≈ $71,370 |
| Aurora storage, primaries | 30 TB × $0.225 | ≈ $6,750 |
Aurora DR secondaries, 68 × db.r7g.large reader, I/O-Optimized | 68 × $0.359/h × 730 h | ≈ $17,810 |
| Aurora storage, DR secondaries | 30 TB × $0.225 | ≈ $6,750 |
| Global Database replicated write I/O | assume ~10 write I/Os per transfer: 100M × 10 × 30.4 days ≈ 30.4B × $0.20 per million | ≈ $6,090 |
Valkey, 3 regions × 3 × cache.r7g.large (one shard, primary + 2 replicas) | 9 × $0.175/h × 730 h | ≈ $1,150 |
MSK, 15 × kafka.m7g.xlarge (6, 6, 3) + storage | 15 × $0.408/h × 730 h + ~3.2 TB × $0.10 | ≈ $4,780 |
| MSK Connect, 68 Debezium connectors × 1 MCU | 68 × $0.11/h × 730 h | ≈ $5,460 |
| Fargate API, 255 tasks × 2 vCPU, 4 GB, Arm | 255 × $0.079/h × 730 h | ≈ $14,710 |
| Fargate compliance, FX, risk, sagas, recon, export, ~60 tasks | 60 × $0.079/h × 730 h | ≈ $3,460 |
| S3 audit archive and lake | ≈ $200 | |
| CloudWatch, logs, alarms | ≈ $2,000 | |
| Total | ≈ $140,500/month at us-east-1 prices, ≈ $155,000 with about 10% regional uplift |
That's about 100M × 30.4 ≈ 3.04 billion transfers a month, so roughly $0.046 per thousand transfers at us-east-1 prices and $0.051 with the uplift, similar to Round 2. Sanctions-screening and identity-verification vendors charge per check and aren't in this AWS bill. The cache sizing: 60M active wallets × ~350 bytes ≈ 21 GB across all regions, about 7 GB per region, which fits one cache.r7g.large (13.07 GiB). Throughput fits too: the Americas' peak of about 93K reads/s spread over a primary and two replicas is about 31K reads/s per node. So one cache shard per region, three nodes, one per AZ. The replicated-write-I/O line is the least certain number here; we'd measure it from the clusters' write I/O metrics in the first week.
R3.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| Quote window length | 60 s | Longer windows are kinder to users and cost a wider spread (more market movement to cover); shorter ones make users re-quote. Some flows (scheduled payments) can't use a locked quote at all and must disclose "rate at execution". |
| Hold vs post-hoc review | Hold in the path for risky flows; post-hoc review for learning patterns | Held transfers delay legitimate customers; pure post-hoc review lets money leave before anyone looks. |
| Home region vs a global ledger | Home region per wallet, asynchronous DR, sagas between regions | A synchronously replicated multi-region database (Aurora DSQL multi-Region, Spanner) would give zero data loss on region failure, at the cost of an inter-region round trip on every write, contention on hot accounts (under Aurora DSQL's optimistic concurrency, conflicting writers abort at commit), and a residency story to solve per region. We keep it in mind for a small, critical table rather than the whole ledger. |
| Build vs buy | Build the ledger; partner for regulated banking | A banking-as-a-service partner can provide licensed accounts and payment rails in countries where we don't hold a licence, at a per-transaction cost and with their limits on our product. Purpose-built ledger databases also exist (TigerBeetle is one open-source example). What we can't outsource is the responsibility: in many countries holding customer money needs a licence (money-transmitter licences in US states, e-money licences in the EU and UK, for example), and the safeguarding and audit duties come with it. |
The opening question was: how do we move money between balances concurrently, at scale, so that every balance is right and every cent is accounted for? The answer is now: one writer per balance (a row lock in Round 1, a shard in Round 2, a home region in Round 3); money between writers always sits in a named account (in_transit, held accounts, payouts_pending), and every such account has a check that it nets to what's in flight; and proof outside the ledger: a daily reconciliation against real bank balances and a hash chain an auditor holds.
R3.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| A region outage mid-saga | Cross-region transfers stuck in PROCESSING; in-transit age alarms in the healthy region | If the recipient's region is down: the sender's region keeps retrying; nothing is cancelled on a timeout; money waits in in_transit. If the sender's region is down: every Try that reached another region was already durable in the DR region (step 3.4), so after promotion the DR region's saga worker finishes those transfers. |
| FX provider down | No fresh rates | The FX service refuses to quote from rates older than a few seconds, so new cross-currency transfers fail with a clear error; quotes already issued are still honoured until they expire, because their rate is locked. Same-currency transfers are unaffected. |
| A safeguarding break | The daily reconciliation doesn't balance | Same-day investigation. Treasury tops up the account from company funds while it's investigated if the break is a shortfall, and compliance decides whether it must be reported. |
| Compliance services down | Risk-score timeouts | Fail closed for risky flows (they're held, not rejected); low-risk domestic transfers continue on the last screening status and are re-scored later. |
| A quote used twice | Two transfers for one quote | Can't happen: the quote row is locked and marked used in the Try transaction; a retry with the same idempotency key returns the first result. |
| An audit file altered or missing | Daily verification fails: a hash or chain link doesn't match, or an account's entry_seq has a gap | Object Lock should make this impossible in the archive; a mismatch means the export itself was wrong. Re-export from the ledger, compare, and escalate as a security event. |
R3.9 Runbook and Incident Response
Golden signals, per region OPS 8 · REL 6
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Ledger imbalance (per transfer, per currency) | any non-zero | P1 | Freeze the affected accounts; stop the deploy that introduced the code path |
| In-transit age (cross-shard and cross-region) | any transfer > 5 min; any > 30 min | P2 / P1 | Check the receiving shard or region; re-drive the saga |
| In-transit net per transfer not in {0, amount} | any | P1 | A confirm and a cancel both happened, or a one-sided entry: freeze both wallets |
| Saga cancels | > 10/min per region | P2 | Look for a shard rejecting confirms in bulk |
| Stripe lock wait P99 | > 50 ms for 5 min | P3 | Raise the merchant's stripe count |
| Cache version lag | > 2 s for 5 min | P3 | Check the cache updater and CDC connector |
| Global Database replication lag | > 5 s for 2 min | P2 | Cross-region confirms and bank payouts slow down by design; check the primary's write load |
| Safeguarding break | any at the daily run | P1 for a shortfall | Treasury and compliance on the same day |
| Held-transfer queue | older than the review SLA | P3 | Add analysts; check whether a rule change is holding too much |
| Audit chain verification | any failure | P1 | Treat as a security event |
Stuck-saga procedure OPS 10
- Find it. The in-transit check lists the transfer ID, the sender's shard and region, and the age.
- Ask the receiver. Read the
inbound_transfersrow on the receiving shard. Confirmed: finish the sender's side (markCOMPLETED). Rejected: run the cancel. No row: continue. - Re-drive. Republish the saga command from the sender's outbox row; the worker retries.
- Give up only through the receiver. If the receiver is healthy but can't process it, ask it to "reject if not yet processed", then cancel. Never post a cancel by hand without the receiver's recorded rejection.
- Record it. Every manual step is itself a transfer with an idempotency key and the operator's identity.
Reconciliation-break procedure. Drill from the break to lines: for each system account in the safeguarding table, compare its movements during the day with the bank statement lines of the matching type. The unmatched item is either a missing ledger path (add a line and a ledger entry type) or a bank error (raise with the bank). Corrections are new transfers, approved by a second person.
Regional failover procedure. Follow Trace 3 in R3.5: stop routing, fence the old writers, fail over every shard's global cluster, scale up the promoted instances, start services, open writes, then reconcile the lost tail from the failover snapshots once the old region returns. Rehearse it twice a year per region. REL 12 · REL 13
Go deeper: CLI playbook
Plain commands an on-call engineer runs, one at a time. Replace the names and ARNs with real ones.
text# 1. Members and roles of one shard's cluster aws rds describe-db-clusters --region eu-central-1 --db-cluster-identifier eu-shard-08 --query "DBClusters[0].DBClusterMembers" # 2. Planned in-region failover to a bigger reader (event scale-up), one shard at a time aws rds failover-db-cluster --region eu-central-1 --db-cluster-identifier eu-shard-08 --target-db-instance-identifier eu-shard-08-2xl # 3. Replication lag from the EU primary to its DR secondary (published by the secondary cluster, in its region) aws cloudwatch get-metric-statistics --region eu-west-1 --namespace AWS/RDS --metric-name AuroraGlobalDBReplicationLag --dimensions Name=DBClusterIdentifier,Value=eu-shard-08-dr --start-time 2026-09-27T10:00:00Z --end-time 2026-09-27T10:15:00Z --period 60 --statistics Maximum # 4. Regional disaster only: promote the DR secondary, accepting the loss of unreplicated writes aws rds failover-global-cluster --region eu-west-1 --global-cluster-identifier eu-shard-08-global --target-db-cluster-identifier arn:aws:rds:eu-west-1:111122223333:cluster:eu-shard-08-dr --allow-data-loss # 5. After the old region returns: find the snapshot Aurora took of the old primary's storage aws rds describe-db-cluster-snapshots --region eu-central-1 --query "DBClusterSnapshots[?contains(DBClusterSnapshotIdentifier, 'unplanned-global-failover')].[DBClusterSnapshotIdentifier,SnapshotCreateTime]" # 6. Consumer lag of the cache updater on the ledger-entries topic aws cloudwatch get-metric-statistics --region eu-central-1 --namespace AWS/Kafka --metric-name MaxOffsetLag --dimensions Name="Cluster Name",Value=wallet-eu Name="Consumer Group",Value=cache-updater Name=Topic,Value=ledger-entries --start-time 2026-09-27T10:00:00Z --end-time 2026-09-27T10:15:00Z --period 60 --statistics Maximum # 7. Confirm an audit file is under compliance-mode retention aws s3api get-object-retention --region eu-central-1 --bucket wallet-audit-eu --key ledger/2026/09/26/shard-08.parquet # 8. Verify a day's signed chain head aws kms verify --region eu-central-1 --key-id alias/wallet-audit-chain --message fileb://head-2026-09-26.json --message-type RAW --signature fileb://head-2026-09-26.sig --signing-algorithm ECDSA_SHA_256
And the two checks every on-call engineer should be able to read, as short SQL over the lake (Athena):
sql-- Transfers whose in-transit entries don't net to 0 or to the amount in flight SELECT transfer_id, currency, SUM(amount_minor) AS in_transit_net FROM ledger_entries_lake WHERE account_role = 'IN_TRANSIT' GROUP BY transfer_id, currency HAVING SUM(amount_minor) <> 0 -- not finished AND SUM(amount_minor) <> MAX(amount_minor); -- and not exactly one Try in flight -- Any transfer whose entries don't sum to zero in some currency SELECT transfer_id, currency FROM ledger_entries_lake GROUP BY transfer_id, currency HAVING SUM(amount_minor) <> 0;
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | Home regions isolate failures; money leaving a region waits for DR durability; a fenced, rehearsed regional failover with lost-tail reconciliation; an honest 99.99% instead of an impossible five nines REL 10 · REL 12 · REL 13 |
| Performance Efficiency | Checks in parallel before the transaction and limits inside it (52 ms budget); regional fleets sized per region from the math; stripes and local merchant payments carried over PERF 1 · PERF 2 |
| Security | Identity data and ledger classified and kept in the home region; append-only ledger; S3 Object Lock and KMS-signed chain heads; step-up authentication and holds against account takeover; failover and tampering treated as security events SEC 4 · SEC 7 · SEC 8 · SEC 10 |
| Cost Optimization | About $155K a month with regional pricing, ≈ $0.05 per thousand transfers; audit archive moved to Deep Archive after 90 days; build the ledger, partner for licences where that's cheaper than holding one COST 4 · COST 11 |
| Operational Excellence | Compliance requirements (residency, safeguarding, retention) as design inputs; runbooks for stuck sagas, reconciliation breaks and regional failover; a daily reconciliation owner OPS 1 · OPS 10 · OPS 11 |
| Sustainability | Regions chosen by where users and their data must be, not duplicated everywhere; the audit archive in the coldest storage that meets the auditors' needs; compute scaled up only for known events SUS 1 · SUS 2 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Treats each currency as its own ledger and designs the FX legs so every currency balances; defines rounding once.
- Puts compliance in the transfer path with a held state that is still the customer's money, and decides where to fail closed.
- Knows the ledger balancing proves nothing about the bank, and designs the safeguarding reconciliation with one line per money path.
- Chooses one writer per wallet across regions, contains region failure by making outbound money wait for DR durability, and plans the lost-tail reconciliation.
- Gives auditors proof that doesn't depend on trusting the operators: a hash chain, write-once storage and a head held outside the team.
- Says no to five nines with arithmetic, and offers what can be promised instead.
- Keeps legal claims general and makes compliance the owner of the rules.
Follow-up questions
-
"Why not let the recipient's region do the FX conversion?" Answer: the sender approved a specific amount in her currency and a specific amount for Bob; that decision, the quote, and the debit must be one transaction on her shard. Converting later in Bob's region would need the quote there too and would split the decision across regions. Putting both currencies' legs on the sender's side keeps the quote check, the limit check and the debit atomic.
-
"A user moves from Germany to the US. Their wallet is in the EU region." Answer: a supervised migration: freeze the wallet briefly, move its balances with cross-region transfers from the EU accounts to new US accounts (so the ledger shows the move), copy the history the user may see, update
home_region, and keep the EU records where EU retention rules require them. It's rare enough to be a procedure, not a feature. -
"Auditors ask for every entry of one merchant in 2028, but 2028 is in Deep Archive." Answer: we restore those days' files (hours, not seconds), verify each file's hash against its chain record and the chain up to a signed head, then query them with Athena. The proof that nothing changed comes from the chain, not from trusting the restored copy.
Interview gotchas from this round
| Gotcha | Why it's wrong |
|---|---|
| "Convert at display time" | Balances change with the market; nobody can say what a customer owns. |
| "One transfer with a USD leg and a EUR leg that sums to zero" | Currencies can't be summed; each currency must balance on its own. |
| "Our ledger balances, so the money is there" | It proves nothing about the bank; only the reconciliation does. |
| "Active-active balances across regions" | Asynchronous replication lets the same money be spent in two regions. |
| "Five nines for transfers" | 26 seconds a month; one database failover spends it. |
| "Use QLDB for tamper evidence" | It reached end of support on 2025-07-31. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: top-ups, overdrafts, instant or not, history, one currency | Restate Round 1 in 60 seconds | Restate Round 2 in 60 seconds |
| 5–15 min | Requirements + API (integer cents, idempotency keys) | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.5: ledger, locks, lock order, idempotency, top-ups | Steps 2.1–2.6: shards, the saga, hot accounts, cache, CDC, withdrawals | Steps 3.1–3.6: FX, compliance, safeguarding, regions, audit, takeover |
| 40–50 min | Numbers + trade-offs (pessimistic vs optimistic) | Numbers (shards from the ceiling), cost, trade-offs (where the ledger lives) | Numbers, five nines honestly, cost, build vs buy |
| 50–60 min | Failures + pillar check | Failures + pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint. For the conversation with card processors and banks, see the payment processing interview loop.
The Two Sentences That Matter Most
- Opening a round: "Before I design, let me ask: can a balance ever go negative, and must a transfer be instant?"
- When the scope is raised: "Here's what breaks, and I'll fix anything that could lose or create money first, then latency and cost."
And the one sentence specific to this system: "Every balance has exactly one writer, every movement is entries that sum to zero, and money between writers always sits in an account we can add up."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What if the app retries a transfer?" (REL 4) | The idempotency key is inserted in the same transaction, so a retry returns the first result. | 1 | Step 1.4 |
| "What if the database fails over mid-transfer?" (REL 11) | Uncommitted work rolls back; committed Try steps survive with their outbox row and the saga finishes. | 1–2 | R1.9, R2.8 | |
| "What if the other shard never answers?" (REL 5) | The money waits in in_transit; we retry, and cancel only on the receiver's recorded rejection. | 2 | Step 2.2 | |
| "What if a region fails?" (REL 13) | Wallets have one home region; money that left it was already in the DR region; we promote after fencing and reconcile the tail. | 3 | Step 3.4 | |
| Performance | "How do you handle a merchant getting 5,000 payments a second?" (PERF 3) | Stripes on every shard, paid from the payer's shard, swept into one account that alone pays out. | 2 | Step 2.3 |
| "How do you serve 50,000 balance reads a second?" (PERF 3) | A cache that only moves forward by version, with min_version for the user's own writes. | 2 | Step 2.4 | |
| Cost | "What does it cost?" (COST 5) | About $530, $24,700 and $155,000 a month; about 4–5 cents per thousand transfers at scale. | 1–3 | R1.7, R2.6, R3.6 |
| "Build or buy?" (COST 11) | Build the ledger; partner for licensed banking where holding our own licence costs more. | 3 | R3.7 | |
| Operations | "How do you know every cent is accounted for?" (OPS 8) | Imbalance, in-transit age and in-transit net alarms, and a daily safeguarding reconciliation. | 2–3 | R2.10, Step 3.3 |
| "What are your compliance constraints?" (OPS 1) | Residency by home region, safeguarding by entity and currency, 7-year tamper-evident retention. | 3 | R3.1 | |
| Security | "Who can change the ledger?" (SEC 3) | Nobody can update or delete entries; services may only insert, and corrections are new transfers. | 1 | R1.6 |
| "How do you prove nothing was changed?" (SEC 8) | A daily hash chain in S3 Object Lock compliance mode, with a KMS-signed head held by the auditor. | 3 | Step 3.5 | |
| Sustainability | "Where is this system's footprint?" (SUS 4) | Mostly always-on databases; we keep 400 days online and move history to cold storage. | 2–3 | R2.6, R3.6 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Ledger | Double-entry, integer cents, balance cached in the same transaction. | System accounts per shard; in-transit and payout accounts; per-transfer net checks. | Per-currency balancing, FX accounts, and reconciliation against real bank balances. |
| Concurrency | FOR UPDATE at Read Committed, sorted lock order, a check constraint; can explain write skew. | Sagas instead of 2PC; never cancel on a timeout; stripes with a correct debit path. | One writer per wallet across regions; limits under the row lock; fenced regional promotion. |
| Idempotency | Key in the same transaction; stored responses, including rejections. | Every saga step idempotent by transfer ID; consumers deduplicate by version. | Manual repairs and re-applied lost-tail transfers go through idempotency keys too. |
| Reads and events | Balance from the account row. | Version-fenced cache; min_version; CDC keyed by account. | Regional caches and streams; no cross-region reads for money decisions. |
| Proof | A nightly check that entries sum to zero and match balances. | In-transit netting; bank-file reconciliation for withdrawals. | Safeguarding reconciliation; hash-chained, write-once audit; head held by the auditor. |
| Well-Architected trade-offs | Pessimistic vs optimistic; says the database is 3% busy and why that's fine. | Shards from the ceiling, instances from the peak; I/O-Optimized from the I/O pattern. | Says no to five nines with numbers; home region vs global ledger; build vs buy. |
| Evolving under new scope | Builds from a balance column, one problem at a time. | Opens with "what breaks" and fixes money-safety first. | Adds currencies, law and proof without weakening any earlier invariant. |