Design a Payment Processing System
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | A marketplace takes card payments through one payment provider | A payments platform: thousands of merchants, three providers, payouts to merchants' banks | Global money: 25 currencies, strong customer authentication, fraud, regulators |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Volume | 50K payments/day; ~10/s peak | 20M payments/day (231.5/s average); 2,500/s peak; 12,500 status reads/s | 200M payments/day; ~24K/s summed regional peaks |
| Storage | ~18 GB of payments and ledger a year | ~7.3 TB a year raw (36.5 TB over 5 years) | ~73 TB a year raw, across 11 ledger shards |
| Footprint | 1 region, 3 AZs | 1 region, 3 AZs; survives losing an AZ | 4 home regions, each with a disaster-recovery copy |
| Targets | No double charges; 99.9% | Charge P99 < 800 ms including the provider; 99.99%; books balance to the cent | Audit-ready; each region 99.99%; no double-spend through a region failover |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is a Payment System?
You Already Use One: the Checkout Button
You click Pay on a shopping site, and a second later the page says "Order confirmed". In that second, money did not move. What happened is that several companies that don't trust each other agreed that it will move, and each of them wrote that down.
| Party | Who it is | What it wants |
|---|---|---|
| Cardholder | The buyer | To pay once, and to get the money back if something goes wrong |
| Merchant | The shop (in Round 1, the marketplace) | To be paid, and never to ship goods for a payment that failed |
| Issuer | The buyer's bank, which issued the card | To approve only real, affordable purchases |
| Card network | Visa, Mastercard and others | To route messages between banks and enforce the rules |
| Acquirer | The merchant's bank for card payments | To collect card money for the merchant |
| PSP (payment service provider) | Stripe, Adyen and similar | One API that hides the acquirer, the networks and the paperwork |
Synthesizing vector architecture diagram...
The solid arrows happen in about a second, while the buyer waits. The dotted arrow happens days later. Everything in between is a promise that somebody has to keep track of.
Five words carry this whole loop. Each gets one line now:
- Authorize: the issuer checks the card and puts a hold on the amount. No money moves yet.
- Capture: the merchant says "take the held money now". Online shops often capture at shipping.
- Settle: the banks actually move the money, usually a day or more after capture, and the PSP pays the merchant, minus its fees.
- Refund: the merchant sends money back to the card on purpose.
- Chargeback (dispute): the buyer asks their bank to take the money back, and the bank does, often weeks later. The merchant can contest it with evidence.
What Makes It Hard
Three things make payments harder than most systems that just store data:
- The network lies. A request to the PSP can time out after the PSP charged the card. From our side, "no answer" looks exactly like "failed", but the money moved.
- Retries charge twice. The easy fix for a timeout, "try again", is exactly how a customer gets charged twice.
- Our books must match the bank's to the cent. A missing cent is not a rounding detail. It means money went somewhere we can't explain, and an auditor will ask.
So every component in this loop defends one of three rules:
| Rule | In one line |
|---|---|
| At most once | Each payment happens at most once, however many times anyone retries. |
| Every cent accounted for | Every movement of money is written down as debits that equal credits. |
| Our books match the bank's | Every day, we prove our records agree with what the PSPs and banks say happened. |
The Question the Whole Loop Answers
How do we make sure every payment happens exactly once in effect, and every cent is accounted for, when the network can't tell us what happened?
The answer gets sharper every round:
- Round 1: idempotency keys, a state machine, "never trust a timeout", and a double-entry ledger.
- Round 2: the machinery of a real processor: sagas, locking, an outbox, routing between providers, reconciliation, payouts and a card vault.
- Round 3: money across borders, rules and adversaries: sharded ledgers, hot accounts, 3-D Secure, fraud scoring, currencies, home regions and audit.
Round 1 · Mid-level · "Card Payments for One Marketplace"
~35 min · SDE II (L5) · 1 region, 3 AZs · ~10 payments/s peak · 99.9% · no double charges
R1.1 Establish Design Scope
The interviewer says: "We run an online marketplace. Design the part that takes the buyer's payment." Before we draw anything, we ask questions, and we say out loud what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Which payment methods? | Cards only, through one PSP (think Stripe). | One integration. We talk to one API, and the PSP talks to the card networks. |
| Do we store card numbers? | No. The PSP's hosted fields collect them. | Hosted fields are the PSP's card-entry boxes, loaded in an iframe on our page. Card numbers go from the browser straight to the PSP, and we get back a token. That keeps our servers out of most card-security rules (R1.8). |
| Refunds? | Yes, full and partial, started by our support team. | A second money-moving operation, with the same retry dangers as a charge. |
| Authorize now and capture later? Partial captures? | Not yet. We charge in full at checkout. | One call per payment: the PSP authorizes and captures together. |
| Payouts to sellers? | Not yet. Finance pays sellers weekly from a report. | We don't move money to sellers. But we must record exactly what we owe each seller. |
| How many payments, and how big? | About 50,000 a day. The average order is about $40, all in US dollars. | Small. Throughput won't be the hard part (R1.7). |
| How fast? | Checkout should complete within about a second. | The PSP's own time dominates. We must never make the buyer wait on our retries. |
Out of scope for this round:
- Several PSPs. One provider, one integration.
- Our own card storage. The PSP keeps every card.
- Other currencies and countries. US dollars only, but we store the currency anyway.
- Fraud checks. The PSP's built-in checks are enough for now.
The interviewer will widen this scope later. Write your out-of-scope list where you can see it: in a multi-round loop, some of it comes back.
R1.2 Functional Requirements, Derived Step by Step
We read the problem one phrase at a time and turn each phrase into a requirement:
| Phrase from the problem | Requirement |
|---|---|
| "Take the buyer's payment" | POST /v1/payments charges a card token for an order |
| "The buyer must know if it worked" | The result is final and correct: succeeded, failed, or "still confirming", never a guess |
| "Support can give money back" | POST /v1/payments/{id}/refunds, full or partial, never more than was charged |
| "Everyone asks: did this order get paid?" | GET /v1/payments/{id} returns the current status |
| "Finance pays sellers from a report" | Every payment and refund records what the seller is owed and what we earned |
Not yet: authorizing now and capturing later, payouts, and more than one PSP. The interviewer may bring these back.
R1.3 Non-Functional Requirements: the Questions
Numbers come in R1.7. For now, the questions, in the order that matters for money:
- Correctness. No double charges, and no lost payments: a buyer who was charged must get their order. This beats every other requirement. A slow payment is an annoyance; a double charge is a support ticket, a refund, and maybe a chargeback.
- Availability. If payments are down, the marketplace sells nothing. But "available" must never mean "guessing".
- Auditability. Finance must be able to explain every dollar: who paid, what we kept, what we owe each seller.
- Latency. About a second at checkout, most of which is the PSP.
R1.4 The API
Create a payment. The checkout makes one idempotency key per payment attempt: a random ID the client generates once and sends again on every retry of the same attempt, so the server can recognize the retry. It is not the order ID: if a card is declined and the buyer tries a different card, that's a new attempt with a new key.
httpPOST /v1/payments HTTP/1.1 Host: payments.internal.example.com Idempotency-Key: 5f1c2a9e-8d4b-4c61-9a3e-2b7f0d6c1e45 Content-Type: application/json { "order_id": "ord_20260927_000184", "amount_minor": 4000, "currency": "USD", "payment_method_token": "pm_tok_8GQ2x", "seller_id": "sel_7731", "description": "Order ord_20260927_000184" }
amount_minor is the amount in the currency's minor unit (cents for US dollars): 4000 means $40.00. Step 1.6 explains why it's an integer.
httpHTTP/1.1 201 Created Content-Type: application/json { "payment_id": "pay_01J9ZK3M7Q", "status": "SUCCEEDED", "amount_minor": 4000, "currency": "USD", "psp_reference": "pi_3Q8xYz", "created_at": "2026-09-27T14:32:01.458Z" }
What a retry with the same key returns. The same status code and the same body as the first time, plus a header Idempotent-Replayed: true. A retry never charges again.
| Status | When |
|---|---|
201 Created | The payment succeeded (first time) |
201 + Idempotent-Replayed: true | A retry of an attempt that already succeeded: the stored result |
202 Accepted with "status": "PROCESSING" | We don't know the outcome yet (step 1.2). Poll GET /v1/payments/{id} |
402 Payment Required | The card was declined (stored and replayed like a success) |
409 Conflict + Retry-After: 1 | The first request with this key is still running |
422 Unprocessable Entity | The same key was sent with a different body: a client bug, never a new payment |
503 Service Unavailable | The PSP is down; nothing was charged; try again later (R1.9) |
Refund:
httpPOST /v1/payments/pay_01J9ZK3M7Q/refunds HTTP/1.1 Host: payments.internal.example.com Idempotency-Key: 0b6e4d1a-7c3f-4e2a-8f5d-91c2e7a4b380 Content-Type: application/json { "amount_minor": 4000, "reason": "ITEM_NOT_RECEIVED" }
It returns 201 with a refund object ("status": "PENDING" until the PSP confirms, then SUCCEEDED or FAILED), and follows the same idempotency rules.
Status: GET /v1/payments/{id} returns the payment with its refunds.
How the card gets to the PSP without touching us:
Synthesizing vector architecture diagram...
The card number crosses from the buyer's browser to the PSP and nowhere else. Our servers only ever see a token that is useless outside our PSP account.
Recap
- One create call with an idempotency key; one refund call; one status call.
- A retry returns the stored result; a different body with the same key is an error.
- "We don't know yet" is a real answer (
202 PROCESSING), not a failure. - Card numbers never touch our servers.
- Correctness first: at most one charge per attempt, and every dollar written down.
Let's build it, starting with the simplest thing that works.
R1.5 Design Evolution: From "Call the PSP" to a Ledger
Every step below follows the same pattern: a problem, your turn to think, the answer, and what the answer costs us. The cost is always the next problem.
Step 1.0: The Baseline
The checkout calls the PSP. If the PSP says "succeeded", we write paid = true on the order.
Synthesizing vector architecture diagram...
What's good about it: it's two lines of code, and on a good day it works.
What it costs us: it trusts every message to arrive exactly once, and every answer to come back. Neither is true. The next six steps are the ways it breaks.
Step 1.1: The Buyer Was Charged Twice
The problem: a buyer double-clicked Pay. Two requests reached the checkout service a few milliseconds apart, each called the PSP, and the buyer was charged $40 twice. What would you do?
The key table, in the same Aurora PostgreSQL database as the payments (so claiming the key and creating the payment can share one transaction):
sqlCREATE TABLE idempotency_keys ( scope_key TEXT PRIMARY KEY, -- 'payments:5f1c2a9e-...' request_hash BYTEA NOT NULL, -- SHA-256 of the request body status TEXT NOT NULL CHECK (status IN ('IN_FLIGHT', 'COMPLETED')), locked_until TIMESTAMPTZ NOT NULL, -- set from the database clock: now() + 60 s payment_id TEXT, -- filled once known response_code INT, response_body JSONB, created_at TIMESTAMPTZ NOT NULL DEFAULT now() );
The claim is one statement:
sqlINSERT INTO idempotency_keys (scope_key, request_hash, status, locked_until) VALUES ($1, $2, 'IN_FLIGHT', now() + interval '60 seconds') ON CONFLICT (scope_key) DO NOTHING;
If it inserted zero rows, somebody else holds the key, and we read the row to decide between replay, 409 and 422. The lock time comes from the database's clock (now()), so two app servers with slightly different clocks can't disagree about whether a lock expired. A request that finds an IN_FLIGHT row whose locked_until has passed takes over by updating the row only if locked_until still has the value it read, and then checks whether a payment already exists before doing anything (step 1.2 explains what it does next). A nightly job deletes rows older than 7 days, in batches of a few thousand so it never locks the table for long.
Step 1.2: The PSP Call Timed Out. Did It Charge?
The problem: we sent the charge to the PSP. After 5 seconds, our HTTP client gave up. We don't know whether the card was charged. What would you do?
Why 5 seconds? The PSP's normal P99 is 300 to 600 ms, so 5 seconds is far above normal and a timeout really means trouble. A shorter timeout would turn slow-but-successful charges into unknowns we then have to resolve; a much longer one leaves the buyer staring at a spinner.
Synthesizing vector architecture diagram...
The replay is safe only because the key is the same. It does not charge again; it reads the first request's saved result.
Primitive: Circuit Breaker, Bulkhead and Fault-Tolerance Patterns (timeouts, retries and backoff with jitter)
Step 1.3: The Statuses Are a Mess of Booleans
The problem: the payments table has paid, failed, refunded and pending_check columns. Support found a payment with paid = true and failed = true, and another that was refunded twice.
What would you do?
Synthesizing vector architecture diagram...
There is no arrow out of FAILED and no arrow from UNKNOWN to anything but a confirmed answer. The only way to leave UNKNOWN is for the PSP to tell us.
Step 1.4: The PSP Tells Us Things Later
The problem: a refund we asked for finished two days later. A buyer disputed a charge three weeks later. A payment we marked UNKNOWN was actually charged, and the PSP knows. How do we hear about these?
What would you do?
Disputes after refunds are allowed. It's tempting to write "a refunded payment ignores dispute events". That's wrong. A buyer can still dispute a partially refunded charge, and Stripe notes that a dispute can arrive while a refund is still on its way (it fails such a refund with the reason charge_for_pending_refund_disputed). So DISPUTED must be reachable from SUCCEEDED, PARTIALLY_REFUNDED and REFUNDED, and the dispute handling must check how much was already refunded so the buyer isn't paid back twice. Round 2 builds the dispute flow; in Round 1 we record the event and alert support.
| Event arrives | Current status | What we do |
|---|---|---|
payment_intent.succeeded | UNKNOWN or PROCESSING | Move to SUCCEEDED |
payment_intent.succeeded | SUCCEEDED | Duplicate in effect: nothing to do |
payment_intent.payment_failed | UNKNOWN or PROCESSING | Move to FAILED |
refund.updated with refund status succeeded | SUCCEEDED or PARTIALLY_REFUNDED | Mark the refund SUCCEEDED, add to refunded_minor, post the refund journal |
refund.updated with refund status succeeded | UNKNOWN | Resolve the payment with the PSP first, then apply |
refund.failed (can come weeks later: Stripe says a failed refund's money can take up to 30 days to return) | refund SUCCEEDED | Mark the refund FAILED, subtract it from refunded_minor, move the payment back (REFUNDED to PARTIALLY_REFUNDED or SUCCEEDED), post a reversing journal, and alert support to repay the buyer another way |
refund.failed | refund PENDING | Mark the refund FAILED; no journal was posted, so none to reverse |
We confirm a refund from the refund's own status (refund.updated moving to succeeded), not from charge.refunded: Stripe sends charge.refunded when a refund is created, before the refund has actually gone through.
| charge.dispute.created | SUCCEEDED, PARTIALLY_REFUNDED or REFUNDED | Record the dispute, alert support |
Step 1.5: Finance Can't Tell Where the Money Went
The problem: finance asks: "How much do we owe seller 7731 right now, and why?" We have a balance column on the sellers table that every payment adds to and every refund subtracts from. Last month a bug ran one refund twice, and the balance is $36 too low. Nobody can say which payment it came from.
What would you do?
Debits and credits in one paragraph. Assets (things we have or are owed, like psp_receivable) go up with a debit. Liabilities (things we owe, like seller_payable) and revenue go up with a credit. Expenses (like psp_fees) go up with a debit. So "the PSP owes us more" is a debit, and "we owe the seller more" is a credit, and in every transaction the two sides are equal.
One $40.00 payment, with the PSP's fee of 2.9% + 30¢ ($1.16 + $0.30 = $1.46) and our 10% commission ($4.00):
| Account | Type | Debit | Credit |
|---|---|---|---|
psp_receivable | Asset | $38.54 | |
psp_fees | Expense | $1.46 | |
seller_payable:sel_7731 | Liability | $36.00 | |
platform_revenue | Revenue | $4.00 | |
| Total | $40.00 | $40.00 |
The PSP keeps its fee and will pay us $38.54; we owe the seller $36.00 and keep $4.00. The fee is the amount the PSP reports in its response, never one we compute ourselves: our entries must match what the PSP actually does.
A full refund of that payment. PSPs generally don't return their processing fee on a refund (Stripe says so), and our policy (a choice) is to return our commission:
| Account | Type | Debit | Credit |
|---|---|---|---|
seller_payable:sel_7731 | Liability | $36.00 | |
platform_revenue | Revenue | $4.00 | |
psp_receivable | Asset | $40.00 | |
| Total | $40.00 | $40.00 |
After both: the seller is owed $0, our revenue is $0, and psp_receivable is down $1.46 net, which is exactly the fee the PSP kept, still visible in psp_fees. Every cent is explained.
sqlCREATE TABLE accounts ( account_id BIGSERIAL PRIMARY KEY, account_code TEXT UNIQUE NOT NULL, -- 'seller_payable:sel_7731' account_type TEXT NOT NULL CHECK (account_type IN ('ASSET','LIABILITY','REVENUE','EXPENSE','EQUITY')), currency CHAR(3) NOT NULL, created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE journal_transactions ( journal_id BIGSERIAL PRIMARY KEY, journal_key TEXT UNIQUE NOT NULL, -- 'payment:pay_01J9ZK3M7Q:capture'; a repeat post fails here payment_id TEXT, kind TEXT NOT NULL, -- 'CAPTURE', 'REFUND', 'DISPUTE', 'REVERSAL' posted_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE ledger_entries ( entry_id BIGSERIAL PRIMARY KEY, journal_id BIGINT NOT NULL REFERENCES journal_transactions(journal_id), account_id BIGINT NOT NULL REFERENCES accounts(account_id), direction CHAR(1) NOT NULL CHECK (direction IN ('D','C')), amount_micros BIGINT NOT NULL CHECK (amount_micros > 0), currency CHAR(3) NOT NULL ); CREATE INDEX ON ledger_entries (account_id, entry_id); CREATE INDEX ON ledger_entries (journal_id);
Two rules make the ledger trustworthy, and both are enforced by the database, not by hoping every code path is right:
- Balanced journals. A deferred constraint trigger runs at commit and rejects the transaction if, for any journal it touched, debits and credits differ in any currency.
- Append-only. The application's database role may only insert into the ledger tables:
REVOKE UPDATE, DELETE ON ledger_entries, journal_transactions FROM payments_app;
A balance is a query:
sqlSELECT SUM(CASE WHEN direction = 'C' THEN amount_micros ELSE -amount_micros END) AS owed_micros FROM ledger_entries WHERE account_id = $1; -- a liability account: credits minus debits
Primitive: Database Isolation Levels, ACID and Concurrency Anomalies (why the whole journal commits or nothing does)
Step 1.6: The Amounts Are Off by a Cent
The problem: a report adds up a day of payments and gets $1,999,999.9999998. Another report splits a $10.00 bundle between three sellers and pays out $10.01. What would you do?
A worked example: a $33.33 order, for a seller whose contract says the seller bears the PSP fee. Our 10% commission is 333.3 cents; half-even to the cent gives $3.33. The PSP reports its fee: 2.9% × $33.33 + $0.30 = $1.26657, which it rounds to $1.27. We record the fee the PSP reports, compute our commission, and derive the seller's share by subtraction: $33.33 − $1.27 − $3.33 = $28.73.
| Account | Debit | Credit |
|---|---|---|
psp_receivable | $32.06 | |
psp_fees | $1.27 | |
seller_payable:sel_7731 | $28.73 | |
platform_revenue | $3.33 | |
psp_fee_recovery (fee charged on to the seller) | $1.27 | |
| Total | $33.33 | $33.33 |
The seller's line is derived by subtraction, so the journal balances by construction. psp_fee_recovery offsets psp_fees, so our net fee cost for this order is zero, and both facts stay visible.
(Who bears the PSP fee is a business rule. In the $40 example above, we absorbed it; here the seller does. Pick one rule per contract and write it in the journal code, not in each report.)
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | Checkout calls the PSP, sets paid | Trusts every message and every answer |
| 1.1 | Double click charged twice | Idempotency key claimed with a conditional insert; stored response; PSP key per charge | A key store with a 7-day window |
| 1.2 | Timeout: did it charge? | UNKNOWN state; resolver asks the PSP (replay with the same key, or look up), backoff with full jitter; review queue after 15 min | Some payments resolve later |
| 1.3 | Flags contradict each other | One status, allowed moves enforced by compare-and-set; payment_events history | Every write names its allowed "from" states |
| 1.4 | The PSP tells us things later | Webhooks: HMAC signature, 300 s replay window, dedupe by event ID, allowed moves only | A public endpoint to secure |
| 1.5 | Where did the money go? | Double-entry ledger; balances derived; append-only; balanced journals enforced at commit | More rows; a new way of thinking |
| 1.6 | Off by a cent | Integer micros in the ledger, minor units in the API, half-even rounding, splits that add up | Discipline in every money code path |
R1.6 Architecture v1
Now the concepts get AWS names.
Synthesizing vector architecture diagram...
Follow a payment: the browser sends only a token; WAF filters abuse; the service claims the key, creates the payment and calls the PSP; the result and the ledger entries commit in one Aurora transaction. Webhooks come in through the same front door and change payments through the same allowed-moves rules. The resolver works only on payments stuck in UNKNOWN.
The pieces:
- Payment service: a stateless container on ECS Fargate (containers without managing servers), three tasks, one per AZ. It serves the API and the webhook endpoint.
- Aurora PostgreSQL: one writer and one replica in another AZ, which Aurora promotes if the writer fails. It holds the key table, payments, refunds, webhook events and the ledger. One database means one transaction can cover "claim key, create payment, post ledger".
- Resolver: a scheduled Fargate task that runs every 30 seconds, picks payments that have been in
UNKNOWNorPROCESSINGfor more than 10 seconds, and asks the PSP with the backoff from step 1.2 (it stores each payment's next try time, so a run skips payments that aren't due). It also runs the nightly key cleanup. - Secrets: the PSP API key and the webhook signing secret live in AWS Secrets Manager, read at start-up.
- Egress: calls to the PSP leave through NAT gateways (one per AZ).
The remaining tables:
sqlCREATE TABLE payments ( payment_id TEXT PRIMARY KEY, -- 'pay_01J9ZK3M7Q' scope_key TEXT UNIQUE NOT NULL, -- permanent guard; key rows are deleted after 7 days order_id TEXT NOT NULL, seller_id TEXT NOT NULL, amount_minor BIGINT NOT NULL CHECK (amount_minor > 0), refunded_minor BIGINT NOT NULL DEFAULT 0, currency CHAR(3) NOT NULL, status TEXT NOT NULL, psp_reference TEXT, next_check_at TIMESTAMPTZ, -- resolver schedule while UNKNOWN version INT NOT NULL DEFAULT 0, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), updated_at TIMESTAMPTZ NOT NULL DEFAULT now(), CHECK (refunded_minor >= 0 AND refunded_minor <= amount_minor) ); CREATE TABLE refunds ( refund_id TEXT PRIMARY KEY, payment_id TEXT NOT NULL REFERENCES payments(payment_id), scope_key TEXT UNIQUE NOT NULL, amount_minor BIGINT NOT NULL CHECK (amount_minor > 0), status TEXT NOT NULL CHECK (status IN ('PENDING','SUCCEEDED','FAILED','UNKNOWN')), psp_reference TEXT, created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE payment_events ( event_id BIGSERIAL PRIMARY KEY, payment_id TEXT NOT NULL REFERENCES payments(payment_id), from_status TEXT, to_status TEXT NOT NULL, cause TEXT NOT NULL, -- 'api', 'resolver', 'webhook:evt_1Q...' created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE TABLE webhook_events ( psp_event_id TEXT PRIMARY KEY, -- dedupe event_type TEXT NOT NULL, received_at TIMESTAMPTZ NOT NULL DEFAULT now(), payload JSONB NOT NULL );
Trace 1: a successful payment.
Synthesizing vector architecture diagram...
The payment row exists before the PSP is called. If our service dies between the PSP call and the second transaction, the resolver finds a PROCESSING payment and asks the PSP; nothing is lost.
Trace 2: a timeout resolved by asking. The PSP charges, but the reply never arrives. At 5 s the service sets UNKNOWN, sets next_check_at, and returns 202 PROCESSING; the page starts polling. On its next run, the resolver replays the charge with the same PSP key and gets the saved "succeeded". In one transaction it moves the payment to SUCCEEDED, posts the journal, and completes the key with a 201 response, so a buyer's retry now replays 201. The page's next poll shows "Payment confirmed". The PSP's payment_intent.succeeded webhook arrives later and finds nothing left to do.
Trace 3: a refund confirmed by webhook.
Synthesizing vector architecture diagram...
The refund journal is posted when the PSP confirms the refund, not when we ask for it, because until then the money hasn't moved. If a refund.failed arrives later, a reversing journal undoes it and the payment moves back.
R1.7 Numbers
Traffic.
For the peak we assume (and would confirm with the interviewer) that 30% of the day's payments land in the busiest two hours, and that a promotion can multiply that five-fold for a few minutes:
Each payment is about ten row writes (key claim, payment insert, status update, two history events, journal, four entries, key completion, rounded), so the peak is about 100 row writes a second. One small database does that easily.
Ledger rows per day. Four entries per payment, and we assume 5% of payments get a refund journal of three entries:
Storage. About 1 KB per payment all in: the payment row (~350 bytes), four ledger entries (~100 bytes each, ~400 bytes), and events and journal header (~250 bytes).
With indexes (about 1.6× the raw data, an assumption), about 29 GB a year.
Idempotency keys. About 52,500 a day (payments plus refunds), kept 7 days, at up to ~1 KB each with the stored response:
Availability. 99.9% of a 30.4-day month is minutes. Note that the PSP is in our path: if the PSP is down, we are down for new payments, whatever our own uptime.
Latency. Dependent steps, so they add: key claim and payment insert (~5 ms), PSP call (300 to 600 ms at P99), commit (~5 ms), network inside the region (~10 ms). That's about 620 ms at P99, inside the "about a second" target, and almost all of it is the PSP.
Monthly cost (us-east-1 on-demand prices, 730 hours a month):
| Item | Math | Monthly |
|---|---|---|
Aurora PostgreSQL, 2 × db.r6g.large | 2 × $0.26/h × 730 h | ≈ $380 |
| Aurora storage and I/O | ~30 GB × $0.10/GB-month, plus a few dollars of I/O | ≈ $10 |
| Fargate, 3 tasks of 0.5 vCPU / 1 GB | 3 × (0.5 × $0.04048 + 1 × $0.004445)/h × 730 h | ≈ $54 |
| Application Load Balancer | $0.0225/h × 730 h, plus a few capacity units | ≈ $22 |
| AWS WAF | $5 per web ACL + ~5 rules × $1 + a few million requests × $0.60/M | ≈ $12 |
| NAT gateways, 3 | 3 × $0.045/h × 730 h | ≈ $99 |
| Secrets Manager, CloudWatch | 3 secrets × $0.40, logs, metrics, alarms | ≈ $20 |
| Total | ≈ $600/month |
Now compare with what the PSP charges. At a common published card rate of 2.9% + 30¢ and our $40 average order, each payment costs $1.46 in fees:
Say it in the interview: the infrastructure is about 0.03% of the payment fees. Saving money on servers is irrelevant here. Every design choice in this round is about correctness.
R1.8 Trade-Offs
| Choice | Option A | Option B | Our pick for Round 1 |
|---|---|---|---|
| Card data | Hosted fields: the PSP's iframe collects the card; we get a token | Handle cards ourselves: card numbers pass through our servers | Hosted fields. The PCI DSS (the card industry's data security standard) applies to any system that stores, processes or transmits card numbers. With hosted fields from a PCI-compliant PSP, a merchant can often use the shortest self-assessment questionnaire (SAQ A), as long as the checkout page is protected against malicious scripts. Handling cards ourselves puts our servers, network and staff into full PCI scope. Confirm the exact category with a qualified assessor. |
| Ledger store | Relational (Aurora PostgreSQL) | NoSQL (e.g. DynamoDB) | Relational. A journal is several rows that must commit together, with foreign keys, check constraints and a balanced-journal check at commit. DynamoDB transactions can write up to 100 items atomically, but constraints like "debits equal credits" and ad-hoc finance queries are natural in SQL. At 10 payments a second, scale doesn't push us anywhere else. |
| Confirmation | Synchronous: the buyer waits for the PSP's answer | Asynchronous: accept the order, charge in the background | Synchronous, with an asynchronous fallback (202 PROCESSING) only when the answer is unknown. Charging in the background means telling buyers "order placed" and emailing "your card was declined" later, and the goods may already be on their way. |
| Where keys live | Same database as payments | Separate key store | Same database. One transaction covers key and payment, so they can never disagree. Round 2 revisits this at 2,500 payments a second. |
R1.9 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| The PSP is down | Connection errors or 5xx from the PSP; error-rate alarm | Fail the checkout clearly: 503, "Payments are temporarily unavailable, you have not been charged". Don't queue charges to run later. A queued charge surprises the buyer hours later, may be declined when they're gone, and the order may already have shipped. Where the error proves the PSP never received the request (a refused connection, say), we're sure nothing was charged; where it doesn't, the payment is UNKNOWN and the resolver asks later. |
| The database fails over mid-payment | A burst of connection errors while Aurora promotes the replica (typically tens of seconds) | Two cases. Before the PSP call: the transaction rolled back, nothing was charged, and the client's retry with the same key starts cleanly. After the PSP call but before our commit: the payment row still says PROCESSING, so the resolver asks the PSP and finishes it. A client retry in the meantime finds the key IN_FLIGHT and gets 409, then the stored result. |
| A service task crashes mid-request | One request's connection resets | Same as the second case above: the key lock expires after 60 s (database clock), and the resolver finishes the payment. |
| Duplicate webhooks | The same event ID twice | The second insert into webhook_events fails on the primary key; we answer 200 and change nothing. |
| A forged or replayed webhook | Signature mismatch or a timestamp older than 300 s | Reject with 400; alarm on the rate, since a burst means someone is probing the endpoint. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Idempotency keys make every retry safe; UNKNOWN plus a resolver instead of guessing; allowed-moves state machine; Aurora replica in a second AZ, service tasks in three REL 5 · REL 11 |
| Security | Card numbers never touch us (hosted fields, smaller PCI scope); PSP key and webhook secret in Secrets Manager; signed webhooks with a 300 s replay window; WAF in front; the app's database role can't update or delete ledger rows SEC 3 · SEC 7 · SEC 9 |
| Performance Efficiency | P99 ≈ 620 ms, dominated by the PSP; our own work is a few indexed writes PERF 1 |
| Cost Optimization | About $600 a month, derived, and about 0.03% of the PSP fees, which is the real cost of payments COST 5 |
| Operational Excellence | Light this round: alarms on PSP error rate, UNKNOWN count and age, webhook signature failures, and the review queue OPS 8 |
| Sustainability | Light this round: three small tasks and two small database instances; nothing idles at scale SUS 2 |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Puts correctness first and says so: no double charge, no lost payment.
- Uses an idempotency key with an atomic claim, stores the response, and passes a key to the PSP.
- Treats a timeout as "unknown", never as "failed", and knows how to find out: replay with the same key inside the PSP's window, or look up.
- Models the payment as a state machine with allowed moves, enforced where the data lives.
- Verifies webhooks, dedupes them, and doesn't assume their order.
- Writes money movements as balanced double entries in integer units.
Follow-up questions
-
"Why not use the order ID as the idempotency key?" Answer: because an order can have several legitimate payment attempts. If the first card is declined and the buyer tries another, an order-ID key would replay the decline forever. The key identifies one attempt; the order can have many, and the service can still refuse a second successful payment for one order by checking the order's paid state inside the same transaction.
-
"The buyer closes the tab while the payment is
UNKNOWN. What happens?" Answer: nothing depends on the tab. The resolver finishes the payment, the order moves to paid or cancelled, and the buyer gets an email. If they come back and click Pay again, the checkout reuses the same key for the same attempt (it keeps the key with the order attempt), so they get the stored result instead of a second charge. -
"Why store the response instead of just a 'done' flag?" Answer: because a retry must get the same answer as the first request, including the payment ID and the PSP reference. With only a flag, the retry would have to rebuild the answer, and anything that changed since (a refund, say) would make it different.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Disable the button" as the fix for double charges | Clients, proxies and load balancers retry on their own; the server must be idempotent. |
| "Check if paid, then charge" | Two requests pass the check together; the check and the claim must be one atomic write. |
| "On timeout, mark failed" | The PSP may have charged; the buyer pays and gets nothing, then pays again. |
| "Retry the PSP call with the same key, any time" | PSPs forget keys (Stripe after at least 24 hours); a late replay can be a new charge. |
| "Refunded payments ignore dispute events" | Buyers can dispute after a partial refund, or while a refund is on its way. |
"Store money as FLOAT" | Binary fractions drift; use integer minor units or micros with a named rounding rule. |
Round 2 · Senior · "20M Payments a Day, Many PSPs, Payouts"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 2,500 payments/s peak · 12,500 reads/s · 99.99% · the books balance to the cent
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built card payments for one marketplace: 50,000 payments a day, about 10 a second at peak, one PSP, US dollars. The browser sends the card to the PSP through hosted fields, so we only see a token. Every request carries an idempotency key; we claim it with a conditional insert in the same Aurora PostgreSQL database as the payment, store the response, and replay it on retries. We pass our own key to the PSP. A timeout is never a failure: the payment goes to
UNKNOWN, the buyer sees 'confirming', and a resolver asks the PSP by replaying with the same key inside the PSP's key window, or by looking the payment up, with backoff and full jitter. Statuses are one state machine with allowed moves enforced by compare-and-set. Webhooks are verified with HMAC, rejected after 300 seconds, deduped by event ID and applied only as allowed moves, and a dispute can follow a refund. Money is a double-entry ledger in integer micros: balanced journals, derived balances, append-only. About $600 a month, against $2.2 million a month in PSP fees. Open costs: one PSP, one database for everything, and nobody checks our books against the PSP's."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: one database holds keys, payments and the ledger, so one transaction keeps them consistent; the resolver turns unknowns into answers.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Double click, double charge | Idempotency key, conditional insert, stored response, PSP key | Key store, 7-day window |
| 1.2 | Timeout: did it charge? | UNKNOWN + resolver (replay same key, or look up) | Late answers |
| 1.3 | Contradicting flags | State machine, compare-and-set | Every write names allowed moves |
| 1.4 | PSP tells us later | Webhooks: HMAC, 300 s window, dedupe, allowed moves | Public endpoint |
| 1.5 | Where did the money go? | Double-entry ledger, append-only | More rows |
| 1.6 | Off by a cent | Integer micros, half-even, splits that add up | Discipline |
Open costs: one PSP is a single point of failure; one database does everything; nothing proves our books match the PSP's and the bank's.
R2.1 The Scope Raise
Interviewer: "The marketplace worked so well that we're turning it into a payments platform. Twenty thousand merchants, 20 million payments a day, 2,500 a second at peak. Many merchants want to authorize at checkout and capture when they ship. We'll use three PSPs, routed by cost and approval rate, and we must keep taking payments when one of them is failing. We now pay merchants out to their bank accounts ourselves. Finance wants the nightly bank and PSP files to match our ledger to the cent. We want our own card vault so we're not locked into one PSP. Disputes are ours to handle. And customers now have a stored wallet balance, which they spend from their phone and their laptop."
As in Round 1, we ask back before fixing anything.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How long between authorize and capture? | Usually one to five days; a few merchants wait longer. | Card authorizations expire. Stripe documents about 7 days for most online card payments (5 for some Visa merchant-initiated ones). A payment's life now spans days, so its workflow must survive restarts and deploys (step 2.1). |
| What share of payments capture later? | About 40%. The rest capture at once. | Only 40% need a multi-day workflow. That drives the orchestration bill (R2.6). |
| Can every PSP charge every saved card? | No. A card saved at one PSP is a token that only that PSP understands. | Failover and routing are impossible for saved cards unless we hold the cards ourselves (step 2.7). |
| When and how are merchants paid? | Daily, by bank transfer (ACH in the US), for money we've actually received and checked. | We move money out, through a bank rail with its own delays and returns (step 2.6). |
| Which files arrive, and when? | Each PSP sends a daily settlement report; our bank sends a statement. Both by 06:00. | A nightly three-way match between our ledger, the PSPs and the bank (step 2.5). |
| Who else needs to know when money moves? | Merchant notifications, analytics, risk, finance. | Every ledger change must be published reliably, with no gaps (step 2.3). |
| How much can a wallet go negative? | Never. | Concurrent spends from two devices must not overdraw it (step 2.2). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Merchants | One marketplace | 20,000 merchants |
| Traffic | 50K/day, ~10/s peak | 20M/day (231.5/s average), 2,500/s peak; 12,500 status reads/s |
| Flows | Charge, refund | + authorize then capture (40%), cancel, payouts, disputes, wallet transfers |
| PSPs | One | Three, routed and failed over |
| Cards | PSP tokens | Our own vault |
| Books | A ledger | + nightly reconciliation to the cent against PSPs and bank |
| Data | ~18 GB/year | ~7.3 TB/year raw; 36.5 TB over 5 years |
| Survive | A task or a DB failover | An AZ, and a PSP that fails slowly |
| Availability | 99.9% (43.8 min/month) | 99.99% (4.4 min/month) for our API |
| Latency | ~1 s | Charge P99 < 800 ms including the PSP |
The "Not yet" list from R1.2 is now mandatory: authorize and capture, payouts, and several PSPs.
About "availability": we promise 99.99% for our API. We can't promise that end to end, because PSPs, card networks and issuers are in the path and publish their own targets. We measure our own availability separately from each PSP's, and routing (step 2.4) is how we stay up when one PSP isn't.
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | What breaks at the new scope |
|---|---|
| One PSP | When it fails or slows down, every payment fails. It's also the only price we pay. |
| The payment service calls the PSP and writes the ledger in one request | Authorize now, capture in three days, post the ledger, notify the merchant: five steps across three systems over days. A request can't hold that, and any step can fail in the middle. |
| No reconciliation | At 20M payments a day, a 0.01% mismatch is 2,000 payments a day that nobody has explained, and we're about to pay merchants from these balances. |
| Wallet balance checked, then debited | Two devices spend at once, both see enough money, both succeed. |
| Card tokens live at the PSP | We can't send a saved card to a different PSP, so routing and failover don't work for repeat customers. |
| Webhooks processed inline, in order of arrival | Three PSPs' webhooks, out of order, at thousands a second; a dispute can arrive before or after a refund. |
| Keys, payments, ledger and 12,500 reads/s on one Aurora writer | Retry storms and status reads compete with ledger writes on the one machine that can't scale out. |
| "We'll publish events after commit" (not built yet) | Downstream teams need every ledger change; a crash between commit and publish loses events silently. |
R2.3 New Requirements and API Additions
Authorize now, capture later
httpPOST /v1/payments HTTP/1.1 Host: api.payments.example.com Authorization: Bearer <merchant API key> Idempotency-Key: 7a0c4b2e-1f6d-4c8a-b3e9-5d2f8e1a6c07 Content-Type: application/json { "amount_minor": 12999, "currency": "USD", "payment_method": "vtok_9f3kQ2", "capture_method": "MANUAL", "merchant_reference": "order-551-2207" }
It returns 201 with "status": "REQUIRES_CAPTURE" and "capture_before": "2026-10-04T09:12:00Z", the deadline after which the authorization lapses (taken from the PSP's answer, never assumed).
httpPOST /v1/payments/pay_01JA2B/capture HTTP/1.1 Idempotency-Key: 3e8d1c7a-cap-2207 Content-Type: application/json { "amount_minor": 11999 }
A capture may be for less than the authorized amount; the rest is released. For most card payments you can capture only once (Stripe documents this), so the API rejects a second capture with 409. POST /v1/payments/{id}/cancel releases an uncaptured authorization.
Routing rules (an admin API, versioned, read by the router every 30 seconds):
json{ "version": 42, "rules": [ { "match": { "card_country": "US", "card_type": "debit" }, "prefer": ["PSP_B", "PSP_A"] }, { "match": { "card_country": "*" }, "prefer": ["PSP_A", "PSP_B", "PSP_C"] } ], "min_approval_rate_to_stay_primary": 0.85 }
Payouts: GET /v1/merchants/{id}/payouts lists payouts with their status (PENDING, SENT, SETTLED, RETURNED) and the payments each one covers. A merchant sets its bank account and schedule through the dashboard.
Reconciliation reports: GET /v1/reconciliation/runs/{date} returns counts and amounts per category (matched, amount mismatch, missing on one side, timing) per PSP, for finance.
Dispute events to merchants (our own webhooks, signed the same way PSPs sign theirs):
json{ "id": "evt_01JA3K9", "type": "payment.disputed", "created": "2026-10-19T08:14:02Z", "data": { "payment_id": "pay_01JA2B", "dispute_id": "dp_01JA3K", "amount_minor": 11999, "reason": "fraudulent", "evidence_due_by": "2026-11-02T23:59:59Z" } }
Wallets: POST /v1/wallets/{id}/debits with an idempotency key, which fails with 402 and insufficient_funds if the balance can't cover it.
R2.4 Design Evolution: The Machinery of a Real Processor
Step 2.1: A Payment Is Five Steps Across Three Systems
The problem: a manual-capture payment is: authorize at the PSP; record the hold; wait days for the merchant to ship; capture at the PSP; post the ledger and notify the merchant. The PSP, our ledger and the merchant's systems are separate, and any step can fail, or our service can be redeployed, in the middle. What would you do?
Synthesizing vector architecture diagram...
The saga for one manual-capture payment. There is no separate notify state: PostLedger writes the journal, the status and an outbox row in one transaction, and the merchant notification flows from the outbox (step 2.3). Two arrows lead to Capture because the merchant may call capture before the execution even starts waiting; the next paragraph shows why that race is safe.
The capture race, made safe by the row lock. The workflow saves its task token on the payment row; the capture API marks the row capture_requested. Each does it with one UPDATE ... RETURNING on the same row, so Postgres runs them one after the other:
- The workflow runs
UPDATE payments SET task_token = $t WHERE payment_id = $1 RETURNING capture_requested. Ifcapture_requestedis already true, it goes straight toCapture. - The capture API runs
UPDATE payments SET capture_requested = true WHERE payment_id = $1 AND status = 'REQUIRES_CAPTURE' RETURNING task_token. If a token is there, it callsSendTaskSuccesswith it; if not, the workflow will see the flag when it registers.
Whichever runs second sees the first one's write, so a capture request can never be missed.
Primitive: Two-Phase Commit and Saga Orchestration
Drill: The flight booking that charged without a seat (a half-failed compensation is answered above; saga versus 2PC is the first wrong answer)
Step 2.2: Two $40 Spends From a $50 Wallet Both Succeeded
The problem: a customer's wallet holds $50. Their phone and their laptop each start a $40 purchase in the same millisecond. Each request sums the wallet's ledger entries, sees $50, and inserts a $40 debit journal. Both commit. The wallet is at −$30. What would you do?
Primitive: Database Isolation Levels, ACID and Concurrency Anomalies
Drill: The on-call doctor anomaly (why snapshot isolation misses write skew, and why not SERIALIZABLE everywhere, are both answered above)
Step 2.3: The Ledger and the Event Stream Disagree
The problem: after committing a payment, the service publishes payment.captured to Kafka for merchant notifications, analytics and risk. One night a network blip made 312 publishes time out after the commits had succeeded. Those payments exist in the ledger, but no merchant was notified, and analytics under-counts the day.
What would you do?
The same stream also starts the manual-capture sagas: a small consumer reads payment.requires_capture events and calls StartExecution with the payment ID as the name, which is idempotent while the execution runs.
Primitive: Change Data Capture and the Outbox Pattern
Drill: The dual-write that broke search consistency (the commit-succeeded-publish-failed case and CDC versus a polling outbox are both answered above)
Step 2.4: PSP A Is Failing, and Approvals Are Dropping
The problem: PSP A carries 60% of our traffic. At 14:05 its median latency climbs from 300 ms to 5 seconds, and about a third of its requests start timing out. It isn't down; it's slow. Within two minutes our checkout latency is terrible for every payment, including those going to PSPs B and C. What would you do?
Primitive: Circuit Breaker, Bulkhead and Fault-Tolerance Patterns
Drill: The slow recommendation service that took down checkout (why a slow dependency exhausts threads and why tiny timeouts aren't the answer are both covered above)
Step 2.5: Our Books Don't Match the Bank's
The problem: PSP A deposited $4,812,337.19 into our bank account this morning. Our ledger says PSP A owed us $4,812,402.19 for that batch. Sixty-five dollars are missing, and we're about to pay merchants from these balances. What would you do?
Synthesizing vector architecture diagram...
Two checks, both required: the line-by-line join proves every payment is on both sides, and the net check proves the money that actually landed in the bank is what the lines add up to.
Step 2.6: Merchants Need Their Money
The problem: twenty thousand merchants expect to be paid daily. The money arrives from three PSPs on their own schedules, some sales will be refunded or disputed later, and bank transfers can bounce days after we send them. What would you do?
Synthesizing vector architecture diagram...
The ledger always changes before money moves out. Settlement is posted on the settlement date for every payout, so payouts_in_transit always empties; a return moves the money back onto the merchant's balance, so the next payout can pay it once the bank details are fixed.
Step 2.7: We're Locked Into One PSP's Card Tokens
The problem: 70% of our payments use cards saved from earlier purchases. Each saved card is a token that only the PSP that saved it can charge. When PSP A's breaker opens, those customers can't pay, and we can't move our volume to a cheaper PSP. What would you do?
sqlCREATE TABLE card_vault ( token_id TEXT PRIMARY KEY, -- 'vtok_9f3kQ2', random fingerprint BYTEA NOT NULL, -- HMAC-SHA256(card number, fingerprint key) kms_key_arn TEXT NOT NULL, key_version INT NOT NULL, encrypted_dek BYTEA NOT NULL, -- data key, encrypted by KMS ciphertext_pan BYTEA NOT NULL, -- card number, AES-256-GCM under the data key last4 CHAR(4) NOT NULL, brand TEXT NOT NULL, exp_month SMALLINT NOT NULL, exp_year SMALLINT NOT NULL, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), rewrapped_at TIMESTAMPTZ ); CREATE INDEX ON card_vault (fingerprint);
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Five steps, three systems, days | Step Functions Standard saga per manual-capture payment; compensations; inline authorize | Eventual consistency between steps; per-transition price |
| 2.2 | Two spends overdraw a wallet | FOR UPDATE on the account row, balance snapshot, per-account check constraint, lock ordering | Contention on busy accounts |
| 2.3 | Ledger and events disagree | Transactional outbox, Debezium CDC on MSK Connect, idempotent consumers | A CDC pipeline; WAL retention to watch |
| 2.4 | A slow PSP stalls everything | Bulkheads, circuit breaker per PSP, BIN/cost/approval routing, fail over only when safe, key per PSP attempt | Routing logic; three reconciliations |
| 2.5 | Books don't match the bank | Nightly three-way full-outer-join reconciliation, every row classified, net check against deposits | A pipeline and reviewers |
| 2.6 | Merchants need paying | Daily batched payouts from reconciled funds minus reserve; ledger first; deterministic batch ID; returns | Payout delay; reserve policy |
| 2.7 | Locked into PSP tokens | Isolated vault, envelope encryption with KMS, fingerprint, forwarding proxy; KMS rotation | A PCI-scoped service |
R2.5 Architecture v2
Now the concepts get AWS names. Two pictures: the request path, then the money and data path.
Synthesizing vector architecture diagram...
The request path. The API claims the key in DynamoDB, authorizes through the router, and commits the payment, its ledger entries and an outbox row in one Aurora transaction. Card numbers exist only inside the vault and on the wire to the PSP.
Synthesizing vector architecture diagram...
The money and data path. Every change leaves the ledger through the WAL, so sagas, notifications, the lake and reconciliation all see the same committed history. Payouts wait for reconciliation.
What moved since Round 1, and why:
- Idempotency keys moved to DynamoDB. The key check is the first thing every request does, including abusive retry loops, and the Aurora writer is the one machine we can't scale out. DynamoDB takes the conditional writes, and its TTL feature deletes 7-day-old keys without costing us write capacity. Aurora stays the source of truth: the payments table keeps a
UNIQUE (merchant_id, scope, idempotency_key)constraint forever, so even a retry after the DynamoDB item is gone can't create a second payment. (R2.8 covers what happens when the two disagree.) - Webhooks are verified and stored by the API tier (insert into
webhook_events, dedupe by event ID, plus an outbox row, then200); the processing happens behind the stream, so a burst of deliveries can't overwhelm payments. - Reads go to three read replicas. Anything that decides something (a balance check before a debit, "is this captured?" before a refund) reads the writer.
The DynamoDB key item:
| Attribute | Example | Meaning |
|---|---|---|
pk (partition key) | m_4410#payments#7a0c4b2e-... | Merchant, endpoint and key: everything that makes two requests different operations |
status | IN_FLIGHT or COMPLETED | |
request_hash | 9c1e... | SHA-256 of the body; a mismatch is 422 |
payment_id | pay_01JA2B | Filled once known |
response | compressed JSON | The stored answer |
lease_until | 1790500061000 | Our clock, ms: claim time + 60 s |
expires_at | 1791104801 | TTL attribute, epoch seconds: claim time + 7 days |
The claim is a PutItem with ConditionExpression: attribute_not_exists(pk). Two details that matter:
- DynamoDB has no "current time" in conditions, so
lease_untilcomes from our servers' clocks. A request that finds anIN_FLIGHTitem treats the lease as expired only whenlease_untilis at least 5 seconds (our clock-skew margin) in the past, and takes it over with anUpdateItemconditioned onlease_untilstill equalling the value it read. Only one taker can win, and no decision depends on DynamoDB's clock. - TTL deletes items eventually, typically within a few days of expiry, not at the instant. So an item can outlive its 7 days; replaying it is harmless because it's the true answer. And we never use TTL as the lease: the lease is
lease_until, checked by us.
Trace 1: authorize, then capture three days later.
Synthesizing vector architecture diagram...
The buyer waits only for the top half. Everything after "three days later" is the saga, and each step is idempotent.
Trace 2: PSP A is failing. At 14:05, PSP A's timeouts pass 20% for 30 seconds and its breaker opens. New payments with a vault token route to PSP B (the vault forwards the card there); payments with a PSP-A-only token (older cards not yet in our vault) get 503 with "try again shortly". The 400 or so payments that timed out at A before the breaker opened are UNKNOWN; the resolver asks A, gets answers as A recovers, and only those A confirms never happened are left for the buyer to retry. At 14:20 the breaker goes half-open, 1% of eligible traffic tries A, succeeds, and the router shifts A back to its normal share over a few minutes.
Trace 3: a reconciliation break. At 06:20, the Glue job finds one PSP A line with no ledger match: du_1QK..., a $50.00 dispute plus a $15.00 fee. The webhook for it had failed signature verification during a secret rotation and was never processed. The net check for A's batch is off by the same $65.00, so the break is explained, not just detected. The merchant is held from today's payout; an operator replays the event from the PSP (Stripe keeps events available to resend for weeks), the dispute journal posts, the merchant's balance drops by $65, and tomorrow's run shows zero breaks for A.
Trace 4: a payout with a return. On Tuesday the payout job debits merchant_payable:m_4410 $8,200.00 and credits payouts_in_transit, then sends the ACH file. On Wednesday, the settlement date, it debits payouts_in_transit and credits bank_cash $8,200.00. On Thursday the return file shows R03 for that entry. The job debits bank_cash, credits merchant_payable:m_4410 $8,200.00, marks the payout RETURNED and blocks the bank account. The merchant updates its bank details on Friday and the next payout includes the $8,200.00.
Ledger entries for one payment and one refund (the $129.99 authorization captured at $119.99; PSP fee 2.9% + 30¢ = $3.78, which PSP A reports after rounding $3.77971; platform fee 1% = $1.20; the merchant bears the PSP fee):
| Journal | Account | Debit | Credit |
|---|---|---|---|
| Capture | psp_receivable:PSP_A | $116.21 | |
| Capture | psp_fees:PSP_A | $3.78 | |
| Capture | merchant_payable:m_4410 | $115.01 | |
| Capture | platform_revenue | $1.20 | |
| Capture | psp_fee_recovery | $3.78 | |
| Capture total | $119.99 | $119.99 | |
| Refund $19.99 | merchant_payable:m_4410 | $19.79 | |
| Refund $19.99 | platform_revenue | $0.20 | |
| Refund $19.99 | psp_receivable:PSP_A | $19.99 | |
| Refund total | $19.99 | $19.99 |
Check the capture: $119.99 − $3.78 − $1.20 = $115.01 to the merchant, and $116.21 + $3.78 = $119.99. The refund returns our 1% on the refunded part ($0.1999, half-even to $0.20) and takes the rest from the merchant; the PSP fee is not returned.
R2.6 Numbers and Cost
Traffic.
The 10× covers the daily peak hour plus flash sales. Status reads are 5 per payment at peak: reads/s.
Storage. About 1 KB per payment all in, as in Round 1 (payment row, entries, events, journal):
With indexes (1.6×, an assumption): about 11.7 TB after one year and 58.4 TB after five, inside one Aurora cluster's limit (128 TiB, or 256 TiB on recent Aurora PostgreSQL versions). Aurora keeps six copies across three AZs but bills storage once, so the storage bill is for 58.4 TB, not three times that. We partition the ledger tables by month, so old months can later be detached and archived to S3 without a giant delete.
Idempotency keys.
That's the steady state for 7 days of keys. Because TTL deletes lag by up to a few days, we budget for somewhat more.
Bandwidth (4 KB per request or response, from Round 1's payloads plus headers):
Out counts status replies plus our calls to PSPs. Averaged over a month: MB/s, about 14.6 TB a month leaving AWS.
Concurrency to the PSPs (Little's law): calls in flight at peak, split across the three bulkheads by traffic share.
Database sizing. Each payment is roughly ten row writes, so peak is about 25,000 row writes a second on one writer. We plan on a db.r6g.8xlarge writer (32 vCPUs) and an identical standby in a second AZ that Aurora promotes on failure. This is a planning assumption we load-test before launch: the writer is the one component that doesn't scale out, and Round 3 shards it. For reads, we plan 8,000 simple indexed reads a second per db.r6g.2xlarge replica (also to be load-tested) and run three, one per AZ: losing an AZ leaves two, (78% busy).
API fleet. We plan 500 requests a second per vCPU (an assumption to load-test): vCPUs at peak. To survive losing one of three AZs, the remaining two-thirds must carry peak: vCPUs. We run 24 tasks of 2 vCPU (48 vCPUs), 8 per AZ.
Step Functions. Standard workflows cost $0.000025 per state transition ($25 per million) in us-east-1. Only manual-capture payments run a saga, 4 transitions each on the happy path (RegisterToken, WaitForCapture, Capture, PostLedger):
Disputes (0.1% of payments, assumed, about 6 transitions each) and payouts (20,000 merchants × about 5) add about 0.22 million a day:
If every payment ran a 10-transition saga, the same math gives $5,000 a day, about $152K a month. That's why auto-capture payments don't use a saga.
Step Functions quotas. One default in us-east-1 is below our peak: StartExecution for Standard refills at 300 a second (bucket 1,300), and our manual-capture starts peak at a second. The others have headroom: state transitions refill at 5,000 a second against roughly 2,300 at peak, and SendTaskSuccess at 500 a second against about 93 (captures arrive spread over days, at about the average rate of ). These are soft quotas that AWS can raise; we request a StartExecution increase, load-test it, and have the saga starter pace StartExecution calls, since a saga that starts a few seconds late only waits for a capture that's hours away anyway (the capture race above is safe in either order).
Monthly cost (us-east-1 on-demand prices, 730 hours a month, storage at the end of year one):
| Item | Math | Monthly |
|---|---|---|
| Aurora instances | 2 × db.r6g.8xlarge ($4.152/h) + 3 × db.r6g.2xlarge ($1.038/h), × 730 h | ≈ $8,340 |
| Aurora storage | ~11.7 TB × $0.10/GB-month, billed once | ≈ $1,170 |
| Aurora I/O | Assumed ~3 billed write I/Os per transaction, ~2 transactions per payment: ~1,400/s average ≈ 3.65 billion × $0.20/M | ≈ $730 |
| DynamoDB keys | ~2.9 writes per payment (claim and complete, plus capture and refund keys) × 20M × 30.4 days ≈ 1.76 billion × $0.625/M; 70 GB × $0.25 | ≈ $1,120 |
| Step Functions Standard | from above | ≈ $24,500 |
| Fargate | 24 API tasks (2 vCPU, 4 GB) ≈ $1,730 + 6 worker tasks (1 vCPU, 2 GB) ≈ $216 | ≈ $1,950 |
| Card vault | 2 × db.r6g.large $380 + ~$20 storage + 6 Fargate tasks $216 + KMS: 20M decrypts/day × 30.4 × $0.03 per 10K ≈ $1,824 | ≈ $2,440 |
| MSK + Debezium | 3 × kafka.m7g.large ($0.204/h) ≈ $447 + 840 GB storage (40 GB/day × 7 days × 3 copies) × $0.10 ≈ $84 + 2 MSK Connect units × $0.11/h ≈ $161 | ≈ $690 |
| S3 lake, Glue, Athena | lake storage grows slowly; Glue ~20 DPUs × 0.75 h × $0.44 nightly ≈ $200 | ≈ $300 |
| ALB | $16 fixed + ~24 capacity units (~23 GB/h processed) × $0.008/h | ≈ $160 |
| AWS WAF | ~1,440 requests/s average ≈ 3.8 billion/month × $0.60/M | ≈ $2,280 |
| Data transfer and NAT | ~14.6 TB out ($0.09/GB first 10 TB, $0.085 next) ≈ $1,290 + NAT hours and processing ≈ $210 | ≈ $1,500 |
| CloudWatch, logs, alarms | an estimate | ≈ $1,000 |
| Total | ≈ $46K/month |
Aurora storage grows by about $1,170 a month each year, to about $5,840 a month at year five (58.4 TB). Watch the I/O line: if Aurora I/O ever passes about a quarter of the Aurora bill, Aurora I/O-Optimized (no per-I/O charge, higher instance and storage prices) becomes cheaper.
Now the number that matters. At 2.9% + 30¢ on a $40 average, PSP fees are 20 \times 10^6 \times \1.46 = $29.2M a day, about **\888M a month**. Our whole platform is about 0.005% of that. Large platforms negotiate far better rates than list price, but the conclusion holds: a 0.1-point better routing decision is worth more than the entire AWS bill. That's why routing by cost and approval rate (step 2.4) is a revenue feature, not an optimization.
Availability budget. 99.99% of a 30.4-day month is minutes. An Aurora failover of tens of seconds uses a real part of that, so we don't spend it on deploys: rolling deploys, one AZ at a time.
Latency budget (P99, dependent steps add): key claim in DynamoDB (~10 ms) + payment insert (~5 ms) + router and vault decrypt (~20 ms) + PSP (≤ 600 ms) + commit with ledger and outbox (~15 ms) + network inside the region (~20 ms) ≈ 670 ms, under the 800 ms target. The PSP is 90% of it.
R2.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| Orchestration vs choreography | An orchestrator (Step Functions) for multi-day sagas | In choreography, each service reacts to the previous one's events and no one is in charge. It scales and decouples, but "where is payment X stuck?" has no single answer, and compensations are spread over many services. For money we want one place that knows the state. |
| Pessimistic locks vs optimistic concurrency | FOR UPDATE on balance-checked accounts | Optimistic concurrency (a version column, retry on conflict) avoids waiting but retries under contention, which on a busy wallet means repeated work. Locks queue instead. Accounts without a balance rule take no lock at all. |
| Our vault vs PSP tokens | Our vault (plus network tokens where offered) | PSP tokens keep us almost out of PCI scope but lock saved cards to one PSP. Our vault gives routing and failover for saved cards at the price of running a PCI-scoped service. |
| Standard vs Express workflows | Standard, only for multi-day flows | Express is far cheaper per execution and has no transition quota, but it's at-least-once (asynchronous) or at-most-once (synchronous), stops at 5 minutes, and has no callbacks. Our saga waits days for a merchant. |
| Keys in DynamoDB vs in Aurora | DynamoDB at the front, Aurora's unique constraint as truth | Two stores that can disagree (R2.8); in exchange, retries and replays stay off the writer. |
Four ways to keep several writes consistent, with verified limits:
| Saga + orchestrator (Step Functions Standard) | Two-phase commit | Choreography (events on Kafka) | DynamoDB TransactWriteItems | |
|---|---|---|---|---|
| Works across an external PSP? | Yes: each step is a separate call with compensation | No: PSPs don't take part in 2PC | Yes, with compensation spread across services | No: only DynamoDB items |
| Consistency | Eventual between steps, each step atomic | Atomic across participants | Eventual | Atomic, up to 100 items and 4 MB in one call |
| Behaviour under timeouts | State persists; waits up to a year; retries per step | Participants hold locks until the coordinator decides | Messages wait in the log; nobody owns the whole flow | Fails as a unit if any condition fails |
| Throughput limits | Default Standard StartExecution 300/s refill and 5,000 transitions/s in us-east-1, both raisable; priced per transition | Bounded by the slowest participant holding locks | Very high; limited by partitions and consumers | Each item's partition limits apply (1,000 write units/s per partition) |
| "Where is payment X?" | One execution history | The coordinator's log | Reassembled from traces | Not applicable |
R2.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| PSP gray failure (slow, not down) | Latency and timeouts rise at one PSP; approval rate drops | Bulkhead caps its in-flight calls; breaker opens on the sustained pattern; router moves new payments only; timeouts go to UNKNOWN and are resolved at that PSP. |
| Out-of-order webhooks | A refund event before the capture event; a dispute after a refund | Events are stored first, then applied as allowed moves. A move not allowed yet waits while we fetch the payment's current state from the PSP. A dispute is allowed from CAPTURED, PARTIALLY_REFUNDED and REFUNDED; the dispute journal checks what was already refunded, and if a refund is still pending when the dispute arrives, we let the dispute decide and don't also refund (the PSP may fail that refund, as Stripe does with charge_for_pending_refund_disputed). |
| Idempotency store and ledger disagree | DynamoDB says IN_FLIGHT long after Aurora committed (the "complete" write failed) | A retry finds the lease expired (by our clock, with the 5 s margin), takes it over conditionally, and checks Aurora by (merchant, scope, key). Committed: it repairs the DynamoDB item from Aurora and replays the answer. Still PROCESSING or UNKNOWN: it returns 202 and leaves the payment to the resolver; it never starts a second PSP attempt. No payment row: it proceeds as a new request, and Aurora's unique constraint stops any zombie of the first request from creating a duplicate. |
| Rounding drift in fee splits | A reconciliation amount mismatch of a few cents on one PSP | We never compute what the PSP deducts; we post the fee the PSP reports. A PSP-side currency conversion on a foreign card can settle a few cents differently from its authorization; the difference posts to an fx_variance account and alarms above $10 a day. Round 3 handles currencies properly. |
| KMS key rotation under load | Nothing, for automatic rotation | KMS keeps old key material and picks the right version on decrypt, so rotation needs no downtime or re-encryption. A switch to a new KMS key runs as a background re-wrap while reads use the key named on each row. |
| Losing an AZ | Tasks, one replica and possibly the writer disappear | Aurora promotes the standby (tens of seconds, with a burst of connection errors; in-flight payments become PROCESSING rows the resolver finishes). Two AZs of API tasks (32 vCPUs of our 48) and two replicas (16,000 reads/s) carry peak. DynamoDB, MSK (three brokers across three AZs), Step Functions and KMS are multi-AZ already. |
R2.9 Production Gotchas
| Gotcha | Why it hurts | What we do |
|---|---|---|
| Floats for money | Binary fractions drift; two systems round differently | Integer micros in the ledger, minor units at the edges, named rounding |
| Updating a balance in place | No history; a bad update can't be explained or undone | Append-only entries; the balance column is a snapshot recomputed nightly |
| No idempotency key on outbound PSP calls | A retry after a lost reply is a second charge | A PSP key per attempt, reused only within that attempt and that PSP |
| Synchronous PSP calls on shared request threads | One slow PSP exhausts every thread | Per-PSP bulkheads, breakers, and non-blocking calls with a timeout |
| Encryption keys without version metadata | A forced key change becomes a stop-the-world migration | KMS key ARN and version on every vault row; automatic KMS rotation |
| Starting a saga right after commit, in the same code | Crash between commit and start loses the saga | Start from the outbox stream, with the payment ID as execution name |
| Treating TTL as a deadline | DynamoDB TTL deletes lag by days | Leases use our own timestamp; TTL only cleans up |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | Sagas with compensation for multi-day flows; per-PSP bulkheads and breakers; failover only when provably safe; outbox + CDC so no event is lost; AZ loss sized (32 of 48 vCPUs, 2 of 3 replicas, standby writer) REL 4 · REL 5 · REL 10 · REL 11 |
| Security | Card numbers only in an isolated vault; envelope encryption with KMS and automatic rotation; no CVC storage; the router never sees a card number; signed webhooks both ways; least-privilege roles (only the vault can decrypt) SEC 3 · SEC 7 · SEC 8 · SEC 9 |
| Performance Efficiency | P99 ≈ 670 ms with the PSP at 90%; keys on DynamoDB, reads on replicas, the writer kept for writes; concurrency to PSPs sized by Little's law PERF 1 · PERF 3 |
| Cost Optimization | ≈ $46K/month, derived; Step Functions only where its guarantees are needed ($24.5K instead of $152K); PSP fees (~$888M/month) are the real cost, so routing by price is where money is saved COST 5 · COST 6 |
| Operational Excellence | Reconciliation every night with every row classified; review queues; payouts held on breaks; alarms on replication-slot lag, UNKNOWN age, breaker state and saga failures OPS 8 · OPS 10 |
| Sustainability | Graviton instances throughout; sagas only for the 40% that need them; ledger partitions archived to S3 as they age SUS 2 · SUS 4 |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Rejects 2PC with the right reason (external PSPs, locks held across network calls) and designs a saga with compensations that can themselves fail safely.
- Chooses between Standard and Express workflows by execution semantics, duration and price, and knows the quotas.
- Explains write skew under snapshot isolation and fixes the balance race with explicit locks where the invariant lives.
- Uses the outbox with CDC instead of dual writes.
- Routes between PSPs with bulkheads and breakers, and never fails over a payment in an unknown state.
- Reconciles line by line with a full outer join, accounts for every row, and holds payouts on breaks.
- Designs payouts ledger-first with returns, and a vault with envelope encryption, knowing what KMS rotation does and doesn't do.
- Says out loud that PSP fees dwarf infrastructure.
Follow-up questions
-
"A merchant captures, then immediately refunds, and the capture's ledger post is still in flight. What happens?" Answer: the refund API checks the payment's status on the writer. While it's
CAPTURING, a refund isn't an allowed move yet, so the API returns409withRetry-After. Once the capture journal commits (seconds), the refund proceeds. We never refund money we haven't recorded as captured. -
"Why not let merchants call PSP B directly when PSP A is down?" Answer: because every payment must go through our idempotency, routing and ledger. A payment made around us is money we don't know about: it would show up in reconciliation as "missing in our ledger", and the merchant's balance and payouts would be wrong.
-
"Your reconciliation says everything matched, but a merchant says they're owed $200 more. Who's right?" Answer: we list the merchant's ledger entries for the period; each links to a payment, refund, dispute or payout, and each of those links to a reconciled PSP or bank line. Either the entries explain the balance (and we show them), or one is missing, which reconciliation would have flagged unless both we and the PSP missed it, and then the bank net check is the tiebreaker. The ledger's value is that the answer is a query, not an argument.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Use 2PC across services and the PSP" | PSPs don't take part in 2PC; locks would be held across slow network calls. |
| "Step Functions Express for the payment saga" | At-least-once (asynchronous) or at-most-once (synchronous), 5-minute limit, no callbacks; the saga waits days. |
| "Repeatable read prevents the overdraft" | Snapshot isolation allows write skew: both transactions insert different rows. |
| "Publish to Kafka after the commit" | A crash in between loses the event; use an outbox read by CDC. |
| "On timeout, fail over to the other PSP" | The first PSP may have charged; resolve it first. |
| "Reconcile with an inner join on totals" | Inner joins drop the one-sided rows that are the actual problems; totals hide cancelling errors. |
| "Rotating the KMS key means re-encrypting every card" | KMS rotation keeps old key material; re-wrapping is needed only when switching to a different key. |
Round 3 · Architect · "Global Money: Currencies, Rules, Fraud, Scale"
~45 min · Principal (L7) · 4 home regions, each with a disaster-recovery copy · 200M payments/day · 25 currencies · no double-spend through a region failover
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Round 2.
Round 2 in 60 seconds. "We built a payments platform in one region: 20,000 merchants, 20 million payments a day, 2,500 a second at peak, 12,500 status reads a second. Idempotency keys are claimed in DynamoDB with a conditional put and a lease timed by our own clock, and Aurora's unique constraint is the source of truth. Payments authorize inline; the 40% that capture later run a Step Functions Standard saga that waits for the merchant with a task token, voids before the authorization expires, and compensates safely. Wallet spends lock the account row, with a per-account never-negative constraint, because snapshot isolation allows write skew. Every ledger change leaves through an outbox read by Debezium into MSK. The router picks among three PSPs by card, cost and live approval rate, with a bulkhead and a circuit breaker per PSP, and never fails over a payment whose outcome is unknown. A nightly Glue job reconciles every line three ways and holds payouts on breaks. Payouts are ledger-first ACH batches with returns handled. Cards live in an isolated vault with KMS envelope encryption. About $46K a month, against about $888 million a month in PSP fees. Open costs: one database writer, accounts that every payment touches, and one region."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: one writer holds all the money; everything else reads its log.
Round 2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Five steps, days apart | Standard saga for manual capture; compensation | Eventual consistency; per-transition price |
| 2.2 | Wallet overdraft race | FOR UPDATE, snapshot balance, per-account check | Contention on busy accounts |
| 2.3 | Ledger and events drift | Outbox + Debezium CDC | A pipeline; WAL retention |
| 2.4 | Slow PSP stalls everything | Bulkheads, breakers, safe routing | Routing logic |
| 2.5 | Books vs bank | Three-way full-outer-join reconciliation | Reviewers |
| 2.6 | Paying merchants | Ledger-first batched payouts, returns | Delay, reserves |
| 2.7 | PSP token lock-in | Isolated vault, envelope encryption | PCI-scoped service |
Open costs: one Aurora writer caps us; accounts that every payment touches serialize on a row lock; one region, one currency, and no defense of our own against fraud.
R3.1 The Scope Raise
Interviewer: "We're going global. Forty countries, 25 currencies, 200 million payments a day. Cards issued in Europe need strong customer authentication. Last month fraudsters ran stolen cards through one of our merchants at a thousand attempts a minute. One of our markets requires payment data to stay in the country. The regulators audit us now. Our own revenue account is written by every single payment. And a whole region can fail, but money must never be double-spent."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Who converts currency, and in which currency are merchants paid? | Buyers pay in their currency; merchants are paid in theirs. We quote the rate. | We own FX: quotes, rates, a currency position, and rounding at every conversion (step 3.4). |
| Which payments need strong customer authentication? | Customer-initiated online payments with cards issued in the European Economic Area, when our acquirer is also in the EEA, under PSD2 (when only one side is in the EEA, it applies on a best-effort basis). Some exemptions exist, and the PSP asks for them. | Checkout can't finish in one call: some payments pause for a bank challenge (step 3.2). |
| Where do the fraud attempts come from? | A script hitting one merchant's checkout with small amounts on many cards, from many IP addresses. | IP blocking won't work. We need scoring on cards, devices and merchants before authorization (step 3.3). |
| Which data must stay in-country, and is there a second AWS Region in that country? | Payment and card data for market M. Assume there are two AWS Regions inside M. | Market M's merchants and cards live in M, including their disaster-recovery copy (step 3.5). |
| How is traffic spread? | Americas 40%, Europe 30%, market M 15%, rest of Asia-Pacific 15%. | Four home regions, sized separately (R3.6). |
| What does "every payment writes our revenue account" mean at this scale? | Every capture credits platform_revenue. | 2,500 captures a second on one shard would serialize on one account (step 3.1). |
| What do auditors need? | Proof that no record was altered after the fact, and reproducible reconciliations, kept for years. Assume 7 years. | Tamper-evident storage and segregation of duties (step 3.6). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Volume | 20M/day, 2,500/s peak | 200M/day; regional peaks 10,000 + 7,000 + 3,500 + 3,500/s |
| Footprint | 1 region, 3 AZs | 4 home regions, each with a disaster-recovery copy |
| Money | USD | 25 currencies, our own FX |
| Checkout | One call | + a 3-D Secure challenge step for European cards |
| Fraud | PSP's checks | Our own scoring before authorization |
| Data | One region | Residency: market M's data stays in M |
| Ledger | One writer | 11 shards, hot accounts handled |
| Audit | Reconciliation | + tamper evidence, segregation of duties, 7-year retention |
| Availability | 99.99% | 99.99% per region; region failover in minutes without double-spend |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| One ledger database | Ten times the writes; one writer planned for 2,500 payments/s can't take 24,000. |
One platform_revenue account | Every capture takes the same row lock: at ~2 ms each, one row serializes at about 500 a second. |
| No fraud checks of our own | A thousand stolen-card attempts a minute at one merchant: network fees, a flagged merchant, and chargebacks weeks later. |
| Synchronous checkout | A European card may need the buyer to approve in their bank's app; there's no "wait for the buyer" step. |
| One region | A region outage stops every payment in the world, and market M's data isn't allowed to live there anyway. |
| FX not modelled | Converting with floats at display time; no record of the rate used, no position, rounding differences everywhere. |
| "Trust the database" for audit | An administrator can change a row; nothing proves it didn't happen. |
| Step Functions for every manual capture | At ten times the volume, about $245K a month for sagas that are mostly waiting. |
R3.3 New Requirements and API Additions
The payment intent with a "requires action" state. A payment intent is one object that tracks a payment from "the buyer wants to pay" to the final result, across however many steps it takes.
httpPOST /v1/payment_intents/pi_01JB7/confirm HTTP/1.1 Host: api.eu.payments.example.com Idempotency-Key: 2c9f0e6b-confirm-1 Content-Type: application/json { "payment_method": "vtok_eu_77Hc", "return_url": "https://shop.example/checkout/return" }
httpHTTP/1.1 200 OK Content-Type: application/json { "id": "pi_01JB7", "status": "REQUIRES_ACTION", "amount_minor": 8900, "currency": "EUR", "next_action": { "type": "three_ds_challenge", "url": "https://acs.issuer.example/challenge/abc" }, "expires_at": "2026-10-02T10:44:00Z", "risk": { "score": 34, "decision": "CHALLENGE" } }
FX quotes:
httpPOST /v1/fx/quotes HTTP/1.1 Content-Type: application/json { "sell_currency": "EUR", "buy_currency": "USD", "sell_amount_minor": 10000 }
json{ "quote_id": "fxq_01JB8", "rate": "1.08500000", "sell_amount_minor": 10000, "buy_amount_minor": 10850, "expires_at": "2026-10-02T10:15:30Z" }
Rates are decimal strings with a fixed number of places, never floats. A quote expires 60 seconds after it's issued (our choice); a payment presenting an expired quote gets 409 quote_expired and a fresh quote.
Risk: every payment carries a risk block (score 0 to 99 and the decision taken). Merchants can see it but not change it.
Residency attributes: every merchant and every customer record has a home_region (for example eu-central-1) and a residency (for example M or NONE), set at onboarding and changed only by a supervised migration.
R3.4 Design Evolution: Money Across Borders, Rules and Adversaries
Step 3.1: Every Payment Writes the Platform Revenue Account
The problem: we need ten times the write capacity, so the ledger must be split across databases. And every capture credits platform_revenue, one account, one row. At 2,500 captures a second on one database, with each transaction holding that row's lock for about 2 ms, the account can only take about 500 a second.
What would you do?
| Strategy | Use it for | Balance read | Can enforce "never negative"? |
|---|---|---|---|
| Append-only, no snapshot | Revenue, fees, clearing accounts with no floor | Rollups + recent entries | No |
| N sub-accounts | A giant merchant's payable | Sum of N rows | Yes, by locking all N in order |
| Batched posting | Very hot accounts where a minute's delay is fine | One row, a minute behind | Yes, at posting time |
| Dedicated shard | A merchant too big to share | One shard's rows | Yes |
Primitive: Database Sharding and Partition Keys
Drill: One customer, one shard, one outage (the giant tenant's move is the "very biggest merchants" paragraph; cross-tenant reports read the lake, in "What it costs us")
Step 3.2: European Cards Need a Challenge Step
The problem: for many cards issued in Europe, the issuer wants the buyer to prove it's really them before approving an online payment: a code in their banking app, a fingerprint. Our checkout is one API call that returns a final answer. The buyer's bank now needs to talk to the buyer in the middle of it. What would you do?
Synthesizing vector architecture diagram...
REQUIRES_ACTION is the only state where we wait for a person, and it has a deadline enforced by the database's own clock.
Step 3.3: Card-Testing Attacks
The problem: a script sends 1,000 payment attempts a minute through one merchant's checkout, $1 each, with a different stolen card every time, from thousands of IP addresses. Most are declined. The ones that succeed tell the fraudster which cards work, and will come back to us as chargebacks. What would you do?
Primitive: Bot Defense, Sybil Resistance and Registration Abuse
Step 3.4: Twenty-Five Currencies
The problem: a buyer pays €100.00; the merchant is paid in US dollars. The current code multiplies by a float rate when displaying the merchant's balance. Finance can't say what rate was used, how much money we hold in euros, or why the dollar total is three cents off. What would you do?
The €100.00 payment at our quoted rate of 1.0850 (the rate already includes our margin):
| Currency | Account | Debit | Credit |
|---|---|---|---|
| EUR | psp_receivable:PSP_A@shard-03 | €100.00 | |
| EUR | fx_clearing:EUR | €100.00 | |
| USD | fx_clearing:USD | $108.50 | |
| USD | merchant_payable:m_9012 | $108.50 |
Each currency balances on its own: €100.00 = €100.00, and $108.50 = $108.50. (Fees are left out to keep the example small.) Later, treasury sells €100.00 at the bank for $109.00:
| Currency | Account | Debit | Credit |
|---|---|---|---|
| EUR | fx_clearing:EUR | €100.00 | |
| EUR | bank_cash:EUR | €100.00 | |
| USD | bank_cash:USD | $109.00 | |
| USD | fx_clearing:USD | $108.50 | |
| USD | fx_gain_loss | $0.50 |
Both clearing accounts are back to zero, and the $0.50 is a realized gain, recorded.
Rounding, worked. €10.00 converted at 1.08235 is $10.8235 = 10,823,500 micros. The merchants are paid in cents, so the outward amount is $10.82 (half-to-even: the digits after the cent are 35, below half). The $0.0035 = 3,500 micros goes to fx_rounding_variance. If the $10.82 is split among three sellers, largest remainder gives $3.61, $3.61 and $3.60. The USD side: debit fx_clearing:USD 10,823,500 micros; credit the three sellers 10,820,000 micros in total and fx_rounding_variance 3,500 micros. It balances to the micro, which is why the ledger stores micros.
Step 3.5: Data Must Stay In-Country, and a Region Can Fail
The problem: market M's payment data must stay in M. The other markets want their data near them for latency. And if a region fails, we must keep taking payments without ever spending the same money twice. What would you do?
Why not synchronous replication for every write? Every commit would wait for a round trip to another region (tens of milliseconds between nearby regions, more across oceans) on the ledger's hottest path, and a slow or unreachable second region would stall payments in the first. We keep strong consistency where it's needed, one writer per account, and accept asynchronous copies for recovery.
Primitive: Cloud Disaster Recovery and Multi-Region Active-Active
Drill: The booking that existed in Frankfurt but not in Virginia (reading a lagging copy for a decision is point 3; synchronous replication for all writes is the paragraph above)
Step 3.6: Auditors Want Proof That Nothing Was Altered
The problem: a regulator asks: "How do you know nobody changed a ledger entry from last March?" Our answer today is "the application can't update or delete rows". The regulator replies: "And your database administrators?" What would you do?
Synthesizing vector architecture diagram...
The chain makes tampering detectable; Object Lock makes the evidence undeletable; the verifier belongs to someone who can't write the ledger.
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Hot revenue account; one writer | Shard by merchant; per-shard sub-accounts; lock-free append for accounts with no floor; N-bucket and batched posting; dedicated shards | Balance reads become sums; reports read the lake |
| 3.2 | European cards need a challenge | Payment intent with REQUIRES_ACTION, 3DS, 30-minute expiry by the database clock | Longer, asynchronous checkout |
| 3.3 | Card testing | Velocity counters in Valkey, rules + model, graduated actions, shadow mode, feedback from disputes | Friction; a model to own |
| 3.4 | 25 currencies | Original amount kept; quotes with expiry; per-currency balancing through FX clearing; rounding variance; realized gain/loss | FX risk; a rate provider |
| 3.5 | Residency and region failure | Home region per merchant and customer; one writer per account; Aurora Global Database; fenced failover; bounded RPO | Cross-region sagas; failover in minutes |
| 3.6 | Proof of no tampering | Hash chain in commit order, S3 Object Lock, independent verifier, segregation of duties | Storage; slower operations |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Each home region is a complete Round 2 stack with its own ledger shards, vault, risk service and PSP connections. The only global piece is the small directory. Dotted arrows are disaster-recovery copies; market M's copy never leaves the country.
Each home region also runs its own Step Functions (disputes, payouts, cross-shard and cross-region transfers), MSK cluster, S3 lake and nightly reconciliation. A global reconciliation job reads every region's lake summaries (totals and break counts, not raw data from market M) to produce the company-wide books. The FX service quotes from a rate provider and caches rates per region; treasury converts balances daily.
What changed for manual captures. At ten times the volume, running a Step Functions saga for every manual-capture payment would cost about $245K a month (Round 2's $24.5K × 10), mostly to pay for waiting. We moved captures onto the same pattern the auto-capture path already uses: the payment row holds the state, the capture API runs the capture inline with its PSP key and the "unknown, then ask" protocol, and a sweeper selects WHERE status = 'REQUIRES_CAPTURE' AND capture_before < now() + interval '6 hours' to void authorizations nobody captured, using the database's clock. Step Functions stays for flows that are rarer and need a visible multi-step history.
Trace 1: a 3DS payment in the EU.
Synthesizing vector architecture diagram...
The webhook happened to arrive first here; if the return had come first, it would have triggered the same fetch and the same compare-and-set, and the webhook would have found nothing left to do.
Trace 2: a card-testing attack blocked. At 10:14 the risk service sees merchant m_4410's distinct-cards-per-minute counter jump from its usual 40 to 900, with 85% declines and $1 amounts. The rule "distinct cards per merchant per minute above 10× its 7-day baseline" fires: new cards at m_4410 now require a bot check on the hosted fields and a 3DS challenge, and the merchant's limit on new cards drops to 60 a minute (429 above it). Attempts that reach the PSP fall from about 1,000 a minute to under 60, most of those with a challenge the script can't pass. The merchant gets an alert; after an hour without the pattern, the limits relax step by step.
Trace 3: a regional failover. eu-central-1 has a severe outage at 09:00.
| Time | What happens | Why it takes that long |
|---|---|---|
| 09:00 | API error rate for EU shards jumps; alarms evaluating 1-minute periods start breaching | |
| 09:03 | Alarms fire (3 of 3 periods breaching) and page on-call | 3 × 60 s alarm evaluation |
| 09:10 | Incident commander decides to fail over EU shards 5-7 | Human decision, up to ~7 min |
| 09:11 | failover-global-cluster for shards 5-7, in parallel; each takes about a minute | Aurora promotion |
| 09:12 | Writer flags for shards 5-7 set to eu-west-1; any EU service still running stops writing within 10 s because it can't renew its flag lease | 10 s lease on our own clock |
| 09:12 + 15 s | eu-west-1 services start writing shards 5-7 | Waits out the old region's 10 s lease plus 5 s margin |
| 09:12 | Directory entries for EU merchants point to eu-west-1; DNS records switch | |
| 09:13 | New traffic arrives in eu-west-1 as DNS caches expire | 60 s record TTL, plus clients that cache longer |
| 09:13 onward | Resolver in eu-west-1 asks PSPs about every PROCESSING, UNKNOWN and CAPTURING payment | Backoff from Round 1 |
| next morning | Reconciliation checks the lag window against PSP reports |
About 13 minutes from failure to traffic in the recovery region, most of it detection and decision.
What is waiting in the recovery region, and what isn't. Each shard's secondary already runs one writer-sized instance, so the promoted database can take the full write load at once. It has no read replicas yet: for the first minutes, status reads go to the writer (a db.r6g.8xlarge has room for them only because writes are what we sized it for; we shed status polling with 429 before it competes with payments), and we add read replicas after promotion, which takes several more minutes each. The API, risk and vault services run a warm fleet at a quarter of the home region's size (about 16 of eu-central-1's 63 API tasks) and scale out at failover. Fargate adds tasks in minutes, so for roughly the first 10 minutes in the recovery region, peak-hour traffic above a quarter of normal gets 503 and retries. We accept that, and we cost the warm fleets in R3.6.
R3.6 Numbers and Cost
Traffic per home region (the interviewer's split of 200M payments a day):
| Region | Share | Payments/day | Average/s | Peak/s (×10, rounded up) | Reads/s at peak (×5) | Ledger shards |
|---|---|---|---|---|---|---|
| us-east-1 | 40% | 80M | 926 | 10,000 | 50,000 | 4 |
| eu-central-1 | 30% | 60M | 694 | 7,000 | 35,000 | 3 |
| Market M | 15% | 30M | 347 | 3,500 | 17,500 | 2 |
| ap-southeast-1 | 15% | 30M | 347 | 3,500 | 17,500 | 2 |
| Total | 200M | 2,315 | 24,000 (not simultaneous) | 11 |
Shards. We keep Round 2's planning figure: one shard (a db.r6g.8xlarge writer) carries up to 2,500 payments a second at peak, the same figure we load-tested. So , , and for each of the smaller regions: 11 shards. Reads per shard stay at or below Round 2's 12,500 (us-east-1: ), so each shard keeps three db.r6g.2xlarge readers.
Storage.
That's about 10.6 TB per shard per year. We keep two years hot in Aurora (a choice) and detach older monthly partitions to S3, so each shard holds up to about 23 TB (the busiest shards take about 11.7 TB a year with indexes), well inside the Aurora limit.
Audit archive. The sealed copy is Parquet, which we assume compresses the ledger about 4×: 50 GB a day. Storage from a retention period is writes per day × days kept:
The newest year (18.3 TB) in S3 Standard at $0.023/GB-month is about $420; the other six years (about 110 TB) in S3 Glacier Deep Archive at about $0.00099/GB-month is about $110. About $530 a month. Object Lock itself adds no storage charge.
Risk-scoring latency budget (checkout P99; dependent steps add, parallel steps take the max):
| Step | P99 | Notes |
|---|---|---|
| Key claim (DynamoDB) | 10 ms | |
| Payment insert | 5 ms | |
| Risk score ‖ vault decrypt | max(50, 20) = 50 ms | They run in parallel; both must finish before the PSP call |
| PSP authorize (frictionless) | 600 ms | Unchanged from Round 2 |
| Commit ledger + outbox | 15 ms | |
| Network in the region | 20 ms | |
| Total | 700 ms | Under 800 ms. A 3DS challenge adds the buyer's own time on top, which no budget covers. |
Inside the 50 ms: about 8 counter reads and increments against Valkey in one pipelined round trip (~2 ms), the rules (~1 ms) and the model (a few ms), with the rest as margin for garbage collection and queueing. Counter load: us-east-1 peaks at Valkey operations a second, spread over 2 shards. We plan each region with 2 shards of a primary and a replica (4 nodes).
Monthly cost (us-east-1 on-demand prices for every region; the other three regions cost more, especially in Asia-Pacific, where both instance and data-transfer prices are higher):
| Item | Math | Monthly |
|---|---|---|
| Ledger shards, primary | 11 × (2 × db.r6g.8xlarge + 3 × db.r6g.2xlarge) = 11 × $8,335 | ≈ $91,700 |
| Ledger shards, DR secondaries | 11 × 1 writer-sized db.r6g.8xlarge = 11 × $3,031 | ≈ $33,300 |
| Aurora storage | ~117 TB after year one × $0.10, billed once per cluster, in both primary and secondary regions | ≈ $23,400 |
| Aurora I/O, primary | ~6 billed write I/Os per payment (assumed) × 200M × 30.4 ≈ 36.5 billion × $0.20/M | ≈ $7,300 |
| Aurora Global replication | replicated write I/O in the secondaries ≈ $7,300, plus ~$100–200 of cross-region transfer | ≈ $7,500 |
| Step Functions Standard | cross-shard and cross-region transfers (2% of payments × 5 transitions = 20M) + disputes (0.1% × 200M × 6 = 1.2M) + payouts (200,000 merchants, assumed, × 5 = 1M) = 22.2M/day × $0.000025 × 30.4 | ≈ $16,900 |
| API fleets (Fargate) | 217 tasks of 2 vCPU across regions, sized as in Round 2 per region ≈ $15,600 + 60 worker tasks ≈ $2,200 | ≈ $17,800 |
| Risk service + Valkey | 18 tasks of 2 vCPU ≈ $1,300 + 16 cache.r7g.large Valkey nodes at roughly $0.175/h ≈ $2,000 | ≈ $3,300 |
| DynamoDB keys | 10 × Round 2 | ≈ $11,200 |
| KMS | 200M vault decrypts/day × 30.4 × $0.03 per 10K | ≈ $18,200 |
| MSK, Debezium | 4 clusters × 3 brokers ≈ $1,800 + ~8.4 TB storage ≈ $840 + 11 connectors × 2 units ≈ $1,770 | ≈ $4,400 |
| AWS WAF | 10 × Round 2 | ≈ $22,800 |
| Data transfer | 10 × Round 2, at US prices | ≈ $15,000 |
| Vaults, lakes, Glue, audit archive | 4 regional vault databases, S3, nightly jobs, $530 archive | ≈ $6,000 |
| Warm fleets in DR regions | a quarter of each region's API, risk and vault tasks, about 54 API tasks plus the rest | ≈ $4,500 |
| CloudWatch, logs, alarms | an estimate | ≈ $8,000 |
| Total | ≈ $291K/month at US prices; budget ≈ $300K |
Against the money. At the same list rate, fees on 200M payments are about $292M a day, roughly $8.9 billion a month. Infrastructure is about 0.003% of that. The costs that matter at this level are processing fees, fraud losses and chargebacks, and FX spread: a 0.1% improvement in authorization rate or fraud loss is worth more than the whole AWS bill. That's why routing, risk and FX are the architect's real levers.
R3.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| Hot-account strategy | Lock-free append for accounts with no floor; buckets or batching only where a check is needed | Balance reads become sums or are a minute behind; each hot account needs a deliberate choice |
| Home region vs multi-writer | One writer per account, in its home region | Cross-region payments need sagas; failover takes minutes; a region's merchants depend on that region |
| Risk friction vs fraud losses | Graduated actions, with 3DS as the step-up | Some real buyers get challenged or declined. We tune thresholds per merchant against measured fraud and conversion, not one global number. |
| Build vs buy | Our own processor on top of several PSPs | A PSP's platform product (for example Stripe Connect or Adyen for Platforms) gives vaults, payouts, 3DS and much of the compliance work out of the box. Below this scale, buying is usually right; we build because routing across PSPs, our own ledger and our own risk models are now worth more than the engineering cost. |
| Step Functions vs row + sweeper for captures | Row + sweeper for captures; Step Functions for rarer multi-step flows | Less out-of-the-box execution history for captures; we rely on payment_events and the ledger instead |
The opening question was: how do we make sure every payment happens exactly once in effect, and every cent is accounted for, when the network can't tell us what happened? The answer is now: at most once by construction (idempotency keys at our edge and at every PSP, one writer per account, compare-and-set state machines), every cent in a balanced journal (micros, per-currency balancing, explicit variance and FX accounts, hot accounts without shortcuts), and proven every day (three-way reconciliation, a hash chain sealed in write-once storage). The network still lies; the design never believes a silence.
R3.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| A region outage mid-saga | Sagas stuck, API errors in one region | Fenced failover (Trace 3). Sagas are restarted from replicated rows; every step is idempotent by its PSP key and journal key. Transfers from other regions into the failed one pause in transfers_in_transit until it's back, and the money stays accounted for. |
| FX provider outage | Stale rates; quote failures | We never quote from a rate older than 5 minutes (our choice). When rates are stale, we stop offering conversion for affected pairs: payments proceed in the merchant's currency (the buyer's issuer converts) or wait. Open quotes keep their promised rate until they expire. |
| Model false-positive storm | Risk declines jump from ~1% to 8%; approval rate falls; merchants complain | An alarm on decline rate against the 7-day baseline per merchant segment. The kill switch reverts to the previous model version, or to rules only, in one configuration change. New models only ever reach production after shadow mode. |
| A ledger shard failover (within a region) | Tens of seconds of errors for one shard's merchants | Standby promoted as in Round 2. Only that shard's merchants are affected; the other 10 shards don't notice. |
| Reconciliation mismatch spike | Breaks jump from dozens to thousands at one PSP | Almost always systemic: a PSP report format change, a missing file, a webhook outage. Payouts for affected merchants are held automatically; we check the file first (row counts, totals versus the bank deposit), then our webhook error logs, before anyone reviews individual rows. |
| Risk service down | Timeouts from the risk call | Fail to "rules only" with conservative limits (lower amounts, more 3DS), not to "allow everything" and not to "block everything". The choice is written down, because either extreme is expensive. |
R3.9 Runbook and Incident Response
Golden signals, per region and per PSP OPS 8 · REL 6
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Authorization rate per PSP | < 85% for 5 min, or 5 points below its 7-day baseline | P1 | Check the PSP's status; shift routing weight away (below) |
UNKNOWN payments, count and oldest age | > 500, or oldest > 15 min | P1 | Look for a PSP gray failure; check the resolver's error logs |
Saga failures (Step Functions ExecutionsFailed) | > 10 in 5 min | P2 | Read the failed executions' last step; redrive once fixed |
| Reconciliation breaks per PSP | > 0 unexplained after 07:00 | P1 over $10K, P2 below | Follow the break procedure (below) |
| Ledger imbalance: debits − credits per currency per shard | ≠0, checked every 5 min | P1 | Freeze payouts on that shard; this should be impossible, so treat it as a bug or tampering |
| Hash-chain verification | any mismatch | P1, security | Security incident process |
| Payout returns | > 1% of a day's payouts | P2 | Look for a bank-details change or a bank-side problem |
| Fraud: risk decline rate, and fraud disputes per 1,000 payments | 2× baseline | P2 | Check for an attack (counters by merchant) or a model problem |
| Aurora Global replication lag | > 5 s for 5 min | P2 | A lagging secondary means a bigger loss in a failover; with rds.global_db_rpo it will soon stall commits |
| Replication-slot lag (Debezium) | > 10 min | P2 | The connector is stuck; WAL is growing on the writer |
PSP failover procedure OPS 10
- Confirm it's the PSP: its error and latency metrics are bad while the other PSPs are fine.
- The breaker usually opens on its own. If the PSP is degraded but not failing enough to trip it, set its routing weight to zero with the parameter below.
- Never move
UNKNOWNpayments to another PSP. Watch the resolver drain them as the PSP recovers. - Tell affected merchants whose customers only have PSP-specific tokens.
- After recovery, bring the weight back in steps (10%, 50%, 100%) while watching its approval rate.
Reconciliation-break procedure
- Check the inputs: did every expected file arrive, with the expected row count and a net total matching the bank deposit?
- Group the breaks by category and cause. One cause behind many rows (a missed webhook batch, a new fee type) gets one fix.
- For "missing in our ledger": fetch the object from the PSP by reference and replay the event through the normal webhook processor, which posts the journal with its usual idempotency.
- For "amount mismatch": the PSP's number wins for money the PSP moved; post a correcting journal, reviewed by a second person.
- Release the merchant's payout hold only when its breaks are zero.
Go deeper: CLI playbook
Plain commands an on-call engineer runs, one at a time. Replace the names and ARNs with real ones.
text# 1. Alarms currently firing for payments in a region aws cloudwatch describe-alarms --region eu-central-1 --state-value ALARM --alarm-name-prefix payments- # 2. Take PSP A out of routing (the router reads this parameter every 30 s) aws ssm put-parameter --region eu-central-1 --name /payments/routing/psp-a/weight --value 0 --type String --overwrite # 3. Failed dispute or payout sagas in the last hour aws stepfunctions list-executions --region eu-central-1 --state-machine-arn arn:aws:states:eu-central-1:111122223333:stateMachine:payout-saga --status-filter FAILED --max-results 50 # 4. Redrive a failed Standard execution after fixing the cause (possible for 14 days after it ends) aws stepfunctions redrive-execution --region eu-central-1 --execution-arn arn:aws:states:eu-central-1:111122223333:execution:payout-saga:po_20261002_m4410 # 5. Replication state of one ledger shard's global database aws rds describe-global-clusters --global-cluster-identifier ledger-shard-05 # 6. Planned move of a shard's writer to another region (no data loss) aws rds switchover-global-cluster --global-cluster-identifier ledger-shard-05 --target-db-cluster-identifier arn:aws:rds:eu-west-1:111122223333:cluster:ledger-shard-05-euw1 # 7. Unplanned failover when the primary region is down (may lose the replication lag) aws rds failover-global-cluster --global-cluster-identifier ledger-shard-05 --target-db-cluster-identifier arn:aws:rds:eu-west-1:111122223333:cluster:ledger-shard-05-euw1 --allow-data-loss # 8. Point shard 5's writer flag at the new region (services poll it every 10 s) aws ssm put-parameter --region eu-west-1 --name /payments/shards/05/writer-region --value eu-west-1 --type String --overwrite # 9. Rerun last night's reconciliation for one date aws glue start-job-run --region eu-central-1 --job-name recon-nightly --arguments '{"--batch_date":"2026-10-01"}' # 10. List yesterday's unexplained breaks aws athena start-query-execution --region eu-central-1 --work-group finance --query-string "SELECT psp, category, COUNT(*) AS n, SUM(amount_micros) AS micros FROM recon.breaks WHERE batch_date = DATE '2026-10-01' GROUP BY psp, category"
The failover in command 7 can lose the last moments of replication. Command 8 flips the writer flag, but it can't reach the old region's copy while that region is down, so the fence doesn't depend on it: services in the old region stop writing shard 5 as soon as they can't refresh their flag lease (within 10 seconds), and services in the new region wait 15 seconds after the flip before writing. Aurora adds its own fence: the old primary can only rejoin as a read-only secondary. (When the old region is reachable, the runbook updates its copy of the parameter too.)
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | 11 shards so one failure touches a fraction of merchants; home regions with Aurora Global secondaries; fenced failover in about 13 minutes; RPO capped at 20 s; idempotent resumption of in-flight work REL 10 · REL 13 |
| Security | 3DS for strong authentication; pre-authorization risk scoring against card testing; residency by home region, including the DR copy; hash-chained ledger sealed with Object Lock; segregation of duties and break-glass SEC 3 · SEC 4 · SEC 7 · SEC 10 |
| Performance Efficiency | Checkout P99 ≈ 700 ms with risk and vault in parallel; lock-free hot accounts; shards sized from load-tested figures; counters pipelined in one round trip PERF 1 · PERF 3 |
| Cost Optimization | ≈ $291K/month derived (≈ $300K budgeted for non-US prices); $245K/month of waiting sagas removed; build versus buy argued with the real levers (fees, fraud, FX) COST 5 · COST 11 |
| Operational Excellence | Golden signals with first actions; PSP-failover and reconciliation-break procedures; ledger-imbalance and hash-chain alarms; shadow-mode model releases; a timed regional-failover runbook OPS 6 · OPS 8 · OPS 10 |
| Sustainability | Data stays in its home region (no global copies); old ledger partitions leave Aurora for S3, and audit data moves to archive storage after a year; Graviton throughout SUS 1 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Shards the ledger so that a payment's journal stays on one shard, and treats hot accounts with a strategy per account type, not one trick.
- Designs checkout as a state machine that can pause for a person, with deadlines enforced by the database's clock.
- Puts fraud scoring on the critical path with a latency budget, graduated actions and a feedback loop, and says what happens when it fails.
- Models FX in the ledger: original amounts, quotes, per-currency balancing, rounding variance and realized gain or loss.
- Chooses home regions with one writer per account, explains exactly why that can't double-spend and what a failover can lose, and bounds it.
- Makes the ledger tamper-evident with checks owned by someone else.
- Names the real economics: fees, fraud and FX, not servers, and can argue build versus buy.
Follow-up questions
-
"A merchant moves from Europe to the US. How do you move their ledger?" Answer: as a planned migration, not a replication trick. Freeze their writes briefly, copy their accounts' entries to a US shard, verify the balances and hash chain match, flip the directory and the routing table, then unfreeze. In-flight payments finish in the old home first. If residency rules forbid the move, their European data stays in Europe and only new activity starts in the US.
-
"Why not keep a global balance for each customer's wallet, readable anywhere?" Answer: we can offer a display balance anywhere, from replicas, clearly marked as possibly a moment old. A spendable balance can only be checked on the one writer for that wallet, in its home region. Two regions deciding about one balance is exactly how money gets double-spent.
-
"Could we use DynamoDB global tables for the ledger instead?" Answer: global tables replicate asynchronously by default and resolve conflicting writes with last-writer-wins, so two regions could both approve a spend and one record would be discarded. The multi-Region strong-consistency mode avoids that, but it runs in exactly three Regions and doesn't support transactions, and a journal is a multi-row transaction. It's a good fit for our small global directory, not for the ledger.
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: payment methods, do we store cards, refunds, volume | Restate the Round 1 design in 60 seconds | Restate the Round 2 design in 60 seconds |
| 5–15 min | Requirements + API (idempotency key, status codes, hosted fields) | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.6: key → unknown → state machine → webhooks → ledger → integers | Steps 2.1–2.7: saga, locking, outbox, routing, reconciliation, payouts, vault | Steps 3.1–3.6: shards and hot accounts, 3DS, fraud, FX, home regions, audit |
| 40–50 min | Numbers, cost against PSP fees, trade-offs | Numbers, cost (Step Functions, Aurora billed once), trade-offs | Numbers, latency budget, cost against fees, build vs buy |
| 50–60 min | Failures + pillar check | Failures + pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening a round: "Before I design, I'll ask what we must never get wrong. For payments that's charging twice or losing a payment, so every piece I add will defend 'at most once', 'every cent accounted for', or 'our books match the bank's'."
- When the network fails: "A timeout isn't a failure, it's an unknown. We don't guess: we ask the PSP with the same idempotency key, and until it answers, the payment is
UNKNOWN."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What if the client retries?" (REL 4) | The idempotency key is claimed atomically and the stored response is replayed; a different body with the same key is 422. | 1 | Step 1.1 |
| "What if the PSP times out?" (REL 5) | The payment is UNKNOWN; a resolver asks the PSP with the same key inside its key window, never fails it and never fails over. | 1–2 | Step 1.2, step 2.4 | |
| "What if one PSP gets slow?" (REL 10) | A bulkhead caps its in-flight calls, a breaker stops sending to it, and only provably unsent payments move to another PSP. | 2 | Step 2.4 | |
| "What if an AZ fails?" (REL 11) | A standby writer is promoted; two AZs of tasks (32 of 48 vCPUs) and two replicas carry peak; in-flight payments are finished by the resolver. | 2 | R2.8 | |
| "What if a region fails?" (REL 13) | Promote each shard's secondary, fence the old region with the writer flag, resume in-flight work by the same keys; RPO capped at 20 s. | 3 | Step 3.5, R3.5 | |
| Security | "Where do card numbers live?" (SEC 7) | Only in hosted fields and an isolated vault with KMS envelope encryption; everything else sees tokens. | 1–2 | R1.8, step 2.7 |
| "How do you trust webhooks?" (SEC 9) | HMAC over timestamp and raw body, constant-time compare, 300 s window, dedupe by event ID. | 1 | Step 1.4 | |
| "Could an admin change a ledger entry?" (SEC 4) | Not undetected: hash-chained batches sealed in Object Lock, verified daily by a separate team. | 3 | Step 3.6 | |
| Performance | "Where does the latency go?" (PERF 1) | About 90% is the PSP; our part is a key claim, a few writes, and a 50 ms risk check in parallel with the vault. | 1–3 | R1.7, R2.6, R3.6 |
| "How do you scale the ledger?" (PERF 3) | Shard by merchant so a journal stays on one shard; hot accounts append without locks or use buckets. | 3 | Step 3.1 | |
| Cost | "What does it cost?" (COST 5) | About $600, $46K and $291K a month, against about $2.2M, $888M and $8.9B a month in fees; the fees are the cost. | 1–3 | R1.7, R2.6, R3.6 |
| "Why not Step Functions for everything?" (COST 5) | Standard is priced per transition; sagas that mostly wait would have cost $152K a month in Round 2 and $245K in Round 3. | 2–3 | R2.6, R3.5 | |
| "Why build this instead of using a PSP's platform?" (COST 11) | Below this scale, buy; we build because routing across PSPs, our own ledger and our own risk models are worth more than the engineering effort. | 3 | R3.7 | |
| Operations | "How do you know the books are right?" (OPS 8) | Nightly three-way reconciliation with every row classified, a net check against bank deposits, and a ledger-imbalance alarm that must read zero. | 2–3 | Step 2.5, R3.9 |
| "What do you do when reconciliation breaks?" (OPS 10) | Hold the merchant's payouts, check the inputs, group by cause, replay missing events, correct with reviewed journals. | 2–3 | R3.9 | |
| Sustainability | "Where is the footprint?" (SUS 4) | Years of ledger and audit data: hot for two years, then S3, then archive storage; data stays in its home region. | 2–3 | R2.6, R3.6 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| At most once | Idempotency key with an atomic claim and stored response; a key to the PSP | Keys in DynamoDB with a lease on our own clock; Aurora's unique constraint as truth; one key per PSP attempt | One writer per account across regions; idempotent resumption after failover |
| Unknown outcomes | UNKNOWN + resolver; replay inside the PSP's key window | Never fail over an unknown; compensations that can fail safely | Unknowns resumed in the recovery region; lag-window losses found by reconciliation |
| Ledger | Double entry, append-only, integer micros, half-even | Row locks where a floor exists; write skew explained; outbox + CDC | Sharded by merchant; hot-account strategies; per-currency balancing; hash-chained |
| External truth | Webhooks verified, deduped, applied as allowed moves; disputes after refunds | Three-way reconciliation; payouts held on breaks; ACH returns | Global reconciliation from regional summaries; reproducible, sealed runs |
| Security | Hosted fields; secrets; signed webhooks | Isolated vault; envelope encryption; correct KMS rotation | 3DS; fraud scoring; residency; segregation of duties |
| Well-Architected trade-offs | Infra ≈ 0.03% of fees; correctness over cost | Standard vs Express by semantics and price; quotas raised and paced | Build vs buy; fees, fraud and FX as the real levers |
| Evolving under new scope | Builds from one PSP call, one problem at a time | Opens with "what breaks", fixes money-safety first | Changes the architecture's shape (shards, regions) without weakening the three rules |