Design a Crowdfunding Platform
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | Crowdfunding for one creative community (board games and indie gadgets) | A national platform: one hit campaign draws a flash crowd, and a campaign with a million backers ends | A global marketplace: many currencies, fraud and scam campaigns, moderation, three regions |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Volume | 3,000 live campaigns; 25,000 pledges/day; a launch peak of ~20 pledges/s; ~100 campaigns end a day | 30,000 live campaigns; 500,000 pledges/day; 3,000 pledges/s on one campaign; one campaign ends with 1M backers; ~1,000 campaigns end a day | 150,000 live campaigns; 3M pledges/day in 3 home regions; ~20 currencies; 2,000 new campaigns to review a day |
| Money | Pledges become charges at the deadline | + platform fees, creator payouts, refunds, failed-card recovery | + FX, mass refunds, chargebacks months after payout |
| Footprint | 1 region, 3 AZs | 1 region, 3 AZs; survives losing an AZ mid-settlement | 3 home regions, each with a disaster-recovery copy; pages served worldwide |
| Targets | Never charge an unfunded campaign's backers; never charge twice; never oversell a tier; 99.9% | Pledge P99 < 200 ms at 3,000/s; displayed total ≤ ~5 s stale; a 1M-pledge settlement in about an hour at an assumed negotiated PSP limit; 99.99% | Pledging up per region; settlement resumes after a region failure with no double charge; every cent auditable |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up. Card payments themselves are designed in Design a Payment Processing System; this page links there instead of re-teaching them, and does the same for the other loops it builds on.
Loop Opener: What Is Crowdfunding?
You Already Know One: a Pre-Order With a Promise
A creator says: "I'll make 1,000 copies of this board game if at least $50,000 is pledged by March 31." People who want the game pledge: they promise to pay, usually in exchange for a reward such as a copy of the game. If $50,000 or more is pledged by the deadline, every backer is charged and the creator gets the money to make the game. If not, nobody pays anything, and the creator gets nothing. That rule is called all-or-nothing.
It's a pre-order with a condition attached. The best-known example is Kickstarter. As of 2026-09-28, its help center says that when a campaign misses its goal "no money changes hands" (article "Why is funding all-or-nothing?"), and its blog says a backer's card "will be charged when the project reaches its funding deadline" ("When is My Card Charged?", 17 May 2024).
| Word | What it means on this page |
|---|---|
| Campaign | A creator's project page: a goal, a deadline, a description and reward tiers. |
| Creator | The person or company raising the money. |
| Backer | Someone who pledges to a campaign. |
| Pledge | A backer's promise to pay a stated amount if the campaign is funded. Not a payment. |
| Reward tier | A perk offered at a minimum pledge, often limited: "Early Bird, $49, 500 available". |
| Goal | The amount that must be pledged by the deadline. |
| Deadline | The instant the campaign ends. No pledges are accepted after it. |
| Funded / unsuccessful | The outcome at the deadline: total pledged ≥ goal, or not. |
| Settlement | Turning a funded campaign's pledges into charges, or releasing an unsuccessful campaign's pledges. |
| Payout | Sending the collected money, minus fees, to the creator. |
| Platform fee | The share of collected money the platform keeps. |
| Chargeback | A cardholder disputes a charge with their bank, and the money is pulled back. |
| Dunning | The process of recovering a payment that failed: telling the payer, retrying, and eventually giving up. |
Synthesizing vector architecture diagram...
The whole product in one picture. Weeks of promises on the left, one decision in the middle, and money that moves only on the right, after the deadline.
What Makes It Hard
- Money moves at one instant, weeks after the promise. A card can't simply be held for a month: card authorizations expire long before a campaign ends. We need a different way to be sure we can charge later.
- The first minute and the last minute. A hit campaign gets thousands of pledges a second when it launches. When it ends, a million charges must each happen exactly once, as fast as the payment provider allows.
- Limited perks. 500 Early Bird slots, 20,000 people who want one.
- Books that must close. Pledged, charged, failed, refunded, fees, paid out: every cent has to be accounted for, long after the campaign ended.
The Question the Whole Loop Answers
How do we collect millions of promises safely, and then, at one instant, turn each one into exactly one charge, or release it, for every campaign on the platform?
The answer gets sharper every round:
- Round 1: pledges are saved payment methods plus a ledger of promises; a fence closes pledging at the deadline; settlement charges each pledge with a per-attempt idempotency key.
- Round 2: a hot campaign that never locks one row, tiers that never oversell under a flash crowd, settlement as a rate-limited, resumable saga, failed-card recovery, and a money ledger with fees, payouts and refunds.
- Round 3: currencies, fraud, moderation, regions, and disputes that arrive months after the money left.
Round 1 · Mid-level · "Crowdfunding for One Creative Community"
~35 min · SDE II (L5) · 1 region, 3 AZs · 25,000 pledges/day, ~20 pledges/s at a launch peak · 99.9% · never charge an unfunded campaign, never charge twice, never oversell a tier
R1.1 Establish Design Scope
The interviewer says: "We run a community for board-game and gadget makers. Design the crowdfunding site: creators launch campaigns, backers pledge." Before we draw anything, we ask questions, and we say out loud what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| All-or-nothing, or does the creator keep whatever is raised? | All-or-nothing. | Money can't move until the deadline decides the outcome. The other model, flexible funding (the creator keeps whatever is raised), simply charges at pledge time like any shop; it's simpler and it's not what we're building. |
| How long do campaigns run? | Up to 60 days, usually 30. | A pledge is a promise that must stay chargeable for up to two months (step 1.1). Kickstarter's limit is the same: its help center says it can't extend a project "beyond the maximum funding period of 60 days" (article "Can a project be edited after launching?", checked 2026-09-28). Ours is our own policy. |
| When is the backer charged? | Only if the campaign is funded, at the deadline. | We must collect a way to charge later without charging now (step 1.1). |
| Reward tiers? | Yes. Some are limited in quantity. | A counter per limited tier that can never exceed its limit (step 1.4). |
| Can backers change or cancel a pledge? | Yes, until the deadline. | Pledges are edited and cancelled; the total must follow (step 1.3). |
| How do backers pay? | Cards, through one payment service provider (PSP). Card numbers never touch our servers. | The PSP's hosted card fields collect the card; we only store the PSP's references. The card mechanics are the payment loop's. |
| Fees, payouts to creators, other currencies? | Not yet. | One currency, no money ledger beyond charges yet. |
Out of scope for this round:
- Fees and payouts. We record what was collected; paying creators comes in Round 2.
- Multiple currencies, fraud, moderation. Round 3.
- Taxes on rewards, shipping and fulfilment. The creator's job, outside our system.
- Pledges after the deadline. Never.
The interviewer will widen this scope later. Write your out-of-scope list where you can see it: in a multi-round loop, some of it comes back.
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Requirement |
|---|---|
| "Creators launch campaigns" | POST /v1/campaigns (draft with goal, deadline, tiers) and POST /v1/campaigns/{id}/publish |
| "Backers see the campaign" | GET /v1/campaigns/{id}: goal, deadline, total pledged, tiers with remaining counts |
| "Backers pledge" | POST /v1/campaigns/{id}/pledges: an amount and an optional tier; the card is saved in a separate PSP step |
| "Change or cancel before the deadline" | PATCH /v1/pledges/{id} and DELETE /v1/pledges/{id} |
| "All-or-nothing at the deadline" | A job that closes pledging, decides funded or unsuccessful from the exact total, then charges every pledge or releases every pledge |
| "Tell people" | Emails to backers and the creator on pledge, outcome and charge |
Not yet: flash crowds, failed-card recovery, fees and payouts, refunds after charging, currencies.
R1.3 Non-Functional Requirements: the Questions
Numbers come in R1.7. For now, the invariants in plain words, most important first:
- Nobody is charged for an unfunded campaign. A charge that should never have happened costs a refund, the fee we don't get back, and the backer's trust.
- Each pledge of a funded campaign is charged at most once. Retries, crashes and duplicate webhooks must never produce a second charge.
- A limited tier never sells more than its limit. The creator promised 500 Early Bird copies, not 504.
- No pledge is accepted after the deadline, and no pledge that was accepted is left out of the funded decision.
- Pledging feels instant. Backers click and see "You're a backer" in well under a second.
- Auditability. Months later, we can explain every change to every total.
- Availability. If pledging is down during a campaign's last hours, the creator can miss their goal because of us.
R1.4 The API
Create and publish a campaign. Amounts are integers in the currency's minor unit (cents), never floats.
httpPOST /v1/campaigns HTTP/1.1 Authorization: Bearer <creator token> Idempotency-Key: 3f1c9a2e-7b44-4d0e-9a51-6c2f8e1d0b77 Content-Type: application/json { "title": "Harbor Lights: a cooperative lighthouse game", "goal_minor": 5000000, "currency": "USD", "deadline": "2027-03-31T23:59:00-04:00", "tiers": [ { "title": "Early Bird", "min_pledge_minor": 4900, "limit": 500 }, { "title": "Standard copy", "min_pledge_minor": 5900, "limit": null } ] }
It returns 201 with "campaign_id": "cmp_01JF3KQ8" and "status": "DRAFT". POST /v1/campaigns/cmp_01JF3KQ8/publish moves it to LIVE. The deadline is stored as an absolute instant (2027-04-01T03:59:00Z); the creator's time zone is only for display.
Read a campaign.
httpHTTP/1.1 200 OK Content-Type: application/json { "campaign_id": "cmp_01JF3KQ8", "status": "LIVE", "goal_minor": 5000000, "pledged_minor": 3712400, "backers": 612, "deadline": "2027-04-01T03:59:00Z", "tiers": [ { "tier_id": "tir_eb", "title": "Early Bird", "min_pledge_minor": 4900, "remaining": 37 }, { "tier_id": "tir_std", "title": "Standard copy", "min_pledge_minor": 5900, "remaining": null } ] }
Pledge. The client makes one idempotency key per pledge attempt: a random ID generated once and sent again on every retry of the same attempt, so the server recognizes the retry.
httpPOST /v1/campaigns/cmp_01JF3KQ8/pledges HTTP/1.1 Authorization: Bearer <backer token> Idempotency-Key: 8b0e5d7c-1a2f-4e93-b6c4-5d9f0a7e3c21 Content-Type: application/json { "amount_minor": 6000, "tier_id": "tir_eb", "payment_method_id": null }
A backer with no saved card gets a pledge that is waiting for one:
httpHTTP/1.1 201 Created Content-Type: application/json { "pledge_id": "plg_01JF4B2M", "status": "PENDING_SETUP", "hold_expires_at": "2027-03-02T15:00:00Z", "setup": { "provider": "psp", "client_secret": "seti_..._secret_..." } }
The browser hands client_secret to the PSP's card form. The card goes from the browser straight to the PSP, which saves it; the PSP's webhook (or the page's return call) then moves the pledge to ACTIVE. A backer who already has a saved card sends its reference in payment_method_id and gets "status": "ACTIVE" at once. A pledge left in PENDING_SETUP for 15 minutes expires and gives its tier slot back (step 1.4).
| Status | When |
|---|---|
201 Created | Pledge created, ACTIVE or PENDING_SETUP |
201 + Idempotent-Replayed: true | A retry of an attempt that already succeeded: the same pledge |
409 Conflict, "reason": "tier_sold_out" | The tier has no slot left. Nothing was created. |
409 Conflict, "reason": "campaign_closed" | The campaign isn't LIVE, or its deadline has passed by the database's clock |
409 Conflict, "reason": "already_backing" | This backer already has an active pledge on this campaign; change it with PATCH |
422 Unprocessable Entity | Amount below the tier's minimum, or the same key with a different body |
Change and cancel. PATCH /v1/pledges/plg_01JF4B2M with { "amount_minor": 10000, "tier_id": "tir_std" } changes the amount or tier; DELETE /v1/pledges/plg_01JF4B2M cancels. Both carry idempotency keys, both return 409 campaign_closed after the deadline.
The deadline is half-open. A pledge is accepted only while now() < deadline, checked by the database's clock inside the pledge's own transaction, never by an app server's clock read earlier. At exactly 03:59:00.000Z pledging is over. Step 1.5 shows why this check alone is not the whole story.
Recap
- Pledging is two steps for a new card: create the pledge, then save the card with the PSP.
- Every call that changes something carries an idempotency key.
- The total a page shows is not the thing we decide on; the ledger is (step 1.3).
- Correctness first: no charge without funding, no double charge, no oversold tier, no pledge after the deadline.
Let's build it, starting with the simplest thing that works.
R1.5 Design Evolution: From "Charge on Click" to a Promise Ledger
Every step below follows the same pattern: a problem, your turn to think, the answer, and what the answer costs us. The cost is always the next problem.
Step 1.0: The Baseline
A pledge charges the card immediately and adds the amount to the campaign's total:
sql-- On "Pledge": charge the card at the PSP first, then: UPDATE campaigns SET raised_minor = raised_minor + $amount WHERE campaign_id = $1; INSERT INTO pledges (campaign_id, backer_id, amount_minor, psp_charge_id) VALUES (...);
Synthesizing vector architecture diagram...
What's good about it: it's how a shop works, and it's two lines.
What it costs us: it ignores the one rule that makes this product crowdfunding. If the goal isn't reached, every backer was charged for nothing. The next steps fix that first, then the races underneath.
Step 1.1: Where Is the Money for 30 Days?
The problem: the campaign runs until March 31. A backer pledges $60 on March 2. If the goal is missed, they must pay nothing; if it's reached, we must be able to take $60 from them on March 31, when they may be asleep. Where is that money, or that right to charge, for 29 days? What would you do?
Synthesizing vector architecture diagram...
The card goes from the browser to the PSP; we only ever see references. Nothing is charged, and the pledge counts toward the total only once the card is saved.
Step 1.2: The Backer Double-Clicked and Has Two Pledges
The problem: a backer on a slow train clicks Pledge, sees a spinner, and clicks again. Two requests arrive. The campaign now shows two $60 pledges from the same person, and at the deadline we'd charge them twice. What would you do?
sqlCREATE UNIQUE INDEX one_live_pledge_per_backer ON pledges (campaign_id, backer_id) WHERE status IN ('PENDING_SETUP', 'ACTIVE');
Loop primitive: Idempotency & Effectively-Once Processing
Step 1.3: The Total on the Page Doesn't Match the Pledges
The problem: support gets a message: "Your page says $37,124 but my spreadsheet of backers adds up to $37,184." Somewhere a pledge change updated the pledge but not the total, or a cancel ran twice. With one raised_minor column edited in place, we can't even say what happened.
What would you do?
The ledger is append-only by construction, not by convention:
sqlCREATE TABLE pledge_ledger ( entry_id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY, campaign_id BIGINT NOT NULL, pledge_id BIGINT NOT NULL, request_id BIGINT NOT NULL, -- the pledge_requests row that caused it delta_minor BIGINT NOT NULL, kind TEXT NOT NULL CHECK (kind IN ('ACTIVATE','CHANGE','CANCEL')), created_at TIMESTAMPTZ NOT NULL DEFAULT now(), UNIQUE (request_id, kind), -- a replayed request can't add a second entry CHECK ((kind = 'ACTIVATE' AND delta_minor > 0) OR (kind = 'CANCEL' AND delta_minor < 0) OR kind = 'CHANGE') ); CREATE INDEX pledge_ledger_by_campaign ON pledge_ledger (campaign_id) INCLUDE (delta_minor); -- The service's role gets INSERT and SELECT only; a trigger also rejects UPDATE and DELETE. REVOKE UPDATE, DELETE ON pledge_ledger FROM pledge_service;
Related loops: the outbox and event-driven ledger loop (step 2.6, an auditable ledger) and the payment loop (step 1.5, double entry).
Step 1.4: The 500-Slot Early Bird Tier Sold 504
The problem: Early Bird is limited to 500. The pledge code counts existing Early Bird pledges and inserts a new one if the count is below 500. On launch morning the creator finds 504 Early Bird backers, and only 500 copies at the early price. What would you do?
sqlCREATE TABLE reward_tiers ( tier_id BIGINT PRIMARY KEY, campaign_id BIGINT NOT NULL REFERENCES campaigns(campaign_id), title TEXT NOT NULL, min_pledge_minor BIGINT NOT NULL CHECK (min_pledge_minor > 0), limit_qty INT CHECK (limit_qty > 0), -- NULL: unlimited, no counting claimed INT NOT NULL DEFAULT 0 CHECK (claimed >= 0), held INT NOT NULL DEFAULT 0 CHECK (held >= 0), CONSTRAINT never_oversell CHECK (limit_qty IS NULL OR claimed + held <= limit_qty) ); -- Card saved in time: the hold becomes a claim. UPDATE reward_tiers SET held = held - 1, claimed = claimed + 1 WHERE tier_id = $1; -- only after the pledge's own -- PENDING_SETUP -> ACTIVE update changed 1 row -- Card saved after the sweeper expired the hold: try to re-claim. UPDATE reward_tiers SET claimed = claimed + 1 WHERE tier_id = $1 AND claimed + held + 1 <= limit_qty; -- 0 rows: ask the backer to choose again
Primitive: Database Isolation Levels, ACID and Concurrency Anomalies
Drill: The on-call doctor anomaly (why snapshot isolation allows write skew is the "count, then insert" answer above; why not Serializable everywhere is the answer below it)
Step 1.5: A Pledge Arrived After the Deadline, and Another Committed After We Added Up the Total
The problem: the campaign ends at 03:59:00Z. At 03:59:00.4 a pledge is accepted. Worse: at 03:58:59.9 a pledge transaction starts, passes every check, and commits at 03:59:00.8, after the job that decides "funded or not" had already summed the ledger at 03:59:00.5. The decision didn't include a pledge we accepted. What would you do?
sql-- Inside every pledge-activation, change or cancel transaction (Read Committed): UPDATE campaigns SET pledged_minor = pledged_minor + $delta, backers = backers + $backer_delta WHERE campaign_id = $1 AND status = 'LIVE' AND now() < deadline; -- 0 rows: ROLLBACK and answer 409 campaign_closed. -- The close job, once now() >= deadline: UPDATE campaigns SET status = 'ENDED', ended_at = now() WHERE campaign_id = $1 AND status = 'LIVE' AND deadline <= now(); -- Waits for in-flight pledges' row locks. 1 row: we closed it. 0 rows: another job instance did.
Synthesizing vector architecture diagram...
Pledge A started before the deadline and committed after it, and it's still counted, because the close waited for it. Pledge B arrived after the close and was refused. Nothing slips between.
Loop primitive: Leases, Fencing Tokens & Distributed Locks (a fence is whatever makes a late writer's write fail where the data lives; here, a row lock and a status the late writer must re-read)
Step 1.6: The Deadline Passed. Who Settles, and How Do We Charge 250 Backers Exactly Once Each?
The problem: a funded campaign with 250 backers ends at 03:59. Someone must decide it's funded, then charge 250 saved cards, each exactly once, even if the worker crashes halfway, the PSP times out, or two workers pick up the same campaign. An unsuccessful campaign must release its pledges and charge nobody. What would you do?
Synthesizing vector architecture diagram...
Round 1's campaign lifecycle. Every arrow is one conditional update on the current status. Round 2 adds what happens after SETTLED (collection, reconciliation, payout); Round 3 adds review and suspension.
Synthesizing vector architecture diagram...
Round 1's pledge lifecycle. Only ACTIVE pledges count toward the total. A timeout leaves the pledge in CHARGING with an attempt the resolver asks the PSP about; it never moves on by guesswork.
The charge attempt, with its guard:
sqlCREATE TABLE charge_attempts ( pledge_id BIGINT NOT NULL, attempt_no SMALLINT NOT NULL, psp_key TEXT NOT NULL UNIQUE, -- 'plg_01JF4B2M:charge:1' status TEXT NOT NULL CHECK (status IN ('IN_FLIGHT','SUCCEEDED','FAILED')), psp_payment_id TEXT, -- stored as soon as the PSP answers decline_code TEXT, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), PRIMARY KEY (pledge_id, attempt_no) ); CREATE UNIQUE INDEX one_success_per_pledge ON charge_attempts (pledge_id) WHERE status = 'SUCCEEDED'; -- Start attempt n, in one short transaction, before calling the PSP: UPDATE pledges SET status = 'CHARGING', charge_attempt = charge_attempt + 1 WHERE pledge_id = $1 AND status = 'TO_CHARGE' RETURNING charge_attempt; -- 0 rows: someone else has it INSERT INTO charge_attempts (pledge_id, attempt_no, psp_key, status) VALUES ($1, $n, $pledge_public_id || ':charge:' || $n, 'IN_FLIGHT');
Primitive: Two-Phase Commit and Saga Orchestration
Drill: The GC pause that corrupted shared storage (why an expiring lease doesn't protect the charge is point 5; the conditional state change on the pledge is exactly "optimistic concurrency at the database instead of a lock")
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | Charge on click, a mutable total | Charges backers of campaigns that fail |
| 1.1 | Where is the money for 30 days? | Save the card now (SetupIntent), charge off-session at the deadline | Deadline failures; authentication_required; card setup is a separate step |
| 1.2 | Double click, two pledges | Idempotency key per request; one live pledge per backer per campaign | Changes need their own path |
| 1.3 | Total doesn't match pledges | Append-only pledge ledger; snapshot updated in the same transaction | Every pledge writes one row |
| 1.4 | 504 of 500 Early Birds | Tier counter row, conditional update, check constraint; 15-min holds with a sweeper; re-claim for a late card | A hot row when a tier opens |
| 1.5 | Pledges after the deadline or after the sum | Conditional on LIVE and now() < deadline, plus a closing fence on the campaign row | A lock per pledge on one row |
| 1.6 | Charge 250 backers exactly once | Campaign state machine; close job; per-pledge charge with a (pledge, attempt) key; one success per pledge; webhooks via an inbox | Settlement takes minutes; attempts to keep |
R1.6 Architecture v1
Now the concepts get AWS names.
Synthesizing vector architecture diagram...
One writer holds every invariant: the tier counters, the ledger, the campaign's status. The PSP is reached for card setup at pledge time and for charges only after the deadline. Emails never run inside a money transaction; they go through the outbox.
The pieces:
- Campaign and pledge service: one stateless container on ECS Fargate (containers without managing servers), three tasks, one per AZ. It serves the API, receives PSP webhooks, and runs the tier-hold sweeper.
- Aurora PostgreSQL: one writer and one replica in another AZ that Aurora promotes if the writer fails. Everything that decides anything is on the writer.
- Close job and settlement workers: two Fargate tasks. The close job polls every 30 seconds for due campaigns; workers claim settlement jobs with
SKIP LOCKED. - Outbox relay → SQS → notification worker → SES. Events ("pledge confirmed", "campaign funded", "you were charged") are written to an outbox table in the same transaction as the state change, relayed to SQS, and sent by a worker through Amazon SES. The outbox pattern and its relay are the outbox loop's steps 1.1 to 1.3. A new SES account starts in the sandbox: "a maximum of 200 messages per 24-hour period" and "1 message per second" (docs.aws.amazon.com/ses, request-production-access), and quotas are "separate for each AWS Region". We request production access before launch.
- Secrets: the PSP's API key and webhook signing secret live in AWS Secrets Manager.
Scheduler per deadline, or a poll? Amazon EventBridge Scheduler can fire a one-time schedule at an exact time ("A one-time schedule will invoke a target only once", docs.aws.amazon.com/scheduler), and its default quota is 10,000,000 schedules per Region, adjustable. One schedule per deadline would be precise to the minute. We choose the poll for Round 1: every 30 seconds, WHERE status = 'LIVE' AND deadline <= now(), on a partial index. It's one query, it can't miss a campaign whose schedule failed to be created or wasn't updated when the creator changed the deadline, and 30 seconds of delay is invisible next to a 30-day campaign. Correctness never depended on the trigger anyway: the fence and the status updates do.
The tables (the tier table is in step 1.4, the ledger in step 1.3, the attempts in step 1.6):
sqlCREATE TABLE campaigns ( campaign_id BIGINT PRIMARY KEY, public_id TEXT NOT NULL UNIQUE, -- 'cmp_01JF3KQ8' creator_id BIGINT NOT NULL, goal_minor BIGINT NOT NULL CHECK (goal_minor > 0), currency CHAR(3) NOT NULL, deadline TIMESTAMPTZ NOT NULL, status TEXT NOT NULL CHECK (status IN ('DRAFT','LIVE','CANCELLED','ENDED', 'FUNDED','SETTLING','SETTLED','UNSUCCESSFUL','RELEASED')), pledged_minor BIGINT NOT NULL DEFAULT 0, -- snapshot of the ledger sum (step 1.3) backers INT NOT NULL DEFAULT 0, final_minor BIGINT, -- exact sum after the fence ended_at TIMESTAMPTZ, published_at TIMESTAMPTZ ); CREATE INDEX campaigns_due ON campaigns (deadline) WHERE status = 'LIVE'; CREATE TABLE pledges ( pledge_id BIGINT PRIMARY KEY, public_id TEXT NOT NULL UNIQUE, -- 'plg_01JF4B2M' campaign_id BIGINT NOT NULL REFERENCES campaigns(campaign_id), backer_id BIGINT NOT NULL, tier_id BIGINT, -- NULL: no reward amount_minor BIGINT NOT NULL CHECK (amount_minor > 0), status TEXT NOT NULL CHECK (status IN ('PENDING_SETUP','ACTIVE','EXPIRED', 'CANCELLED','RELEASED','TO_CHARGE','CHARGING','CHARGED','FAILED')), psp_customer_id TEXT NOT NULL, -- the backer at the PSP psp_setup_id TEXT, -- the SetupIntent, if a new card psp_payment_method TEXT, -- set when the card is saved hold_expires_at TIMESTAMPTZ, -- only while PENDING_SETUP with a limited tier charge_attempt SMALLINT NOT NULL DEFAULT 0, created_at TIMESTAMPTZ NOT NULL DEFAULT now() ); CREATE INDEX pledges_hold_due ON pledges (hold_expires_at) WHERE status = 'PENDING_SETUP'; CREATE INDEX pledges_to_charge ON pledges (campaign_id, pledge_id) WHERE status = 'TO_CHARGE'; CREATE TABLE pledge_requests ( -- idempotency keys and stored replies backer_id BIGINT NOT NULL, idempotency_key TEXT NOT NULL, request_hash BYTEA NOT NULL, response JSONB NOT NULL, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), request_id BIGINT GENERATED ALWAYS AS IDENTITY UNIQUE, PRIMARY KEY (backer_id, idempotency_key) ); CREATE TABLE settlement_jobs ( campaign_id BIGINT PRIMARY KEY, kind TEXT NOT NULL CHECK (kind IN ('CHARGE','RELEASE')), lease_owner TEXT, lease_until TIMESTAMPTZ, done BOOLEAN NOT NULL DEFAULT false ); CREATE TABLE webhook_inbox ( -- dedupe PSP events by their ID event_id TEXT PRIMARY KEY, received_at TIMESTAMPTZ NOT NULL DEFAULT now() );
Synthesizing vector architecture diagram...
A pledge has many ledger entries (activate, change, cancel) and, if funded, one or more numbered charge attempts, at most one of them successful.
Trace 1: a new backer pledges for a limited tier. The sequence in step 1.1: one transaction writes the pledge PENDING_SETUP and held + 1 on Early Bird (stored with its idempotency key); the service creates a SetupIntent and returns its client secret; the backer types the card into the PSP's form; the setup_intent.succeeded webhook arrives, is deduplicated in the inbox, and one transaction moves the pledge to ACTIVE, moves the slot from held to claimed, appends ACTIVATE +6000 to the ledger and adds 6,000 to the campaign's snapshot, conditional on the campaign still being LIVE before its deadline.
Trace 2: a pledge after the deadline. At 03:59:00.3 a backer clicks Pledge on a saved card. The pledge transaction's conditional update on the campaign row waits behind the close job's update (step 1.5's diagram), re-reads ENDED, changes 0 rows and rolls back: 409 campaign_closed, and no ledger entry, no tier change, nothing to undo.
Trace 3: a funded campaign settles.
Synthesizing vector architecture diagram...
Each pledge is its own small transaction on each side of one PSP call. A crash anywhere leaves at most one pledge in CHARGING, which the next worker resolves by re-sending the same key.
Trace 4: an unsuccessful campaign. The exact sum is $41,300 against a $50,000 goal. One transaction moves the campaign to UNSUCCESSFUL, all ACTIVE pledges to RELEASED, and writes an outbox event; the notification worker tells 250 backers "not funded, you weren't charged". The campaign moves to RELEASED. The PSP is not called at all.
R1.7 Numbers
Pledges. From the interviewer's numbers plus our assumptions:
The average is useless here. Pledges bunch at launches and deadlines. We assume the community's most popular creator launches to a big mailing list and gets 1,200 pledges in the first minute: pledges a second on one campaign, about 70 times the whole platform's average. That's the rate we design for.
Campaigns ending. campaigns end a day. We assume 60% are funded (an assumption, not an industry figure):
- Funded: charges a day.
- Unsuccessful: pledges released a day, with no PSP calls.
- Creators like round deadlines, so we assume a quarter of the day's deadlines fall in one hour: 25 campaigns, 15 funded, charges. Our workers charge at 10 a second (our choice, far below the PSP's limit): s, about 6 minutes.
PSP calls against the PSP's limits. Stripe documents live-mode limits of "100 requests per second" per account and "25 requests per second" for individual API endpoints unless noted (docs.stripe.com/rate-limits, checked 2026-09-28). We assume 40% of pledges come from backers without a saved card, so each needs a SetupIntent: a day, and at the launch peak a second. Settlement adds 10 charges a second at worst. Both are well under 25 a second per endpoint and 100 a second overall. In Round 1 the PSP's limit is a non-issue; in Round 2 it becomes the constraint.
Ledger and table rows. We assume 10% of pledges are changed once and 5% cancelled: ledger entries a day. A ledger entry, from PostgreSQL's row layout (an estimate):
| Part | Bytes |
|---|---|
| Row header (23, padded to 24) | 24 |
entry_id, campaign_id, pledge_id, request_id, delta_minor: 5 × 8 | 40 |
created_at | 8 |
kind (short text, 1-byte header + up to 8 characters), padded to 8 | 16 |
| Line pointer | 4 |
| Heap row | 92 |
Primary-key and (request_id, kind) index entries, the covering campaign index: about 20 + 28 + 28 | 76 |
| Total per entry | ≈ 170 |
MB a day, about 1.8 GB a year. Pledges (≈ 400 B each with indexes) add 10 MB a day, charge attempts (15,000 a day at ≈ 300 B) 4.5 MB, and idempotency rows (kept 30 days) about 1 GB in steady state. Everything fits in the memory of the smallest instance for years. Size is not our problem; correctness is.
Availability. 99.9% of a 30.4-day month is minutes.
Latency. A pledge activation is about five statements on the writer (request row, pledge update, tier update, ledger insert, campaign snapshot update) at about 1 ms each plus a durable commit of a few milliseconds: about 15 ms at the median. The card entry at the PSP takes as long as the backer takes, and isn't ours to measure.
Monthly cost (us-east-1 on-demand prices, 730 hours a month):
| Item | Math | Monthly |
|---|---|---|
Aurora PostgreSQL, 2 × db.r6g.large | 2 × $0.26/h × 730 h | ≈ $380 |
| Aurora storage and I/O | a few GB × $0.10/GB-month, plus a few dollars of I/O | ≈ $10 |
| Fargate, 5 tasks of 1 vCPU / 2 GB | 5 × ($0.04048 + 2 × $0.004445)/h × 730 h | ≈ $180 |
| Application Load Balancer | $0.0225/h × 730 h, plus a few capacity units | ≈ $25 |
| AWS WAF | $5 per web ACL + 5 rules × $1 + ~5M requests × $0.60/M | ≈ $13 |
| NAT gateways, 3 | 3 × $0.045/h × 730 h | ≈ $99 |
| SES | ~51,000 emails a day (pledge receipts, outcome emails, creator notices) × 30 = 1.53M × $0.10 per 1,000 | ≈ $153 |
| SQS, Secrets Manager, CloudWatch | an estimate | ≈ $25 |
| Total | ≈ $885/month |
Against the money moved. We assume an average pledge of $60: 15{,}000 \times \60 = $900{,}000 collected a day, about \27M a month. At a common card list rate of 2.9% + 30¢ (an assumption; our contract decides), each $60 charge costs $2.04 in PSP fees: about $918K a month. The infrastructure is about 0.003% of the money it moves. And the funding-model choice is worth real money: under charge-now-and-refund, the 10,000 pledges released each day would have been charged and refunded, losing about 10{,}000 \times \2.04 = $20{,}400 a day in fees, about \612K a month, about 690 times this whole platform's infrastructure bill.
R1.8 Trade-Offs
| Choice | Option A | Option B | Our pick for Round 1 |
|---|---|---|---|
| Funding model | Save the card, charge at the deadline | Charge now, refund on failure (or authorize and capture) | Save and charge later. Charge-now is simpler (no deadline failures, no dunning, trivial settlement) and fine for a small platform with few failed campaigns, but it pays PSP fees on every refund (about $612K a month at our numbers) and ties up backers' money for weeks. Authorize-and-capture can't cover a 30- to 60-day campaign. The operational cost of our choice is a settlement pipeline and failed charges, which Round 2 handles. |
| Deadline trigger | Poll every 30 s | A one-time schedule per deadline (EventBridge Scheduler) | The poll: one indexed query, can't miss a campaign whose schedule wasn't created or updated, and 30 s of delay doesn't matter. Schedules win when there are millions of timers with no cheap query to find the due ones; Round 2 revisits it. |
| The total | Derived from the ledger, with a snapshot in the same transaction | A counter column only, or a SUM on every read | Derived plus snapshot. A counter alone has no history; a SUM on every page view is correct but wasteful. The snapshot is exact because it's written with the entry, and checkable because the ledger exists. |
| Settlement control | Orchestrated: a job table driven by one worker per campaign | Choreographed: modules react to each other's events ("campaign failed" → "refund contributions") | Orchestrated. One place knows how far a settlement got, which Round 2 needs for checkpoints, rate limits and an ETA. Choreography keeps modules decoupled (the reference implementation uses it for refunds) but spreads a settlement's progress across several event handlers. |
R1.9 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| A settlement worker crashes mid-campaign | One pledge left CHARGING, the job's lease stops being renewed | After the lease expires, another worker claims the job. The resolver finds the CHARGING pledge with an IN_FLIGHT attempt and re-sends the same key: the PSP returns the stored result (or charges for the first time if the first request never reached it). No double charge. |
| The PSP is down at the deadline | Charge errors; the settlement's progress stalls | The campaign's outcome is already decided and recorded; only the charging waits. Workers back off and retry. We tell backers and the creator "funded; charges are delayed". A timeout is an unknown, not a failure: nothing moves to a new attempt until the PSP answers about the old one (the payment loop's step 1.2). |
| The database fails over during a pledge | Connection errors for tens of seconds | An uncommitted pledge rolled back completely: no pledge, no hold, no ledger entry. The client retries with the same idempotency key once the new writer is up; if the commit had succeeded and only the reply was lost, the retry replays the stored response. |
| A duplicate or out-of-order webhook | The same setup_intent.succeeded twice, or a charge's failure event after its success | The inbox's primary key drops the duplicate. State moves are conditional, so an event that isn't allowed from the current state is recorded and ignored, or resolved by asking the PSP first. |
| The close job is down at the deadline | Due campaigns stay LIVE past their deadline | No pledge can be accepted after the deadline anyway: every pledge write checks now() < deadline by the database's clock. The campaign is closed late, and all pledges that were accepted were accepted in time. We alarm when a due campaign is more than 5 minutes past its deadline. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Idempotency keys on every write; per-attempt PSP keys; conditional state moves; a closing fence; Aurora replica in a second AZ, tasks in three REL 4 · REL 11 |
| Security | Card details go browser-to-PSP, never to us; webhooks verified by signature and timestamp; PSP keys in Secrets Manager; the app's database role can insert into the ledger but not update or delete it SEC 3 · SEC 9 |
| Performance Efficiency | ~15 ms pledge activations; the snapshot total avoids a SUM per page view PERF 3 |
| Cost Optimization | About $885 a month, derived, and about 0.003% of money moved; the funding model avoids about $612K a month of refund fees COST 5 |
| Operational Excellence | Light this round: alarms on campaigns past their deadline and still LIVE, pledges stuck in CHARGING, and snapshot-versus-ledger differences OPS 8 |
| Sustainability | Skipped this round: a handful of small tasks and two small instances. |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Says "a pledge is a promise, not a payment", and compares the three funding models with what each costs, including why authorizations don't last.
- Uses idempotency keys plus a business-level unique rule (one live pledge per backer per campaign).
- Keeps an append-only ledger of promises and derives the total from it.
- Protects a limited tier with a counter row, a conditional update and a check constraint, and names write skew in "count, then insert".
- Closes pledging with a fence that waits for in-flight pledges, not with a timestamp comparison.
- Charges each pledge with a (pledge, attempt) key, retries the same attempt with the same key, and allows one success per pledge.
Follow-up questions
-
"Why not keep the pledge as a PaymentIntent at the PSP and let the PSP be our database?" Answer: the PSP knows about cards and charges, not about goals, tiers, deadlines or who changed their pledge from $60 to $100 on day 12. The funded decision needs an exact sum under our own fence. We store the PSP's references on our pledges and treat the PSP as the truth for one thing only: whether money moved.
-
"A backer's card was saved at 15:02, but their Early Bird hold expired at 15:00. What happens?" Answer: the activation tries to re-claim a slot with the same conditional update. If one is free, the pledge activates normally. If not, the backer is asked to choose another tier or no reward. Unlike the hotel loop's version of this race, nothing was charged, so there's nothing to refund.
-
"Can a backer cancel one minute before the deadline and sink the campaign?" Answer: in this design, yes: cancels are allowed until the deadline. That's a real complaint from creators, and it has a known fix: a rule that in the last 24 hours a backer can't cancel or lower a pledge if it would drop the campaign below its goal. It hides a write-skew trap of its own. Round 2's step 2.1 builds it.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Authorize now, capture at the deadline" | Online card authorizations usually last about 7 days; extended ones reach 30 days at best, on some cards and merchant types. Campaigns run 30 to 60. |
| "Count the tier's pledges, then insert" | Write skew: two transactions count 499 and both insert. Materialize the count into a row and update it conditionally. |
| "Check the deadline in the app" | Clocks differ, and the check and the insert are two moments. The database decides, inside the write. |
"now() is when the pledge committed" | In PostgreSQL it's the transaction's start time; a pledge can commit after the decision summed the ledger. Close with a fence. |
| "One idempotency key per pledge for every retry" | The PSP replays the first result, a decline included; retries after a failure need a new attempt and a new key. |
| "Charge everyone in one big transaction" | Hours of locks, and a rollback that forgets charges which really happened at the PSP. |
Round 2 · Senior · "A Hit Campaign, a Million Backers, and Settlement Day"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 500,000 pledges/day, 3,000 pledges/s on one campaign · a 1M-pledge settlement · pledge P99 < 200 ms · 99.99%
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built all-or-nothing crowdfunding for one community: 3,000 live campaigns, 25,000 pledges a day, a launch peak of about 20 a second, and about 100 campaigns ending a day. A pledge is a promise, not a payment: the PSP saves the backer's card with a SetupIntent, and we charge it off-session only if the campaign is funded, because card authorizations last about 7 days and campaigns run 30 to 60. Idempotency keys stop retries, and a partial unique index allows one live pledge per backer per campaign. Every change is an entry in an append-only pledge ledger, with a snapshot total on the campaign row updated in the same transaction. Limited tiers are counter rows with a conditional update and a check constraint; a new card holds a slot for 15 minutes. Pledging closes with a fence: every pledge updates the campaign row conditionally on
LIVEandnow() < deadline, so the close job's status update waits for in-flight pledges, and only then do we sum the ledger. Settlement charges each pledge with a key per (pledge, attempt), one success per pledge, webhooks through an inbox. One Aurora writer, a few Fargate tasks, about $885 a month. Open costs: every pledge writes one campaign row, a tier row is contended when it opens, settlement is a simple loop, and failed cards just fail."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: the writer owns every invariant, the PSP saves cards at pledge time and charges them only after the deadline.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Money for 30 days | Save the card now, charge off-session at the deadline | Deadline failures; setup is a separate step |
| 1.2 | Double click | Idempotency key; one live pledge per backer | Changes need their own path |
| 1.3 | Total doesn't match | Append-only pledge ledger + snapshot in the same transaction | One row written by every pledge |
| 1.4 | Oversold tier | Counter row, conditional update, check constraint; 15-min holds | Hot tier row at opening |
| 1.5 | Pledge after the deadline or the sum | DB-clock condition + a fence on the campaign row | A lock per pledge on one row |
| 1.6 | Charge exactly once | State machine; close job; (pledge, attempt) keys; one success per pledge | Minutes of settlement; attempts to keep |
Open costs: the campaign row is written by every pledge; tier rows at opening; settlement with no rate plan; no recovery for failed cards; no fees, payouts or refunds.
R2.1 The Scope Raise
Interviewer: "The community site grew into a national platform: about 30,000 live campaigns and half a million pledges a day. Next month a famous designer launches at 10:00. Last time they launched, the first minute saw about 3,000 pledges a second on that one campaign, and the 1,000 Early Bird slots went in seconds. That campaign will end in 30 days with about a million backers, all to be charged, and the creator wants their money soon after. Many campaigns end at the same hour, because creators pick round-number deadlines. Backers try to cancel in the final minutes, and creators complain they were sunk at the last moment. A few percent of cards fail at the deadline. We now take a fee and pay creators out, and we sometimes refund backers. And we must survive losing an AZ in the middle of settlement."
We ask back before fixing anything.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How fresh must the total on the page be? | Seconds are fine. | The displayed total can come from a stream instead of a row every pledge writes (step 2.1). |
| Does the "goal reached" moment need to be exact? | Only at the deadline. The "we're funded!" banner can be a few seconds late. | Two numbers: a display total that lags, and an exact total for the decision (step 2.1). |
| How fast must settlement finish? | Hours are fine; days are not. | Settlement is planned against the PSP's rate limit, not our database (step 2.3). |
| What's our PSP rate limit? | The documented defaults, until we ask for more. | Stripe's defaults are per account and per endpoint, and a flash launch spends the same budget as settlement (steps 2.2, 2.3). |
| Can a backer cancel at the last minute? | Not if it would sink a campaign that has reached its goal. | A last-24-hours rule, and the write skew it hides (step 2.1). |
| What if a card fails at the deadline? | Backers get 7 days to fix it, then we try once more, then the pledge is dropped. | Dunning, with a new attempt per retry (step 2.4). |
| When is the creator paid, and how much do we keep back? | After collection closes and the books match; we hold a reserve for refunds. Creators must be verified by the PSP before they launch. | A money ledger, reconciliation before payout, a reserve, and a verification gate at launch (steps 2.5, 2.6). |
Two of those answers are Kickstarter's documented policies, which we adopt as our own. Kickstarter's help center (article "What happens when there is a problem with one of my backer's pledges?", checked 2026-09-28) says: "Backers with errored pledges will have 7 days from the project's deadline to update their payment method. At the end of this period, we'll automatically try to collect their pledge again. If this attempt fails, the backer will be dropped from the campaign." And its blog ("The Simple Steps to Cancel Your Kickstarter Pledge", Kickstarter Staff, 21 Mar 2024) says: "It is not possible to decrease or cancel a pledge during the last 24 hours of the campaign if doing so would drop the project below its funding goal."
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Live campaigns | 3,000 | 30,000 |
| Pledges | 25,000/day; ~20/s launch peak | 500,000/day (5.8/s average); 3,000/s on one campaign |
| Largest settlement | 250 pledges | 1,000,000 pledges |
| Campaigns ending | ~100/day | ~1,000/day, ~250 in the busiest hour |
| Money | Charges | + platform fee, creator payouts, a reserve, refunds, failed-card recovery |
| Survive | A task or a database failover | Losing an AZ mid-settlement |
| Targets | 99.9% | Pledge P99 < 200 ms at 3,000/s (at our load balancer, card entry excluded); displayed total ≤ ~5 s stale; a 1M-pledge settlement in about 70 minutes at an assumed negotiated PSP limit (11 hours at the default); 99.99% (4.4 min/month) |
The "Not yet" list from R1.2 comes back: flash crowds, failed-card recovery, fees, payouts and refunds are now in scope. Currencies are still out.
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | What breaks at the new scope |
|---|---|
| Every pledge updates the campaign's snapshot row | 3,000 pledges a second queue on one row lock. If each holds the lock for about 3 ms (the update plus a durable commit; a rough figure), one row allows about 1,000 / 3 ≈ 330 pledges a second. We'd need a tenth of that lock time. |
| The closing fence is that same row lock | Take the lock away and the fence goes with it. A FOR SHARE lock instead would let pledges run side by side, but thousands of transactions sharing one row lock go through PostgreSQL's multi-transaction lock machinery, a known contention point. |
| One tier row takes every Early Bird request | 20,000 people in 10 seconds for 1,000 slots: 2,000 requests a second on one row, and 95% of them only learn "sold out". |
| A SetupIntent per new card, "far below the limit" | At 3,000 pledges a second, the new-card share alone is far above Stripe's default of 25 requests a second per endpoint. |
| Settlement as a loop at 10 charges a second | A million charges would take 100,000 s, 28 hours, and there's no plan for 429s, a crash halfway, or 250 other campaigns ending that hour. |
| Failed charges just fail | 40,000 backers who wanted to pay are silently dropped, and the creator's funding shrinks. |
| No money ledger | No place for fees, payouts, reserves or refunds; finance can't reconcile. |
| Emails from a simple worker | A million outcome emails at one deadline, against an email quota. |
The order we fix it in: the hot row and the closing fence (2.1), hot tiers and the PSP budget at launch (2.2), settlement at scale (2.3), failed charges (2.4), webhooks and reconciliation (2.5), money, fees and payouts (2.6), notifications (2.7).
R2.3 New Requirements and API Additions
The pledge response no longer waits for the total. A pledge is exact and durable when 201 returns; the campaign's displayed total catches up within seconds. The backer's own pages ("You back this campaign: $60, Early Bird") read their own pledge from the writer, so they always see their own write (loop primitive Replication, Quorums & Read-Your-Writes).
Campaign totals, labeled.
json{ "campaign_id": "cmp_09HZ", "pledged_minor": 4130927400, "backers": 688211, "as_of": "2027-05-02T10:14:03Z", "exact": false, "tiers": [ { "tier_id": "tir_eb", "remaining": 0 }, { "tier_id": "tir_std", "remaining": null } ] }
After the deadline the same call returns "exact": true and the fenced final_minor.
Settlement status, per campaign.
json{ "campaign_id": "cmp_09HZ", "status": "SETTLING", "charged": 412000, "failed": 9310, "remaining": 578690, "unknown": 12, "rate_per_s": 250, "eta": "2027-06-01T05:02:00Z" }
Fix your payment (a backer in dunning): POST /v1/pledges/{id}/payment-method with a new saved-card reference, or POST /v1/pledges/{id}/confirm which returns the declined payment's client secret when the bank asked for authentication (step 2.4).
Creators' money. GET /v1/creators/{id}/verification returns the PSP's verification status; POST /v1/campaigns/{id}/publish returns 409 creator_not_verified until it's complete. GET /v1/campaigns/{id}/balance returns collected, fees, reserve, paid out and the payout date.
Refunds. POST /v1/pledges/{id}/refunds with an Idempotency-Key, an amount and a reason (CREATOR_REQUEST, DUPLICATE, CAMPAIGN_SUSPENDED, SUPPORT), for support staff and creators; backers ask through support.
Money-ledger entry types (step 2.6): CHARGE, PSP_FEE, REFUND, PLATFORM_FEE, FEE_RECHARGE, RESERVE_HOLD, RESERVE_RELEASE, PAYOUT.
R2.4 Design Evolution: The First Minute and the Last Minute
Step 2.1: 3,000 Pledges a Second, and Every One Updates the Same Row
The problem: the designer's campaign launches at 10:00. Every pledge transaction updates the campaign's snapshot row, which is also our closing fence. At about 3 ms of lock time per pledge, that row takes about 330 pledges a second, and 3,000 arrive. What would you do?
Synthesizing vector architecture diagram...
The pledge path only inserts. The total is recomputed from the ledger and written as an absolute value, so a replayed event can cost a recompute but never a double count.
Synthesizing vector architecture diagram...
The Round 2 fence. It's the same promise as Round 1's ("the sum includes everything accepted") bought with a database-enforced time limit instead of a lock every pledge takes.
Primitives: Change Data Capture and the Outbox Pattern · loop primitive Sharding, Hot Keys & Rebalancing
Drill: The on-call doctor anomaly (the last-minute cancel is the same write skew as step 1.4's "count, then insert"; why not Serializable for everything is in step 1.4)
Step 2.2: 1,000 Early Bird Slots, 20,000 People in the First 10 Seconds
The problem: at 10:00:00 the Early Bird tier opens: 1,000 slots, and 20,000 people click within 10 seconds, 2,000 a second on one tier row. And 40% of the 3,000 backers a second have no saved card, so each needs a SetupIntent from the PSP, whose documented default is 25 requests a second per endpoint. What would you do?
Synthesizing vector architecture diagram...
The database takes both lanes at full speed; only the new-card lane is metered, because only it spends the PSP's budget. Holds are taken before the queue, so the queue's length must stay well inside the hold.
Loop primitive: Retries, Timeouts, Backpressure & Load Shedding (admission before the expensive work, and a client-side token bucket in front of someone else's rate limit)
Step 2.3: The Campaign Ends: a Million Charges, 250 Other Campaigns, and a Rate Limit
The problem: at 23:59 the designer's campaign ends with 1,000,000 backers. In the same hour, about 250 other campaigns end (we assume a quarter of the day's 1,000 deadlines land in the busiest hour). Every funded pledge needs one PSP call. The PSP allows N requests a second, shared by everything we do. What would you do?
Synthesizing vector architecture diagram...
The checkpoint is the pledges' own states. Nothing needs a separate progress counter that could disagree with them.
Primitive: Two-Phase Commit and Saga Orchestration · loop primitives Idempotency & Effectively-Once Processing and Retries, Timeouts, Backpressure & Load Shedding · related loop: Design a Distributed Job Scheduler
Step 2.4: 40,000 Cards Failed at the Deadline
The problem: we assume 4% of cards fail at the deadline: expired cards, insufficient funds, banks that want the cardholder to authenticate. On the million-backer campaign that's 40,000 backers who wanted to pay. What would you do?
Synthesizing vector architecture diagram...
Round 2's charging lifecycle. Every retry is a new attempt with a new key, and every arrow into CHARGING is a conditional update, so at most one attempt is ever in flight for a pledge.
Step 2.5: Our Records, the PSP's Webhooks and the PSP's Reports Disagree
The problem: after the big settlement, our database says 958,000 charged; the PSP's balance report lists 958,014 successful payments for the campaign; and a few hundred webhooks arrived for payments we still show as CHARGING.
What would you do?
Step 2.6: Where Do the Fees Go, When Is the Creator Paid, and What About Refunds?
The problem: the platform keeps 5% of what it collects, the PSP keeps its processing fees, the creator gets the rest, and some backers are refunded. Finance asks where every dollar of the campaign went, and when the creator will be paid. What would you do?
One campaign's money, end to end. "Harbor Lights" collects 1,000 charges of $60. We assume a PSP fee of 2.9% + 30¢ ($2.04 per charge), our 5% platform fee on collected money net of refunds, and one refund before payout.
| # | Event | Debit | Credit | Amount |
|---|---|---|---|---|
| 1 | 1,000 charges succeed (summed) | PSP balance | Campaign clearing | $57,960.00 |
| 1 | Processing fees | Campaign clearing | $2,040.00 | |
| 2 | One refund before payout | Campaign clearing | PSP balance | $60.00 |
| 3 | Split at RECONCILED: clearing holds $59,940 | Campaign clearing | Platform fee revenue | $2,997.00 |
| 3 | Campaign clearing | Processing fees (charged to the creator) | $2,040.00 | |
| 3 | Campaign clearing | Creator payable | $54,903.00 | |
| 4 | Reserve, 10% of the creator's share | Creator payable | Creator reserve | $5,490.30 |
| 5 | Payout | Creator payable | PSP balance | $49,412.70 |
Check: clearing gets $60,000 in and $60 + $2,997 + $2,040 + $54,903 = $60,000 out, so it's back to zero. The PSP balance holds \57{,}960 - $60 - $49{,}412.70 = $8{,}487.30, which is exactly our fee (\2,997) plus the creator's reserve ($5,490.30). The processing-fee account is back to zero: the fees were charged to the creator. Each row pair balances on its own.
sqlCREATE TABLE money_entries ( entry_id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY, journal_id BIGINT NOT NULL, -- one business event; its entries sum to zero account TEXT NOT NULL, -- 'clearing:cmp_09HZ', 'creator_payable:crt_77', 'platform_fee' amount_minor BIGINT NOT NULL, -- positive debit, negative credit currency CHAR(3) NOT NULL, kind TEXT NOT NULL CHECK (kind IN ('CHARGE','PSP_FEE','REFUND','PLATFORM_FEE', 'FEE_RECHARGE','RESERVE_HOLD','RESERVE_RELEASE','PAYOUT')), source_ref TEXT NOT NULL, -- 'plg_..:charge:2', 'payout:cmp_09HZ' created_at TIMESTAMPTZ NOT NULL DEFAULT now(), UNIQUE (source_ref, account, kind) -- a replayed event can't post twice ); -- A deferred constraint trigger checks that each journal_id sums to zero at commit; no UPDATE or DELETE.
Step 2.7: A Million Emails at the Deadline
The problem: when the big campaign ends, a million backers should hear "funded", and then each should get a receipt when charged. Creators also post updates to all their backers. What would you do?
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | 3,000 pledges/s on one row | Append-only pledge path; total recomputed from the ledger via CDC, set as an absolute value; exact SUM after close; close + quiesce window beyond a 2 s transaction timeout; last-day cancels serialized per campaign | Seconds-stale display; two numbers; a 3 s longer close |
| 2.2 | Hot Early Bird tier and the PSP at launch | Sold-out short-circuit; hotel-style gate for big launches; a new-card lane metered by the PSP budget | "Sold out" a moment too long; minutes of waiting for new cards |
| 2.3 | 1M charges under a rate limit | Resumable saga, batches, pledge states as checkpoints; global token bucket; round-robin across campaigns; lookups after the key window | ~72 min at an assumed 250/s (11 h at the default) |
| 2.4 | 40,000 failed cards | Dunning: 7-day window, new attempt per retry, on-session confirm for authentication_required, final attempt, DROP; stays funded | Money arrives over a week; a fraud incentive |
| 2.5 | Records disagree | Inbox under the burst; per-campaign reconciliation before payout | Payout waits |
| 2.6 | Fees, payouts, refunds | Double-entry money ledger; clearing → fee, recharge, payable; 10% reserve; verified creators; refunds as compensating entries | Delayed reserve; negative balances possible |
| 2.7 | A million emails | Outbox events → queue → idempotent, quota-paced senders | Emails lag by minutes |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
Follow a pledge down the left: gate, pledge service, one insert-only transaction. Everything that shows a number (totals, remaining slots) is recomputed from the ledger off the change stream. Everything that calls the PSP takes a token from the same buckets first.
The pieces that are new since Round 1:
- CloudFront serves campaign pages. The page's static parts are cached for minutes; the small totals-and-tiers JSON is cached for 2 seconds. That short TTL is on purpose: a total or a "37 left" is wrong within seconds, and at the deadline the status flips to
ENDED. Caching those forever would show a live "Pledge" button on a closed campaign. When a hot page's entry expires, CloudFront's request collapsing sends one request per edge location to the origin while the others wait for its response (CloudFront developer guide, "Request and response behavior for custom origins"). If Valkey itself restarts empty, the origin doesn't let every miss recompute totals: concurrent misses in one task share one read (single-flight), a short lock key (SET lock:<campaign> <task> NX PX 2000) lets one task per campaign refill it, and the rest serve "total updating" for a second instead of hitting the database.
Drill: The product page that melted Redis (the cold-cache stampede is the request-collapsing and single-flight answer above; why not cache forever is the 2-second TTL and the status flip at the deadline)
- Waiting room and gate for scheduled launches (step 2.2), scaled up before the launch time.
- Valkey (ElastiCache): display totals, tier
remainingcounts, and the PSP token buckets. A primary and a replica in two AZs. Nothing in it is a source of truth: totals are recomputed, and a lost token bucket just starts full, which the PSP's own 429s correct. - Debezium on MSK Connect → MSK, topics keyed by
campaign_id, as in the hotel loop's step 2.1. We choose MSK over Kinesis because the Debezium connector writes Kafka natively and we reuse the topics for notifications. - Aggregator: marks campaigns dirty from the stream and recomputes totals from a replica at most every 2 seconds each.
- Settlement orchestrator, workers, dunning scheduler and resolver: Fargate tasks that scale out before deadline hours and back to a minimum after.
- Webhook endpoint and inbox consumer, reconciliation, payouts.
- Aurora: a
db.r6g.4xlargewriter and a replica in another AZ (R2.6 shows why).
Trace 1: a pledge at 3,000 a second. A returning backer, saved card, Standard tier (unlimited). The waiting room admits them; the pledge service runs one transaction: insert the request row, the pledge (ACTIVE) and its ACTIVATE ledger entry, with a plain read of the campaign's status and deadline. No row is shared with other backers. Commit, 201. About 1.5 seconds later the aggregator's next recompute includes it, and within 2 more seconds CloudFront's copy of the totals expires and the page shows it.
Trace 2: an Early Bird sell-out. Slot 1,000 is held at 10:00:04. The tier row's change event reaches the aggregator, which sets remaining = 0 in Valkey; from then on, Early Bird requests get 409 from the pledge service without a database call. At 10:15:07 a new-card backer's hold expires, the sweeper returns the slot, the next recompute sets remaining = 1, and the next person admitted gets it.
Trace 3: a 1M-pledge settlement with a worker crash. Step 2.3's sequence: the close, the 3-second quiesce, the exact sum, the batched move to TO_CHARGE, workers charging at 250 a second through the bucket. At minute 31 one worker's AZ is lost; its lease expires after 60 seconds; the other workers pick up its batches (the pledges that were still TO_CHARGE), and the resolver re-sends the same keys for the 32 pledges that worker left CHARGING (its 32 calls in flight). The PSP replays 31 stored successes and charges one that had never reached it. No backer is charged twice.
Trace 4: a failed card through dunning. Step 2.4's diagram: declined at 00:41 with insufficient_funds; emailed at once; retried on day 1 (declined) and day 3 (succeeded) as attempts 2 and 3 with new keys.
Trace 5: a payout, then a chargeback. The campaign is COLLECTED on day 7, RECONCILED on day 9, and paid out: step 2.6's entries 3 to 5. A dispute arrives four months later; step 3.6 posts it against the reserve.
R2.6 Numbers and Cost
Pledges.
Campaigns ending. a day; a quarter in the busiest hour (an assumption): 250. At 60% funded: 600 funded campaigns a day × 500 = 300,000 charges on a normal day, plus a giant campaign's million when one ends.
Database writes at the launch peak. Returning backers (1,800/s): request row + pledge + ledger entry = 3 inserts. New-card backers (1,200/s): request row + pledge = 2 inserts, plus a tier update for those taking a limited tier; their activation later (200/s) is a pledge update, a tier update and a ledger insert. We assume a third of pledges take a limited tier:
The hotel loop plans about 7,500 row writes a second for a db.r6g.2xlarge writer; we take a db.r6g.4xlarge and plan 15,000 (an assumption to load-test), so the launch uses about 63% of it. The tier row itself sees only about one hold per slot plus returns, thanks to the gate.
The PSP call budget. Stripe's documented live defaults are 100 requests a second per account and 25 per endpoint (payouts: "15 create requests per second"). Our design assumes a negotiated account limit of 1,000 a second and 400 per endpoint; the rows below are how we'd spend it at the worst moment, a launch and a big settlement in the same hour:
| Use | Endpoint | Peak need | Share we meter | At the defaults |
|---|---|---|---|---|
| New backers' PSP customers | Customers, create | up to 200/s at a launch | 200/s | 25/s |
| New cards at a launch | SetupIntents, create | 1,200/s arriving | 200/s (the rest queue) | 25/s |
| Settlement charges | PaymentIntents, create with confirm=true (one call per attempt) | 1,075,000 in the busiest hour | 250/s | 25/s → 11.1 h for 1M |
| Dunning retries | PaymentIntents, create and confirm | 40,000 × 3 retries over 7 days ≈ 0.2/s | 10/s | shares the 25/s above |
| Lookups of unknown outcomes | PaymentIntents, list and retrieve | rare | 10/s | 25/s |
| Refunds | Refunds, create | bursts from support | 20/s | 25/s |
| Transfers to creators | Transfers, create | ~600 campaigns/day | 5/s | 25/s |
| Everything else | detach, retrieve, reports | 50/s | ||
| Account-wide | 745/s of an assumed 1,000/s | 100/s |
The 25% left over is room for retries after 429s. At the defaults, one caller alone hits its endpoint's 25 a second first (that's the 11-hour settlement), but as soon as a launch, a settlement and refunds run together, the account-wide 100 a second is the binding limit: four endpoints at 25 each already use all of it.
Settlement time. From step 2.3: 1M ÷ 25/s ≈ 11.1 hours at the default; 1,075,000 ÷ 250/s ≈ 72 minutes in the busiest hour at the assumed limit, small campaigns done in about 5. Each call takes about 0.5 s (an assumption), so 250 a second means calls in flight: four worker tasks with 32 concurrent calls each.
Webhooks. About one event per charge we subscribe to: a second during the busiest hour, each one inbox insert plus one conditional update.
Emails per deadline. A million outcome emails at an assumed 1,000/s SES quota: about 17 minutes; the same again as receipts over the settlement hour.
Ledgers and rows. We assume 15% of pledges are changed or cancelled: pledge-ledger entries a day at ≈ 170 B (R1.7's row) = 98 MB a day, about 36 GB a year. Pledges: MB a day, about 73 GB a year. Charge attempts: about 330,000 a day on average (normal days plus a giant a month) at ≈ 300 B, about 36 GB a year. Money entries: 3 per charge (the charge's two lines, plus its share of splits and payouts, amortized) at ≈ 170 B: MB a day, about 61 GB a year. About 210 GB a year in Aurora, most of it rarely read after a campaign is paid out.
Cache. Totals and tier counts for 30,000 live campaigns: a few MB. The token buckets: a few keys. Valkey is sized for availability, not memory.
Latency budget, pledge P99 (at our load balancer, card entry excluded; dependent steps add):
| Step | P99 |
|---|---|
| WAF, load balancer, waiting-room token check | 5 ms |
| Sold-out check in Valkey | 2 ms |
| Pledge transaction: 3 inserts and a status read at ~1 ms, a durable commit of a few ms, up to ~30 ms of lock wait on a hot tier row | 40 ms |
SetupIntent creation for a new card, when its turn comes (not counted: the backer is already in the queue with a 201) | – |
| Service work and serialization | 10 ms |
| Total | ≈ 57 ms, under 200 ms |
Availability. 99.99% of a 30.4-day month is minutes.
Monthly cost (us-east-1 on-demand list prices, 730 hours a month):
| Item | Math | Monthly |
|---|---|---|
Aurora, 2 × db.r6g.4xlarge | 2 × $2.076/h × 730 h | ≈ $3,030 |
| Aurora storage and I/O | ~210 GB × $0.10 ≈ $21, plus I/O (estimate) ≈ $300 | ≈ $320 |
| Fargate | 16 tasks of 1 vCPU / 2 GB (6 pledge service, 4 settlement, 2 aggregator, 2 webhook, 2 dunning and notifications) × $0.0494/h × 730 h ≈ $577, plus launch scale-outs ≈ 520 task-hours ≈ $26 | ≈ $600 |
| MSK and Debezium | 3 × kafka.m7g.large ($0.204/h) ≈ $447, storage ≈ $30, 1 connector × 1 MSK Connect unit × $0.11/h ≈ $80 | ≈ $560 |
| Valkey | 2 × cache.r7g.xlarge at about $0.35/h × 730 h | ≈ $510 |
| CloudFront | we assume 20M page and totals views a day: 600M requests, 10M free then $0.0100 per 10,000 ≈ $590; 600M × 30 KB = 18 TB: 1 TB free, 9 TB × $0.085 + 8 TB × $0.080 ≈ $1,405 | ≈ $2,000 |
| AWS WAF | ~620M requests × $0.60/M + ACL and rules | ≈ $385 |
| Amazon SES | ~50M emails (15M pledge receipts, 15M outcome emails, 9M charge receipts, ~1M dunning, ~10M update digests) × $0.10 per 1,000 | ≈ $5,000 |
| Load balancer, NAT gateways | an estimate | ≈ $205 |
| CloudWatch, logs, alarms | an estimate | ≈ $800 |
| Total | ≈ $13.4K/month |
The biggest line is email, not databases: the system's real volume is people being told what happened.
Against the business. 300,000 charges a day at an assumed $60 is $18M a day, about $540M a month. Our 5% platform fee is about $27M a month; PSP fees at 2.9% + 30¢ are 300{,}000 \times 30 \times \2.04 \approx $18.4$M a month. PSP fees are about 1,400 times the infrastructure. For comparison, Kickstarter's blog describes a 5% platform fee plus payment processing of "roughly 3-5%" ("Kickstarter Fees: A Comprehensive Guide for Creators", 13 Mar 2024); our 5% is an assumption for this design.
R2.7 Trade-Offs
Three ways to show a hot campaign's total, figures rough:
| Recompute from the ledger, triggered by the stream (chosen) | Sharded sub-counters | Exact SUM on every read | |
|---|---|---|---|
| Pledge path cost | Inserts only | One update on one of N counter rows | Inserts only |
| Read cost | One cache read | Sum N rows | Up to a million entries per view |
| Freshness | ~2–4 s | Exact at read time | Exact |
| Replay safety | Absolute values; replays only recompute | Each counter update must be written once, in the pledge's transaction | Nothing to replay |
| Fits | Hot and cold campaigns alike; reuses the stream | When reads must be exact and N is small | Small campaigns only |
| Choice | We chose | What we give up |
|---|---|---|
| Job table vs Step Functions Distributed Map | Job table with workers and a shared token bucket | Step Functions' built-in visibility and retries. We'd still need the global bucket (MaxConcurrency is per Map), Express children are at-least-once, and Standard adds about $100 per million pledges. |
| Funded on the deadline's pledges vs on collected money | The deadline's pledges; the campaign stays funded after drops | A creator can end below their goal. The other rule would refund paying backers a week later, and it doesn't remove the fraud incentive either, only moves it. |
| A 7-day dunning window vs longer | 7 days, one final attempt | Some recoverable pledges are dropped. A longer window delays every payout and reconciliation by the same amount. |
| Save-and-charge-later vs charge-now at this scale | Save and charge later | Deadline failures and dunning. Charge-now would pay fees on every pledge of every unsuccessful campaign (400 campaigns × 500 pledges × $2.04 ≈ $408K a day at our numbers), tie up backers' money for weeks, and turn every failed campaign into a refund saga. |
| Quiesce window vs a lock | A 3 s wait after close, backed by a 2 s transaction timeout | A pledge transaction that runs past 2 s is aborted. A lock every pledge takes would serialize the launch. |
R2.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| An AZ is lost mid-settlement | A third of the workers and possibly the Aurora writer disappear | Aurora promotes the replica (tens of seconds; charges in flight retry with the same keys). The workers' leases expire and others take their batches; the pledges' states are the checkpoint, and every CHARGING pledge is resolved with its own key. |
| The PSP returns 429s or slows down during a big settlement | 429s with Stripe-Rate-Limited-Reason; rising latency | The token bucket lowers its rate; retries back off with jitter. If error rates pass a threshold, a circuit breaker pauses charging for that PSP and we probe it, instead of piling on (Circuit Breaker, Bulkhead and Fault Tolerance). The ETA on the settlement dashboard moves; nothing is lost. |
| A webhook storm | 250+ events a second, duplicates, events for unknown payments | The endpoint only verifies and inserts into the inbox; the consumer applies at its own pace. Unknown payment IDs are matched by pledge ID in metadata or parked for the resolver. |
| The display total drifts | The drift checker (sampling campaigns every minute) finds a cached total different from the ledger's | Recompute and overwrite; totals are always rebuildable from the ledger, and the decision never used them. |
| 1,000 campaigns end at 00:00 UTC | A spike of closes, sums and settlement jobs | Close jobs run in parallel per campaign (each fence is independent); round-robin keeps small campaigns moving; workers were scaled up before the hour because deadlines are known in advance. |
| A creator tries to cancel during settlement | DELETE /v1/campaigns/{id} after the deadline | Refused once the campaign is ENDED: cancelling is a pre-deadline action. After funding, only a suspension by trust and safety can stop it (step 3.4). |
| A suspension lands while charges are in flight | Some charges return success after the campaign was suspended | The orchestrator stops creating new attempts at once. A charge that succeeds afterwards is recorded, then immediately refunded with a compensating entry. The reference implementation handles the same race in its charge-now model: a payment that succeeds after a campaign was cancelled is marked refunded in the same transaction that records it. |
R2.9 Production Gotchas
| Gotcha | Why it hurts | What we do |
|---|---|---|
| Updating a shared counter on every pledge | One row caps the launch at a few hundred pledges a second | Insert-only pledges; totals recomputed off the stream |
total += amount per event | At-least-once delivery double-counts on every replay | Recompute from the ledger and set the absolute value |
| Deciding "funded" from the display counter | It lags and can drift | Exact SUM after the close |
| Summing before the closing fence | A pledge that read LIVE commits after the sum | Close, wait out the transaction timeout, then sum |
| Checking the last-day cancel rule with a plain sum | Write skew: two cancels each see enough headroom | Serialize last-day cancels and decreases per campaign |
| Reusing one idempotency key for dunning retries | Every retry replays the first decline | A new attempt and key, only after the last one failed |
| Relying on the PSP's key after its retention window | After 24 hours the same key makes a new charge | Store the payment ID; look up by customer and metadata before retrying |
| Forgetting the pledge path spends the PSP budget | A launch can starve a settlement, or the reverse | One token bucket per limit, for every caller |
| Settlement with no rate plan | 429 storms and an 11-hour settlement nobody expected | Negotiate limits weeks ahead; plan time from the limit |
| Editing ledger rows for refunds | History disappears; reconciliation can't explain the gap | Compensating entries only |
| Paying out before reconciliation | Money leaves before drops, refunds and unknowns are settled | COLLECTED → RECONCILED → PAID_OUT |
| Sending notifications inside transactions | A slow provider slows settlement; rollbacks can't unsend | Outbox events and idempotent senders |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | An AZ lost mid-settlement resumes from the pledges' own states with the same keys; the PSP's limit treated as a quota we plan, meter and negotiate; waiting room and token buckets shed load before the database and the PSP REL 1 · REL 5 · REL 11 |
| Security | Webhook signatures and timestamps; creators verified by the PSP before launch; only the settlement and payout services' roles can post to the money ledger, and no role can update or delete it SEC 3 · SEC 9 |
| Performance Efficiency | No shared row on the pledge path; hot tiers answered from the cache when sold out; pledge P99 ≈ 57 ms derived PERF 3 |
| Cost Optimization | ≈ $13.4K/month; PSP fees are about 1,400 times that; job table instead of per-pledge Step Functions transitions COST 5 · COST 6 |
| Operational Excellence | A settlement dashboard per campaign (charged, failed, unknown, rate, ETA); a rehearsal before a known big deadline (scale workers, confirm the PSP limit, dry-run the close on a copy) OPS 8 · OPS 10 |
| Sustainability | Light this round: settlement, dunning and email workers scale down between deadline hours SUS 2 |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Takes the shared row off the pledge path and says which number is allowed to lag and which never is.
- Rebuilds the closing fence without a lock, with a database-enforced transaction limit and a quiesce window.
- Finds the write skew in the last-24-hours cancel rule and serializes only that rare path.
- Treats the PSP's rate limit as the settlement's real bottleneck, derives the hours from it, and meters every caller, pledges included, through one bucket.
- Designs dunning with a new attempt per retry, handles
authentication_requiredon-session, and states the "stays funded" policy with its fraud risk. - Posts fees, reserves, payouts and refunds as double-entry journals, and pays out only after reconciliation.
Follow-up questions
-
"The PSP won't raise our limit before the big deadline. Now what?" Answer: we tell the creator the honest timeline: at 25 charges a second, about 11 hours for the first pass. We keep round-robin so the other campaigns ending that day aren't held hostage, and we spend the account-wide budget on settlement by pausing non-urgent calls (bulk card detaches, report downloads). The one thing we don't do is exceed the limit and rely on retries.
-
"Why not keep a Serializable transaction for the last-day cancel and skip the advisory lock?" Answer: it would work: Serializable would detect the two cancels' read-write conflict and abort one. But every cancel would have to be ready to retry, and on a campaign with a flurry of last-day cancels the aborts pile up. The advisory lock makes them wait in line for a few milliseconds instead, and it's scoped to the only path that can break the rule.
-
"A backer fixed their card on day 6 while the day-5 retry was still in flight. Can we charge them twice?" Answer: no. The day-5 retry is attempt 3, in
CHARGING. The fix-your-payment path can only start attempt 4 fromPAYMENT_FAILED, which the pledge isn't in; it waits for attempt 3's result. And the partial unique index allows one successful attempt per pledge, whatever the code does.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "A bigger database fixes the hot row" | One row lock is one queue; lock time is the commit, not the CPU. |
"total += amount from the event stream" | Replays double-count; recompute or set absolute values. |
| "Deciding funded from the counter" | It lags and can drift; the decision is the exact ledger sum after the fence. |
| "Our database is the settlement bottleneck" | The PSP's rate limit is; 1M charges at 25/s is 11 hours. |
| "Drop failed cards at the deadline" | Many are fixable; a 7-day dunning window recovers them. |
| "Pay creators as charges succeed" | Refunds, drops and unknowns haven't settled; pay after reconciliation. |
Round 3 · Architect · "A Global Marketplace: Currencies, Fraud, and Trust"
~45 min · Principal (L7) · 3 home regions, each with a disaster-recovery copy · 150,000 live campaigns · 3M pledges/day · ~20 currencies · pledging up per region · no double charge across a region failure
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Round 2.
Round 2 in 60 seconds. "We scaled to a national platform: 30,000 live campaigns, 500,000 pledges a day, 3,000 a second on one campaign at launch, and a million-backer campaign to settle. Pledges only insert; the campaign's total is recomputed from the ledger off the change stream and set as an absolute value, and the funded decision is an exact ledger sum after the close. The close sets
ENDED, waits 3 seconds past a 2-second database-enforced transaction timeout, then sums. Last-day cancels and decreases are serialized per campaign to stop write skew. Hot tiers get a sold-out short-circuit and, for big launches, a gate; new cards wait in a lane metered by the PSP's budget. Settlement is a resumable saga paced by a global token bucket, round-robin across campaigns: 11 hours for a million charges at Stripe's default, about 72 minutes at an assumed negotiated 250 a second. Failed cards get 7 days of dunning with a new attempt per retry, and the campaign stays funded after drops. A double-entry money ledger splits clearing into our fee, recharged processing fees and the creator's payable, holds a 10% reserve, and pays out only after reconciliation. About $13.4K a month. Open costs: one currency, no fraud checks, no moderation, one region, and chargebacks that arrive after payout."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: pledges insert, totals are recomputed off the stream, and settlement is paced by the PSP.
Round 2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Hot campaign row | Insert-only pledges; totals from the stream; exact sum after close + quiesce; serialized last-day cancels | Seconds-stale display |
| 2.2 | Hot tier and PSP at launch | Sold-out short-circuit, gate, new-card lane metered by the PSP | Minutes of waiting for new cards |
| 2.3 | 1M charges | Resumable saga, global token bucket, round-robin | ~72 min at an assumed limit |
| 2.4 | Failed cards | 7-day dunning, new attempt per retry; stays funded | Money over a week; fraud incentive |
| 2.5 | Records disagree | Inbox under burst; reconcile before payout | Payout waits |
| 2.6 | Fees, payouts, refunds | Double-entry ledger, reserve, verified creators | Negative balances possible |
| 2.7 | A million emails | Outbox → queue → idempotent senders | Minutes of lag |
Open costs: one currency; no risk or moderation; one region; disputes after payout.
R3.1 The Scope Raise
Interviewer: "We're global now: backers and creators in about 40 countries and 20 currencies. A backer in Japan backs a campaign priced in euros. Scam campaigns appear, and some creators take the money and disappear. Some campaigns get pushed over their goal by fake backers, and small pledges are being used to test stolen cards. Regulators and the card networks expect us to act on reports quickly. Sometimes we must suspend a campaign, live or already funded, and refund every backer. Chargebacks arrive months later, often because rewards never shipped, after the creator was paid. We'll run in three home regions, and a region can fail during a settlement. And creators want our campaign events in their own tools."
We ask back before fixing anything.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Which currency is the goal in, and can the backer pay in theirs? | The goal is in the campaign's currency. Backers should see their own currency, at least as an estimate. | The campaign currency is the unit of the goal and the ledger; FX never decides an outcome (step 3.1). |
| How do we know a pledge is fake? | We don't, for sure. Velocity, linked accounts and card patterns are signals. | Risk checks on the pledge path and a review before settlement for suspicious campaigns (step 3.2). |
| Must every campaign be reviewed before launch? | Yes, and reports after launch must be handled within a day. | A moderation pipeline with automated checks and a human queue (step 3.3). |
| If we suspend a funded campaign, who pays for the refunds? | The creator's balance and reserve first; we absorb what can't be recovered. | A mass-refund saga that takes money in a stated order (step 3.4). |
| Where do campaigns live, and what if a region fails? | Each campaign in its creator's home region. Pledging may pause for a region's campaigns, but no double charges, ever. | A home region per campaign, fenced failover, settlement resumed with PSP lookups (step 3.5). |
| How long after payout can money come back? | Months. Disputes for undelivered rewards can come long after. | Chargebacks as compensating entries and a rolling, delivery-aware reserve (step 3.6). |
| What do creators want from us? | Signed webhooks for pledges and settlement events, to their own URLs. | Outbound webhooks with HMAC signatures, retries and SSRF protection (R3.3). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Live campaigns | 30,000 | 150,000 |
| Pledges | 500,000/day | 3M/day (150,000 × 600 / 30), 40% / 35% / 25% across us-east-1, eu-central-1 and ap-southeast-1 |
| Campaigns ending | ~1,000/day | ~5,000/day; ~1.8M charges/day |
| Currencies | 1 | ~20 |
| New campaigns to review | none | ~2,000/day |
| Footprint | 1 region, 3 AZs | 3 home regions, each with a disaster-recovery copy; pages served worldwide |
| Money after payout | Refunds | + chargebacks months later, mass refunds, negative creator balances |
| Targets | 99.99% in one region | Pledging up per region; settlement resumes after a region failure with no double charge; RPO ≤ 20 s for pledges; payouts and refunds auditable end to end |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| One currency | A euro campaign's goal and ledger can't mix yen amounts; converting at pledge time would let exchange rates change a campaign's outcome. |
| No risk checks | Fake pledges on bad cards push campaigns over their goal, and "stays funded" rewards the fraud; card testers use pledges to check stolen cards. |
| Campaigns go live on publish | Scams and prohibited items launch and collect pledges before anyone looks. |
| Refunds are one-off support actions | Suspending a funded campaign means refunding 500,000 backers, in a stated order, against a rate limit. |
| One region | A regional outage stops every campaign on the platform, possibly in the middle of a settlement. |
| A flat 10% reserve | A creator whose rewards ship in a year carries far more dispute risk than one shipping next month. |
R3.3 New Requirements and API Additions
Currency on the pledge. The campaign page and the pledge carry the campaign's currency and an estimate in the backer's:
json{ "pledge_id": "plg_02K1XZ", "amount_minor": 6000, "currency": "EUR", "estimate": { "currency": "JPY", "amount_minor": 9780, "rate": "163.00", "as_of": "2027-09-02T08:00:00Z", "binding": false } }
Risk. Every pledge and every creator carries a risk score and the signals behind it (internal API only). A campaign can be in UNDER_REVIEW after its deadline.
Moderation. GET /v1/moderation/cases/{id} with states PENDING → APPROVED | REJECTED; POST /v1/campaigns/{id}/reports for anyone to report a live campaign; POST /v1/campaigns/{id}/suspend for trust and safety, with a reason and an appeal deadline.
Outbound webhooks for creators. Creators register HTTPS URLs; we send pledge.created, campaign.funded, settlement.progress and payout.sent, each signed with an HMAC over the timestamp and the raw body, retried with backoff, and deduplicated by event ID on their side. The sender resolves the URL's address itself and refuses private, loopback and link-local ranges (including the instance metadata address) after every DNS resolution and redirect, so a creator's URL can never reach our internal network (SSRF, server-side request forgery).
Home region. Every campaign has a home_region attribute, fixed at creation from the creator's country, and its public ID carries a region prefix so any region can route to it without a lookup.
R3.4 Design Evolution: Money, Trust and Regions
Step 3.1: A Backer Pays in Yen for a Campaign Priced in Euros
The problem: a Munich creator's campaign has a €50,000 goal. A backer in Osaka pledges what the page shows as "about ¥9,780". At the deadline, the euro is 4% stronger than on the day they pledged. What would you do?
Related loop: Design a Payment Processing System, step 3.4 (FX).
Step 3.2: Fake Backers Pushed a Campaign Over Its Goal, and Small Pledges Are Testing Stolen Cards
The problem: a campaign sat at 70% of its goal for 28 days, then crossed it in the last 6 hours on 900 pledges from brand-new accounts with brand-new cards. After the deadline, 80% of those cards failed in dunning. Under our Round 2 policy, the campaign stays funded. Separately, thousands of $1 pledges a day come from a few devices, each with a different card. What would you do?
Primitive: Bot Defense, Sybil Resistance and Registration Abuse · related loop: the payment loop's step 3.3 (card testing)
Step 3.3: 2,000 New Campaigns a Day, Some of Them Scams or Prohibited
The problem: creators submit about 2,000 campaigns a day. Some sell prohibited items, some use stolen photos of someone else's product, some are outright scams. Once a campaign is live, it collects pledges from real people. What would you do?
Synthesizing vector architecture diagram...
The campaign lifecycle as it stands after Round 3 (Round 1's CANCELLED branch is unchanged and left out). SUSPENDED can be reached from any state where backers could still be harmed. It ends in a refund saga when money was collected; a campaign suspended while still LIVE has charged nobody, so its pledges are simply released. Every arrow is one conditional status update.
Step 3.4: A Funded Campaign Turns Out to Be a Scam: Refund 500,000 Backers
The problem: three weeks after payout, a funded campaign with 500,000 backers (€30M collected) is confirmed as a scam: the product never existed. We decide to refund every backer. What would you do?
Primitive: Two-Phase Commit and Saga Orchestration
Drill: The flight booking that charged without a seat (a compensation that fails halfway is point 1: every refund is an idempotent, recorded step resumed from the pledges' states, and a refused refund becomes a case; why a saga and not 2PC is step 2.3's third wrong answer)
Step 3.5: A Region Fails at 23:59 With Three Settlements in Flight
The problem: eu-central-1 becomes unreachable at 23:59:30. It's the home of 35% of campaigns. Three campaigns are mid-settlement, one campaign's deadline is 23:59:59, and backers everywhere are pledging to European campaigns. What would you do?
Synthesizing vector architecture diagram...
The deterministic key is what makes lost records safe: the new writer can only ask the PSP the same question again, and the PSP answers from memory.
Primitive: Cloud Disaster Recovery and Multi-Region Active-Active · loop primitive Multi-Region Failover · related loop: the hotel reservation loop, step 3.6
Step 3.6: Chargebacks Arrive Months After Payout Because Rewards Didn't Ship
The problem: a gadget campaign was paid out in March, promising delivery in December. In January, backers start filing disputes with their banks for "goods not received". Each dispute pulls the charge back from us, plus a fee. The creator's reserve was 10% and has already been released. What would you do?
A chargeback, posted. One €60 pledge from the gadget campaign, disputed in January; we assume a fee equivalent to the US $15 for illustration and that the reserve still holds money.
| Event | Debit | Credit | Amount |
|---|---|---|---|
| Dispute opened: PSP takes amount + fee | Creator reserve | PSP balance | €60.00 |
| Dispute fees (charged to the creator) | PSP balance | €15.00 | |
| Fee recharged to the creator | Creator payable | Dispute fees | €15.00 |
| Dispute won (reverse the amount; the dispute fee stays lost) | PSP balance | Creator reserve | €60.00 |
Related loop: the payment loop's step 2.6 (payout returns and negative balances).
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Yen for a euro campaign | Campaign currency is the unit of goal and ledger; estimates shown; charge in the campaign currency; FX can't flip outcomes | Final amount differs from the estimate |
| 3.2 | Fake backers and card testing | Velocity limits, linked-account checks, UNDER_REVIEW before the outcome, card-testing defenses on card setup, rules + a model | Friction; a review queue |
| 3.3 | Scams and prohibited campaigns | Pre-launch moderation (Comprehend, Rekognition, rules, humans), post-launch reports, its own module | Launch waits for review |
| 3.4 | Refund 500,000 backers | Mass-refund saga: batches, per-pledge keys, token bucket, money taken in a stated order, refused refunds as cases | Platform losses beyond the reserve |
| 3.5 | A region fails mid-settlement | Home region per campaign; fenced Global Database failover; resume with the same keys; bucket shares move; reconcile the loss window before closing | 10–15 min pause for a third of campaigns |
| 3.6 | Chargebacks after payout | Compensating journals in a stated order; rolling, delivery-aware reserve; evidence from updates; negative-balance recovery | Slower payouts for risky creators |
R3.5 Global Architecture
One home region in detail; the other two are the same shape.
Synthesizing vector architecture diagram...
Everything that decides money for a campaign is in its home region, on one writer. Everything browsers read is a copy. The PSP budget is split between regions and lent back and forth, so no region depends on another to charge.
- Regional stacks as in Round 2, one per home region, each with a Global Database secondary in a nearby region.
- Global pages: CloudFront serves each campaign page; the origin is the nearest region, which holds a read copy of every campaign's page data built from the other regions' change streams, seconds stale.
- Risk service per region: rules and a model scoring pledges and creators on the pledge path.
- Moderation module per region, with a shared reviewer tool.
- FX: the payment platform's quotes and rates, used for estimates (and quotes if we ever charge in the backer's currency).
- Outbound creator webhooks: a sender fleet with HMAC signing, retries and an egress path that refuses internal addresses.
- Audit export: each region's ledgers exported nightly to S3 with Object Lock, as in the payment loop's step 3.6 (tamper evidence); we don't re-teach it.
Trace 1: a cross-currency pledge and charge. The Osaka backer opens the Munich campaign from ap-southeast-1's cache ("about ¥9,780"), pledges €60, and the request is routed to eu-central-1, where the risk service scores it (low) and the pledge is inserted in euros. At the deadline, the charge is €60; the backer's bank converts it at the stronger euro, about ¥10,170, plus any fee the bank adds. Our ledgers only ever saw euros.
Trace 2: a scam campaign suspended and refunded. Step 3.4: reports and a risk signal open a case; trust and safety suspend the campaign (payouts and charges stop); after the appeal window the decision is final; the refund saga runs 500,000 refunds at 100 a second in about 83 minutes, taking money from reserve, then balance, then transfer reversal, then platform loss.
Trace 3: a region failover mid-settlement. Step 3.5's sequence: the fence, the promotion, the same keys re-sent, attempt 3 replayed from the PSP's memory, the late deadline closed only after the loss window is reconciled.
R3.6 Numbers and Cost
Pledges.
By home region: us-east-1 40% (1.2M), eu-central-1 35% (1.05M), ap-southeast-1 25% (0.75M). Launch peaks stay per campaign, as in Round 2.
Settlements. campaigns end a day; at 60% funded, M charges a day. Each region's busiest hour, with a quarter of its day's deadlines (an assumption):
| Region | Ending/day | Busiest hour | Funded × 600 | Time at 250/s |
|---|---|---|---|---|
| us-east-1 | 2,000 | 500 | 300 × 600 = 180,000 | 720 s, 12 min |
| eu-central-1 | 1,750 | ~438 | ~263 × 600 ≈ 158,000 | ≈ 630 s |
| ap-southeast-1 | 1,250 | ~313 | ~188 × 600 ≈ 113,000 | ≈ 450 s |
The regions' busiest hours are in different time zones, so their peaks rarely overlap; if all three regions share one PSP account, the coordinator lends each region the others' idle share, and a region can charge faster than its fixed share when the others are quiet. A giant campaign still takes about an hour, as in Round 2.
Moderation. 2,000 submissions a day, each checked automatically: its text (a few DetectToxicContent calls) and, we assume, 12 images and one short video of about a minute. That's 24,000 image checks a day. With 15% flagged (an assumption) and 10 minutes per human review, that's minutes, 50 reviewer-hours a day. Add post-launch reports, which we assume at 500 a day at 5 minutes each, about 42 hours: about 92 reviewer-hours a day, 12 reviewers on 8-hour shifts. People, not servers, are the moderation cost.
Refund saga. 500,000 refunds: about 83 minutes at an assumed 100/s refund budget, 5.6 hours at the default 25/s.
Ledgers and audit. Pledge-ledger entries: MB a day. Money entries: MB a day. Pledges and charge attempts add about 1.7 GB a day. In Aurora across the three regions: about 1.2 TB a year, plus the same again in the DR secondaries. The nightly audit export compresses the two ledgers, about 1.5 GB a day, to roughly a quarter (an assumption): about 140 GB a year in S3, kept for years at a few dollars a month per year kept.
Availability and recovery. RPO normally under a second, capped at about 20 seconds by rds.global_db_rpo. RTO 10 to 15 minutes, mostly the human decision (the hotel loop's step 3.6 derives it).
Monthly cost (us-east-1 on-demand prices for every region first; the regional adjustment follows):
| Item | Math | Monthly |
|---|---|---|
| Aurora instances | 3 regions × (writer + replica + DR secondary) = 9 × db.r6g.4xlarge × $2.076/h × 730 h | ≈ $13,640 |
| Aurora storage, I/O, replication | ~1.2 TB primary + 1.2 TB DR in the first year × $0.10 ≈ $240; I/O and replicated writes (estimate) ≈ $1,500 | ≈ $1,750 |
| Fargate | 3 × Round 2's $600, doubled for the larger regions ≈ $3,600; risk, moderation and webhook senders ≈ $500 | ≈ $4,100 |
| MSK and Debezium | 3 × Round 2's $560 | ≈ $1,680 |
| Valkey | 3 × Round 2's $510 | ≈ $1,530 |
| CloudFront | 6 × Round 2's views: 3.6B requests ≈ $3,590; 108 TB: 1 TB free, 9 × $85 + 40 × $80 + 58 × $60 per TB ≈ $7,445 (US list tiers; other continents cost more) | ≈ $11,035 |
| AWS WAF | ~3.7B requests × $0.60/M | ≈ $2,250 |
| Amazon SES | 6 × Round 2's 50M = 300M emails × $0.10 per 1,000 | ≈ $30,000 |
| Rekognition and Comprehend | a month is 720,000 images × $0.001 = $720, 60,000 one-minute videos × $0.10/min = $6,000, and ~300,000 text checks of about 1,000 characters (10 units × $0.0001) ≈ $300 | ≈ $7,000 |
| Cross-region replication and page copies | an estimate | ≈ $800 |
| Load balancers, NAT, CloudWatch | 3 × Round 2's ≈ $1,005 | ≈ $3,015 |
| Audit archive in S3 | ~140 GB a year | ≈ $50 |
| Total at us-east-1 prices | ≈ $76.9K/month |
Other regions cost more. Frankfurt and Singapore charge more than us-east-1 for instances and transfer; they carry about 60% of this spend. We add about 15% to that share (an estimate, to be replaced by the pricing calculator), about $6.9K, for a budget of about $84K a month. Email is still the biggest line.
Against the business. 1.8M charges a day at an assumed $60 equivalent is $108M a day, about $3.2B a month. A 5% platform fee is about $162M a month. The infrastructure is about 0.05% of the platform's fee revenue. The expensive things in Round 3 are people (reviewers, analysts, support) and losses (fraud, unrecovered chargebacks), which is why the design spends its effort on the reserve, the review states and the audit trail.
R3.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| Hold funds at the platform vs in the PSP's platform accounts | The PSP's platform product (Connect-style charges, then transfers to verified creator accounts) | Some control over timing and fees. Holding customer money ourselves brings money-transmission and safeguarding obligations that vary by country; we keep that general and let the PSP's licensed product carry it. |
| Reserve size vs creator cash flow | A rolling reserve sized by risk and delivery date | Risky creators get less money up front, when they need it most. A zero reserve maximizes creator happiness and makes us the insurer of every campaign. |
| Strict pre-launch review vs speed to launch | Every campaign reviewed; 85% automatically, the flagged 15% by people | Launches wait minutes to hours. Publishing first and reviewing later is faster and lets scams collect pledges for a day. |
| Home region vs multi-writer | One writer per campaign, in its home region | Far backers' pledges cross an ocean (a few hundred ms), and a region failure pauses a third of campaigns. Multi-writer would be faster for them, but the funded decision needs one exact sum under one fence. |
| Charge in the campaign's currency vs the backer's | The campaign's currency | Some backers pay their bank's FX fee. Charging in their currency would add conversion fees and an FX account to every campaign's books. |
Back to the opening question: how do we collect millions of promises safely, and then, at one instant, turn each one into exactly one charge, or release it, for every campaign on the platform? The answer is now: a pledge is a promise, not a payment, and the ledger, not a counter, decides. We save cards instead of charging them, record every promise in an append-only ledger owned by one home region, and close pledging with a fence so the exact sum includes everything accepted and nothing after. Then settlement turns promises into money as a resumable saga, paced by the PSP's limit, with keys per (pledge, attempt) so a retry, a crash or a failed-over region can only ask the PSP the same question again. Around it sit the things that make the money trustworthy: dunning, reconciliation before payout, a double-entry money ledger, reviews before outcomes, and reserves for the disputes that come after.
R3.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| The PSP has a regional outage during a global deadline hour | Charges fail or time out for one region's campaigns | Outcomes are already decided and recorded; only charging waits. The breaker pauses charging, unknowns stay unknown (never retried with a new key), and the settlement ETA shows the delay. If the outage passes 24 hours, resolution switches to lookups before any retry (step 2.3). |
| The FX rate provider is down | Stale reference rates | Estimates older than our limit (for example 1 hour) are hidden ("amount in EUR"), never shown as current. Pledges and charges don't depend on FX, because they're in the campaign's currency. |
| A moderation false positive suspends a legitimate hit campaign | A suspended campaign with 200,000 backers and an angry creator | Suspension stops charges and payouts, not refunds: refunds start only when the decision is final. The creator's appeal goes to a senior reviewer within a stated time; if upheld, the campaign resumes where it stopped (settlement from the pledges' states). A refund issued too early could never be taken back. |
| A mass-refund saga fails halfway | Refund progress stalls at 212,000 of 500,000 | It resumes from the pledges' states; every refund in flight is resolved with its own key; refused refunds are cases, not retries. |
| A region outage | Health checks fail; pledges for that region's campaigns get 503 | Step 3.5: pause, then a fenced failover if it lasts, settlement resumed with the same keys, bucket share moved, the loss window reconciled before closing deadlines that fell inside it. |
R3.9 Runbook and Incident Response
Golden signals, per region OPS 8 · REL 6
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Pledge P99 | > 200 ms for 5 min | P2 | Check tier-row lock waits and the gate; a launch without a waiting room |
| Total-aggregation lag (last recompute age for hot campaigns) | > 10 s | P3 | Check the connector and the aggregator; decisions are unaffected |
| Tier sold-out mismatches (cache says 0, database says free, or the reverse) | > 1% of checks | P3 | Recompute tier counts; the database is the truth |
Campaigns past their deadline and still LIVE | any, > 5 min | P2 | The close job is down; no late pledges can be accepted anyway |
| Settlement progress and ETA per campaign | ETA slips > 30 min | P2 | PSP 429 rate and latency; worker count; bucket share |
| PSP 429 rate | > 1% of calls for 5 min | P2 | Lower the bucket rate; check who is spending the budget |
Unknown-outcome charges (CHARGING older than 10 min) | > 100, or any older than 12 h | P1 at 12 h | Resolve by key now; after 24 h only by lookup |
| Dunning backlog | failures > 2× the 30-day baseline | P3 | A PSP or network problem, or a fraud pattern (step 3.2) |
| Reconciliation breaks per campaign | any open for > 1 day | P2 | Payout held; match by metadata |
| Payout holds | campaigns COLLECTED > 3 days without RECONCILED | P3 | Work the breaks list |
| Chargeback rate per creator | > 1% of their charges | P2 | Raise their reserve; review the campaign |
| Moderation queue age | oldest > 12 h | P3 | Add reviewers; launches are waiting |
| Notification lag | > 30 min after a deadline | P3 | SES quota, sender count |
Procedure: a stalled settlement REL 11
- Look at the campaign's settlement dashboard: charged, failed, unknown, rate, ETA.
- If the rate is zero with no 429s, the workers are down or not holding leases: check the task count and the job's lease owner.
- If 429s dominate, someone else is spending the budget: the bucket metrics per caller show who.
- Resolve stuck
CHARGINGpledges with their own keys; never create a new attempt for an unknown.
Procedure: a PSP incident at a big deadline OPS 10
- Confirm on the PSP's status page and our error rates. Open the breaker for charging; keep pledging up (card setup may also be failing: show a clear message).
- Tell creators whose settlements are affected: "funded; charging is delayed".
- When the PSP recovers, close the breaker gradually through the bucket; resolve unknowns first.
Procedure: a campaign suspension SEC 10
- Trust and safety suspend with a reason; charges and payouts stop at once.
- Freeze the creator's other payouts and linked accounts pending review.
- After the appeal window, if confirmed: start the refund saga; watch its progress like a settlement.
- Afterwards: add the pattern to the risk rules; report as regulators and card networks require (keep to our legal team's guidance).
Go deeper: CLI playbook
Plain commands an on-call engineer runs, one at a time. Replace the names and ARNs with real ones.
text# 1. Alarms currently firing for settlement in a region aws cloudwatch describe-alarms --region eu-central-1 --state-value ALARM --alarm-name-prefix settlement- # 2. Depth of the notification queue after a deadline aws sqs get-queue-attributes --region eu-central-1 --queue-url https://sqs.eu-central-1.amazonaws.com/111122223333/notify-outcomes --attribute-names ApproximateNumberOfMessages ApproximateAgeOfOldestMessage # 3. The region's SES sending quota and today's usage aws sesv2 get-account --region eu-central-1 # 4. Replication state of a home region's global database aws rds describe-global-clusters --global-cluster-identifier cf-eu # 5. Unplanned failover to the DR secondary (may lose up to about 20 s of commits) aws rds failover-global-cluster --global-cluster-identifier cf-eu --target-db-cluster-identifier arn:aws:rds:eu-west-1:111122223333:cluster:cf-eu-euw1 --allow-data-loss # 6. State of the region's CDC connector aws kafkaconnect describe-connector --region eu-central-1 --connector-arn arn:aws:kafkaconnect:eu-central-1:111122223333:connector/cf-eu-cdc/1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d-2 # 7. State of the region's cache (totals, tier counts, token buckets) aws elasticache describe-replication-groups --region eu-central-1 --replication-group-id cf-cache
Audit queries (read-only, against a region's replica):
sql-- A decided campaign's promise entries must add up to its decision. Should return nothing. SELECT c.campaign_id, c.final_minor, SUM(l.delta_minor) AS promises FROM campaigns c JOIN pledge_ledger l ON l.campaign_id = c.campaign_id WHERE c.status IN ('FUNDED','SETTLING','SETTLED','COLLECTED','RECONCILED','PAID_OUT') AND l.kind IN ('ACTIVATE','CHANGE','CANCEL','REMOVE') GROUP BY c.campaign_id, c.final_minor HAVING SUM(l.delta_minor) <> c.final_minor; -- No pledge may have two successful charges (the unique index makes this impossible). SELECT pledge_id, count(*) FROM charge_attempts WHERE status = 'SUCCEEDED' GROUP BY pledge_id HAVING count(*) > 1; -- Journals must balance. Should return nothing. SELECT journal_id, SUM(amount_minor) FROM money_entries GROUP BY journal_id HAVING SUM(amount_minor) <> 0; -- Charges in flight too long: resolve by key, and by lookup after 24 hours. SELECT pledge_id, attempt_no, created_at FROM charge_attempts WHERE status = 'IN_FLIGHT' AND created_at < now() - interval '10 minutes' ORDER BY created_at;
The first query needs no time filter: after the fence, no transaction can add an ACTIVATE, CHANGE or CANCEL entry, and REMOVE entries (step 3.2's fake pledges) are written only before the decision. Post-close entries (DROP) are a different kind. A row in this result means the fence failed, which is exactly what an audit should catch.
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | One writer per campaign in its home region; fenced failover with a majority-read lease and Aurora's 20 s RPO cap; settlement resumed with the same keys and PSP lookups; bucket shares that move with campaigns; fault isolation by home region REL 10 · REL 13 |
| Security | Risk scoring, velocity limits and linked-account checks on pledges; pre-launch moderation; outbound webhooks HMAC-signed with SSRF protection; a suspension procedure for scam campaigns; backers' personal data stays in the campaign's home region and its DR copy SEC 5 · SEC 7 · SEC 10 |
| Performance Efficiency | Campaign pages from the nearest region's caches through CloudFront; only pledges travel to the home region PERF 3 · PERF 4 |
| Cost Optimization | ≈ $77K/month at US prices, budget ≈ $84K; email the largest line; three regional stacks against about $162M a month of fee revenue; replication and page copies between regions priced as their own line COST 5 · COST 8 |
| Operational Excellence | The golden signals with first actions; settlement ETAs per campaign; procedures for stalled settlements, PSP incidents and suspensions OPS 8 · OPS 10 |
| Sustainability | Home regions chosen by where creators and backers are, not a full copy of everything everywhere; settlement and email fleets scale to near zero between deadline hours instead of idling warm in every region; old ledgers archived compressed SUS 1 · SUS 2 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Keeps one currency per campaign's goal and ledger, and shows why FX can't flip an outcome.
- Connects the "stays funded" policy to the fraud it invites, and puts a review before the outcome instead of changing the policy.
- Treats moderation as its own module with its own states, automated first and human for what's flagged.
- Designs mass refunds as the mirror of settlement, with a stated order for where the money comes from and a loss account for what's left.
- Fails over a home region with fences, resumes settlement from the pledges' states, and explains why deterministic keys make lost records harmless.
- Owns the money after payout: chargebacks as journals, a delivery-aware reserve, evidence from the campaign's own records.
Follow-up questions
-
"Why not DynamoDB global tables for pledges, so every region can accept them?" Answer: global tables check condition expressions against the local copy and resolve conflicting writes with last-writer-wins by default, so two regions could each take the last Early Bird slot, and the funded decision would depend on which copy you summed. The multi-Region strong-consistency mode runs in exactly three Regions and doesn't support transactions, and a pledge is a transaction over a pledge, a ledger entry and a tier counter. One writer per campaign is simpler and correct.
-
"A creator moves countries. Can their campaign move home regions?" Answer: only between campaigns, or as a planned migration while it isn't settling: freeze pledges briefly, copy the campaign's rows, verify the ledger sum, flip the routing, and have the old region refuse that campaign's writes for good. We never move a campaign mid-settlement.
-
"The review found 900 fake pledges and the campaign falls below its goal without them. What now?" Answer: the fake pledges are removed with compensating entries while the campaign is
UNDER_REVIEW, before any outcome was decided, so the exact sum without them decides: unsuccessful, and every real backer is released without being charged. The creator is investigated for the linked accounts.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Convert every pledge to one currency at pledge time" | Rates move between pledge and charge; the ledger and the goal must share one currency. |
| "Block fraud by IP" | Rings rotate addresses; honest backers share them. Use card, device and account signals. |
| "Publish first, review later" | A scam collects a day of pledges and trust. |
| "Refund with a loop script" | Rate limits, crashes and re-runs: refunds need the same saga as charges. |
| "Restart settlement after a failover" | Lost records plus new keys can charge twice; resume with deterministic keys. |
| "Absorb chargebacks" or "freeze payouts until shipping" | One makes us everyone's insurer; the other defeats the purpose of crowdfunding. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: all-or-nothing, length, when charged, tiers, cancels, payment | Restate Round 1 in 60 seconds | Restate Round 2 in 60 seconds |
| 5–15 min | Requirements and API (pledge with idempotency key, the separate card step, the half-open deadline) | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.6: funding models → idempotency → promise ledger → tier counters → closing fence → per-attempt settlement | Steps 2.1–2.7: insert-only pledges and the quiesce fence → hot tiers and the PSP budget → rate-limited saga → dunning → reconciliation → money ledger → notifications | Steps 3.1–3.6: currency → fraud → moderation → mass refunds → region failover → chargebacks |
| 40–50 min | Numbers, the funding-model cost comparison, trade-offs | Numbers (PSP budget, settlement hours, rows), cost, trade-offs | Numbers per region, reviewers, cost with regional uplift |
| 50–60 min | Failures and pillar check | Failures and pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening a round: "A pledge is a promise, not a payment: I collect promises for weeks and move money once, at the deadline, as a resumable saga."
- When someone points at the counter: "Totals can lag; decisions can't: the page shows a counter, but the funded decision and every charge come from the ledger."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What if the backer double-clicks, or the worker retries a charge?" (REL 4) | Idempotency keys on every request, and a PSP key per (pledge, attempt) so a retry replays and a new attempt only follows a known failure. | 1 | Steps 1.2, 1.6 |
| "What's your real throughput limit?" (REL 1) | The PSP's: 25 requests a second per endpoint by default, so a million charges take 11 hours unless we negotiate more weeks ahead; every caller meters through one token bucket. | 2 | Steps 2.2, 2.3, R2.6 | |
| "What if 3,000 people pledge in one second?" (REL 5) | A waiting room and a sold-out short-circuit in front, insert-only pledges behind, and a new-card lane metered by the PSP budget. | 2 | Steps 2.1, 2.2 | |
| "What if an AZ fails mid-settlement?" (REL 11) | Leases expire, other workers continue from the pledges' states, and stuck charges are resolved with their own keys. | 2 | R2.8 | |
| "What if a region fails?" (REL 13) | Its campaigns pause, fail over behind a lease fence and a 20-second RPO cap, and settlement resumes with the same keys, so lost records become PSP replays. | 3 | Step 3.5 | |
| Security | "Do you ever see card numbers?" (SEC 9) | No: cards go browser-to-PSP; we store references and verify every webhook's signature and timestamp. | 1 | Steps 1.1, 1.6 |
| "Who can change the money ledger?" (SEC 3) | Only settlement and payout roles can insert; no role can update or delete; journals must balance. | 2 | Step 2.6 | |
| "What do you do about a scam campaign?" (SEC 10) | Review before launch, a risk review before the outcome, suspension, and a mass-refund saga once the decision is final. | 3 | Steps 3.2–3.4, R3.9 | |
| Performance | "Why doesn't one hot campaign melt the database?" (PERF 3) | Pledges only insert; totals are recomputed off the change stream; the exact sum is taken once, after the fence. | 2 | Step 2.1 |
| "How is a campaign page fast in Tokyo?" (PERF 4) | CloudFront and a read copy in the nearest region; only the pledge goes to the campaign's home region. | 3 | R3.5 | |
| Cost | "What does it cost?" (COST 5) | About $885, $13.4K and $84K a month; PSP fees are about 1,400 times Round 2's infrastructure, and email is the biggest AWS line. | 1–3 | R1.7, R2.6, R3.6 |
| "Why not charge now and refund failed campaigns?" (COST 5) | The PSP keeps its fee on refunds: about $612K a month at Round 1's numbers, $408K a day at Round 2's. | 1–2 | Step 1.1, R1.7, R2.7 | |
| "What about traffic between regions?" (COST 8) | Only page copies and database replication cross regions, priced as their own line; pledges go straight to one home region. | 3 | R3.6 | |
| Operations | "How do you know a settlement is healthy?" (OPS 8) | A dashboard per campaign: charged, failed, unknown, rate and ETA, plus alarms on 429s and unknowns. | 2–3 | R2.10, R3.9 |
| "What happens when the PSP goes down at a big deadline?" (OPS 10) | Outcomes are already decided; a breaker pauses charging; creators are told; unknowns are resolved first on recovery. | 3 | R3.9 | |
| Sustainability | "Where does the footprint go?" (SUS 2) | Settlement and email fleets scale with deadline hours instead of idling, and only three home regions hold the full stack. | 2–3 | R2.10, R3.10 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Funding model | Save and charge later; compares charge-now and authorize-capture with their costs; knows authorizations expire | Knows the model's price at scale (dunning, authentication_required) and still defends it with numbers | Adds currency: goal and ledger in one currency, FX never decides an outcome |
| Hot campaign | Tier counter with a conditional update; write skew named | Insert-only pledges, totals off the stream, sold-out short-circuit, a launch metered by the PSP | The same per home region, with pages served globally |
| Closing and deciding | A fence on the campaign row; exact sum after it | A quiesce window backed by a transaction timeout; last-day cancels serialized | A risk review before the outcome; late deadlines closed only after a failover's loss window is reconciled |
| Settlement | Per-pledge charges with (pledge, attempt) keys; one success per pledge | A rate-limited, resumable saga with a global token bucket and fair scheduling; lookups after the key window | Resumed across a region failover with the same keys; bucket shares move with campaigns |
| Money after the deadline | Charges recorded | Dunning, reconciliation before payout, double-entry money ledger, reserve, refunds as compensating entries | Mass refunds in a stated order, chargebacks as journals, a delivery-aware reserve |
| Trust and safety | Out of scope | Creators verified before launch | Velocity and linked-account checks, moderation pipeline, suspension procedure |
| Evolving under new scope | Builds from "charge on click" one problem at a time | Opens with "what breaks", removes the shared row without losing the fence | Changes the shape (currencies, regions, review states) without weakening "the ledger decides, once" |
See It Built: the Reference Implementation
~5 min.
CrowdFundingHub is a C# (.NET) modular monolith that implements Round 1's design and parts of Round 2: one deployable application, one PostgreSQL database with a schema per module, modules that talk to each other only through published contracts and events carried by a transactional outbox, and inboxes that make every consumer idempotent. It's the codebase behind an upcoming course on clean architecture; the repository is private, so this section describes it rather than linking to it.
A modular monolith is a sensible shape for Round 1: one database and one deploy keep the fence and the ledger simple, while the module boundaries mean a piece like Moderation could later be moved into its own service without untangling shared tables.
| Module | What it owns | Where it appears on this page |
|---|---|---|
| Campaigns | The campaign lifecycle (draft, published, successful or failed, cancelled); an append-only campaign ledger in which a unique key on contribution and entry kind makes a redelivered pledge or refund a no-op, a check constraint ties each entry's sign to its kind, and a trigger rejects updates and deletes (refunds are negative compensating entries); reward tiers with a claimed + reserved ≤ capacity invariant, 15-minute reservations and a scavenger that releases expired ones, serialized by a per-tier advisory lock; a background job that resolves expired campaigns to succeeded or failed under a per-campaign lock | Steps 1.3, 1.4, 1.6 |
| Contributions | Pledges and their payment states; a signed payment webhook reconciled through an inbox claimed in the same transaction as the state change; a payment that succeeds after its campaign was cancelled is marked refunded in the same transaction; the refund saga that reacts to a failed or cancelled campaign; per-pledge metering of platform and processing fees | Steps 1.2, 1.6, 2.5, R2.8 |
| Identity | Accounts, roles and permissions; tokens signed with asymmetric keys and verified offline against a published key set | R1.10 Security |
| Moderation | A review per campaign (pending, approved or rejected) created from the campaign-created event, with asynchronous, signed media analysis | Step 3.3 |
| Notifications | Email on campaign and pledge events, with preferences and unsubscribe | Step 2.7 |
| CampaignUpdates | Creators' updates and creator-registered outbound webhooks, HMAC-signed, sent from a background loop with SSRF protection | R3.3, step 3.6 |
Where it differs from this page's design (checked against the code on 2026-09-28; the codebase is under active development):
- It charges at pledge time and refunds if the campaign fails: the charge-now model from step 1.1, a legitimate simpler choice, rather than save-and-charge-later.
- It recomputes the raised total as a sum over the ledger under a per-campaign lock on every pledge: exactly right at Round 1's scale, and exactly what step 2.1 replaces.
- Fees are metered, not yet posted to a double-entry money ledger, and there are no creator payouts yet.
- Its refunds record state only: the handler that reacts to a failed or cancelled campaign marks contributions refunded in batches of 100, and the late-success case marks a payment refunded, without yet calling the payment provider's refund API.
- Its 15-minute tier reservations match step 1.4's hold, but they exist for its charge-now checkout, not for card setup.