Design Bot Defense and Sybil-Resistant Registration
This page is one interview loop in three rounds. All three rounds design the same system: the part of a consumer platform that decides whether a signup (and later a login, a post or a payment) comes from a real person or from someone pretending to be many people. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | One app's signup form is flooded by scripts | Fake-account farms on residential proxies and real browsers | One trust system for signup, login, posting and payments |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Traffic we plan for | 50,000 real signups a day (peak ≈ 3/s); attacks of 1,000 attempts/s (assumptions) | 10M real signups a month (peak 50/s); 100,000 malicious requests/s at the edge (planning figures) | ≈ 3,150 scored actions/s on average, ≈ 29,500/s at the peak of a credential-stuffing wave (derived) |
| Target | A real user signs up in under 30 s; a fake account needs a real inbox, and a challenge when it looks risky | Edge decisions in under 100 ms; at most 500 requests/s reach the origin; a false-positive budget | A measured fake-account rate and false-positive rate; a friction budget per action |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
How to read this case study. Unlike the other case studies in this section, this one isn't about one company. Bot defense is an architecture pattern that many platforms build from the same parts, so we teach it as a design loop grounded in publicly documented practice: vendor documentation (AWS, Cloudflare, Google), standards (NIST, the IETF, the W3C) and public security guidance. Every claim about a product says what the vendor documents, with the year we checked it, and the page ends with a Sources list. Everything else is marked a design that fits (ours), and numbers marked assumption are ours, chosen to make the arithmetic concrete.
On attackers. We describe what attackers do only at the level a defender needs, to explain why a defense exists. This page is about building defenses, not about getting past them.
Loop Opener: Why Bot Defense?
A Bouncer Who Can't See Faces
Picture a club with a bouncer at the door, but the bouncer is behind a wall. All they get is a slip of paper pushed through a slot: a name, an email, a password. Most slips come from real people. Some come from one person with a printing press, pushing through ten thousand slips with ten thousand made-up names. The bouncer can't see faces. They can only look at the slips, how fast they arrive, what they're written on, and what happens after each person gets inside.
That is signup on the internet. A bot is a program that uses our site the way a person would, but automatically. A Sybil attack (named after a case study of a person with many identities) is one actor creating many fake identities so they count as many people: many votes, many free trials, many referral bonuses, many accounts to send spam from.
A few words we'll use all page:
| Word | What it means on this page |
|---|---|
| Bot | A program that drives our site or API automatically. Some are welcome (search engines); here we mean the unwelcome kind. |
| Sybil / fake account | One of many accounts controlled by one actor, created to look like separate people. |
| Headless browser | A real browser engine run by a script, with no screen. It runs our JavaScript, so it looks much more like a person than a simple script does. |
| Residential proxy | A service that relays an attacker's traffic through home internet connections, so each request comes from an ordinary-looking home IP address. |
| Fingerprint | A value computed from how a client behaves technically (for example how its TLS handshake is built), used to recognize similar clients. It identifies a kind of client, not a person. |
| CAPTCHA | A challenge meant to be easy for people and hard for programs, such as picking images. |
| Managed challenge | A vendor service (Cloudflare Turnstile, Google reCAPTCHA, AWS WAF's Challenge and CAPTCHA actions) that decides for itself whether to run invisible checks or show the user something. |
| Proof-of-work (PoW) | A puzzle that costs the client a measurable amount of computing to solve and costs the server almost nothing to check. |
| False positive | A real person we treated as a bot: challenged, slowed down, or blocked. |
What Makes It Hard
- The attacker adapts. Every rule we ship gets probed. What works this month may not work next month, so the design must be cheap to change and must measure itself.
- Identities are cheap. IP addresses, email addresses, browsers and phone numbers can all be rented or generated. Anything we key a limit on, the attacker tries to rotate.
- Every check costs real users something. A challenge adds friction. A strict limit catches a whole office behind one IP address. A visual puzzle can lock out a blind user. Collecting device signals raises privacy questions.
- Decisions are made on a live request path. Signup has to answer in well under a second, while the strongest evidence (what the account does next) arrives much later.
The Question the Whole Loop Answers
How do we make each fake account cost the attacker more than it's worth to them, while keeping signup nearly free for real people?
The idea underneath every round is an inequality. For an attacker, a fake account is worth it when
We can't lower the value much (that's the product), so we raise the cost, and we raise it most for the traffic that looks least like a person. A rough estimate of the starting point, to show the scale: a script on a small cloud machine at $0.05 an hour (assumption) sending 10 signups a second makes 36,000 attempts an hour, about $0.0000014 per attempt. With a free throwaway inbox, a fake account costs the attacker around a millionth of a dollar. That figure is our order-of-magnitude estimate, not a measured market price, but the point stands: before we defend, fakes are almost free.
Synthesizing vector architecture diagram...
Each defense adds to the attacker's column. The design goal is to add a lot to the left for the traffic that looks automated, while keeping most real users in the top row on the right.
The answer grows every round:
- Round 1: rate limits keyed on what's cheap to check, a honeypot and timing check, email verification with a disposable-domain screen, and a managed challenge only when something looks wrong.
- Round 2: many signals combined into a risk score, a proof-of-work puzzle sized by risk, a ladder of challenges that escalates only as far as needed, defenses against SMS toll fraud, and edge filtering for volume.
- Round 3: detection after signup (clusters of accounts that share devices or payment cards), login defenses against stolen passwords, a feedback loop that retrains as attackers change, appeals for people we got wrong, and privacy and accessibility built in.
Round 1 · Mid-level · "Stop Scripted Signups on One App"
~35 min · SDE II (L5) · one app · 50,000 real signups a day, attacks of 1,000 attempts/s (assumptions) · real users sign up in under 30 s
R1.1 Establish Design Scope
The interviewer sets the scene: "We run a community app. New accounts get to post, message other members, and receive 50 free credits. For a few months, scripts have been creating accounts and using them to spam. Fix signup without making it painful for real people."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| What's the attack? | Scripted signups. Spam posts and messages start within minutes of account creation. | Accounts must not get full powers until they prove something (step 1.3). |
| How many? | It started at thousands of fake accounts an hour. Last Tuesday one script sent 3.6 million signup attempts in one hour, from about 50 IP addresses in cloud data centers. | 3,600,000 ÷ 3,600 s = 1,000 attempts/s from ~50 addresses: a crude attacker. Per-IP limits will help this round (step 1.1). |
| How many real signups? | About 50,000 a day, with an evening peak. | ≈ 0.58/s on average, ≈ 3/s at the peak (R1.7). The attack is about 345 times the real peak. |
| What must stay easy? | Signup for real users: an email, a password, a name. Marketing says no puzzles for everyone. | Challenges only when something looks wrong (step 1.4). |
| How do we verify people? | By email. There's no phone number in the flow today. | Email verification is our main proof (step 1.3). Phone comes in Round 2. |
| Clients? | A website and mobile apps that call the same API. | Anything browser-only (honeypots, some challenges) needs an app equivalent. |
Out of scope for this round: device fingerprints and risk scores (Round 2), proof-of-work (Round 2), phone verification (Round 2), detecting fakes after signup (Round 3), login attacks (Round 3).
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Operation |
|---|---|
| "Create an account" | POST /v1/auth/register with email, password and display name |
| "Prove the email is yours" | We email a one-time link; clicking it moves the account from PENDING to ACTIVE |
| "Block obvious abuse" | Rate limits, a honeypot field, a minimum fill time, a disposable-domain screen |
| "Only annoy people when it looks risky" | Return CHALLENGE_REQUIRED; the client shows a managed challenge and retries with its token |
| "Scripts spam right after signup" | A PENDING account can't post, message or receive credits |
Not yet: device signals, proof-of-work, phone verification, graph detection.
R1.3 Non-Functional Requirements: the Questions
We name each quality first; the numbers come in R1.7.
- Friction. How many real users see a challenge? We set a budget: at most 5% of real signups (assumption).
- False positives. How many real users do we block outright? Blocks must be rarer than challenges, and a blocked person needs a way forward.
- Latency. The server's part of a signup (not counting the user's typing or a challenge) should finish in well under a second; we budget p99 < 500 ms (R1.7).
- Cost of verification. Every verification email and every challenge costs money, and an attacker can make us send them.
- Availability. If a dependency is down (the rate-limit cache, the challenge provider, the email sender), signup should degrade to a safer mode, not fail open and not fail closed for everyone.
R1.4 The API
Register
httpPOST /v1/auth/register HTTP/1.1 Host: auth.example.com Content-Type: application/json Idempotency-Key: 5b0c2f0e-7a1d-4b8e-9f3a-2d6c1e8b4a70 { "email": "ana@example.org", "password": "<sent over TLS; the server hashes it with Argon2id>", "display_name": "Ana", "form_token": "1790500000.8f3c...hmac", "website": "", "challenge_token": null }
form_tokenis issued when the form is rendered: the render time plus an HMAC (a keyed hash only our servers can make). It lets us measure how long the form took to fill without trusting the client's clock. Because a token is only a timing signal, a count of submissions perform_tokenover its lifetime is a signal too: one render, many posts means a script.websiteis the honeypot: a field people never see and never fill (step 1.2). Its name is ordinary on purpose.challenge_tokenis empty the first time. It carries a managed-challenge token when we asked for one.Idempotency-Keynames one attempt, so a retry after a lost response gets the same answer instead of a second account or a second email.
Responses
httpHTTP/1.1 202 Accepted Content-Type: application/json { "status": "PENDING_VERIFICATION", "message": "Check your inbox to finish signing up." }
httpHTTP/1.1 403 Forbidden Content-Type: application/json { "error": "CHALLENGE_REQUIRED", "challenge": { "provider": "turnstile", "site_key": "0x4AAAA..." } }
| Code | Meaning | Client does |
|---|---|---|
202 Accepted | Account created (or already exists: same answer, see below); email sent | Show "check your inbox" |
400 Bad Request with DISPOSABLE_EMAIL | The email's domain is a known throwaway service | Ask for another address |
403 Forbidden with CHALLENGE_REQUIRED | Something looked risky | Show the challenge, then resend with challenge_token |
409 Conflict with IN_PROGRESS | The same Idempotency-Key is still being processed | Retry in a second |
429 Too Many Requests with Retry-After | Too many attempts from this source | Wait; show a friendly message |
Why no 409 EMAIL_TAKEN? Because "this email already has an account" tells anyone which addresses are registered with us, which helps attackers who test stolen email lists. We answer 202 either way and send the existing owner an email ("someone tried to sign up with your address; if it was you, sign in here"), rate-limited per address. Some products accept that leak for usability; it's a product decision to make on purpose.
Why 403 and not 428 Precondition Required for the challenge? 428 is defined for requests that should be conditional (with If-Match and friends), not for "prove you're a person". A 403 with a clear error code is easier for clients to handle.
An account's states
Synthesizing vector architecture diagram...
A PENDING account exists but can't post, message or receive credits. That single rule removes most of the reward from scripted signups that can't read a real inbox.
Recap
- One write endpoint, idempotent by key, with an explicit challenge error the client knows how to handle.
- Accounts start
PENDINGand earn their powers by verifying email. - We don't reveal whether an email is registered.
R1.5 Design Evolution: From an Open Form to Layered Checks
Each step is a problem, your turn to think, the answer, and what it costs us.
Step 1.0: The Baseline
Synthesizing vector architecture diagram...
An open form: every request that reaches the signup service creates an account and sends an email. The scripts and the people look the same.
At 1,000 attempts a second, this creates up to 3.6 million accounts an hour and asks our email provider to send 3.6 million emails, many to addresses that don't exist. The bounces alone threaten our ability to send email at all: Amazon SES places an account under review at a bounce rate of 5% and may pause sending at 10% (SES documentation, checked September 2026).
Step 1.1: One Script Creates Thousands of Accounts
The problem: one script sends 20 signup attempts a second from each of about 50 cloud IP addresses, 1,000 a second in all. Meanwhile real people sign up about 3 times a second at the evening peak, some of them from an office, a university or a mobile carrier where hundreds of people share one IP address. What would you do?
Primitive: Distributed Rate Limiting · Drill: The Partner Whose Retry Loop Took Down Everyone Else (answered here: counters kept locally on N servers let a source through N times; a fixed window allows a double burst at the boundary) · Loop: Design a Distributed Rate Limiter (Round 1, steps 1.1–1.3, in depth)
Step 1.2: Scripts Fill the Form Instantly
The problem: the simple scripts post the form in 40 milliseconds, with every field filled, including fields a person would never touch. They stay under the rate limits by spreading across addresses. What would you do?
Primitive: Honeypot Fields and Canary Traps · Drill: The Headless Scraper That Filled the Database with Garbage (answered here: why hidden fields fail against headless browsers that see what's visible, and why invisible traps beat a CAPTCHA for everyone)
Step 1.3: Throwaway Emails
The problem: the scripts now slow down and pass the traps. Every fake account uses an address at a disposable email service: a site that hands out inboxes that last minutes, often readable through an API, so a script can click our verification link without a person. What would you do?
Primitive: Disposable Email and Domain Blocklists · Primitive: Bloom Filters and Counting Filters · Drill: The 100,000 Free Trials That Vanished Overnight (answered here: why a hard-coded list goes stale within days, and why SMTP probing is the wrong tool)
Step 1.4: Still Too Many
The problem: the attacker now uses real webmail inboxes and paces signups under our limits. Fake accounts still get through, a few hundred an hour. Marketing still refuses a puzzle for everyone. What would you do?
Primitive: Cloudflare Turnstile Managed Challenges · Primitive: Google reCAPTCHA v2 and Enterprise Risk Analysis · Drill: The Flash Sale That Blocked 40% of Paying Mobile Shoppers (answered here and in R1.9: why challenges hit real users behind shared addresses, inside in-app browsers or with blocked scripts, and why a managed challenge beats an image puzzle for everyone)
Why challenges hit real people, and what we do about it. A challenge that fires on the wrong signal catches the wrong people:
- Shared addresses. A mobile carrier's shared IP address can hit a per-IP limit that no single person hit. That's why the per-IP rule challenges rather than blocks, and why its limit counts accounts, not requests.
- Scripts that can't run. In-app browsers (inside social apps), strict content blockers and some privacy settings can stop the challenge's script from loading or from keeping its token. We treat "the challenge failed to load" as a reason to offer another route (a code to the email address), never as proof of a bot.
- Expired tokens. A Turnstile token lasts 300 s. A person who solves the challenge, then spends six minutes choosing a display name, submits an expired token. The client refreshes the token when it expires instead of sending the dead one.
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | An open form | Unlimited fakes; email reputation at risk |
| 1.1 | One script, thousands of accounts | WAF volume rule per IP; shared sliding-window limits on accounts per IP and per domain; over the limit means challenge | A cache on the path; challenges for shared IPs |
| 1.2 | Scripts fill the form instantly | Honeypot field and signed timing token | Stops only naive scripts |
| 1.3 | Throwaway emails | PENDING until verified; a self-refreshing disposable-domain screen | A list pipeline; rare false flags |
| 1.4 | Still too many | Turnstile, only when a signal fires; verdicts stored per request | A vendor dependency; accessibility work |
R1.6 Architecture v1
Synthesizing vector architecture diagram...
Requests pass the edge's volume and reputation rules, then the signup service runs its cheap checks, its limits and (only for risky requests) the challenge check before writing anything. The email goes out afterwards, from an outbox item written in the same transaction as the account.
The data model (a design that fits)
One DynamoDB table holds several item types; the key design makes each check one read or one conditional write.
| PK | SK | Attributes | Purpose |
|---|---|---|---|
USER#<user_id> | PROFILE | email, display_name, password_hash, status, created_at | The account |
EMAIL#<normalized_email> | UNIQUE | user_id | Makes the email unique: written with attribute_not_exists(PK) in the same transaction as the user |
IDEM#<idempotency_key> | REQ | request_hash, status, owner, lease_until, challenge_verdict, response, ttl | The durable guard for retries |
VERIFY#<token_hash> | EMAIL | user_id, expires_at, used, ttl | Email verification links |
OUTBOX#<user_id> | VERIFY_EMAIL | email, token reference | Picked up by the email sender from the table's stream |
Synthesizing vector architecture diagram...
The EMAIL# item, not a secondary index, is what makes two concurrent signups with the same address impossible: a global secondary index can't enforce uniqueness.
The durable guard, placed before any outside call. A signup calls two outside systems: the challenge check (sometimes) and the email sender. The guard goes first:
textregister(request, key): 1. stateless checks: form_token HMAC (constant-time compare) and age, honeypot, email syntax, disposable screen 2. rate-limit reservations (step 1.1); a reservation over the limit -> CHALLENGE_REQUIRED. A request whose claim finds a DONE record releases its reservation, so a retry never counts twice. 3. claim: PutItem IDEM#key {status=IN_PROGRESS, owner=my_id, lease_until=now()+10 s, request_hash=h} if attribute_not_exists(PK) - claim lost: read the item with a strongly consistent read (an eventually consistent read can miss the item just written or show an older status) status DONE and same request_hash -> return the stored response different request_hash -> 422, key reused with another body IN_PROGRESS and lease_until > now() -> 409 IN_PROGRESS (the original is still working) IN_PROGRESS and lease expired -> take over: UpdateItem set owner=my_id, lease_until=now()+10 s if lease_until = old value 4. if a challenge is needed: if challenge_verdict is recorded, use it; else call siteverify (with its own idempotency_key) and record the verdict on IDEM#key 5. TransactWriteItems, all or nothing: Put USER#id (new) Put EMAIL#email if attribute_not_exists(PK) Put VERIFY#token_hash, OUTBOX#id Update IDEM#key set status=DONE, response=202 if owner = my_id 6. EMAIL# condition failed: the address has an account -> Update IDEM#key DONE and queue a "someone tried to sign up" notice; respond 202 either way
now()is the signup task's clock. DynamoDB conditions have no server time; the writer supplies the value it compares against. A lease of 10 s against tasks whose clocks agree to well under a second is safe.- An original request can outlive its retry. Suppose the original stalls for 15 s after claiming. Its lease expires; the retry takes over and finishes. When the original wakes, its transaction carries
owner = my_idon theIDEM#item, which no longer matches, so the whole transaction is cancelled. TheEMAIL#condition would stop a duplicate account anyway; the owner check also stops a duplicate email. - TTL only tidies up.
IDEM#andVERIFY#items carry a TTL so old ones disappear, but DynamoDB deletes expired items "within a few days of their expiration time". Every decision comparesexpires_atorlease_untilin code.
Trace 1: a real person signs up
Synthesizing vector architecture diagram...
No challenge, no delay: Ana never learns there were checks. The email leaves after the commit, so a crash can't send an email for an account that doesn't exist.
Trace 2: the 4th signup from one IP in an hour
Synthesizing vector architecture diagram...
A request that passed a challenge is allowed past the per-IP limit. The verdict is stored on the idempotency record before the account is written, so a retry after a lost response doesn't need the single-use token again.
The email sender. The stream delivers each outbox item at least once, so an email can occasionally go out twice. That's harmless: both emails carry the same link. SES sends at $0.10 per 1,000 emails.
R1.7 Numbers
Figures marked published come from vendor documentation; the rest are assumptions or derived from them. Every monthly figure on this page uses a 30-day month: 720 hours, 2,592,000 seconds.
Traffic
| Quantity | Arithmetic | Result |
|---|---|---|
| Real signups | Assumption | 50,000 a day |
| Average | 50,000 ÷ 86,400 s | 0.58/s |
| Evening peak | × 5 (assumption) | ≈ 2.9/s |
| Attack | 3.6M attempts in an hour ÷ 3,600 s | 1,000/s |
| Per attacking IP | 1,000 ÷ 50 IPs | 20/s = 6,000 per 5 minutes |
| Attack vs real peak | 1,000 ÷ 2.9 | ≈ 345× |
What the limits let through
| Quantity | Arithmetic | Result |
|---|---|---|
| Edge rule, per IP | 1,000 requests per 5 min (our setting) | 3.3/s per IP |
| Before the rule reacts | Each IP sends 20/s, so it crosses 1,000 in 5 min after 50 s; AWS then usually reacts within 30 s: (50 + 30) s × 1,000/s | ≈ 80,000 requests reach the service |
| An attacker who paces just under the rule | 3.3/s × 50 IPs | ≈ 167 requests/s |
| Accounts without a challenge | 3 per IP per hour × 50 IPs | 150 an hour, all PENDING until a real inbox clicks |
| Everything above that | ≈ 167/s × 3,600 s | ≈ 600,000 challenges an hour, each one a cost for the attacker |
Rate-limit memory
| Quantity | Arithmetic | Result |
|---|---|---|
| Distinct IPs in a peak hour | 2.9/s × 3,600 (real) + 50 (attacker) | ≈ 10,500 |
| Keys, counter design | (10,500 IPs + ~1,500 domains) × 2 windows | 24,000 |
| Memory, counter design | 24,000 × ~90 B per key (assumption, including overhead) | ≈ 2.2 MB |
| Memory, log design, attack hour | 3.6M entries × ~70 B (assumption) | ≈ 252 MB, growing with whatever the attacker sends |
Challenges and email
| Quantity | Arithmetic | Result |
|---|---|---|
| Real users challenged | 5% of 50,000 (our budget) | 2,500 a day |
| Turnstile cost | Free plan, unlimited challenges (published) | $0 |
| If we used the WAF CAPTCHA instead | Real: 75,000 a month × $0.40/1,000; attack: 5 attack hours × 600,000 challenges, if the attacker submits answers to all of them | $30 for real users; up to $1,200 that the attacker chooses for us |
| Verification emails | 50,000 × 1.1 (resends) × 30 | 1.65M a month |
| SES | 1.65M × $0.10/1,000 | $165 a month |
Latency: the server's part of a signup, every hop counted
| Hop | p99 (assumption) | How it adds |
|---|---|---|
| CloudFront and WAF inspection | 5 ms | Sequential |
| Edge to origin Region and back, over AWS's network | 70 ms (a far-away user; about 10 ms nearby) | Sequential |
| ALB | 2 ms | Sequential |
| Stateless checks and the in-memory domain screen | 1 ms | Sequential |
| Two counter reservations in ElastiCache | 2 ms each, sent together: max = 2 ms | Parallel |
| Idempotency claim in DynamoDB | 10 ms | Sequential |
| Account transaction in DynamoDB | 25 ms | Sequential |
| Without a challenge | 5 + 70 + 2 + 1 + 2 + 10 + 25 | 115 ms |
siteverify, only when challenged | 300 ms (assumption; an outside service) | Sequential |
| With a challenge | 115 + 300 | 415 ms, inside the 500 ms budget |
Adding each hop's p99 overstates the true p99 of the whole path, because they rarely all hit their worst case at once. That makes it a safe upper bound for a budget. The user's own time (typing, solving the challenge, reaching the edge) is on top, and the 30-second target leaves plenty of room for it.
Cost (AWS list prices, us-east-1, checked September 2026)
| Item | Arithmetic | Monthly |
|---|---|---|
| AWS WAF | $5 web ACL + 3 rules × $1 + $10 Bot Control + (normal 50 requests/s × 2,592,000 s = 129.6M, plus 5 attack hours × 3.6M = 18M) × $0.60/M + Bot Control common scoped to POST /v1/auth/register: (18M attack + 1.5M real − 10M free) × $1/M | $116.06 |
| Signup service | 3 tasks × 0.5 vCPU, 1 GB: 1.5 × $0.04048 × 720 + 3 × $0.004445 × 720 | $53.32 |
| ElastiCache (Valkey) | 2 × cache.m7g.large (primary and replica in two AZs) × $0.1264 × 720 | $182.02 |
| DynamoDB (on-demand) | 1.5M signups × 10 write units × $0.625/M + 15M read units × $0.125/M + 27 GB × $0.25 | $18.01 |
| SES | above | $165.00 |
| Email sender (Lambda) | 1.65M invocations × $0.20/M, plus a little duration | ≈ $1 |
| ALB | $0.0225 × 720 + 2 LCUs × $0.008 × 720 | $27.72 |
| Turnstile | Free plan | $0 |
| Total | ≈ $563 a month |
That's about 38 cents per 1,000 real signups ($563 ÷ 1.5M signups), dominated by the cache and the email. The attacker's side of the ledger is what changed: every account beyond 150 an hour now costs them a solved challenge and a real inbox.
R1.8 Trade-Offs
Always-on CAPTCHA vs risk-based challenges
| CAPTCHA for everyone | Challenge only when a signal fires (chosen) | |
|---|---|---|
| Real users who see one | 100% | Our budget: at most 5% |
| Attackers stopped | Those not paying a solving service | The same, plus everything the free checks catch |
| Accessibility exposure | Every user | Only the few challenged; they get another route |
| What the attacker learns | Nothing: everyone's treated the same | Which signals fire; they'll adapt to avoid them |
| Vendor outage impact | All signups | Only the challenged few (R1.9) |
Build vs buy the challenge
| Build our own puzzle | Buy a managed challenge (chosen) | |
|---|---|---|
| Detection quality | Only what we can see | The vendor sees traffic across many sites |
| Cost | Engineering time, forever | Free or per-use |
| Accessibility | Ours to design and test | The vendor's alternatives, plus our fallback route |
| Dependency | None | A third party on the signup path |
| Privacy | We control the data | The vendor's scripts run in our users' browsers; review what they collect |
Where to put the limits
| Edge only (WAF) | Service only | Both (chosen) | |
|---|---|---|---|
| Sees | IPs, headers, TLS fingerprints, reputation | Accounts, domains, outcomes | All of it |
| Precision | Approximate, reacts in about 30 s | Exact, immediate | Floods stopped cheaply, decisions made precisely |
| Can challenge | Yes (its own CAPTCHA or Challenge) | Yes (our choice of provider) | Yes |
R1.9 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Turnstile is unreachable | siteverify times out (we wait 3 s) | Fail to a safer mode, not open: signups that needed a challenge are created as PENDING with a review flag and stricter limits (1 per IP per hour); they can't do anything until they verify email, and the flagged ones get a second look before their credits unlock. Signups that didn't need a challenge are unaffected. |
| The rate-limit cache is down | Reservations fail | Don't fall back to per-task counters with the full limit: with our 3 tasks, a source spread across them would get 3 times the limit. Each task instead enforces the limit divided by the task count (3 ÷ 3 = 1 account per IP per hour, 20 ÷ 3 → 6 per domain), and we pause scale-out while the cache is down, so the divisor stays right and the sum never exceeds the shared limit. More real users hit a challenge for a while; nothing opens up. |
| A cache failover | A few seconds of errors, then counters missing their latest increments | Counters may read a little low for up to an hour: the limits undercount briefly, and the edge rule and challenges still stand. Acceptable for limits; we never keep money or identity decisions only in the cache. |
| DynamoDB throttles or errors | Transactions fail | Return 503 with Retry-After; the client retries with the same key, and the durable guard makes that safe. |
| The domain list job fails | The file in S3 stops changing | Tasks keep the last good copy; an alarm fires when the file is more than 36 hours old. |
| SES bounce rate climbs | Our sending reputation drops | Alarm at 2% bounces (SES recommends staying under 2%); tighten the domain screen and limits before SES's 5% review threshold. |
| An attacker floods the challenge path | Many CHALLENGE_REQUIRED answers, many widget loads | Turnstile has no per-challenge fee; our cost is one 403 each. The edge rule still caps each IP. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Security | Edge volume rules and IP reputation before the service SEC 5; email verification before an account gets powers, passwords hashed with Argon2id, no account-existence leak SEC 2 |
| Reliability | Idempotent signups with a durable guard claimed before any outside call REL 4; safer modes, not open ones, when the challenge provider or the cache fails REL 5 |
| Performance Efficiency | Cheap in-memory checks first, a counter design whose memory doesn't grow with attack volume, an exact set chosen over a Bloom filter at 10 MB PERF 3 |
| Cost Optimization | A free challenge provider chosen with the per-attempt pricing of the alternative in view COST 5 |
| Operational Excellence | Alarms on bounce rate, list freshness and challenge rate OPS 8 |
| Sustainability | Skipped this round: a small fleet; the relevant gain, dropping floods before they use compute, comes in Round 2. |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Keys limits on what's cheap to check (IP, email domain) and knows shared addresses make "block" the wrong action: over the limit means challenge.
- Counts accounts created, not requests, and keeps counters shared and atomic.
- Treats honeypots and timing as signals, not verdicts.
- Verifies email, and keeps a disposable-domain screen fresh automatically.
- Challenges only risky traffic, verifies tokens on the server, and handles single-use tokens across retries.
- Says what happens when each dependency fails, and chooses a safer mode over open or closed.
Follow-up questions
-
"Why not block every data-center IP address?" Answer: some real users come through corporate VPNs, privacy relays and cloud desktops, which use data-center addresses. A data-center label is a strong signal for a challenge, not proof of a bot. We block only when several signals agree.
-
"Your honeypot says bot but the challenge says human. Who wins?" Answer: the challenge, for that request: the honeypot's job was to decide whether to ask. We log the disagreement; a rising rate of honeypot hits that pass challenges means either autofill is filling the field (rename it) or attackers are paying for solves.
-
"Why store the challenge verdict instead of just verifying again on retry?" Answer: tokens are single use (Turnstile: "Each token can only be validated once"). A retry after a lost response would fail verification and block a real person who already proved themselves.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Block the IP forever" | Attackers rotate; innocent users inherit addresses. |
| "5 signups per IP per day, then block" | Offices and carriers share addresses. Challenge instead. |
| "Count requests in each server's memory" | N servers let through N times the limit. |
| "CAPTCHA for everyone" | All users pay; solving services don't. |
| "Silent fake success for honeypot hits" | Strands the real user whose autofill filled the field. |
"409 EMAIL_TAKEN" | Tells attackers which emails have accounts. |
| "Verify the token again on retry" | Single-use tokens fail the second time. |
Round 2 · Senior · "Botnets on Residential Proxies at 100K Requests/s"
~40 min · Senior SDE (L6) · 10M real signups a month, peak 50/s · 100,000 malicious requests/s at the edge (planning figures) · edge decisions under 100 ms, at most 500 requests/s reach the origin
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "Our app's signup was flooded by scripts: 1,000 attempts a second from about 50 cloud IPs. We layered cheap checks. At the edge, an AWS WAF rate-based rule stops floods per IP. In the signup service, shared sliding-window counters in ElastiCache limit accounts created per IP and per email domain, and going over a limit means a challenge, not a block, because offices and carriers share addresses. A honeypot field and a signed render-time token catch naive form-fillers, as signals, not verdicts. Every account starts
PENDINGand gets no powers until its owner clicks an emailed link, and a self-refreshing screen rejects disposable email domains. When any signal fires, we ask for a Cloudflare Turnstile challenge and check its token on the server; tokens are single use, so the verdict is stored on the request's idempotency record, which is claimed before any outside call. If the challenge provider is down we fail to a safer mode, not open. It costs about $563 a month. Open costs: everything keys on IP addresses and email domains, and a patient attacker with many addresses and real inboxes still gets through."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: cheap checks, limits keyed on IP and domain, a challenge only when a signal fires, and nothing works until the email is verified.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | One script, thousands of accounts | Edge volume rule; shared limits on accounts per IP and domain | Challenges for shared IPs |
| 1.2 | Scripts fill forms instantly | Honeypot and signed timing token | Stops only naive scripts |
| 1.3 | Throwaway emails | Email verification; disposable-domain screen | A list pipeline |
| 1.4 | Still too many | Managed challenge only when risky | A vendor dependency |
Open costs: every limit is keyed on something the attacker can rent in bulk, and the only proof of a person is an inbox.
R2.1 The Scope Raise
Interviewer: "We grew. Ten million real signups a month now, and we added phone verification for some markets. The attackers grew too. They rent residential proxies, drive real browsers, buy CAPTCHA answers, and one weekend someone ran up an SMS bill we're still arguing with finance about. At peak we see 100,000 malicious requests a second at the edge. And our growth market is people on cheap Android phones: they must not be blocked."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How many addresses does the attack use? | About 500,000 distinct residential IPs in a big wave. | 100,000 ÷ 500,000 = 0.2 requests/s per IP, 60 per 5 minutes: under any per-IP limit a real carrier address would need. IP limits alone are finished (step 2.1). |
| Are they real browsers? | Many are headless browsers that run our JavaScript and send browser-like headers. | Header checks and honeypots see a browser. We need signals that are costly to fake in bulk (step 2.1). |
| What happened with SMS? | Bots requested verification codes to phone numbers in expensive destinations; almost none of the codes were ever entered. | SMS must not be a free action at the top of the funnel (step 2.4). |
| Our challenges? | Challenge pass rates on attack traffic went up: they're paying solving services. | One challenge type isn't enough; friction must escalate with risk (step 2.3). |
| Real users' devices? | Half of new users are on low-end Android phones, often on shared carrier addresses. | Any client-side work must be sized for the slowest phones we support (step 2.2). |
| What's the budget for friction? | Challenge at most 5% of real signups; hard-block at most 0.1% of them, measured. | A false-positive budget we measure by sampling (R2.6). |
| What reaches our servers today? | Too much. The origin fleet autoscaled to 10 times its size during the last wave. | Decide at the edge for volume; the origin should see at most 500 requests/s (step 2.5). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Real signups | 50,000 a day, peak ≈ 3/s | 10M a month, peak 50/s |
| Attack | 1,000/s from ~50 cloud IPs | 100,000/s from ~500,000 residential IPs |
| Attacker tools | Scripts | Headless browsers, solving services, rented phone numbers |
| Verification | Email, or phone in some markets | |
| Decisions | Per-IP and per-domain limits | A risk score from many signals, and an escalating ladder |
| Where decided | Mostly in the signup service | Volume at the edge; precise decisions at the origin |
The "Not yet" list from R1.2 comes back: device signals and proof-of-work are in scope now.
R2.2 What Breaks in the Round 1 Design
| Round 1 piece | What breaks at the new scale |
|---|---|
| Per-IP limits | 0.2 requests/s per IP is below any sane limit. The per-IP rule sees nothing wrong, while real users behind carrier addresses still trip it. |
| One challenge type | A solving service answers the same challenge for a fraction of a cent. The challenge is now a price, not a barrier, and one price for every risk level is either too low or too high. |
| SMS as a verification step | Each code costs us money, the attacker chooses where it goes, and they gain when it goes to numbers they profit from. A free action that costs us money is an attack surface. |
| Decisions only at the origin | 100,000 requests a second reach our fleet before anything says no. Autoscaling lags the wave, and we pay for the compute that rejects them. |
| Honeypots and timing | Real browsers leave hidden fields empty and can wait as long as they like. |
R2.3 New Requirements and API Additions
1. A risk score, computed for every signup. An internal call inside the signup service's region (we show it as an API so the contract is clear):
httpPOST /internal/v1/risk/score HTTP/1.1 Content-Type: application/json { "action": "signup", "request_id": "req_01J9ZQ3C6X8R2M4T7V1B5N0K9D", "network": { "ip": "203.0.113.24", "asn_type": "residential", "waf_labels": ["bot-control:signal:automated_browser"] }, "tls": { "ja4": "t13d1516h2_8daaf6150882_e5627efa2ab1" }, "device": { "device_id": "dv_7c1e...", "waf_token_fingerprint": "fp_91a2..." }, "form": { "fill_ms": 41000, "honeypot_filled": false }, "identity": { "email_domain": "example.org", "domain_type": "freemail", "phone_type": null } }
httpHTTP/1.1 200 OK Content-Type: application/json { "score": 64, "action": "POW_AND_CHALLENGE", "pow_difficulty_bits": 17, "reasons": ["ja4_cluster_velocity_high", "device_new", "asn_residential"], "model_version": "rules-2026-09-21" }
2. A proof-of-work challenge. The client asks for a puzzle, solves it in a background thread (a Web Worker in the browser), and sends the answer with the signup.
httpPOST /v1/auth/pow-challenge HTTP/1.1 Content-Type: application/json { "action": "signup", "request_id": "req_01J9ZQ3C6X8R2M4T7V1B5N0K9D" }
httpHTTP/1.1 200 OK Content-Type: application/json { "seed": "Jq3v8pX0cM4r2tQy9wLb1A", "difficulty_bits": 17, "expires_at": 1790500120, "signature": "hmac-sha256:4be1...c90a" }
The signup then carries "pow": { "seed": "...", "nonce": 88213, "difficulty_bits": 17, "expires_at": 1790500120, "signature": "..." }.
3. Phone verification, behind the ladder.
httpPOST /v1/auth/phone/start HTTP/1.1 Content-Type: application/json Idempotency-Key: 0d2f4c1a-9b7e-4e3a-8f61-5c2b7a9d1e04 { "account_id": "usr_8842", "phone": "+15550100123", "channel": "sms" }
| Response | Meaning |
|---|---|
202 Accepted | Code sent (or already sent under this key) |
400 with UNSUPPORTED_DESTINATION | We don't send SMS to that country or number type; the client offers email instead |
429 with Retry-After | Too many codes for this account, number or destination prefix |
403 with VERIFY_ANOTHER_WAY | The SMS budget breaker is open; email or an authenticator app instead |
POST /v1/auth/phone/verify takes the code; five wrong codes invalidate it.
The risk actions, as a ladder
Synthesizing vector architecture diagram...
A request only climbs as far as its score demands. Most real users stop at the first arrow and never see anything.
Recap
- Every signup gets a score with reasons and a model version, so decisions can be explained and replayed.
- Proof-of-work and phone verification are rungs on a ladder, not steps for everyone.
- SMS is gated, limited and can be switched off without breaking signup.
R2.4 Design Evolution: Signals, Cost and Escalation
Step 2.1: IPs Rotate Endlessly
The problem: 100,000 requests a second arrive from 500,000 residential addresses, 0.2 a second each. Each IP looks like a household. The browsers run our JavaScript. Our per-IP limits see nothing. What would you do?
Primitive: Bot Defense, Sybil Resistance and Registration Abuse
Synthesizing vector architecture diagram...
Five families of signals feed one score. The score picks the rung; the reasons explain it.
Step 2.2: Make Each Attempt Cost Something
The problem: scoring sorts traffic, but a request with a middling score still costs the attacker nothing to send. We want every suspicious attempt to cost them something, without asking real people to do anything, and without charging money. What would you do?
Primitive: Hashcash and Client-Side Proof of Work · Drill: The Proof-of-Work Challenge That Melted Mobile Batteries (answered here and in R2.8: a fixed difficulty takes a phone many times longer than a desktop, and proof-of-work prices attempts without needing an identity, which is what a token bucket keyed on IPs can't do against rotation)
Synthesizing vector architecture diagram...
Each bit doubles the wait. Two bits past our cap turn a half-second puzzle into a two-second one; four bits turn it into a battery drain.
Step 2.3: One Challenge for Everyone
The problem: our one challenge is now a fixed price the attacker pays to a solving service. Real users on shared carrier addresses see it more than we'd like. We need friction that grows with suspicion, and a way to tell "paid for a solve" from "is a person". What would you do?
Primitive: Google reCAPTCHA v2 and Enterprise Risk Analysis · Drill: The Silent Fraud Ring That Maintained a 0.9 reCAPTCHA Score (answered here: why human-operated farms on real devices score as human, and why a continuous score beats a pass-or-fail checkbox)
Step 2.4: SMS Costs Exploded
The problem: one weekend, bots requested verification codes for phone numbers in a handful of expensive destinations, about 50 a second at the peak. Almost none of the codes were ever entered. The attacker profits from traffic to those numbers (this is SMS pumping, or SMS toll fraud; AWS calls it artificially inflated traffic). At $0.10 a message (assumption for an expensive destination), that's $5 a second: $18,000 an hour. What would you do?
Loop: Design a Notification System (Round 3, step 3.3: SMS routes and providers)
Step 2.5: Decide at the Edge
The problem: during a wave, 100,000 requests a second reach our origin fleet. It autoscales, but minutes late, and we pay for thousands of cores that exist only to say no. Our target is that at most 500 requests a second reach the origin. What would you do?
Primitive: API Gateway and Reverse Proxy · Loop: Design a Distributed Rate Limiter (Round 2, step 2.5: trusting only the edge's view of the client IP, and the layers at the edge) · Loop: Shopify: Pods, Flash Sales and Payment Resilience (Round 2, step 2.5: bots at checkout)
Synthesizing vector architecture diagram...
The edge funnel we plan for, in requests per second during a 100,000/s wave. The 30% and 99.3% are planning targets to be measured, not published rates; if the edge lets more through, the origin's scoring and proof-of-work still apply, at a higher compute cost.
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | IPs rotate endlessly | Network, TLS, device, behavior and identity signals in a points model with reasons | A privacy review; tuning |
| 2.2 | Attempts are free | Stateless Hashcash-style PoW, sized by risk and capped by the slowest phone | Seconds on phones; battery |
| 2.3 | One challenge for everyone | A ladder: invisible → PoW → challenge → verify → block | More parts to tune |
| 2.4 | The SMS bill | SMS gated by the ladder; country rules; limits per prefix and per number; a budget breaker; conversion alarms | Some users verify another way |
| 2.5 | Decide at the edge | A separate auth distribution; WAF rules cheapest first; JA4 rate rule; Bot Control targeted on auth POSTs | Per-request edge fees |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
The edge absorbs volume with rules that need no state of ours; the signup service scores what's left and picks a rung. SMS goes through its own gate with its own limits, and every decision is also written to a stream for tuning and for Round 3.
The device data (a design that fits)
| PK | SK | Attributes | Notes |
|---|---|---|---|
DEVICE#<device_id_hash> | SUMMARY | recent_accounts (list of {user_id, created_at}, at most 5), version, first_seen, ttl | The device's account count, updated with a version check |
DEVICE#<device_id_hash> | USER#<user_id> | linked_at, ja4, ip_prefix | One link per account; kept 90 days |
PHONE#<e164_hash> | SUMMARY | recent_accounts, version | Distinct accounts verified with this number |
Device IDs and phone numbers are stored as keyed hashes (an HMAC with a secret key), so the table alone doesn't reveal them.
Keeping the device limit honest under concurrency. "Read the count, check it, create the account" lets two concurrent signups from one device both see 2 and both succeed. Instead, the account transaction includes the device summary: read it with a strongly consistent read (an eventually consistent read or a global secondary index can lag and show an older count), drop entries older than 30 days, refuse if 3 remain, and write the new list with the condition version = <what we read>. If another signup changed it first, the transaction is cancelled and we retry once with a fresh read.
Trace 1: a real user on a phone, score 12
Synthesizing vector architecture diagram...
Nothing visible happened. The score came from counters read in parallel; the only writes are the idempotency claim and one transaction.
Trace 2: a risky signup, score 64
Synthesizing vector architecture diagram...
The cheap checks run before any cache call. A bot that pays for both the puzzle and a solved challenge gets through this rung, which is why the next rungs, and Round 3, exist.
Trace 3: an SMS pumping attempt
Synthesizing vector architecture diagram...
The attacker's 31st code in that prefix never leaves. The conversion alarm tells us to mark the prefix high-risk or block the country before the budget breaker is ever needed.
R2.6 Numbers and Cost
All figures are assumptions unless marked published or derived. Monthly figures use a 30-day month (720 hours, 2,592,000 s).
Traffic
| Quantity | Arithmetic | Result |
|---|---|---|
| Real signups | 10M a month (planning figure) | 10,000,000 ÷ 2,592,000 s = 3.86/s average |
| Real peak | Planning figure | 50/s |
| Attack at the edge | Planning figure | 100,000/s |
| Attack per IP | 100,000 ÷ 500,000 IPs | 0.2/s = 60 per 5 min |
| Allowed to the origin | Target | ≤ 500/s = the edge stops ≥ 99.5% |
| Burst before the edge reacts | ≈ 30 s (the WAF delay; the 0–10 s counter flush runs in parallel, so the longer one counts), assuming 10% of the wave leaks until the rules act | Up to 10,000/s for about 30 s |
| Origin tasks for that burst | 10,000/s ÷ 250 cheap rejections/s per task (assumption) | 40 tasks at most; 12 on average |
Proof-of-work, both sides
| Quantity | Arithmetic | Result |
|---|---|---|
| Server verify cost | ≈ 5 µs per check (assumption: one HMAC, one SHA-256) × 500/s | 2.5 ms of CPU a second |
| Even if all 100,000/s reached us | 100,000 × 5 µs | 0.5 of one core |
| Replay records | 500/s × 150 s × ~100 B | ≈ 75,000 keys, 7.5 MB |
| Attacker at for the whole wave | 100,000 × 1,048,576 | 1.05 × 10¹¹ hashes/s |
| At browser speed | ÷ 2M per core | ≈ 52,400 cores |
| At GPU speed | ÷ 10⁹ to 10¹⁰ per GPU (assumption) | ≈ 10 to 105 GPUs |
| Attacker cost per attempt | $21 to $210 an hour (at $2 per GPU-hour) ÷ 360M attempts an hour | ≈ $0.00000006 to $0.0000006 |
Fingerprint and signal stores
| Quantity | Arithmetic | Result |
|---|---|---|
| Device links kept 90 days | 10M a month × 3 months × 250 B | 30M links, 7.5 GB |
| Device summaries | 25M × 300 B | 7.5 GB |
| JA4 cluster counters | 50,000 fingerprints × 2 windows × 90 B | ≈ 9 MB |
| Phone summaries | 2M phone signups a month × 3 months × 300 B | 1.8 GB |
SMS
| Quantity | Arithmetic | Result |
|---|---|---|
| Signups verified by phone | 20% of 10M (assumption) | 2M a month |
| Messages | 2M × 1.2 (resends) | 2.4M a month |
| Normal spend | 2.4M × $0.02 blended (assumption) | $48,000 a month |
| Normal hour | 2.4M ÷ 720 × $0.02 | ≈ $67 average; ≈ $200 at the peak hour (× 3) |
| Budget breaker | ≈ 2 × the peak hour | $400 an hour |
| Pumping, undefended | 50/s × $0.10 × 3,600 | $18,000 an hour |
| Pumping, on 20 high-risk prefixes | 20 prefixes × 30 an hour × $0.10 | $60 an hour |
| One day of capped leak | $60 × 24 | $1,440 |
The biggest lever is the 20%. Offering phone verification only on rung 3, to perhaps 5% of signups (assumption), would cut normal SMS spend to $12,000 a month.
False positives, measured
| Quantity | Arithmetic | Result |
|---|---|---|
| Blocked real signups, budget | 0.1% of 10M | ≤ 10,000 a month |
| Audit sample | 500 blocked signups a week, reviewed by people | If 5% of them are real: ±1.9 percentage points at 95% confidence () |
Latency at the origin (p99, every hop, assumptions)
| Hop | p99 | How it adds |
|---|---|---|
| CloudFront and WAF with Bot Control | 10 ms | Sequential |
| Edge to origin and back | 70 ms | Sequential |
| ALB | 2 ms | Sequential |
| In-memory checks, PoW verify, scoring | 3 ms | Sequential |
| ElastiCache reads (counters, JA4 share, seed claim) | 3 ms | Parallel: max |
| DynamoDB: idempotency claim and device summary read | 10 ms each | Parallel: max = 10 ms |
| DynamoDB: account transaction | 25 ms | Sequential |
| Total, no challenge | 10 + 70 + 2 + 3 + 3 + 10 + 25 | 123 ms |
The edge's own decision is inside the first 10 ms (assumption; AWS doesn't publish WAF's added latency), well under the 100 ms target.
Cost (AWS list prices, us-east-1, checked September 2026)
AWS WAF: $5 a web ACL, $1 a rule, $0.60 per million requests. Bot Control: $10 a month per web ACL; the common level includes 10M requests a month then costs $1 per million; the targeted level includes 1M then costs $10 per million. We assume normal auth traffic of 300 requests/s at the edge (50/s of them POSTs to the scoped paths), and 10 attack hours a month at 100,000/s = 3.6 billion requests, 30% of which cheaper rules block before Bot Control.
| Item | Arithmetic | Monthly |
|---|---|---|
| WAF fixed | $5 + 10 rules × $1 + $10 Bot Control | $25.00 |
| WAF requests | (777.6M normal + 3,600M attack) × $0.60/M | $2,626.56 |
| Bot Control targeted, normal | (129.6M − 1M free) × $10/M | $1,286.00 |
| Bot Control targeted, attack | 3,600M × 70% × $10/M | $25,200.00 |
| WAF Challenge responses | Challenges are served only on page loads (never on POSTs), at $0.40 per million served, billed whether or not the client attempts them. An attacker must load a page for each token; if it fetches a fresh token for every one of the 2,520M posts past the cheap rules (to dodge the token-reuse rules), that's 2,520M × $0.40/M. Reusing each token for its 300 s would cut it to about 60M (500,000 IPs × 12 an hour × 10 h) = $24. We budget the upper bound | ≤ $1,008.00 |
| Signup service | 12 tasks × (1 vCPU × $0.04048 + 2 GB × $0.004445) × 720 | $426.56 |
| ElastiCache (Valkey) | 3 × cache.r7g.large × $0.1752 × 720 | $378.43 |
| DynamoDB | 120M write units × $0.625/M + 120M read units × $0.125/M + 195 GB × $0.25 | $138.75 |
| Kinesis (security events) | 1 shard × $0.015 × 720 + 259.2M records × $0.014/M | $14.43 |
| SES | 8.8M emails × $0.10/1,000 | $880.00 |
| SMS | $48,000 normal + $1,440 one day of capped leak (plus per-message fees for Filter mode, not included) | $49,440.00 |
| ALB | $0.0225 × 720 + 10 LCUs × $0.008 × 720 | $73.80 |
| Cross-AZ traffic | 100 requests/s × 10 cache calls × 200 B × ⅔ crossing = 133 KB/s → 346 GB × $0.02 | $6.91 |
| Total | ≈ $81,500 a month |
Synthesizing vector architecture diagram...
Two lines are about 92% of the bill ($74,640 of $81,504): text messages, and inspecting attack traffic at the edge. The servers and databases are rounding error.
Two lessons from the table. First, attack traffic that reaches a paid inspection rule runs up the bill: each million requests Bot Control's targeted level inspects costs $10, so rule order (cheap blocks first) and scope-down (only auth POSTs) are cost decisions, not only security ones. Rejecting the same requests at the origin would cost less in compute, about 40 tasks for a few hours, but the origin can't interrogate browsers and scales minutes late; we pay the edge for capability and speed, not price. Second, SMS is the most expensive way we verify people, and the attacker chooses when we use it.
Per real signup, the whole system costs about $81,500 ÷ 10M = $0.008.
R2.7 Trade-Offs
PoW vs CAPTCHA
| Proof-of-work | Managed challenge (CAPTCHA-style) | |
|---|---|---|
| Real user sees | Nothing, or a short spinner | Sometimes a checkbox or task |
| Needs an identity to work | No | No |
| Beaten by | Buying compute (GPUs make it cheap) | Paying people to solve |
| Accessibility | No perception or cognitive task | Needs a working non-visual path |
| Cost to us | Microseconds a check | Free to per-attempt, by vendor |
| Hurts | Slow phones, batteries | Everyone who sees it |
We use both, on different rungs: they're beaten by different kinds of spending, so stacking them raises the attacker's cost more than doubling either.
Friction vs catch rate
| Rung thresholds | Real users challenged | Fakes caught at signup | What happens to the rest |
|---|---|---|---|
| Strict | More than our 5% budget | Most | Real users leave; support tickets rise |
| Our setting | About 5% | Many | Round 3 catches survivors after signup |
| Loose | Under 1% | Few | More fakes to clean up later, and more harm first |
Buy vs build the score
| Vendor score only (reCAPTCHA Enterprise, Bot Control labels) | Our own points model, using vendor scores as inputs (chosen) | |
|---|---|---|
| Sees | Cross-site traffic we can't see | Our own accounts, devices, outcomes |
| Explains decisions | A score and some reason codes | Our reasons, versioned and replayable |
| Adapts to our abuse | On the vendor's schedule | On ours |
| Effort | Low | A team |
R2.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| PoW difficulty oscillation under bursts | A naive controller sees a spike and jumps from to : M hashes, 8.4 s on a desktop and about 56 s on a slow phone. Real signups collapse, traffic falls, the controller drops back, the flood returns. | Difficulty is set per rung, not globally, so a wave raises it only for traffic already scored as risky. It moves at most 1 bit per 60-second epoch, only when that rung's fast abuse measure (share of its new accounts that fail verification or get flagged within 10 minutes) stays outside a deadband, and it never passes the device cap: beyond on phones and on desktops, we add a challenge instead. A decision waits for the next epoch boundary, 0 to 60 s. From 16 to 20 takes at least 4 minutes. |
| A residential proxy botnet | Traffic from 500,000 home addresses, 0.2 requests/s each | IP limits are ignored as evidence; JA4 cluster share, device and identity signals and the ladder carry it (step 2.1). |
| Battery drain on low-end phones | Long solves, complaints, abandoned signups | Phones never get more than ; puzzles run in a background thread while the user types; a solve that passes 5 s stops and the client asks for a challenge instead. We watch P95 solve time by device class. |
| SMS toll fraud | Spend per hour climbs; conversion per prefix falls toward zero | Prefix caps, conversion alarms, the budget breaker and country rules (step 2.4). |
| ElastiCache fails over | Seconds of errors; some counters missing recent increments | Limits undercount briefly (acceptable); PoW seed claims fail closed (a real user retries with a new puzzle); SMS fails closed to email. |
| Bot Control's ML rules start fresh | After enabling, the coordinated-activity rules need up to 24 hours of baseline (AWS) | Enable them before a known risk window, not during an attack. |
| An edge rule change during an attack | AWS documents that changing a rate-based rule's settings "resets the rule's rate limiting counts" and can pause it "for up to a minute" | Add a new, stricter rule instead of editing the live one. |
R2.9 Production Gotchas
| Gotcha | Why it hurts | What we do |
|---|---|---|
| IP-only defenses | Addresses are rented by the hundred thousand; real users share them | IP is one weak signal among many |
| Solving or checking PoW in the wrong place | A backend that solves puzzles "for" an app pays the attacker's cost itself; verifying after database work wastes the savings | Clients solve; the server checks first, in memory, before any cache or database call |
| A PoW replay record keyed on seed and nonce | One seed can be solved over and over | Claim the seed alone, and keep the record longer than the seed is valid |
| Disposable-domain lists in slow storage | A database lookup per signup, and a list that's only as fresh as its last manual update | In memory, refreshed from S3 automatically |
| Counting events for thresholds | A person asking for a resend looks like an attacker | Count distinct accounts for "how many people"; count messages only for cost |
| A global PoW difficulty | A wave slows every real user | Difficulty per rung, capped per device class |
| Paid edge rules on everything | Attack traffic inspected at $10 a million | Cheap rules first; scope paid rules to the paths that need them |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Security | Edge filtering with reputation, rate rules by IP and JA4, and Bot Control SEC 5; a ladder that asks for stronger proof of a person only as risk rises SEC 2; device IDs and phone numbers stored as keyed hashes SEC 8 |
| Reliability | Safer modes per dependency: SMS fails closed to email, seed claims fail closed, limits undercount briefly on failover REL 5; the origin sized for the burst before the edge reacts REL 7 |
| Performance Efficiency | Parallel reads for scoring, in-memory verification before any state PERF 3 |
| Cost Optimization | A separate auth distribution so edge fees apply only to auth; cheapest rules first; paid inspection scoped down COST 5; SMS spend watched per hour with a breaker COST 3 |
| Operational Excellence | Signals for challenge rate, solve time by device class, conversion per prefix, spend per hour OPS 8; thresholds versioned in config, rolled back in seconds OPS 6 |
| Sustainability | Floods dropped at the edge before they use origin compute, and proof-of-work kept to the lowest difficulty that does the job, because its energy is spent on our users' devices SUS 3 |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Explains why IP limits fail against residential proxies with arithmetic (0.2 requests/s per IP), and moves to many weak signals compared with baselines.
- Does the proof-of-work math: , the geometric tail, a difficulty cap from a device budget, and an honest attacker cost.
- Builds a ladder where each rung is beaten by a different kind of spending.
- Treats SMS as a cost an attacker controls, and caps it with the right kinds of counters.
- Splits work between edge and origin, and knows what the edge costs per request under attack.
Follow-up questions
-
"Why not make JA4 a hard block list?" Answer: a fingerprint identifies a TLS stack, not a person. Real users of the same browser build share it, and attackers can move to a stack that matches a popular browser. We use it for velocity against a baseline, with a Challenge action, so real users whose page or app already holds a WAF token pass.
-
"The attacker claims every request is from a phone to get . Now what?" Answer: that's expected. Claiming "phone" never lowers the score; it only swaps extra puzzle difficulty for a managed challenge at the same rung. The claim also becomes a signal: a JA4 or device profile that doesn't fit the claimed device raises the score.
-
"Why keep email verification if we have all these signals?" Answer: the signals decide how hard to make signup; verification decides what an account may do. A fake account that can't prove a real channel stays
PENDINGand harmless.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Stricter IP limits" | 0.2 requests/s per IP is below any limit a household needs. |
| "A hard puzzle for everyone" | 10 seconds on a slow phone, half a second on a GPU rig. |
| "Proof-of-work makes bots uneconomical" | GPUs make each solve cost well under a millionth of a dollar. It's a tax. |
| "Rate limit SMS per phone number" | Pumpers use a new number each time; cap per prefix and per budget. |
| "Everything at the origin" | Scaling lags the wave; the origin can't interrogate browsers. |
| "Edit the live rate rule during an attack" | Its counts reset and it can pause for up to a minute. |
Round 3 · Architect · "Trust and Safety Across the Whole Platform"
~45 min · Principal (L7) · 200M accounts, 50M daily users (assumptions) · ≈ 3,150 scored actions/s on average, ≈ 29,500/s at the peak of a login attack (derived) · a measured fake-account rate, a measured false-positive rate, a friction budget per action
R3.0 Where We Left Off
Round 2 in 60 seconds. "Attackers moved to 500,000 residential IPs at 0.2 requests a second each, so IP limits stopped meaning anything. We combined weak signals, network, TLS fingerprint velocity against a baseline, device account counts, behavior and identity, into a points score with reasons and a version. The score picks a rung on a ladder: invisible, a Hashcash-style proof-of-work sized by risk and capped at 17 bits on phones so the slowest supported phone finishes in 2 s at P95, a managed challenge, a verified email or phone, or a block. Proof-of-work is a tax, not a wall: GPUs make each solve cost well under a millionth of a dollar, so the expensive rungs sit above it. SMS is gated by the ladder, blocked by country where we don't serve, capped per destination prefix and per hour of spend, and watched for conversion. The auth endpoints have their own CloudFront distribution, with WAF rules cheapest first and Bot Control's targeted level scoped to auth
POSTs. It costs about $81,500 a month, about 92% of it SMS and edge inspection of attack traffic. Open costs: fakes that pass signup are never looked at again, logins are unprotected, rules are hand-tuned, and nobody can appeal a wrong decision."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: a strong front door. Everything after the door, logins, posts, payments, is still unguarded.
Round 2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | IPs rotate | Many weak signals, one explained score | Privacy review, tuning |
| 2.2 | Attempts are free | Proof-of-work sized by risk, capped per device | Seconds on phones |
| 2.3 | One challenge for all | A ladder of different costs | Tuning |
| 2.4 | The SMS bill | Gated SMS, prefix caps, budget breaker | Some users verify differently |
| 2.5 | Decide at the edge | WAF cheapest first; Bot Control scoped | Edge fees under attack |
Open costs: signup-only defense, hand-written rules, no feedback, no appeals.
R3.1 The Scope Raise
Interviewer: "Signup is much better. But the fakes that got through are now posting scams and messaging members. Last month, a wave of logins with stolen passwords took over thousands of real accounts. The attackers change tactics every week, and we find out from users. Legal wants us to show we handle personal data carefully, and we've had complaints from blind users about challenges. And people we blocked by mistake have nowhere to go. We want one trust system for the whole platform."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Which actions matter? | Signup, login, posting, messaging people you don't know, and payments. | Score every sensitive action, not just signup (R3.3). |
| How big? | 200M accounts, 50M daily users (assumptions for this round). | ≈ 3,150 scored actions a second on average (R3.6). |
| How do fakes show up after signup? | Groups of accounts that share phones, cards and devices, and act together. | Look at accounts as a graph and act on clusters (step 3.1). |
| The login attack? | 20,000 attempts a second from about 100,000 residential IPs, each using a different leaked email and password pair. About 0.5% worked. | Per-IP and per-account limits both miss it; we need other keys (step 3.2). |
| What do we learn from? | User reports, chargebacks, what reviewers decide, appeals. None of it feeds back into the rules today. | A labeled feedback loop and safe rollout of new rules (step 3.3). |
| Mistakes? | Real people get blocked and email the CEO. | Appeals with human review, and false positives measured, not guessed (step 3.4). |
| Privacy and accessibility? | We operate in many countries. Accessibility rules apply to us. | Minimize and bound what we keep; accessible alternatives for every challenge (step 3.5). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Protected actions | Signup | Signup, login, posting, messaging, payments |
| When we decide | At signup | At every sensitive action, and again later as evidence arrives |
| Unit of judgment | A request | An account, and the cluster it belongs to |
| Rules | Hand-tuned points | Trained on labeled outcomes, shadow-tested before enforcing |
| Mistakes | Not tracked | Appeals; false-positive rate measured by audit |
| Privacy and accessibility | Case by case | Designed in: purpose, retention, alternatives |
R3.2 What Breaks in the Round 2 Design
| Round 2 piece | What breaks at the new scale |
|---|---|
| Signup-only defenses | A fake that passes once is trusted forever; a real account taken over at login bypasses signup entirely. |
| Static, hand-written rules | Attackers find the thresholds in days. Every change is a guess, shipped straight to enforcement. |
| No feedback loop | We never learn which decisions were right. Reports, chargebacks and reviews sit in other systems. |
| No appeals | Every false positive is permanent, and we can't even count them. |
| Each product for itself | Payments sees a card used by 40 accounts; messaging sees those 40 accounts spamming; neither knows about the other. |
R3.3 New Requirements and API Additions
1. A trust score per account, updated over time. Every product asks the same service.
httpGET /internal/v1/trust/accounts/usr_8842?action=message_stranger HTTP/1.1
httpHTTP/1.1 200 OK Content-Type: application/json { "account_id": "usr_8842", "trust_tier": "NEW", "action_decision": "ALLOW_LIMITED", "limits": { "messages_to_strangers_per_day": 5 }, "reasons": ["account_age_under_7_days", "email_verified"], "cluster_risk": "NONE", "model_version": "trust-2026-09-24.2", "evaluated_at": "2026-09-28T10:02:11Z" }
2. Enforcement actions. Products don't ban accounts themselves; they ask the trust service to act, so every action is recorded, explained and appealable.
httpPOST /internal/v1/enforcement/actions HTTP/1.1 Content-Type: application/json Idempotency-Key: enf-cluster-77120-usr_8842 { "account_id": "usr_8842", "action": "RESTRICT", "scope": ["post", "message_stranger", "payout"], "reason_codes": ["cluster_shared_payment_instrument"], "evidence_ref": "cluster/77120/2026-09-28", "appealable": true }
3. Appeals.
httpPOST /v1/appeals HTTP/1.1 Content-Type: application/json Idempotency-Key: 3c1b9e0a-44f2-4d61-b8e2-7a0c5d19f6b3 { "enforcement_id": "enf_55120", "statement": "This is my only account. I share a laptop with my brother." }
Synthesizing vector architecture diagram...
A reviewer's claim on an appeal is a lease: if they walk away, it returns to the queue. Both outcomes become labels for the feedback loop (step 3.3).
4. A signal-sharing contract. Every product emits events in one schema, and each field says why we're allowed to use it and for how long.
json{ "event_id": "evt_01J9ZT8Q2M", "event_seq": 1840021, "account_id": "usr_8842", "action": "payment_attempt", "occurred_at": "2026-09-28T10:02:10Z", "signals": { "device_id_hash": { "value": "d8c1...", "purpose": "abuse_prevention", "retention_days": 90 }, "payment_instrument_hash": { "value": "p41a...", "purpose": "abuse_prevention", "retention_days": 180 }, "ip_prefix": { "value": "203.0.113.0/24", "purpose": "abuse_prevention", "retention_days": 30 } }, "outcome": "declined_by_issuer" }
event_seq increases per account, so a consumer that sees an event twice can tell (step 3.3).
Recap
- One trust service answers "may this account do this action now?" for every product, with reasons.
- Enforcement goes through one API, so it's recorded and appealable.
- Signals carry their purpose and retention with them.
R3.4 Design Evolution: From a Front Door to a Trust System
Step 3.1: Fakes Got Through Signup
The problem: accounts that passed every signup check are now posting scams. Individually each looks fine: a verified email, a normal device, a residential address. But reviewers keep noticing that the scam accounts share things: the same few phones, the same payment cards, the same devices. What would you do?
Synthesizing vector architecture diagram...
Accounts 1 to 4 never share one thing, but a chain of shared devices, cards and phones links them. The carrier address links them to a stranger too, which is exactly why weak links don't join clusters.
Step 3.2: Credential Stuffing on Login
The problem: 20,000 login attempts a second from about 100,000 residential IPs, 0.2 a second each. Every attempt tries a different email and password pair from someone else's breach. Each account is tried once. About 0.5% of pairs work: undefended, that's 100 taken-over accounts a second, 6,000 a minute. What would you do?
Drill: The Credential Stuffing Wave That Bypassed IP Rate Limits (answered here: why per-IP limits fail at 0.2 attempts per IP, and why an SMS code on every login is the wrong trade) · Loop: Design a Distributed Rate Limiter (Round 3, step 3.3: the same attack seen from the gateway)
Synthesizing vector architecture diagram...
The password was right, and it didn't matter: a correct password from breach data on an unknown device earns a step-up, not a session.
Step 3.3: Attackers Adapt Weekly
The problem: every rule we ship works for about a week. We change thresholds by hand, straight in production, and sometimes a change blocks thousands of real users before anyone notices. We never learn which decisions were right. What would you do?
Synthesizing vector architecture diagram...
A rule or model earns enforcement in stages, and either guardrail trip sends it back. Nothing goes from a draft straight to blocking people.
Step 3.4: A Real User Was Blocked
The problem: a real user shares a laptop with a brother whose account was banned for spam. Our cluster rule restricted her too. She has no way to tell us, and we have no idea how often this happens. What would you do?
Step 3.5: Privacy and Accessibility
The problem: the signal store holds device IDs, IP prefixes, card fingerprints and phone numbers for 200M accounts, forever, and every product can read all of it. Separately, blind users report they can't get past a challenge. What would you do?
Step 3.6: Build or Buy?
The problem: the CFO asks why we run a trust team when vendors sell bot management and fraud scoring. The security team wants a single vendor. The product teams want control. What would you do?
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Fakes got through | A graph of strong links, hub pruning, clusters, graded action; an instant neighbor lookup; limits for new accounts | A pipeline; cluster false positives |
| 3.2 | Credential stuffing | Breached-password checks, velocity on distinct keys, device binding, step-up, edge layers | Step-ups on new devices during attacks |
| 3.3 | Attackers adapt | Labels, audit samples, weekly retraining, shadow → canary → enforce with guardrails; replay-safe features | Slower changes; ML operations |
| 3.4 | A real user was blocked | Appeals with leased review; precision measured by audit | Reviewers |
| 3.5 | Privacy and accessibility | Purpose and retention per field, keyed hashes, decisions not raw data, Privacy Pass where supported; non-visual paths; no puzzles at login | Fewer signals |
| 3.6 | Build or buy | Buy vantage at the edge; build the account-level core | A team |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Products ask the trust service before sensitive actions and stream what happened. Features update continuously; clusters update daily; labels from review and appeals flow back into training, and new versions reach the trust service only through shadow and canary.
The trust data (a design that fits)
| PK | SK | Attributes | Notes |
|---|---|---|---|
TRUST#<account_id> | CURRENT | tier, score, reasons, model_version, cluster_id, last_event_seq, version | Updated only if event_seq > last_event_seq |
FEAT#<account_id> | H#<yyyymmddhh> | posts, reports_received, stranger_messages, ttl | Absolute values per hour, safe to rewrite |
ENF#<account_id> | ACT#<time>#<id> | action, scope, reason_codes, evidence_ref, status | Written with an idempotency key |
APPEAL#<appeal_id> | STATE | status, claimed_by, claim_until, decision, version | Decisions conditional on the current claim |
CLUSTER#<cluster_id> | SUMMARY | size, risk, reasons, computed_at | From the daily job |
Synthesizing vector architecture diagram...
Strong links (device, card, phone) build clusters; IP prefixes are kept as weak context only. Every enforcement can carry one appeal.
Per-action safe defaults decide what happens when the trust service can't answer (R3.8):
| Action | If the trust service is down |
|---|---|
| Signup | Allow into PENDING with proof-of-work and the Round 1 limits: safer, not open |
| Login | Known device: allow. Unknown device: require the second factor |
| Post, message a stranger | Accounts older than 30 days: allow. Newer: hold the post for later scoring, block messages to strangers |
| Payment | Under $50: allow with the payment provider's own fraud checks. Above: hold for asynchronous review |
Trace 1: a fake cluster is removed
Synthesizing vector architecture diagram...
Restriction first, bans only for accounts a person reviewed: the sample found 1 real account in 20, so the 220 unreviewed ones stay restricted and appealable rather than banned. The one real account is restored, and her link is cleared so the next run doesn't catch her again.
Trace 2: the credential-stuffing wave is step 3.2's sequence: the edge challenges most attempts, the login service steps up correct-but-breached passwords, and the failure-ratio rule then steps up every unknown-device login until the wave ends.
Trace 3: an appeal
Synthesizing vector architecture diagram...
The decision write is conditional on the reviewer still holding the claim, so a reviewer whose claim expired and was taken over can't overwrite the newer decision.
R3.6 Numbers and Cost
All figures are assumptions unless marked published or derived. Monthly figures use a 30-day month.
Scoring load
| Action | Per day | Average per second | Peak (× 3) |
|---|---|---|---|
| Signup | 333,333 (10M a month) | 3.86 | 11.6 |
| Login | 20M | 231.5 | 694 |
| Post | 200M | 2,314.8 | 6,944 |
| Message to a stranger | 50M | 578.7 | 1,736 |
| Payment | 2M | 23.1 | 69 |
| Total | 3,152 | 9,456 | |
| Plus a credential-stuffing wave before the edge reacts | 20,000/s | 29,456 |
The scoring fleet
| Quantity | Arithmetic | Result |
|---|---|---|
| CPU per score | Assumption | 2 ms |
| At the wave peak | 29,456 × 2 ms = 58.9 vCPU ÷ 0.6 target utilization | ≈ 98 vCPU → 50 tasks of 2 vCPU |
| On average | 3,152 × 2 ms = 6.3 vCPU | 8 tasks (16 vCPU), about 39% busy |
| Latency inside an action | Parallel reads from ElastiCache (3 ms p99) and one DynamoDB read on a cache miss (10 ms p99) + scoring (2 ms) | ≤ 15 ms p99 (assumption) |
Scaling from 8 to 50 tasks takes minutes; until it does, the edge sheds most of the wave and actions fall back to their safe defaults (R3.8). That's a choice: 50 tasks all month would cost about $3,000 more (42 extra tasks × $0.09874 × 720).
The graph
| Quantity | Arithmetic | Result |
|---|---|---|
| Nodes | 200M accounts + 250M devices + 40M cards + 60M phones + 50M IP prefixes | ≈ 600M |
| Edges | 400M account–device + 50M account–card + 60M account–phone + 1B account–IP prefix (30 days) | ≈ 1.5B |
| Size | 1.5B × ~50 B | ≈ 75 GB in S3 |
| Daily job | 400 vCPU-hours and 3,200 GB-hours (assumption) | ≈ 2 hours |
| Detection delay for a new cluster | Uniform 0–24 h wait + 2 h run | 14 h average, 25.8 h at p99 |
Labels and people
| Quantity | Arithmetic | Result |
|---|---|---|
| Enforcement actions | Assumption | 100,000 a day |
| Appeals | 2% of actions | 2,000 a day |
| Overturned | 20% of appeals | 400 a day (0.4% of actions) |
| Audit sample | Fixed | 500 a day |
| Prevalence sample (new accounts labeled fake or real) | 1,000 a week | ≈ 143 a day; at 3% fake, the week's 1,000 give ±1.1 points at 95% |
| Reviews a day | 2,000 + 500 + 143 | 2,643 |
| Reviewers | 2,643 ÷ 60 = 44.05 reviews' worth per day | 45 people a day (not an AWS cost) |
Cost (AWS list prices, us-east-1, checked September 2026): what Round 3 adds to Round 2's $81,500
| Item | Arithmetic | Monthly |
|---|---|---|
| Trust service, base | 8 tasks × (2 vCPU × $0.04048 + 4 GB × $0.004445) × 720 | $568.74 |
| Trust service, waves | 4 waves × 2 h × 42 extra tasks: 672 vCPU-h × $0.04048 + 1,344 GB-h × $0.004445 | $33.17 |
| Hot features (ElastiCache Valkey) | 3 shards × 2 nodes cache.r7g.large × $0.1752 × 720 | $756.86 |
| Kinesis | 32 shards × $0.015 × 720 + 8.17B records × $0.014/M | $459.98 |
| DynamoDB | 817M write units × $0.625/M + 2.45B read units × $0.125/M + 150 GB × $0.25 | $854.38 |
| Graph job (EMR Serverless) | (400 × $0.052624 + 3,200 × $0.0057785) × 30 | $1,186.22 |
| S3 for the graph | 75 GB × $0.023 | $1.73 |
| WAF on login waves | 4 × 2 h × 20,000/s = 576M requests × $0.60/M + 70% × 576M × $10/M (Bot Control targeted) | $4,377.60 |
| WAF Challenge responses on login waves | At most one fresh token per wave request past the cheap rules: 403.2M × $0.40/M (an upper bound; real clients reuse a token for its 300 s immunity) | ≤ $161.28 |
| Round 3 additions | $8,399.96 | |
| Platform total | $81,504.44 + $8,399.96 | ≈ $89,900 a month |
Notes on the rows: the feature cache holds about 30 GB of hot features (50M daily users' recent features plus device and IP aggregates) on three shards with a replica each (about 29 GiB usable, so we add a fourth shard at the first sign of evictions); 32 Kinesis shards take 32,000 records a second, enough for the wave; DynamoDB reads are eventually consistent (half a read unit each) because trust features tolerate a second of lag, while enforcement and appeal decisions use conditional writes, which always act on the latest item.
Why not AWS's account-takeover rules on every login? Fraud Control's per-request tiers run from $1,000 down to $50 per million requests analyzed (the first 10,000 free; then $0.001, $0.0007, $0.0004 and $0.0002 per request up to 30M, and $0.00005 beyond). The 576M wave requests alone would cost about $38,400 a month at the tiered prices ($28,800 even if every request were billed at the lowest tier), and attackers decide how many requests there are. If we use it, it goes behind the cheaper rules, scoped to login, where it sees only what they didn't stop.
R3.7 Trade-Offs
Friction vs abuse, per action
| Action | Friction budget | Why |
|---|---|---|
| Signup | Challenge ≤ 5% of real users | First impression; losing a signup is final |
| Login | Step-up ≤ 2% of real logins, mostly new devices | Repeated every session |
| Message a stranger | New accounts limited for 7 days | Most harm per fake account happens here |
| Payment | Holds on high value only | Money has its own fraud checks too |
False positives vs catch rate
| Tuned for catch rate | Tuned for precision (our default) | |
|---|---|---|
| Fakes removed quickly | More | Fewer; clusters and limits catch the rest later |
| Real users hurt | More, and most never appeal | Fewer, measured by audit |
| Where we move the line | During an active attack, temporarily, with a rollback time | Normal operation |
Centralized trust vs per-product (closing the loop)
| Each product its own | One trust service (chosen) | |
|---|---|---|
| Sees cross-product patterns (a card on 40 accounts that spam) | No | Yes |
| One enforcement and appeals path | No: five | Yes |
| Blast radius of a bad model | One product | Every product: hence shadow, canary and per-action safe defaults |
| Team autonomy | High | Products own their actions and friction budgets; the service owns scores |
The whole loop, in one line: Round 1 made fakes need a real inbox, Round 2 made them pay per attempt and per rung, and Round 3 makes them pay again after signup, because a fake that looks perfect at the door still has to behave like a person, alongside the other fakes it shares a device, a card or a phone with.
R3.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| The trust service is down | Products' calls time out (we wait 50 ms) | Each action falls back to its safe default (R3.5 table): nothing opens fully, and nothing important closes for everyone. Decisions made in fallback are re-scored when the service returns. |
| A new model blocks many real users | Canary's challenge or block rate doubles; appeals spike | The guardrail rolls back automatically within 15 minutes; if it slipped past the canary, one config change pins the previous version. Actions taken by the bad version are queued for re-evaluation and lifted in bulk. |
| A coordinated attack during the outage | A signup wave while the trust service is down | The edge doesn't depend on our services, so its rules keep working. Fallback signup is PENDING with proof-of-work and Round 1 limits, and new accounts can't message strangers, so fakes created in the window can't do much before the service returns and re-scores them. |
| The graph job fails | No new clusters today | Yesterday's clusters keep applying through the neighbor lookup; an alarm fires when cluster data is more than 36 hours old. |
| Labels are poisoned | A spike of reports against one group of real users | Reports are a weak label; we train on reviewer decisions and audit labels, and cap how much any single reporter's reports count. |
| Kinesis consumer falls behind | Features lag minutes behind | Scores use slightly older features; last_event_seq keeps late events from being applied twice; alarm on iterator age. |
R3.9 Runbook and Incident Response
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Challenge rate on real-looking traffic (signup) | Above 5% for 30 min | P2 | Which rule or model version? Roll back if new. |
| Challenge pass rate | Drops 20 points in 15 min | P2 | Vendor issue, or a real-user segment failing (browser, country)? |
| Edge block rate, auth distribution | 10× baseline | P2 | Attack starting: check sampled requests and top JA4s. |
| Login failure ratio | Above 20% over 10 s windows for 1 min | P1 | Credential stuffing: confirm the adaptive rule engaged. |
| SMS spend per hour | Above $400 | P1 | Breaker should be open; find the prefixes. |
| SMS conversion per prefix | Below 20% with ≥ 50 sends in 15 min | P2 | Mark the prefix high-risk or block the country. |
| Fake-account rate (weekly prevalence sample) | Up 1 point week over week | P3 | Which signup path let them in? |
| Appeal overturn rate | Above 30%, or 2× last week | P2 | A rule or model is too aggressive. |
| PoW P95 solve time, phones | Above 2 s | P3 | Difficulty cap or client regression. |
| Trust service p99 latency | Above 15 ms for 5 min | P2 | Cache misses? Scale out. |
Procedure: an attack is in progress SEC 10 · OPS 10
- Name it. Signup wave, login wave, or SMS pumping? Which distribution, which paths?
- Contain at the edge first. Add a new, stricter rate-based rule (by IP or by JA4) with a Challenge action (on
POSTs, a check for a valid WAF token) rather than editing a live rule, which would reset its counts. - Raise the ladder, not the wall. Lower the score thresholds for the affected action for a fixed period (1 hour) with an automatic end, so a temporary change can't become permanent by accident.
- Protect the money. For SMS, block the pumped countries or prefixes in the protect configuration; for logins, step up every unknown device.
- Watch real users. Challenge pass rate and completion rate by segment: are we hurting a country, a browser, assistive technology users?
- After the attack: revoke sessions opened from unknown devices during the wave, ask their owners to reset, remove the temporary thresholds, and write a correction-of-error review: which signal would have caught it sooner?
Go deeper: CLI checks (replace names and IDs with real ones)
text# 1. Which IPs a rate-based rule is limiting right now (CloudFront web ACLs live in us-east-1) aws wafv2 get-rate-based-statement-managed-keys --scope CLOUDFRONT --region us-east-1 --web-acl-name auth-edge --web-acl-id 11111111-2222-3333-4444-555555555555 --rule-name register-rate-by-ip # 2. A sample of requests one rule matched in a 15-minute window aws wafv2 get-sampled-requests --scope CLOUDFRONT --region us-east-1 --web-acl-arn arn:aws:wafv2:us-east-1:111122223333:global/webacl/auth-edge/11111111-2222-3333-4444-555555555555 --rule-metric-name register-rate-by-ja4 --time-window StartTime=2026-09-28T10:00:00Z,EndTime=2026-09-28T10:15:00Z --max-items 100 # 3. The SMS protect configurations aws pinpoint-sms-voice-v2 describe-protect-configurations --region us-east-1 # 4. Block SMS to one country during pumping (XX is the two-letter ISO country code) aws pinpoint-sms-voice-v2 update-protect-configuration-country-rule-set --protect-configuration-id protect-0123456789abcdef --number-capability SMS --country-rule-set-updates '{"XX":{"ProtectStatus":"BLOCK"}}' --region us-east-1 # 5. The account's monthly SMS spend limits aws pinpoint-sms-voice-v2 describe-spend-limits --region us-east-1 # 6. The feature cache's shards and which nodes are primaries now aws elasticache describe-replication-groups --replication-group-id trust-features # 7. Alarms currently firing for the trust system aws cloudwatch describe-alarms --state-value ALARM --alarm-name-prefix trust-
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Security | Breached-password checks, device binding and step-up at login SEC 2; detection after signup from events and the graph SEC 4; signals classified by purpose and retention in the event contract SEC 7; keyed hashes for identifiers SEC 8; an attack procedure SEC 10 |
| Reliability | Per-action safe defaults when the trust service fails REL 5; replay-safe features with last_event_seq and idempotent enforcement REL 4; guardrail rollbacks REL 8 |
| Performance Efficiency | Hot features in memory, one read on a miss, a scoring budget under 15 ms inside every action PERF 3 |
| Cost Optimization | Managed per-request rules scoped behind cheaper ones, with the arithmetic that justifies it COST 5; build vs buy decided by capability and team cost COST 11; the scoring fleet scales for waves instead of sitting at peak COST 9 |
| Operational Excellence | Shadow, canary and automatic rollback for every rule and model OPS 6; signals and an attack procedure OPS 10; audits, appeals and correction-of-error reviews feed the next version OPS 11 |
| Sustainability | A daily batch for the graph instead of a continuously running cluster, and a scoring fleet that scales with waves SUS 2; attack traffic dropped at the edge and proof-of-work kept minimal, so less energy is spent on requests we'll refuse SUS 3 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Moves from judging requests to judging accounts and clusters, and knows which links are strong and which are hubs.
- Defends logins without lockouts: breach data, velocity on distinct keys, device binding, step-up, with arithmetic for what gets through.
- Builds a feedback loop with labels, an unbiased audit sample, and a safe rollout path with automatic rollback.
- Treats false positives as a measured number, and knows appeals undercount them.
- Designs privacy and accessibility in, with claims kept to what the standards actually say.
- Decides build vs buy by capability, and prices vendor rules under attack.
Follow-up questions
-
"Why not run the graph job every hour?" Answer: it would cut the average delay for a brand-new cluster from 14 hours to about 2.5 (0.5 h average wait plus 2 h run), for 24 times the compute. The neighbor lookup already catches accounts joining known clusters in seconds, so the batch's job is finding new clusters. We'd first try an incremental job on just the last hour's new links.
-
"A card is shared by 12 accounts. Is that a fraud ring?" Answer: maybe a family, a small business, or a ring. The card crossed our hub threshold of 10, so it no longer merges accounts into one cluster, but it's a signal on each account. Their behavior decides: 12 accounts sharing a card that all message strangers on day one is a ring; 12 that buy groceries isn't.
-
"The trust service is a single point of failure for every product. Isn't that worse?" Answer: it is a shared dependency, so each action has a safe default that needs no trust service, and the edge keeps working without it. The alternative, five separate systems, fails more quietly: none of them sees the ring.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Only tighten signup" | Fakes that look perfect at the door are caught by what they do after. |
| "Cluster on shared IPs" | A carrier address links millions of strangers. |
| "Lock accounts after 3 failures" | Stuffing tries each account once, and lockouts let attackers lock out real users. |
| "SMS code on every login" | $400,000 a day, pumping, and friction on every session. |
| "Ship rule changes straight to enforcement" | You find out from angry users. Shadow, then canary. |
| "Appeals tell us our false-positive rate" | Most wrongly blocked people just leave. Audit a random sample. |
| "Invisible challenges are accessible" | Until they escalate to a visual task. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scope: what's attacked, how much, what must stay easy | Restate Round 1 in 60 seconds | Restate Round 2 in 60 seconds |
| 5–15 min | Requirements; register API with idempotency and a challenge error; account states | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.4: limits that challenge → invisible traps → email verification and domain screen → risk-based challenge | Steps 2.1–2.5: signals and score → PoW math → ladder → SMS pumping → edge vs origin | Steps 3.1–3.6: clusters → credential stuffing → feedback loop → appeals → privacy and accessibility → build vs buy |
| 40–50 min | Attack vs real traffic; what the limits let through; memory of counter vs log; cost | PoW tables and cap; attacker cost; SMS spend and caps; edge fees under attack | Scoring load; graph size and delay; audit math; cost of vendor rules |
| 50–60 min | Failures and pillar check | Failures, gotchas, pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint. For rate limiting in depth, see the rate limiter loop; for bots at a checkout rather than a signup, the Shopify case study; for idempotency and leases underneath all of this, the loop primitives on idempotency and leases and fencing.
The Two Sentences That Matter Most
- Opening any round: "Fake accounts are an economics problem: I'll raise the attacker's cost per account above what it's worth to them, spending friction only on traffic that looks automated, so real people see nothing."
- When scale arrives: "No single signal is proof, so I'll combine many weak ones, escalate friction one rung at a time, keep judging accounts after signup as evidence arrives, and measure my own false positives with an audit sample, because the attacker adapts and so must the system."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Security | "How do you stop bots at the front door?" (SEC 5) | Volume and reputation rules at the edge, cheapest first, then a risk score and a ladder of friction at the origin. | 1, 2 | Steps 1.1, 2.5 |
| "How do you know a new account is a real person?" (SEC 2) | We don't at first: it starts PENDING, earns powers by verifying a channel, and is limited for its first week. | 1, 3 | Steps 1.3, 3.1 | |
| "How do you stop credential stuffing?" (SEC 2) | Breached passwords never become sessions, velocity on distinct keys, known-device tokens, and step-up instead of lockout. | 3 | Step 3.2 | |
| "Do you store fingerprints forever?" (SEC 7) | No: each signal carries a purpose and a retention period, identifiers are keyed hashes, and products see decisions, not raw data. | 3 | Step 3.5 | |
| Reliability | "Your challenge vendor is down. Is signup down?" (REL 5) | No: challenged signups become PENDING with stricter limits; a safer mode, never open. | 1, 3 | R1.9, R3.8 |
| "A retry hits your signup. Two accounts?" (REL 4) | No: an idempotency record is claimed before any outside call, the email item is unique, and the owner check fences a late original. | 1 | R1.6 | |
| Performance | "How fast is the decision?" (PERF 3) | Around 120 ms at p99 for a signup without a challenge, every hop counted; scoring inside other actions under 15 ms. | 1, 2, 3 | R1.7, R2.6, R3.6 |
| Cost | "What does an attack cost you?" (COST 5) | Edge inspection of attack traffic at $10 a million requests is our second-biggest line, so cheap rules go first and paid ones are scoped to auth POSTs. | 2 | R2.6 |
| "How do you stop an SMS bill from exploding?" (COST 3) | SMS only after the ladder, blocked by country, capped per prefix and per hour, with conversion alarms: $60 an hour on high-risk prefixes and never more than the $400-an-hour breaker, instead of $18,000. | 2 | Step 2.4 | |
| "Why not buy it all?" (COST 11) | Buy vantage and browser checks at the edge; build what needs our accounts and graph. | 3 | Step 3.6 | |
| Operations | "How do you change rules safely?" (OPS 6) | Shadow for a week, canary at 5% with guardrails, automatic rollback. | 3 | Step 3.3 |
| "How do you know you're not hurting real users?" (OPS 8) | Challenge and pass rates by segment, and a daily audit sample that measures precision; appeals undercount. | 2, 3 | Steps 3.3, 3.4 | |
| Sustainability | "Isn't proof-of-work wasteful?" (SUS 3) | It spends our users' energy, so it's the lowest difficulty that works, only on risky traffic, and floods are dropped at the edge first. | 2 | Step 2.2 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Limits | Shared, atomic counters; over the limit means challenge | Many signals against baselines; distinct counts | Velocity on keys attackers can't rotate cheaply (credentials, devices) |
| Friction | Challenge only when a signal fires | A ladder with different costs per rung | A friction budget per action, measured |
| Proof of a person | Email verification | Gated phone verification; PoW as a tax | Behavior after signup; clusters |
| Failures | Safer modes per dependency | Fail closed where money is at stake (SMS) | Safe defaults per action; automatic rollback of bad models |
| Measurement | Challenge rate | False-positive budget with a sample | Audit-based precision, prevalence sampling, appeals |
| Honesty | Labels assumptions | Says what PoW can't do | Keeps legal and vendor claims to what's documented |
Sources
All sources used on this page, checked September 2026 unless a date is given.
- Back, A., Hashcash, proposed 1997: the original computational postage idea.
- Cloudflare, Turnstile: server-side validation: the
siteverifyendpoint, 300 s token validity, single use,timeout-or-duplicate,idempotency_key. - Cloudflare, Turnstile widget types and Turnstile plans: managed, non-interactive and invisible modes; the free plan's 20 widgets and unlimited challenges.
- Cloudflare, Eliminating CAPTCHAs on iPhones and Macs using new standard, June 8, 2022: Private Access Tokens, supported OS versions, attester and issuer roles.
- Google, reCAPTCHA: verifying the user's response and reCAPTCHA v3: two-minute tokens verified once, the 0.0–1.0 score.
- Google Cloud, Interpret assessments for websites: 11 score levels, four without a billing account.
- AWS, AWS WAF Developer Guide: rate-based rule settings, rate-based rule caveats, aggregation options, request components including JA3 and JA4, Bot Control, the Bot Control rule group, CAPTCHA and Challenge and immunity times.
- AWS, AWS WAF pricing: web ACL, rule and request fees; Bot Control; CAPTCHA; Fraud Control.
- AWS, Adding CloudFront request headers:
CloudFront-Viewer-JA3-FingerprintandCloudFront-Viewer-JA4-Fingerprint. - AWS, Amazon Cognito threat protection: compromised credentials, adaptive authentication, the Plus plan, no rate limits.
- AWS, What is Amazon Fraud Detector?: closed to new customers as of November 7, 2025.
- AWS, AWS End User Messaging SMS: country rule modes (artificially inflated traffic; Allow, Block, Monitor, Filter) and spending quotas.
- AWS, Amazon SES enforcement FAQ: bounce rate below 2% recommended; review at 5%, possible pause at 10%.
- AWS pricing pages and price list API for Fargate, ElastiCache, DynamoDB, Kinesis Data Streams, SES, Lambda, EMR Serverless, S3 and Elastic Load Balancing (
us-east-1). - FoxIO, JA4+ fingerprinting.
- Have I Been Pwned, API v3: Pwned Passwords: the k-anonymity range API, padding, no key, no rate limit.
- NIST, SP 800-63B-4, Digital Identity Guidelines: Authentication and Authenticator Management, 2025: password blocklists, no composition rules or forced periodic changes, minimum lengths, the 100-attempt ceiling and throttling techniques.
- IETF, RFC 9576: The Privacy Pass Architecture, June 2024, with RFC 9577 (the HTTP authentication scheme) and RFC 9578 (issuance protocols).
- W3C, Understanding Success Criterion 3.3.8: Accessible Authentication (Minimum), WCAG 2.2, and success criterion 1.1.1 Non-text Content (its CAPTCHA provision).