Design a Distributed Email Service
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | Email for one company's 5,000 employees on its own domain | A consumer webmail service like Gmail | Business email sold to many companies, like Google Workspace or Microsoft 365 |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Volume | 500K messages/day (5.8/s, 17.4/s at peak); 50 GB/day | 100M mailboxes; 1B messages/day (11,574/s, 34.7K/s at peak); about 83 PB stored after 3 years | 10M seats in 25,000 companies; 1B messages/day, 60% in the US and 40% in the EU |
| Footprint | 1 region | 1 region, 3 AZs | 4 regions: a home and a standby region in each of 2 jurisdictions (US, EU) |
| Targets | Inbox view < 500 ms; never lose accepted mail; 99.9% | New mail visible < 2 s P95; folder view < 40 ms P99; search < 300 ms P95 on recent mail; 99.99% | Retention and legal holds; EU mail stays in the EU; customer-held keys; survive a region loss |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is an Email Service?
You Already Know One: a Post Office, a Filing Cabinet and a Librarian
Think of three jobs in one building.
- The post office accepts letters from anyone in the world, checks they aren't dangerous, and hands them to the right person. It also sends your letters out, and when an address doesn't exist, it brings the letter back with a note.
- The filing cabinet keeps every letter a person ever received, sorted into folders: Inbox, Sent, Spam, Trash, and whatever folders they made.
- The librarian can find any old letter in a second: "the invoice from Acme last spring."
An email service is those three jobs, run for millions of people, on a network where anyone can send you anything. A few words we'll use all page:
| Word | What it means on this page |
|---|---|
| SMTP | Simple Mail Transfer Protocol: the text protocol mail servers use to hand a message to each other, on TCP port 25. |
| MTA | Mail transfer agent: a server that speaks SMTP, sending or receiving. Postfix and Exim are well-known ones. |
| MX record | A DNS record that says which servers accept mail for a domain: "mail for acme.example goes to mx1.acme.example." |
| Envelope vs headers | The envelope is what SMTP carries outside the message (MAIL FROM, RCPT TO). The headers (From:, To:, Subject:) are inside the message. They can differ: a Bcc recipient is in the envelope, not the headers. |
| MIME | The format that lets one message hold plain text, HTML and attachments as nested parts, each with its own type and encoding. |
| SPF, DKIM, DMARC | The three checks that tell a receiver whether a message really comes from the domain it claims. We define each one in step 1.4. |
| Bounce | A report that a message couldn't be delivered. Soft (a 4xx SMTP code: try again later) or hard (a 5xx code: stop). |
| IMAP | Internet Message Access Protocol: how a mail app on a laptop or phone reads and syncs a mailbox kept on the server. |
What Makes It Hard
- The internet is hostile. Much of the mail that arrives at a big provider is spam, phishing or malware. Some of it is built to crash parsers: a zip file that unpacks to petabytes, or MIME nested a thousand levels deep.
- An accepted message may never be lost. Once our server says "250 OK" in SMTP, the sender deletes its copy. From that moment we are the only holder.
- Every mailbox must feel instant. Opening the inbox, the unread count, a thread, a search: all have to answer in tens of milliseconds, over a pile of mail that only grows.
- Sending is judged by others. Whether our mail reaches an inbox or a spam folder is decided by Gmail, Microsoft and Yahoo, based on our authentication and our reputation.
The Question the Whole Loop Answers
How do we accept, store, find and send mail reliably, for millions of mailboxes, when much of the internet is hostile?
The answer grows every round:
- Round 1: point MX at a managed receiver, store the raw message first and parse it later, keep a small metadata record per message apart from the bytes, and authenticate what we send with SPF, DKIM and DMARC.
- Round 2: run our own mail servers, store each attachment once, keep counters and threads exact, feed a search index from the database's change stream, parse hostile mail safely, and send to Gmail and Microsoft the way their rules require.
- Round 3: serve thousands of companies: tenant routing, retention and legal holds, anti-phishing, customer-held keys, EU residency, and importing ten years of old mail.
The notification loop covered the sending side of email as a customer of a provider (templates, sender reputation basics, bounces from SES). See Design a Notification System. Here we are the mail system: we receive from the internet, keep mailboxes, search them, and send as a first-class mail server.
Round 1 · Mid-level · "Hosted Email for One Company"
~35 min · SDE II (L5) · 1 region · 5,000 employees · 500K messages/day, 17.4/s at peak · 50 GB/day · inbox view < 500 ms · never lose accepted mail · 99.9%
R1.1 Establish Design Scope
The interviewer says: "Acme Corp has 5,000 employees and wants to stop paying for a hosted mail product. Build them email on their own domain, acme.example." Before drawing anything, we ask.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Their own domain? | Yes, acme.example, and they control its DNS. | We publish the MX record and the SPF, DKIM and DMARC records ourselves (steps 1.1 and 1.4). |
| Which clients? | A web app, plus desktop and phone mail apps over IMAP. | Two read paths: a REST API for the web app and an IMAP server, both over the same mailbox data. |
| Attachment limit? | 25 MB per message. | Base64 encoding makes that about 34 MB on the wire (25 × 4/3), which fits a managed receiver's 40 MB limit (step 1.1). |
| Spam? | Basic filtering is enough for now. | Authentication checks and a managed spam verdict; no custom models yet (step 1.5). |
| Retention? | Keep everything forever, for now. | Storage only grows. Round 3 brings retention policies. |
| Sending volume? | Normal business mail: people writing to people. No newsletters. | A managed sending service is enough; no dedicated IPs yet (step 1.4). |
Out of scope for this round: search at scale, conversation threading, many companies.
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Operation |
|---|---|
| "Customers and partners email us" | Receive mail from the internet for @acme.example |
| "Colleagues email each other" | Deliver internally without leaving our system |
| "My inbox" | List a folder, newest first, a page at a time |
| "Open that email" | Read one message: headers, body, attachments |
| "Write to a customer" | Send to the internet |
| "File it, delete it" | Move between folders, delete to Trash, mark read or unread |
| "My laptop's mail app" | Sync over IMAP: what's new, what changed since I last looked |
Not yet: search at scale, threading, many companies.
R1.3 Non-Functional Requirements: the Questions
We name each quality in words first; the numbers come in R1.7.
- Never lose an accepted message. SMTP makes this concrete: once a receiver answers
250, the sender forgets the message. So we answer250only when the message is safe on durable storage. This is the correctness bar of the round. - Deliverability of what we send. A message that lands in a customer's spam folder has failed, even though no error appears anywhere. Our domain must prove who it is.
- Fast mailbox views. Opening the inbox should take well under half a second, no matter how many years of mail the folder holds.
- Security. Mail carries contracts and passwords. Encrypt it at rest and in transit, and keep one employee from reading another's mailbox.
- Availability: 99.9%. Mail tolerates short outages better than most systems: senders retry for days (R1.9). 99.9% is about 44 minutes a month.
R1.4 The API
Three interfaces touch the mailbox: SMTP from the internet, a REST API for the web app, and IMAP for mail apps.
1. SMTP: how a message arrives. Another company's mail server looks up our MX record, connects on port 25 and talks like this (C: is the sender, S: is our server; a typical exchange, and the exact greeting and extension lines vary by server):
textS: 220 inbound-smtp.us-east-1.amazonaws.com ESMTP C: EHLO mail.partner.example S: 250-SIZE 41943040 S: 250 STARTTLS C: STARTTLS S: 220 Ready to start TLS (TLS handshake; the client sends EHLO again inside TLS) C: MAIL FROM:<bob@partner.example> SIZE=2641920 S: 250 2.1.0 Ok C: RCPT TO:<alice@acme.example> S: 250 2.1.5 Ok C: DATA S: 354 End data with <CR><LF>.<CR><LF> C: (headers and MIME body, about 2.6 MB) C: . S: 250 2.0.0 Ok: queued as 0100019a2c1e... C: QUIT S: 221 Bye
Three things to notice. The recipient is checked at RCPT TO, before any bytes arrive, so mail for an unknown user can be refused cheaply. The three-digit code carries the meaning (2xx done, 4xx try later, 5xx give up), and the 2.1.5 style enhanced status code (RFC 3463) adds detail. And the final 250 after the dot is the moment responsibility passes to us.
2. What the message looks like inside (MIME, shortened)
textFrom: Bob Lee <bob@partner.example> To: Alice Chen <alice@acme.example> Subject: Q3 pricing Date: Mon, 28 Sep 2026 09:14:02 +0000 Message-ID: <7f3a.1759050842@partner.example> MIME-Version: 1.0 Content-Type: multipart/mixed; boundary="b1" --b1 Content-Type: multipart/alternative; boundary="b2" --b2 Content-Type: text/plain; charset=utf-8 Hi Alice, the price sheet is attached. --b2 Content-Type: text/html; charset=utf-8 <p>Hi Alice, the price sheet is attached.</p> --b2-- --b1 Content-Type: application/pdf; name="q3.pdf" Content-Disposition: attachment; filename="q3.pdf" Content-Transfer-Encoding: base64 JVBERi0xLjcKJeLjz9MKMSAwIG9iago8PC9UeXBlL0NhdGFsb2cv... --b1--
A message is a tree. Here the top part is multipart/mixed with two children: a multipart/alternative (the same text as plain and HTML) and a PDF. Binary attachments travel as base64, which turns every 3 bytes into 4 characters: a 2 MB PDF becomes about 2.7 MB of text.
3. Web API
httpGET /v1/folders/INBOX/messages?limit=50&cursor=eyJ0cyI6MTc1OTA1MDg0Mn0 HTTP/1.1 Authorization: Bearer <session token>
json{ "folder": "INBOX", "unread": 14, "total": 1250, "messages": [ { "id": "01K6F3Q9M2ZC4T8RVKJ0W5H7NA", "from": { "name": "Bob Lee", "email": "bob@partner.example" }, "subject": "Q3 pricing", "snippet": "Hi Alice, the price sheet is attached.", "received_at": "2026-09-28T09:14:02Z", "unread": true, "has_attachments": true, "size": 2641920 } ], "next_cursor": "eyJ0cyI6MTc1OTA1MDExMH0" }
GET /v1/messages/{id}returns headers and body parts;GET /v1/messages/{id}/attachments/{n}streams one attachment.PATCH /v1/messages/{id}with{ "unread": false }or{ "folder": "Archive" }marks or moves it.POST /v1/messages/sendsends:
httpPOST /v1/messages/send HTTP/1.1 Authorization: Bearer <session token> Idempotency-Key: 5b0d2f8e-3c41-4c55-9a1e-7f0e2b6d9c10 Content-Type: application/json { "to": ["dan@customer.example"], "cc": [], "subject": "Re: Q3 pricing", "text": "Thanks, Dan. Details below.", "in_reply_to": "<a91c@customer.example>", "attachment_upload_ids": ["up_3xk9"] }
The answer is 202 Accepted with the new message's id and its Message-ID, because delivery happens later (step 1.6). The Idempotency-Key is made by the web app once per click of "Send" and reused on retries, so a retried request can't send twice.
4. IMAP. Mail apps speak IMAP4rev2 (RFC 9051) to our IMAP server on port 993 (TLS). IMAP needs three things from our data model, and we build them in from the start:
- UIDs: every message in a folder has a number that only grows (
UID 4812). An app asks "anything above UID 4812?" to find new mail. - UIDVALIDITY: a number per folder that changes only if UIDs had to be reassigned; if it changes, the app throws away its cache for that folder.
- MODSEQ: with the CONDSTORE and QRESYNC extensions (RFC 7162), every change (a new message, a flag, a delete) gets a higher modification sequence, so an app can ask "what changed since modseq 90,211?" instead of re-reading every flag.
And IDLE (RFC 2177) lets an app keep a connection open and be told when new mail arrives.
| Status | Meaning |
|---|---|
202 Accepted | Send queued; delivery status comes later on the Sent item |
400 Bad Request | A malformed address |
404 Not Found | No such message, or not in your mailbox: we never reveal which |
413 Payload Too Large | Attachments over 25 MB |
429 Too Many Requests | Over the per-user sending limit (step 1.6) |
Recap
- SMTP decides the contract:
250after the dot means the message is ours;4xxmeans "try later";5xxmeans "give up". - A message is a MIME tree; attachments are base64, a third bigger on the wire.
- The web API lists folders by cursor and sends with an idempotency key; IMAP needs UIDs, UIDVALIDITY and MODSEQ from the data model.
R1.5 Design Evolution: From One Mail Server to a Managed Pipeline
Each step is a problem, your turn to think, the answer, and what it costs us.
Step 1.0: The Baseline
One virtual machine runs an open-source mail server. It accepts SMTP on port 25, writes each message as a file on its disk (one directory per user, one file per message, the classic Maildir layout), serves IMAP from the same disk, and sends outbound mail straight to the internet.
It works for a demo. One disk holds the company's entire mail history, and one machine failure stops mail for everyone.
Step 1.1: How Does Mail Reach Us at All?
The problem: a customer at partner.example writes to alice@acme.example. Their mail server has never heard of us. And whatever answers must never say "OK" for a message it could still lose.
What would you do?
Step 1.2: Listing a Folder Reads Every Message File
The problem: Alice has 40,000 messages. Her inbox page shows 50 lines: sender, subject, date, a snippet, unread or not. In the baseline, building that page means opening and parsing message files. What would you do?
Step 1.3: Parsing Slows Down Accepting Mail
The problem: someone has to read each raw message and write its metadata: decode the headers, find the text for the snippet, list the attachments. A 25 MB message with an odd structure takes a while. And a burst of 2,000 newsletters at 9 a.m. arrives all at once. What would you do?
Primitive: Message Queues vs Event Streams · Drill: The Warehouse System That Fell a Day Behind. Its two questions are the two above: a consumer that throws halfway (the visibility timeout returns the message, retries are idempotent, the DLQ catches poison messages) and why a queue instead of a direct call (it decouples the SMTP accept rate from the parse rate and absorbs the burst).
Step 1.4: Our Mail Lands in Spam Elsewhere
The problem: Alice replies to a customer at Gmail. The reply lands in the customer's spam folder. Anyone on the internet can write From: alice@acme.example on a message, so Gmail has no reason to trust that this one really came from Acme.
What would you do?
Step 1.5: Spam and Fake Senders Reach Our Inboxes
The problem: Acme's staff start getting mail "from" their own CEO asking for gift cards, plus the usual flood of spam and the occasional infected attachment. What would you do?
Step 1.6: The Recipient's Server Is Down
The problem: Alice sends a proposal to dan@customer.example. Their mail server is being restarted and answers 421 4.3.2 Service not available. Separately, she mistypes an address, and the other side answers 550 5.1.1 User unknown.
What would you do?
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | One server, messages as files on its disk | One disk, one machine |
| 1.1 | How mail reaches us | MX → SES receiving → raw message in S3 | SES fees; less control over SMTP |
| 1.2 | Listing reads every file | ~1 KB metadata per message in DynamoDB; bytes in S3; counters in the same transaction | Two stores; bytes first |
| 1.3 | Parsing slows accepting | SNS → SQS → idempotent parser; DLQ | Mail appears a moment later |
| 1.4 | Our mail lands in spam | Custom MAIL FROM + SPF, Easy DKIM, DMARC rolled out slowly | DNS and rollout care |
| 1.5 | Spam and forgeries | SES verdicts, our rules in the parser, quarantine, no backscatter | False positives |
| 1.6 | Recipient server down | Send queue for SES errors; SES retries 4xx up to 14 h; events update the Sent item | Delivery state; rare duplicates |
R1.6 Architecture v1
Synthesizing vector architecture diagram...
Inbound: SES owns the SMTP conversation, stores the raw bytes, and announces them. Everything after that is our idempotent parser, driven by a queue.
Synthesizing vector architecture diagram...
Reading and sending: both clients read the same metadata and bytes; sending goes through a queue to SES, and the results come back as events.
Tables
The mail table (partition key mailbox_id, sort key sk):
sk | What it is | Main attributes |
|---|---|---|
HEAD | One per mailbox | modseq: the last change number used for this mailbox |
S#<folder> | One per folder | unread, total, uid_next, uidvalidity |
F#<folder>#<ulid> | The message, as listed | uid, from, to, subject, snippet, received_at, flags, size, raw_key, has_attachments, modseq |
M#<ulid> | Where a message is now | folder, uid; after a delete it stays as a tombstone (deleted: true) |
The changes table (partition key mailbox_id, sort key modseq, a Number so it sorts numerically): one item per change, {op: NEW | FLAGS | MOVE | DELETE, ids, folder}, with a TTL of 30 days (TTL deletes are eventual, typically within a few days; a reader never trusts an item's presence to mean it's recent, it compares modseq). IMAP's "what changed since modseq 90,211" is one query on this table.
The sends table (partition key idempotency_key): message_id, state, created_at.
The delivery transaction (one TransactWriteItems per new message, 5 items):
- Put
M#<ulid>if it doesn't exist (this makes a retried parse a no-op: the whole transaction is cancelled, even if the user has since moved the message or deleted it, because the locator outlives both). - Put
F#INBOX#<ulid>. - Update
S#INBOX:unread + 1,total + 1,uid_next + 1(the olduid_nextis the message's UID; the parser read it just before and makes the update conditional on it being unchanged, retrying on conflict). - Update
HEAD:modseq + 1, conditional on the value the parser read. - Put the
changesitem for thatmodseq.
Tracing an inbound message
Synthesizing vector architecture diagram...
Everything after SES's 250 is ours. Alice's app hears about the message within about 30 seconds this round, because the IMAP server polls; Round 2 pushes.
Tracing a send that hits a temporary failure (numbered steps)
- 10:02:00: Alice clicks Send. The API writes the MIME to S3, the Sent item (
delivery_status: QUEUED), thesendsrecord, and a message onsend-queue:202. - 10:02:01: the sender worker calls SES; SES accepts. Status
SENT_TO_SES. - 10:02:03:
customer.example's MX answers421 4.3.2. SES reports DeliveryDelay; the Sent item shows "Delayed, still trying." - SES retries on its own schedule. At 10:19 the server is back and accepts: a Delivery event; the Sent item shows "Delivered."
- Had the server stayed down for 14 hours, SES would have reported a Bounce, and Alice would find a non-delivery report in her inbox.
R1.7 Numbers
Targets
| Quality | Target | Why this number |
|---|---|---|
| Availability | 99.9% | 0.1% of a 30.4-day month is 43,776 min × 0.001 ≈ 44 minutes. Senders retry for days, so an outage delays inbound mail rather than losing it. |
| Inbox view | < 500 ms at the browser | The chain below is about 110 ms. |
| Durability | No accepted message lost | S3 is designed for 99.999999999% (11 nines) durability of objects (AWS's design figure); we add bytes-before-metadata and an idempotent parser. |
Traffic (assumptions: each employee receives 80 messages a day, 60 from outside and 20 from colleagues, and sends 20, half of them outside)
| Item | Math | Result |
|---|---|---|
| Mailbox entries a day | 5,000 × 80 received + 5,000 × 20 sent copies | 500,000/day |
| Average rate | 500,000 ÷ 86,400 s | 5.8/s |
| Peak | we assume 3× in the morning | 17.4/s |
| External inbound (through SES) | 5,000 × 60 | 300,000/day |
| External outbound | 5,000 × 10 messages × 2 recipients on average | 100,000 recipients/day (SES counts each recipient as an email) |
| Internal | 5,000 × 10 sent to colleagues | delivered by our API directly, no SMTP |
Storage (assumption: 100 KB per mailbox entry on average; each entry has its own raw object this round)
| Item | Math | Result |
|---|---|---|
| Bytes | 500,000 × 100 KB | 50 GB/day, 18.25 TB/year |
| After 3 years | 18.25 × 3 | 54.75 TB |
| Metadata items | 500,000/day × (1 KB message item + about 100 B locator) | 182.5M messages and about 201 GB a year |
| Change log | about 1.25M changes/day (new mail plus 1.5 flag changes per message) × 200 B × 30 days | 7.5 GB |
Inbox view chain (dependent steps add): TLS to the ALB and routing 5 ms, session check 2 ms, one DynamoDB query of 50 items 15 ms, building the JSON 8 ms, the internet round trip to the browser 80 ms: about 110 ms.
IMAP polling, sized. 5,000 mailboxes × about 2 connected apps each = 10,000 connections. Each IMAP server polls HEAD for its connections every 30 s: 10,000 ÷ 30 ≈ 333 reads a second, eventually consistent, 0.5 read units each: 14.4M read units a day. Cheap now; Round 2 has 20 million active users, and this becomes push.
Monthly cost at the end of year 1 (us-east-1 list prices, 730 hours a month, ignoring free tiers; check the AWS Pricing Calculator before quoting):
| Item | Math | Monthly |
|---|---|---|
| SES receiving | 300,000 × 30.4 = 9.12M messages × $0.10/1,000 ≈ $912; 256 KB chunks, we assume 1.3 per message on average: 11.9M × $0.09/1,000 ≈ $1,067 | ≈ $1,979 |
| SES sending | 100,000 × 30.4 = 3.04M recipients × $0.10/1,000 ≈ $304; attachment data 3.04M × 79 KB ≈ 240 GB × $0.12 ≈ $29 | ≈ $333 |
| S3 | 18.25 TB × $0.023 ≈ $420; 15.2M PUTs × $0.005/1,000 ≈ $76; GETs ≈ $10 | ≈ $506 |
| DynamoDB (on-demand) | writes: 500,000 × (10 units for delivery + 1.5 flag changes × 8) = 11M a day × 30.4 × $0.625/M ≈ $209; reads, IMAP polling included, ≈ $85; storage 205 GB × $0.25 ≈ $51; point-in-time recovery 205 GB × $0.20 ≈ $41 | ≈ $386 |
| Compute | API 3 × (1 vCPU, 2 GB ARM Fargate ≈ $28.84) ≈ $87; IMAP 3 × (2 vCPU, 4 GB ≈ $57.67) ≈ $173; parser Lambda 15.2M × ($0.20/M + 0.2 s × 1 GB × $0.0000133334) ≈ $44; sender Lambda ≈ $5 | ≈ $309 |
| ALB and NLB | ALB $16 + about 5 LCUs ≈ $29; NLB $16 + ≈ $9 | ≈ $70 |
| SQS, KMS, Route 53, CloudWatch | estimates | ≈ $125 |
| Total | ≈ $3.7K/month |
About $0.74 per employee a month. SES is 62% of the bill ($2,312), and receiving alone is 53%. Remember that ratio; it decides a build-or-buy question in Round 2.
R1.8 Trade-Offs
SES vs our own MTA
| SES receiving and sending (chosen) | Our own MTAs on EC2 | |
|---|---|---|
| SMTP control | SES's checks; we filter after the conversation | Full: reject during the conversation, greylist, custom limits |
| Port 25 outbound | Not our problem | Blocked by default on EC2; request to lift it, set reverse DNS |
| Reputation | SES's shared IPs (dedicated ones are optional) | Ours to build from zero |
| Retries to recipients | SES, up to 14 hours | Ours; RFC 5321 suggests at least 4 to 5 days |
| Cost at 500K/day | About $2.3K a month | A few instances, plus the time to run them |
| Cost at 1B/day (Round 2) | About $4.6M a month to receive alone | Servers in the tens of thousands of dollars, plus a deliverability team |
Where metadata lives
| DynamoDB (chosen) | PostgreSQL | Files on disk (Maildir) | |
|---|---|---|---|
| Folder list | One query on a sorted key | An index scan; fine at this size | A directory listing plus header parsing |
| Counters and changes | Same TransactWriteItems | Same transaction | Separate files, easy to get wrong |
| Operations | None to patch; on-demand scaling | A writer to size and patch | Disks to replicate and back up |
| Next round | Scales out by partition | One writer becomes the ceiling | Doesn't scale |
A relational database would also work at 17 messages a second. We pick DynamoDB because the access pattern is simple (by mailbox, sorted by time) and Round 2's 35,000 a second doesn't force a migration.
Parsing synchronously vs asynchronously
| In the SMTP session | After 250, from a queue (chosen) | |
|---|---|---|
| Sender waits for | Our parser | One durable write |
| A parser crash | A failed delivery, then a retry and maybe a duplicate | A message reappears in the queue |
| A burst | Timeouts | A deeper queue |
| When the user sees mail | Immediately after 250 | A moment later |
R1.9 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| The parser falls behind (a bug, a burst, a Lambda concurrency limit) | Queue age grows; new mail appears late | Mail is still accepted, because SES and S3 don't depend on the parser. We alarm on the age of the oldest message (over 60 s), fix, and let the queue drain. Poison messages go to the DLQ after 5 attempts. |
| Our S3 bucket can't be written (a broken bucket policy or KMS key policy after a change) | SES can't complete the receipt rule's S3 action; SES already answered 250, so the messages are lost | Prevent it: every change to the bucket policy or the KMS key policy is gated behind a test delivery through SES in a staging rule before it's applied. Limit it: a canary message goes through the whole path every 5 minutes, and we alarm if it doesn't appear in a test mailbox. The canary bounds the loss to minutes of mail; it doesn't prevent it. With our own receivers in Round 2, a failed S3 write means answering 451 so the sender keeps the message and retries: the SMTP contract. |
| The API can't write to S3 on send | POST /send fails | 503; the web app retries with the same Idempotency-Key. Nothing was queued, so nothing is duplicated. |
| Our domain's reputation drops (a compromised account sends spam through us) | Bounce and complaint rates rise in SES; SES may put the account under review | A per-user limit (1,000 external recipients a day, 429 above it) caps the damage. SES publishes its own thresholds: reviews above about a 5% bounce rate or 0.1% complaint rate, and may pause sending above about 10% or 0.5%. We alarm at half those, suspend the user's sending, and reset the password. |
| A receiving region problem | SES in us-east-1 impaired | Senders get no answer and retry, for days. A second MX in another region is Round 3's topic. |
| A bad deploy corrupts metadata | Wrong folders or counts | Point-in-time recovery on the tables (35 days); raw bytes are untouched in S3, so any message can be re-parsed. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | 250 only once the message is durable; bytes before metadata; an idempotent parser behind a queue with a DLQ; SES's retries for temporary failures REL 4 · REL 5 · REL 9 |
| Security | SPF, DKIM and DMARC out; authentication and malware verdicts in; quarantine instead of backscatter; SSE-KMS at rest; TLS for IMAP and the web, STARTTLS for SMTP; every read and every attachment reference checked against the caller's own mailbox SEC 3 · SEC 8 · SEC 9 |
| Performance Efficiency | Folder views from small sorted metadata items, never from the bytes PERF 3 |
| Cost Optimization | About $3.7K a month; the managed services are the bill, which is right at this size COST 5 |
| Operational Excellence | Light this round: alarms on queue age, DLQ depth, the canary, and SES bounce and complaint rates OPS 8 |
| Sustainability | Skipped this round: serverless parsing and a few small tasks. |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Knows how mail finds a domain (MX with preferences) and what the SMTP reply codes mean.
- Says out loud that
250transfers responsibility, and stores the message durably before answering. - Separates small metadata from large immutable bytes, and writes bytes first.
- Parses asynchronously from a queue, idempotently, with a DLQ.
- Explains SPF, DKIM and DMARC as three different questions, and rolls DMARC out gradually.
- Separates temporary (
4xx) from permanent (5xx) failures, and knows who retries what.
Follow-up questions
-
"Why not store messages in a relational database with the body in a column?" Answer: a message averages about 100 KB and can be 25 MB. Bodies in rows bloat the buffer cache, backups and replication, while the thing we query (folder, date, flags) is about 1 KB. S3 stores bytes for about $0.023 per GB-month; database storage costs several times that and is sized for random access we don't need.
-
"A colleague sends to 200 people inside Acme. How many copies?" Answer: this round, 200 metadata items and 200 raw objects, one per mailbox, written by our API without SMTP. At 50 GB a day nobody minds. Round 2 stores an attachment once and points to it.
-
"How does Alice's phone learn about new mail without polling?" Answer: this round it partly does poll: the IMAP server checks the mailbox's
modseqevery 30 seconds and tells apps in IDLE. For phones that aren't connected, a push through APNs or FCM. Round 2 replaces polling with a notification sent right after the transaction.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Answer 250, then store it" | A crash in between loses a message the sender already deleted. |
| "Parse during the SMTP session" | The sender waits on our slowest code; bursts cause timeouts and duplicates. |
| "Send straight from EC2" | Port 25 is blocked by default, and the IP has no reputation or reverse DNS. |
| "Bounce the spam we accepted" | The envelope sender is forged: that's backscatter to innocent people. |
| "Retry every failure" | Retrying 5xx wastes reputation; 4xx and 5xx mean different things. |
"DMARC p=reject on day one" | Forgotten systems that send as the domain get rejected everywhere. |
Round 2 · Senior · "100M Mailboxes, a Billion Messages a Day"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 100M mailboxes, 20M active each day · 1B messages/day, 34.7K/s at peak · about 83 PB stored after 3 years · new mail visible < 2 s P95 · folder view < 40 ms P99 · search < 300 ms P95 on recent mail · 99.99%
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built email for one 5,000-person company: 500,000 messages a day, 17 a second at peak, 50 GB a day. Mail arrives through an MX record at SES receiving, which stores the raw message in S3 and announces it; an idempotent parser behind an SQS queue writes about 1 KB of metadata per message to DynamoDB, with the folder's counters, the next IMAP UID and the mailbox's change number in the same transaction. Bytes first, metadata second. The web app and an IMAP server read the same data; IMAP apps sync with UIDs and modseqs, and learn about new mail from a 30-second poll. Outbound mail goes through a queue to SES, signed with DKIM, with SPF on a custom MAIL FROM domain and DMARC rolled out slowly; SES retries temporary failures for up to 14 hours and reports bounces as events. Inbound, we act on SES's authentication and malware verdicts and quarantine forgeries instead of bouncing them. About $3.7K a month, two thirds of it SES. Open costs: SES's per-message price at scale, one raw copy per recipient, no search or threads, and polling."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: a managed receiver, raw bytes in S3, metadata in DynamoDB, and a managed sender.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | How mail reaches us | MX → SES receiving → S3 | SES fees |
| 1.2 | Listing reads every file | Metadata in DynamoDB, bytes in S3, counters in the same transaction | Two stores |
| 1.3 | Parsing slows accepting | Queue + idempotent parser + DLQ | A short delay |
| 1.4 | Our mail lands in spam | SPF, DKIM, DMARC | DNS care |
| 1.5 | Spam and forgeries | Verdicts, quarantine, no backscatter | False positives |
| 1.6 | Recipient down | Send queue; SES retries 4xx | Delivery state |
Open costs: SES pricing at scale; a copy per recipient; no search, threads or push.
R2.1 The Scope Raise
Interviewer: "Congratulations, we're now a consumer webmail service: 100 million mailboxes and a billion messages a day, with a big morning surge. Attachments up to 25 MB, and the same newsletter PDF goes to hundreds of thousands of people. Users want conversations grouped, and search over years of mail. Last month an attacker sent a zip bomb that took our parsers down. And Gmail has started deferring our outbound mail."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| What's the split between received and sent? | 70% received from the internet, 30% sent by our users. | 700M inbound a day, 8,102/s on average; 300M outbound, 3,472/s, 10,416/s at peak (R2.6). SES would cost millions a month; we run our own mail servers (R2.2). |
| How many users are active on a given day? | About 20M, each with 1 or 2 connected apps. | Push instead of polling; a recent-headers cache sized for 20M mailboxes (R2.6). |
| How much of the attachment volume is duplicated? | Measure it; plan for about 30% of attachment bytes being repeats. | Store each attachment once, by hash (step 2.1). |
| How should threads work? | Like Gmail: replies group with what they answer, even if the subject changes a little. | Threading by Message-ID, In-Reply-To and References (step 2.3). |
| Does 300 ms apply to ten-year-old mail? | P95 of 300 ms for searches. Most searches are about recent mail; older mail may take a couple of seconds, but it must be found. | A hot index for recent mail and a cheaper warm tier for the rest (step 2.4). |
| How exact must unread counts be? | The number next to "Inbox" must match what the user sees in the list, on every device. | Counters in the same transaction as every flag change, plus reconciliation (step 2.2). |
| What exactly happened with the zip bomb? | A 45 KB zip expanded without limit and killed every parser that touched it, one after another. | Streaming parsing with hard limits and isolated scanning (step 2.5). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Mailboxes | 5,000 | 100M, 20M active a day |
| Messages | 500K/day, 17.4/s peak | 1B/day, 11,574/s average, 34,722/s peak |
| Bytes | 50 GB/day | 100 TB/day at 100 KB per message |
| Receiving and sending | SES | Our own MTA fleets |
| Attachments | A copy per recipient | Stored once per hash, garbage-collected |
| Views | Folders | + threads, search, exact unread counts |
| Client updates | 30 s polling | Push |
| Targets | 99.9%; inbox < 500 ms | Visible < 2 s P95; folder view < 40 ms P99 at the server; search < 300 ms P95 on recent mail; 99.99% |
The "Not yet" list from R1.2 comes back: search at scale and threading are in scope now. Many companies wait for Round 3.
A note on the availability target: we promise 99.99% (4.4 minutes a month), not "five nines". 99.999% allows 5.26 minutes a year.
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | How it fails at the new scope |
|---|---|
| SES receiving | 700M messages a day × 30.4 days = 21.3B a month. At $0.10 per 1,000 plus about 1.3 chunks at $0.09 per 1,000, that's about $4.6M a month just to receive, and we still can't reject during the SMTP conversation. |
| SES sending | 300M recipients a day is about $0.9M a month, before attachment data, and our users' reputation is mixed with everyone's. |
| One raw object per recipient | A newsletter PDF sent to 300,000 of our users is stored 300,000 times. |
| Counters kept right by "every path is a transaction" | True for single changes. "Mark all 10,000 as read" and repair tools are paths that touch thousands of items; one missed path and the badge drifts forever. |
| No threads, no search | Nothing to find mail with except scrolling. Searching 100 TB a day of new mail with a database LIKE is impossible. |
| A Lambda parser that buffers the message | One hostile message crashes every worker that receives it, and it keeps coming back until the DLQ. |
| IMAP polling every 30 s | 20M active users × 1.5 apps = 30M app connections, about half of them open at peak: 15M ÷ 30 s = 500,000 HEAD reads a second to learn mostly nothing. |
| SES's shared retry schedule | When Gmail defers, we can't control how and when we retry, which IP retries, or how fast we push to each provider. |
The order we fix it in: attachments first (2.1), because they decide how bytes are stored; then counters (2.2) and threads (2.3), which shape the metadata; then search (2.4), which reads everything before it; then the two hostile edges: inbound content (2.5) and outbound reputation (2.6).
R2.3 New Requirements and API Additions
Threads
httpGET /v1/threads/th_01K6F3Q9?limit=100 HTTP/1.1 Authorization: Bearer <session token>
json{ "thread_id": "th_01K6F3Q9", "subject": "Q3 pricing", "message_count": 4, "unread": 1, "messages": [ { "id": "01K6F3Q9M2ZC...", "from": "bob@partner.example", "unread": false }, { "id": "01K6G1B7Y4TA...", "from": "alice@ourmail.example", "unread": false } ] }
The folder list now returns one row per thread with its newest message, count and unread count.
Search
httpGET /v1/search?q=from:acme%20invoice%20has:attachment&scope=all&limit=20 HTTP/1.1 Authorization: Bearer <session token>
json{ "results": [ { "id": "01JZ8...", "thread_id": "th_01JZ8...", "subject": "Invoice 4471", "received_at": "2025-03-02T10:11:00Z", "highlight": "...your <em>invoice</em> for March..." } ], "searched": "recent", "older_pending": true, "next": "c2VhcmNoOjI" }
searched: "recent" with older_pending: true means the first page came from the hot index (the last 90 days) and the older tiers are still being searched; the client asks for next to get them (step 2.4).
Attachment download by signed URL
GET /v1/messages/{id}/attachments/2 checks that the message is in the caller's mailbox and that the attachment's malware verdict is clean, then answers 302 Found with a pre-signed S3 GET URL for the attachment's object, valid for 5 minutes, with response-content-disposition set to the attachment's original file name. The bytes never pass through our API servers. A verdict still pending gets 409 SCAN_PENDING and the client retries in a second.
Delivery status of a sent message
json{ "id": "01K6G1B7Y4TA...", "recipients": [ { "address": "dan@customer.example", "status": "DEFERRED", "attempts": 3, "last_reply": "451 4.7.1 Greylisted, try again later", "next_attempt_at": "2026-09-28T10:52:00Z" }, { "address": "eve@other.example", "status": "DELIVERED", "at": "2026-09-28T10:02:04Z" } ] }
Sync with change tokens (JMAP-style)
Every change to a mailbox already gets the next modseq (Round 1). We expose it to web and mobile clients the way JMAP (RFC 8620 and RFC 8621) does: the mailbox's current modseq is its state string, and a client with an old state asks what changed.
httpGET /v1/changes?since=90211&max=500 HTTP/1.1
json{ "old_state": "90211", "new_state": "90219", "has_more": false, "created": ["01K6H..."], "updated": ["01K6F3Q9M2ZC..."], "destroyed": ["01K5Z..."] }
If since is older than the 30 days of change log we keep, the answer is the JMAP error cannotCalculateChanges, and the client does a full resync. IMAP clients get the same data through CONDSTORE/QRESYNC: HIGHESTMODSEQ, CHANGEDSINCE and VANISHED.
Push. Web and mobile clients keep one connection open to our gateways (WebSocket or EventSource) and receive a tiny frame, {"state":"90219"}, when their mailbox changes; they then call /v1/changes. Phones that aren't connected get an APNs or FCM push. IMAP apps get EXISTS and FETCH updates in IDLE; RFC 2177 has clients re-issue IDLE at least every 29 minutes so servers don't time them out.
Recap
- Threads and search are new read APIs; search says which tiers it has covered.
- Attachments download by 5-minute signed URLs, only after a clean verdict.
- Clients sync by a change token (the mailbox
modseq), pushed as a tiny "something changed" frame.
R2.4 Design Evolution: Hash-Addressed Attachments, Exact Counts, Threads and Search
Step 2.1: The Same 10 MB Attachment, Stored 300,000 Times
The problem: a store sends its 10 MB catalogue PDF to 300,000 of our users. Round 1 keeps one raw copy per recipient: 3 TB for one newsletter. Across all mail, we expect about 30% of attachment bytes to be repeats. What would you do?
Step 2.2: Unread Counts Drift, or Need Scans
The problem: the badge says Inbox (3), and the list shows no unread mail. Users notice immediately. Meanwhile a user clicks "mark all as read" on 10,000 messages from their phone while their laptop is marking a few as unread. What would you do?
Primitive: Distributed Cache Patterns and Eviction for the cached counters in R2.5.
Step 2.3: Group Replies into Conversations
The problem: Bob writes "Q3 pricing". Alice replies, Dan replies to Alice, someone forwards it with "Fwd:", and Bob's answer to Dan arrives before Dan's message does (mail isn't delivered in order). Meanwhile a hundred unrelated messages are titled "Hello" or "Invoice". What would you do?
Step 2.4: Search Years of Mail in 300 ms
The problem: "the invoice from Acme last spring." A billion new messages a day, three years kept. A first estimate indexes 10 KB per message with one replica: 20 TB a day, 1.8 PB for just 90 days, and still nothing older. What would you do?
Primitive: Trie Data Structure and Inverted Index · Change Data Capture and the Outbox Pattern · Drill: The Dual-Write That Broke Search Consistency. Its first question is the dual write above (the database committed, the publish failed, the index drifted forever); the stream fixes it because the change and its record are the same write. Its second asks why tail a log instead of polling an outbox table every 500 ms: DynamoDB has no single, ordered outbox to poll unless every change also goes into one global index partition, which would become the hottest key in the system; polling also adds up to 500 ms of delay and costs a read each time, mostly finding nothing. The stream's records are read by Lambda at no charge and arrive in order per item.
Step 2.5: A Zip Bomb Crashed the Parsers
The problem: a 45 KB zip attachment unpacks, layer by layer, into petabytes. Another message nests multipart/mixed a thousand levels deep. A third has a single 400 MB header line. Each one crashed every parser that picked it up, and SQS kept handing it to the next one.
What would you do?
Step 2.6: Gmail Is Deferring Our Mail
The problem: our outbound mail to Gmail starts coming back with 421 4.7.28 and 451 4.7.1 replies, and more of it is landing in recipients' spam. At the same time, a batch of new accounts created yesterday is sending thousands of messages each.
What would you do?
Primitive: Distributed Rate Limiting · Drill: The Partner Whose Retry Loop Took Down Everyone Else. Its two questions are the ones in point 2: local counters on many instances let through many times the limit, so the buckets are shared in Valkey; and a token bucket beats a fixed window because a window allows a double burst at its boundary.
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | The same attachment stored 300,000 times | SHA-256-addressed attachments; reference items in the message transaction; ownership checks; doom, resurrect, delete by generation; scan once per file | The collector; a dedup scope |
| 2.2 | Unread counts drift | Conditional counter updates in every transaction; batched "mark all" up to a mark taken at the click; a sparse unread index; repair only if modseq is unchanged | Index writes; a job |
| 2.3 | Conversations | Header threading with an ID map and placeholders; narrow subject fallback; 100-message cap; loop-safe tree building | ID-map items; a hot thread item |
| 2.4 | Search in 300 ms | OpenSearch fed from DynamoDB Streams, each record re-read from DynamoDB and versioned by modseq; a repair queue for poison records; routed by mailbox; 2 KB documents; hot 90 to 120 days + UltraWarm | About $0.50M a month; seconds of index lag |
| 2.5 | Zip bombs and nested MIME | Streaming parser with hard limits; isolated scanners; downloads wait for the verdict | Some odd mail partly shown |
| 2.6 | Gmail defers our mail | Own MTAs on EC2; 5-day retry schedule; greylist-friendly retries; shared token buckets per destination; IP pools; warm-up; providers' rules | A deliverability team; slow warm-ups |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
The inbound path: our MTAs answer 250 only after the spool write and the queue entry; everything after is asynchronous, and the search index follows the database's own stream.
Synthesizing vector architecture diagram...
The read path: metadata from DynamoDB with a cache checked against modseq, bodies by range from packs, attachments straight from S3.
Synthesizing vector architecture diagram...
The outbound path: the MIME is written once to the spool; the MTAs pace themselves per destination, park 4xx failures in minute buckets, and send failure notices back through the normal inbound pipeline.
The inbound MTAs, in order of what they check. Each check sits as early as it can, so the cheap refusals happen before the expensive work.
- At connect, before TLS. Per-IP connection limits, and a DNS blocklist lookup on the connecting IP (cached). A listed IP gets
554and the connection closes before any TLS handshake: the handshake is the most expensive thing a spammer can make us do, so we shed load in front of it. - At
MAIL FROM. SPF for the envelope domain. - At
RCPT TO. Is the recipient one of our 100M mailboxes? A cached directory answers; unknown users get550 5.1.1. Over-quota mailboxes get452 4.2.2, which tells the sender to retry later. - At the end of
DATA, before answering. DKIM and DMARC. A message that fails DMARC for a domain publishingp=rejectgets550 5.7.1: rejected inside the conversation, so the sender's own server tells its user, and we send no backscatter. A message over 35 MiB was already refused duringDATA. - Store, then answer. Each MTA collects the messages that finish in a 50 ms window and writes them as one spool object (a group commit), then sends one SQS entry per message, pointing at its byte range. Only then does each sender get
250. If the S3 write fails, every message in that window gets451 4.3.0; if a queue entry fails, that message gets451and its bytes in the spool are simply never parsed (the sender's retry brings a fresh copy, and the spool expires after 7 days).
We also publish an MTA-STS policy (RFC 8461) so that senders who support it refuse to deliver to us without valid TLS, and our outbound MTAs honor other domains' MTA-STS policies.
Duplicates the internet sends us. If our 250 is lost on the network, the sender retries and we accept the message twice. The parser writes a dedup item D#<SHA-256 of the body>#<Message-ID> in the mailbox, only if it doesn't exist, in the same transaction as the message; the second copy's transaction fails, and the parser just re-sends the push for the first. The dedup items live 7 days (a TTL, deleted eventually), which covers any realistic retry by a sender.
Tables (DynamoDB, provisioned with auto scaling; storage-heavy tables in the Standard-IA table class)
| Table | Key | Main attributes | Class |
|---|---|---|---|
messages | mailbox_id + sk | F#<folder>#<ulid>: headers summary, flags, uid, modseq, thread_id, body pointer (pack key, offset, length), attachment hashes; M#<ulid>: folder, kept as a tombstone after a delete; I#<hash of Message-ID>: thread_id; D#...: dedup, TTL 7 days. A sparse GSI on unread_key, keys only (step 2.2) | Standard-IA |
state | mailbox_id + sk | HEAD: modseq; S#<folder>: unread, total, uid_next, uidvalidity; T#<thread_id>: count, unread, newest date, participants | Standard |
changes | mailbox_id + modseq (Number) | op, IDs; TTL 30 days | Standard |
refs | sha256 + <mailbox>#<ulid or draft> | nothing else: the item is the reference | Standard-IA |
blobs | sha256 | state, gen, size, created_at, doomed_at, verdict | Standard-IA |
packs | pack_key | state (LIVE, DOOMED, DELETED), doomed_at, created_at, key of its manifest object | Standard-IA |
deferred | due#<minute>#<shard> + delivery_id | attempt; also holds the collector's "check this new file in a day" items | Standard |
deliveries | message_id + recipient_domain | state, attempt, next_at, lease_owner, lease_until, last reply | Standard |
The delivery transaction for one new message (one TransactWriteItems, 7 items plus two per attachment):
- Put the locator
M#<ulid>if it doesn't exist (it outlives moves and deletes as a tombstone, so a redelivered queue message can never bring a message back), and putF#<folder>#<ulid>. - Put the dedup item
D#...if it doesn't exist. - For each attachment: put its
refsitem, and a condition check that itsblobsitem isLIVE. - Update
S#<folder>:unread + 1,total + 1,uid_next + 1, conditional on theuid_nextthe parser read. - Update
HEAD:modseq + 1, conditional on the value read; put thechangesitem. - Update
T#<thread>: count, unread, newest date.
Before it, the parser claims the ID-map items (step 2.3). After it: the notify router, and nothing else. The search index learns from the stream.
The body packs. Parsers don't write one object per message body: each parser task collects the bodies it produced in the last 500 ms into one pack object (about 660 KB on average), writes it, and only then commits the metadata that points into it (bytes first). Reading a body is an S3 ranged GET of offset to offset + length. Section R2.6 shows why: at this scale, the number of objects costs more than their bytes.
Compacting packs safely. Deleted mail leaves dead bytes in packs, so a monthly job rewrites packs whose live fraction has fallen under 70%. A pack's bytes are referenced from many messages, so the job follows the same rules as the attachment collector (step 2.1):
- Every pack has a manifest: the
(mailbox, ulid)of each body in it, written with the pack. - The compactor resolves each manifest entry through its
M#locator to the currentF#item. Only entries whose item still exists and still points into this pack are live; everything else is dead bytes. - It writes the live bodies into a new pack (registered
LIVEinpacks), then moves each pointer in a transaction: update theF#item's pointer only if it still points at the old pack, with a condition check that the new pack isLIVE. - It sets the old pack
DOOMED(only ifLIVE). Anything that needs the old pack again (a point-in-time restore of the tables, for example) sets it back toLIVEfirst. - Just before the physical delete, it resolves the whole manifest again. If any item still points at the old pack, it returns the pack to
LIVE. If none does, it setsDELETED(only if stillDOOMEDwith thedoomed_atit wrote), and deletes the object. - That delete happens 35 days after doom, the same window as point-in-time recovery, so a restored table never points at a pack we've already removed. About 30% of packs are rewritten once, so retired packs waiting out their 35 days add about 210 TB (30% × 20 TB a day × 35 days), counted in R2.6.
The recent-headers cache. Valkey keeps the first page (50 rows) of each active folder and the folder counters, each stamped with the mailbox modseq it was built from. The API reads HEAD.modseq from DynamoDB (one small, strongly consistent read) and uses the cached page only if the stamps match; otherwise it queries DynamoDB and refills. A stale cache entry can never be served, and a slow refill that writes an older page is simply ignored by the next reader.
Pushing new mail. After the transaction, the parser calls the notify router with (mailbox, modseq). The router looks up which gateway tasks hold that mailbox's connections in a Valkey registry and sends each a frame; it also asks SNS mobile push to wake phones without a connection. Each gateway refreshes its registry entries every 5 minutes (a 12-minute TTL): 15M connections ÷ 300 s = 50,000 registry writes a second, sized into the Valkey cluster. A parser that retries after a crash re-sends the push; clients ignore a state they've already seen, and phone notifications carry the message's ULID as their collapse ID (apns-collapse-id on Apple, collapse_key on Android), so a re-sent push replaces the first instead of showing twice.
Outbound deliveries hold a lease. An MTA worker claims a delivery by setting lease_owner and lease_until = now + 15 min only if the lease is free or lease_until < now − 30 s, where now is the worker's own clock and the 30 s absorbs clock differences between workers (DynamoDB conditions have no server time). A delivery to a slow server can take minutes (RFC 5321 lets a client wait up to 10 minutes for the reply after the final dot), so the worker extends the lease every 5 minutes; at 10,416 deliveries a second lasting about 2 seconds, about 21,000 are in flight at peak, and only the rare slow ones ever extend. If the worker times out after sending the final dot but before reading the reply, it can't know whether the message was delivered; retrying may deliver it twice. We retry: a duplicate beats a lost message.
The send API's idempotency is the durable item in sends from Round 1 (conditional on the key not existing, kept 7 days), not SQS: a FIFO queue's deduplication window is only 5 minutes. A retried request that finds its record in state ACCEPTED (the first attempt crashed before enqueueing) enqueues it now; a delivery worker's attempt claim makes an extra enqueue harmless.
Tracing an inbound message, end to end
Synthesizing vector architecture diagram...
About 1.2 seconds from the final dot to the phone showing the message (R2.6), most of it the body-pack window. The scanner's verdict on 9c1f arrives separately and unlocks the download.
Tracing a search
- Alice types
invoice acme. The API sends one query to the hot indices for the last 90 days (4 monthly indices), withroutingset to her mailbox: 4 shards, one per index, searched in parallel. - The top 20 message IDs come back in about 120 ms. The API reads their metadata with one
BatchGetItemand their text by rangedGETs, in parallel, and builds highlights. - It answers with
searched: "recent"andolder_pending: true, about 170 ms after the request arrived. - Alice scrolls. The next request searches the 33 older monthly indices in UltraWarm, one shard each; segments not in the warm cache come from S3, so this page takes about 1 to 2 seconds.
Tracing a greylisted send (numbered steps)
- 10:02:00: Alice sends to
dan@customer.example. The API writes the MIME to the spool, the Sent item, thesendsrecord and adeliveriesitem, and enqueues it:202. - 10:02:01: an MTA claims the delivery (attempt 1), takes a token from the
customer.examplebucket for the established pool, connects from198.51.100.23, and gets451 4.7.1 Greylisted, try again later. - It writes the
deferreditemdue#10:07#17(5 minutes, with jitter) and sets the delivery toDEFERRED, attempt 2. - 10:07: the scheduler reads bucket 10:07 and enqueues it. An MTA claims attempt 2 and connects from
198.51.100.41, another IP in the same /24 pool, which the greylister tracks as one sender.250. The Sent item shows "Delivered".
Losing an AZ. MTAs, parsers, scanners and gateways run in three AZs; DynamoDB, S3, SQS and Lambda are regional. The NLB stops sending connections to the lost AZ's MTAs; senders whose connections broke retry. Sections R2.6 and R2.8 size the survivors.
R2.6 Numbers and Cost
Traffic
| Item | Math | Result |
|---|---|---|
| Messages | given | 1B/day |
| Average rate | 10⁹ ÷ 86,400 s | 11,574/s |
| Peak | we assume 3× in the morning surge | 34,722/s |
| Inbound (70%) | 700M ÷ 86,400 | 8,102/s, 24,306/s at peak |
| Outbound (30%) | 300M ÷ 86,400 | 3,472/s, 10,416/s at peak |
| Folder views | 20M active users × 40 a day | 800M/day |
Message size and bandwidth
| Item | Math | Result |
|---|---|---|
| Average message | 1 KB metadata + 20 KB body + 79 KB attachments | 100 KB |
| Raw bytes | 10⁹ × 100 KB | 100 TB/day, 36.5 PB/year |
| Bandwidth at 100 KB | 11,574/s × 100 KB = 1.157 GB/s | 9.26 Gbps average, 27.8 Gbps at peak |
| On the wire, inbound | base64 makes attachments a third larger: 20 + 79 × 4/3 ≈ 125 KB; 8,102/s × 125 KB | 8.1 Gbps average, 24.3 Gbps at peak |
Storage after 3 years
| Item | Math | Result |
|---|---|---|
| Without dedup, all bytes kept | 10⁹ × 99 KB × 1,095 days | 108.4 PB |
| Attachments after dedup (30% repeats, an assumption) | 79 KB × 0.7 = 55.3 KB per message × 10⁹ × 1,095 | 60.6 PB |
| Bodies | 20 KB × 10⁹ × 1,095 | 21.9 PB |
| Spool | 10⁹ × 125 KB × 7 days | 0.875 PB |
| Stored | about 83.3 PB of live data, plus about 0.2 PB of retired packs waiting out their 35 days | |
| Metadata | 10⁹ × 1 KB | 1 TB/day, 1.1 PB of message items after 3 years |
Where it sits (lifecycle rules by age):
| Class | What | Size |
|---|---|---|
| S3 Standard | Spool (875 TB), body packs under 30 days (600 TB), attachments under 30 days (1,659 TB) | 3,134 TB |
| S3 Standard-IA | Body packs 30 to 180 days (3,000 TB), retired packs (210 TB) | 3,210 TB |
| S3 Glacier Instant Retrieval | Body packs over 180 days (18,300 TB), attachments over 30 days (58,900 TB) | 77,200 TB |
Two rules to know. Lifecycle transitions skip objects smaller than 128 KB by default (since September 2024), and Standard-IA and Glacier Instant Retrieval bill any object as at least 128 KB. That's why bodies go into packs: a 20 KB body alone would stay in Standard forever. Attachment objects average about 263 KB (55.3 KB per message ÷ 0.21 new attachment objects per message), so most of them move; the ones under 128 KB stay in Standard, which makes our transition count below an upper bound.
Search index
| Item | Math | Result |
|---|---|---|
| Per message | about 4 KB of text, indexed without _source (an assumption to measure) | 2 KB |
| Per day | 10⁹ × 2 KB | 2 TB |
| Hot, up to 120 days with 1 replica | monthly indices leave when their newest mail is 90 days old: 2 TB × 120 × 2 | 480 TB |
| Hot nodes | i4g.8xlarge.search, 7.5 TB of NVMe each, filled to at most 75%: 480 ÷ 5.6 = 85.7; per AZ 28.6, rounded up to 29 per AZ | 87 nodes, 5.5 TB each (74%) |
| Shards | a monthly index of 60 TB at about 40 GB per shard | 1,500 primaries per month; about 12,000 hot shards over 87 nodes |
| Warm, one copy | 2 TB × 1,005 days, at most | 2,010 TB |
| Warm nodes | ultrawarm1.large.search, 20 TiB (22 TB) each: 2,010 ÷ 22 = 91.4; 34 per AZ to stay at about 90% full | 102 nodes, 90% full |
The default quota is 80 data nodes per domain; 87 needs an increase (the limit for this instance family is 400 on OpenSearch 2.17 and later). The 102 UltraWarm nodes fit the default of 150. Weekly indices would keep the hot tier close to 90 days and within 80 nodes, at the price of about four times as many indices for an old-mail search to touch; we keep monthly indices and raise the quota.
Metadata writes (transactions cost 2 write units per item)
| Operation | Items | Write units |
|---|---|---|
New message: F#, M#, D#, S#, HEAD, change, T# | 7 × 2 | 14 |
Attachment references and their blobs condition checks | 0.3 × (2 + 2) | 1.2 |
| ID-map items, outside the transaction | about 1.5 | 1.5 |
Flag changes and moves: about 1.5 per message × (message, S#, HEAD, change) | 1.5 × 8 | 12 |
| Unread index updates | about 1 | 1 |
| Per message | about 29.7 |
29.7 × 10⁹ = 29.7B write units a day = 343,750 a second on average, about 1.03M at peak. That's far above the default 40,000 per table, so the quotas are raised (they're adjustable) and the tables pre-warmed. A single DynamoDB partition serves up to 1,000 write units a second; a mailbox receiving more than about 60 messages a second (a mailing-list archive under attack, say) gets 451 at RCPT TO until it slows.
Fleets, with per-AZ rounding
| Fleet | Sizing | Per AZ | Total |
|---|---|---|---|
Inbound MTAs, c7g.xlarge | We assume 400 messages a second each (to load-test). Peak 24,306 ÷ 400 = 60.8; sized so two AZs carry the peak: 30.4 → 31 | 31 | 93 |
Outbound MTAs, c7g.xlarge | We assume 250 a second each (DKIM signing and slower remote servers). 10,416 ÷ 250 = 41.7; two AZs: 20.8 → 21 | 21 | 63 |
| Parsers, 4 vCPU | 34,722/s × 20 ms = 695 vCPUs; 173.75 tasks → 58 per AZ | 58 | 174 |
| Scanners, 4 vCPU | new attachments 34,722 × 0.21 = 7,292/s × 100 ms = 729 vCPUs → 61 per AZ | 61 | 183 |
| Spam scoring, 4 vCPU | 34,722 × 5 ms = 174 vCPUs → 15 per AZ | 15 | 45 |
| Gateways and IMAP, 4 vCPU / 16 GB | 20M active users × 1.5 apps = 30M app connections, about half open at peak = 15M, ÷ 50,000 each (an assumption to load-test) | 100 | 300 |
| Sending IPs | We assume about 1M messages a day per warmed IP to the large providers: 300M ÷ 1M | 300 |
The parsers are sized for peak across three AZs, not for surviving an AZ at peak. Losing one leaves 116 tasks: 464 vCPUs, 23,200 messages a second. At peak the queue then grows by about 11,500 a second until replacement tasks start (a few minutes), so new mail runs a few minutes late while none is refused. The MTAs, which decide whether mail is refused, are the fleets sized for two AZs.
Cache sizes. Recent headers: 20M active mailboxes × 50 rows × 1 KB = 1 TB. Folder counters for every mailbox: 100M × 4 folders × 16 bytes = 6.4 GB.
Latency budgets (P95 unless noted; dependent steps add, parallel steps count once as their maximum)
| New mail visible, from the final dot | ms |
|---|---|
| Group-commit window | 50 |
Spool PUT | 80 |
| Queue entry | 15 |
| Parser receive | 50 |
Ranged GET of the raw message | 40 |
| Parse | 100 |
| In parallel: attachment upload 80, spam score 10, ID-map read 10 | 80 |
| Body-pack window | 500 |
Body-pack PUT | 60 |
TransactWriteItems | 25 |
| Router, gateway, frame to the phone | 100 |
| Client asks for changes and the list | 60 |
| Total | 1,160, under 2,000 |
| Folder view, P99, at the server | ms |
|---|---|
| ALB and WAF | 3 |
| Session check | 1 |
HEAD read, strongly consistent | 8 |
| Cache hit 2, or on a miss a 50-item query 20 | 20 |
| Build the response | 3 |
| Total, worst case | 35, under 40 |
| Search, recent mail, at the server | ms |
|---|---|
| ALB, WAF, session | 4 |
| OpenSearch, 4 shards in parallel | 120 |
In parallel: metadata BatchGetItem 15, text for highlights 40 | 40 |
| Highlights and response | 5 |
| Total | 169, under 300 |
A message becomes searchable about 1.5 to 2.5 seconds after it's visible: the stream's own delay (AWS publishes no bound; we measure it), up to 250 ms for Lambda's poll, a 40 ms GET, a 100 ms bulk request, and up to 1 s for OpenSearch's refresh.
Availability. 99.99% of a 30.4-day month is 43,776 min × 0.0001 ≈ 4.4 minutes.
Monthly cost at the 3-year point (us-east-1 list prices, decimal GB, 730 hours; at this size real contracts are negotiated, so read these as proportions):
| Item | Math | Monthly |
|---|---|---|
| S3 storage | Standard 3,134 TB: 50 × $23 + 450 × $22 + 2,634 × $21 per TB ≈ $66.4K; Standard-IA 3,210 TB (including retired packs) × $12.5 ≈ $40.1K; Glacier Instant Retrieval 77,200 TB × $4 ≈ $308.8K | ≈ $415.3K |
| (for comparison) 108.4 PB in Standard, no dedup, no tiering | 50 × $23 + 450 × $22 + 107,900 × $21 | (≈ $2.28M) |
| S3 requests | spool PUTs: up to 20 a second from each of 93 MTAs and 100 API tasks (one per 50 ms window) = 3,860 a second × 86,400 × 30.4 = 10.14B × $0.005/1,000 ≈ $50.7K; body packs 914M × $0.005/1,000 ≈ $4.6K; new attachments 6.38B × $0.005/1,000 ≈ $31.9K; attachments into Glacier IR 6.38B × $0.02/1,000 ≈ $127.7K (an upper bound: attachments under 128 KB aren't transitioned); packs into Standard-IA and later Glacier IR 914M × ($0.01 + $0.02)/1,000 ≈ $27.4K; GETs (parser 30.4B, outbound 9.1B, indexer 30.4B, reads 18.2B, at $0.0004/1,000) ≈ $35.3K; old-mail reads from Glacier IR 1.2B × $0.01/1,000 ≈ $12.0K + 24 TB × $0.03/GB ≈ $0.7K = $12.7K; attachment downloads, scans, compaction ≈ $4.5K | ≈ $294.8K |
| DynamoDB | writes: 343,750 average, provisioned at 70% target utilization = 491,100 units; 43% in Standard-IA tables at $0.00081/h ≈ $124.9K and 57% in Standard at $0.00065/h ≈ $132.8K; reads ≈ $12K; storage 1,487 TB in Standard-IA × $0.10 ≈ $148.7K and 146 TB in Standard × $0.25 ≈ $36.5K; point-in-time recovery 1,633 TB × $0.20 ≈ $326.6K; small tables (packs, deferred, deliveries) ≈ $5K | ≈ $786.5K |
| OpenSearch | hot 87 × $3.954/h × 730 ≈ $251.1K; UltraWarm 102 × $2.68/h × 730 ≈ $199.6K; warm managed storage 2,010 TB × $0.024/GB ≈ $48.2K; dedicated masters ≈ $1K | ≈ $499.9K |
| Compute | inbound MTAs 93 × $0.145/h × 730 ≈ $9.8K; outbound MTAs 63 ≈ $6.7K; 300 public IPv4 addresses × $3.65 ≈ $1.1K; Fargate ARM at $0.03238 per vCPU-hour and $0.00356 per GB-hour: parsers 174 × $115.34 ≈ $20.1K, scanners 183 ≈ $21.1K, spam 45 ≈ $5.2K, gateways and IMAP 300 × $136.13 ≈ $40.8K, API 100 × $57.67 ≈ $5.8K; indexer Lambda ≈ $3.2K; collector, reconciler, compaction, scheduler ≈ $5K | ≈ $118.8K |
| Valkey | headers cache: r7g.4xlarge nodes hold 105.8 GiB, about 79 GiB usable; 1 TB (931 GiB) ÷ 79 = 11.8, so 12 shards would leave no headroom; 14 shards (about 1,100 GiB) × 2 nodes ≈ $1.40/h each × 730 ≈ $28.6K; registry, token buckets, counters ≈ $1.5K | ≈ $30.1K |
| Internet egress | outbound SMTP 300M × 125 KB = 1,140 TB, plus reads and downloads 685 TB = 1,825 TB: 10 TB × $0.09 + 40 × $0.085 + 100 × $0.07 + 1,675 × $0.05 per GB | ≈ $95.1K |
| Cross-AZ | cache pages 800M × 50 KB = 40 TB/day, two thirds cross an AZ, × $0.02/GB ($0.01 each way) ≈ $16.2K; indexer to OpenSearch and parser to Valkey ≈ $2.0K | ≈ $18.2K |
| ALB, NLB, WAF | WAF 3B requests a day × 30.4 × $0.60/M ≈ $54.7K; ALB ≈ $10K; NLBs ≈ $5K | ≈ $70.0K |
| KMS, SNS mobile push, CloudWatch and logs | Bucket Keys keep KMS near $1K; 10.6B pushes × $0.50/M ≈ $5.3K; logs and metrics ≈ $50K | ≈ $56.3K |
| Total | ≈ $2.39M/month |
About $0.024 per mailbox a month. Keeping data is most of the bill: S3 storage, DynamoDB storage and backups, and the search index together are about 60%. Object counts are the next surprise: request charges are $295K a month, 71% as much as all the bytes S3 stores. The levers we already pulled: dedup and lifecycle tiers ($2.28M → $0.42M for S3 bytes), packs and group commits instead of one object per message (R2.7), and indexing 2 KB per message instead of 10 KB with a replica.
R2.7 Trade-Offs
Search index layout
| One index per mailbox | Shared monthly indices, routed by mailbox (chosen) | Per-mailbox index files in S3, loaded on demand | |
|---|---|---|---|
| Count | 100M indices | 36 monthly indices | 100M small index files |
| Cluster metadata | Far beyond what a cluster can track | Small | None; our own search service |
| Query cost | One small index | One shard per monthly index | Load the file (tens of MB for an average mailbox), then search |
| Old mail | Same as new | UltraWarm, seconds | Seconds on a cold load |
| Build or run | Impossible at this count | Managed | We build and operate a search engine |
At a much larger scale, per-mailbox files on object storage become attractive (the index lives next to the mail and costs S3 prices), but they mean building a search service. We start managed.
Counters: exact vs reconciled
| Recount on every view | Counter in every transaction (chosen) | Counter from the stream | |
|---|---|---|---|
| Read cost | Reads every unread item | One item | One item |
| Correct under retries | Yes | Yes: conditional on the message's current state | No: replays double-count |
| Drift | None | Only from bugs; repaired only if modseq is unchanged | Grows with every replay |
Our own MTAs vs SES, at this scale
| SES | Our MTAs (chosen) | |
|---|---|---|
| Receiving 700M a day | 21.3B × $0.10/1,000 + 27.7B chunks × $0.09/1,000 ≈ $4.6M a month | 93 instances ≈ $9.8K |
| Sending 300M a day | 9.12B × $0.10/1,000 ≈ $0.91M + attachment data 720 TB × $0.12/GB ≈ $86K | 63 instances and 300 IPs ≈ $7.8K |
| Control | SES's checks, retries (up to 14 h) and shared reputation | Rejection in the conversation, our retry schedule (5 days), our pools and pacing |
| People | None | A deliverability and abuse team, on call |
The instances are almost free; the real cost of our own MTAs is the team and the reputation work. Against $5.6M a month of SES, that's easily paid for. SES still has a place: new, unwarmed pools could overflow to it, and Round 3 revisits it per company.
One object per message vs packs
| One object per body | Packs of about 660 KB (chosen) | |
|---|---|---|
PUTs | 30.4B a month × $0.005/1,000 ≈ $152K | 914M ≈ $4.6K |
| Tiering | 20 KB objects never leave Standard: 21.9 PB × $0.021 ≈ $460K a month | Packs move to Standard-IA and Glacier IR: ≈ $123K for the same bytes |
| Deleting one message | Delete the object | Mark it; bytes stay until compaction |
| Reading one message | GET | Ranged GET |
The cost of packs is a delete that isn't immediate: a deleted message's body bytes stay in its pack until the monthly compaction. We say so in the privacy policy, and Round 3's legal requirements build on it.
Point-in-time recovery. It's 42% of the DynamoDB bill ($327K). We keep it: flags, folders, thread membership and references exist nowhere else, and a bad deploy that corrupts them needs a restore to a minute. The alternative, rebuilding metadata by re-parsing 83 PB, would take weeks and still lose every flag and folder move.
R2.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Greylisting and deferral curves | A new destination or pool gets 4xx on first attempts; deferrals climb when we push too fast | Retries at 5 and 15 minutes from the same /24 pass greylisting; the destination's token bucket halves its rate when deferrals pass 5% and grows 10% an hour while they stay under 1%; after 4 hours the user gets a "still trying" notice, after 5 days a DSN. |
| A bounce spike (a user imported an old contact list) | Hard bounces from one account rise | Each 5xx sends the user a DSN, and bounces feed the account's signals: above 10% hard bounces over 100 recipients in a day, the account moves to the suspect pool. We never retry 5xx. |
| A MIME bomb | A parser hits a limit | The limit stops that message; it's delivered as MALFORMED with what's safe; nothing crashes, nothing retries. A scanner over a limit marks the file blocked. |
| A thread cycle | References that point in a loop | The write path only looks up IDs, so it can't loop; the thread view skips any link that would make a message its own ancestor and stops at 30 levels. |
| Counter drift | A user reports "Inbox (3)" with nothing unread | The reconciler counts through the sparse index twice, fixes the counter only if modseq didn't move, and logs the correction as a bug. |
| OpenSearch degraded (queries slow, then failing) | Search latency climbs, then errors | Search goes through its own bulkhead: a separate, small pool of connections and threads in the API. A slow dependency holds every in-flight request for its full timeout, so without the bulkhead the API's threads would all end up waiting on search and folder views would fail too. A circuit breaker opens when errors or slow calls pass a threshold, answers "search is temporarily unavailable" at once, and lets one trial request through every few seconds (half-open) to detect recovery. Mail keeps flowing: indexing falls behind and the stream catches up (it holds 24 hours). |
| Why not just 50 ms timeouts on search? | – | Normal P99 jitter would fail good queries, retries would multiply the load on a struggling cluster, and a timeout alone keeps calling it; the breaker stops calling and probes gently. |
| An AZ is lost | A third of the MTAs, parsers and connections | Senders' connections to the lost MTAs break and they retry; the other 62 MTAs carry the peak. Parsing runs a few minutes late while tasks are replaced. Clients reconnect with jitter; surviving gateways hold 75,000 connections each until more start. |
| Valkey loses a shard | Cache misses for some mailboxes | The modseq check sends those reads to DynamoDB; latency rises for them, nothing is wrong. Token buckets on that shard restart at conservative rates. |
| S3 errors on the spool | Spool PUTs fail | MTAs answer 451 and senders retry later, for up to days. Accepting without storing is never an option. |
| A hot mailbox | One mailbox's partition throttles | 451 at RCPT TO for that mailbox only, until its rate falls. |
Primitive: Circuit Breaker, Bulkhead and Fault Tolerance Patterns · Drill: The Slow Recommendation Service That Took Down Checkout. Its two questions are the OpenSearch rows above: why a slow optional dependency exhausts the caller's threads, and why aggressive timeouts alone are worse than a breaker.
R2.9 Production Gotchas
| Gotcha | Why it hurts | What we do |
|---|---|---|
| Sending straight from EC2 without the port 25 process | Outbound port 25 is blocked by default; even when lifted, an IP without reverse DNS or history goes to spam or is refused | Request the lift, set PTR records per IP, pool the IPs, warm them up |
| Message bodies in database columns | A 1 KB access pattern dragging 100 KB rows through cache, backups and replication | Bodies in packs, attachments by hash, metadata in DynamoDB |
| Parsing MIME inside the SMTP session | The sender waits on our slowest code; bombs become failed deliveries and retries | Store, answer, then parse from a queue |
| Unread counts by table scans | Every view reads every unread item | Counters in the same transaction; a sparse index for recounts |
| Reference counts from an at-least-once stream | Replays double-count; attachments leak or vanish | Reference items plus a collector that re-checks before deleting |
Retrying 5xx, or retrying 4xx immediately | Burns reputation; looks like a spammer | 5xx stops; 4xx waits on a schedule with jitter |
| Local rate counters per MTA | 63 MTAs × the limit | Shared token buckets per destination |
| One object per small message | Request charges and the 128 KB rules dominate | Group commits and packs |
R2.10 Pillar Check
| Pillar | What Round 2 covers |
|---|---|
| Reliability | 250 after the spool write and the queue entry; dedup items for senders' retries; conditional transactions for counters and UIDs; a collector that can't delete a referenced file; fleets sized per AZ; raised quotas; a 5-day retry schedule REL 1 · REL 4 · REL 5 · REL 10 |
| Performance Efficiency | Folder views from a modseq-checked cache in 35 ms; search routed to one shard per month; attachments served straight from S3 PERF 1 · PERF 3 |
| Security | Admission control before TLS; DMARC rejection in the conversation; MTA-STS; ownership checks on every attachment reference; isolated parsers and scanners; downloads only after a clean verdict SEC 3 · SEC 5 · SEC 6 |
| Cost Optimization | About $2.39M a month; dedup and tiering cut S3 bytes from $2.28M to $0.42M; packs and group commits instead of billions of small objects; provisioned capacity for a steady load; our own MTAs instead of $5.6M of SES COST 5 · COST 6 · COST 7 |
| Operational Excellence | Signals for queue age, parse errors, deferrals and bounces per destination, index lag and counter corrections OPS 8 |
| Sustainability | Light this round: each attachment stored once; cold mail in colder classes; 2 KB index documents SUS 4 |
R2.11 Round 2 Rubric and Follow-Ups
What a strong senior (L6) answer shows
- Runs its own MTAs at this scale and knows why (cost, control), and puts cheap checks before expensive ones in the SMTP conversation.
- Stores attachments by hash, builds a collector that can't race a new reference, and checks that every reference belongs to the caller.
- Keeps counters exact with conditional transactions and repairs them only when nothing changed.
- Threads by headers with placeholders for missing parents, and knows cycles are a read-time problem.
- Feeds search from the change stream, never by dual writes, and sizes the index honestly (what's indexed, hot vs warm).
- Parses hostile input as a stream with limits, and isolates scanning.
- Treats
4xxas "later" with a schedule, handles greylisting, paces per destination with a shared limiter, separates IP pools, and warms IPs up.
Follow-up questions
-
"A user deletes a message with a 20 MB attachment that 50,000 others also received. What happens to the bytes?" Answer: only that user's reference item goes. The hash lands on the candidate queue; the collector finds 49,999 references and leaves the file
LIVE. The user's body bytes stay in their pack until the next compaction. When the last reference goes, the file is doomed, re-checked an hour later, deleted by generation, and kept 35 days as a non-current version. -
"Why not deliver to Gmail as fast as it will accept?" Answer: Gmail decides its acceptance rate partly on our history; bursts raise deferrals and spam placement for everyone in the pool. We keep a steady pace per destination, raise it slowly while signals are clean, and cut it quickly when deferrals rise.
-
"A message arrives with
Referencesto a thread in someone else's mailbox." Answer: it can't join it: the ID map lives in each mailbox's own partition. At most it joins a thread in the recipient's own mailbox, which is what a legitimate reply does.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Delete the attachment when its count reaches zero" | Races a new reference; at-least-once counts drift. |
"last_read_at makes mark-all-read O(1)" | Per-message flags change without their modseq, so sync clients can't see it. |
| "Thread by subject" | Every "Invoice" becomes one thread; edited subjects split real ones. |
| "Write the index right after the database" | A dual write: failures and reordering silently corrupt search. |
| "An index per mailbox" | 100M indices is beyond any cluster's metadata. |
| "More memory for the parsers" | Bombs expand past any memory; limits and streaming stop them. |
| "Retry deferrals immediately, from any IP" | Looks like spam, and fails greylisting that tracks the sender's IP range. |
Round 3 · Architect · "Business Email for Many Companies Worldwide"
~45 min · Principal (L7) · 4 regions: a home and a standby in the US and in the EU · 10M seats in 25,000 companies · 1B messages/day · retention, eDiscovery and legal holds · anti-phishing · EU residency · customer-held keys · imports · survive a region loss
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Rounds 1 and 2.
Round 2 in 60 seconds. "We run consumer webmail for 100 million mailboxes: a billion messages a day, 34,700 a second at peak, 83 PB stored after three years. Our own inbound MTAs check cheap things first (blocklists before TLS, recipients at
RCPT TO, DMARC before answering), group-commit each 50 ms of mail into one S3 spool object, queue one entry per message, and only then answer250. Parsers stream each message with hard limits, write bodies into 660 KB packs and attachments once per SHA-256, and commit one DynamoDB transaction per message: the message, its counters, its IMAP UID, the mailbox's change number, the thread summary and a reference item per attachment. A collector deletes an attachment only after it's been doomed, re-checked, and found unreferenced. Threads come fromMessage-IDandReferencesthrough a per-mailbox ID map. Search is OpenSearch fed from DynamoDB Streams, routed by mailbox: 90 days hot, the rest in UltraWarm. Clients sync by change tokens and get pushed a tiny frame. Outbound, our MTAs retry4xxfor 5 days on a greylist-friendly schedule, pace each destination with shared token buckets, and keep reputations apart in IP pools. About $2.39M a month, most of it keeping data. Open costs: one tenant, deletion with no legal holds, filters tuned for consumers, one region, and no way to bring in old mail."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: store, answer, parse; metadata transactions; an index that follows the stream; our own senders.
Round 2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Same attachment everywhere | Hash-addressed files, reference items, a safe collector | The collector |
| 2.2 | Drifting counts | Conditional counters, sparse index, careful repair | A job |
| 2.3 | Conversations | Header threading with placeholders | ID-map items |
| 2.4 | Search | Index fed from the stream, routed, two tiers | $0.50M a month |
| 2.5 | Hostile MIME | Streaming parser with limits; isolated scanners | Odd mail partly shown |
| 2.6 | Deferrals | Own MTAs, retry schedule, token buckets, pools, warm-up | A deliverability team |
Open costs: tenants; legal deletion rules; phishing aimed at businesses; residency and keys; imports.
R3.1 The Scope Raise
Interviewer: "We're launching business email. Companies bring their own domains and their own admins. Their legal teams need retention rules, eDiscovery and legal holds. Their executives are being phished. European customers need their mail kept in the EU, and some want to hold their own encryption keys. And our first big deal is a 5,000-person company moving over with ten years of mail."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How many companies and seats? | 25,000 companies and 10M seats in the first years; the largest has 200,000 seats. Plan 100 messages per seat a day. | 1B messages a day again, but in 25,000 separate tenants: routing, quotas and isolation per tenant (step 3.1). |
| What do admins control? | Their domains, users, groups, sending limits, spam and phishing policies; sign-in through their own identity provider. | Domain verification, a DKIM key per domain, per-tenant policies (step 3.1). |
| What do the legal teams need? | Keep mail N years by policy, then delete it; search everyone's mail, including what users deleted, for a lawsuit; holds that stop any deletion. | Retention as policy, holds as a separate layer, eDiscovery over a tenant-wide corpus (step 3.2). |
| What kind of phishing? | Mail "from" the CEO asking for wire transfers, lookalike domains, credential-stealing links, malicious attachments. | Impersonation and lookalike detection, link checks at click time, sandboxing (step 3.3). |
| Where must EU data stay? | In the EU: mail, metadata, indexes, backups and logs. | A home and a standby region per jurisdiction; no worldwide CDN for mail content (step 3.4). |
| What does "their own keys" mean? | Keys in their own AWS account, which they can revoke, with every use logged on their side. | Per-tenant envelope encryption; dedup only inside a tenant; a clear list of what still works after revocation (step 3.4). |
| How big is the migration? | 5,000 seats, 10 years, from another provider; they want it done in a week, with dates, folders and read flags intact. | A bulk import path that bypasses SMTP (step 3.5). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Customers | Consumers, one service | 25,000 companies, 10M seats |
| Messages | 1B/day | 1B/day: 600M US, 400M EU |
| Deletion | User deletes, trash, compaction, collector | + retention policies and legal holds that override everything |
| Search | Each user's own mail | + tenant-wide eDiscovery, including deleted and held mail |
| Security | Spam and malware | + anti-phishing for business targets |
| Keys | Ours | Per tenant, optionally held by the customer |
| Regions | 1 | 4: US home + standby, EU home + standby |
| Targets | 99.99%; visible < 2 s | + per-tenant isolation; residency; survive a region; no bounced mail because of our outage |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| One tenant: every mailbox equal | Nothing maps acme.example to a company; one company's spam run damages everyone's IP reputation; one 200,000-seat import can starve everyone's parsers. |
| Dedup across all users | An attachment shared by two companies can't be encrypted under both companies' keys, and one company's deletion rules would reach another's bytes. |
| Deletion paths: user delete, trash, pack compaction, the collector | Each one would destroy mail a court ordered kept. |
| Filters tuned for consumer spam | A well-written, correctly authenticated message from acrne.example (with "rn" posing as "m") asking the CFO for a wire transfer passes every spam check. |
| One region, one key | EU mail sits in a US region; customers can't revoke our access; a region outage stops every company. |
| Mail arrives only by SMTP | Re-sending 10 years of mail through SMTP gives every message today's date, runs it through spam filters, and triggers notifications and bounces. |
R3.3 New Requirements and API Additions
Adding a domain
httpPOST /v1/admin/domains HTTP/1.1 Authorization: Bearer <tenant admin token> Content-Type: application/json { "domain": "acme.example" }
json{ "domain": "acme.example", "status": "PENDING_VERIFICATION", "records": [ { "type": "TXT", "name": "acme.example", "value": "mailcloud-verify=5b1f0c2e9a" }, { "type": "MX", "name": "acme.example", "value": "10 mx.eu-central-1.mailcloud.example" }, { "type": "MX", "name": "acme.example", "value": "20 mx.eu-west-1.mailcloud.example" }, { "type": "TXT", "name": "acme.example", "value": "v=spf1 include:_spf.mailcloud.example -all" }, { "type": "TXT", "name": "mc2026a._domainkey.acme.example", "value": "v=DKIM1; k=rsa; p=MIIBIjANBg..." }, { "type": "TXT", "name": "_dmarc.acme.example", "value": "v=DMARC1; p=none; rua=mailto:dmarc@acme.example" }, { "type": "TXT", "name": "_mta-sts.acme.example", "value": "v=STSv1; id=20260928" } ] }
The domain stays PENDING until our checker sees the TXT token in public DNS; only then do we accept or send mail for it. A domain already verified by another tenant can't be claimed without the old tenant releasing it (or a support process with proof of ownership). The two MX records point at the tenant's home region first and its standby second (step 3.4).
Retention and legal holds
json{ "policy_id": "p_finance7", "scope": { "groups": ["finance@acme.example"] }, "keep_for_days": 2555, "then": "PURGE", "user_deletes": "HIDE_ONLY" }
httpPOST /v1/admin/holds HTTP/1.1 Content-Type: application/json { "hold_id": "h_2291", "matter": "Case 2026-114", "scope": { "mailboxes": ["dana@acme.example", "raj@acme.example"], "received_between": ["2024-01-01", "2026-09-28"] } }
DELETE /v1/admin/holds/h_2291 releases it. A hold has no expiry; only a release ends it.
eDiscovery
POST /v1/admin/ediscovery/searches with { "query": "\"project falcon\" AND from:*@rival.example", "custodians": [...], "dates": [...], "include_deleted": true } returns a search ID; results arrive in minutes, not milliseconds. POST /v1/admin/ediscovery/exports packages the results as EML files with a manifest into the tenant's own S3 bucket. Both need the ediscovery-manager role and are written to the audit log.
Import
json{ "import_id": "im_7Q2", "source": { "type": "IMAP", "host": "imap.oldprovider.example", "auth": "OAUTH2", "credential_ref": "secret:im_7Q2" }, "mapping": [ { "from": "dana@acme-old.example", "to": "dana@acme.example" } ], "options": { "keep_dates": true, "keep_flags": true, "folders": "ALL" } }
POST /v1/admin/imports with this body starts a job; GET /v1/admin/imports/im_7Q2 reports progress per mailbox. An archive upload (MBOX or EML files into a staging bucket in the tenant's home region) is the other source type.
Recap
- Domains are verified by DNS before use; MX points at the home region, then the standby.
- Retention is policy; holds are separate and never expire on their own.
- eDiscovery is asynchronous and exported to the tenant; imports run as jobs, never through SMTP.
R3.4 Design Evolution: Tenants, Law, Attackers and Borders
Step 3.1: Thousands of Tenants, Each with Domains and Admins
The problem: mail for acme.example, globex.example and 24,998 other domains arrives at the same MTAs. Each company has its own users, admins, sending limits and policies. One of them, a 200,000-seat retailer, runs a quarterly all-staff mailing, and another has a compromised account spraying spam.
What would you do?
Primitive: Circuit Breaker, Bulkhead and Fault Tolerance Patterns (cells are bulkheads at the scale of a whole stack) · Drill: One Customer, One Shard, One Outage. Its first question (the biggest tenant grows tenfold) is the dedicated cell above; its second (a query across all tenants) is the usage warehouse fed by the cells' streams.
Step 3.2: Retention, eDiscovery and Legal Hold
The problem: Acme keeps finance mail 7 years and everything else 3. Legal places a hold on Dana's and Raj's mailboxes for a lawsuit. Next week Dana deletes a thread and empties her trash, the 3-year retention purge runs, pack compaction rewrites her packs, and the attachment collector sweeps unreferenced files. Then Acme's lawyers ask for every message mentioning "Project Falcon", deleted ones included. What would you do?
Step 3.3: Executives Are Being Phished
The problem: the CFO receives "From: Jane Park (CEO) jane.park.ceo@gmail.com": please wire $240,000 today, I'm in a meeting. The next day, a message from billing@acrne.example (not acme) with a "view invoice" link to a perfect copy of the sign-in page. Both pass SPF, DKIM and DMARC: they really come from those domains.
What would you do?
Primitive: Bot Defense, Sybil Resistance and Registration Abuse, for the same idea applied to accounts: judge the sender by its history and its signals, not by a single check.
Step 3.4: EU Mail Stays in the EU; Some Customers Want Their Own Keys
The problem: EU customers' contracts require mail, metadata, search indexes, backups and logs to stay in the EU. Some also want the data encrypted under keys in their own AWS accounts, which they can revoke. And today a regional outage stops every company's mail. What would you do?
Primitive: Cloud Disaster Recovery and Multi-Region Active-Active · Drill: The Booking That Existed in Frankfurt but Not in Virginia. Its first question is the collision above: two regions accepting writes to the same record during replication lag, which one owner region per tenant prevents. Its second (why not replicate synchronously everywhere) is answered by a cross-region round trip on every delivery, by deliveries failing whenever any region is down, and by multi-Region strong consistency, which runs in exactly three regions and supports no transactions, while our delivery is a transaction.
Step 3.5: Import Ten Years of Mail
The problem: a 5,000-seat company arrives with 10 years of mail at its old provider: 5,000 × 100 messages a day × 3,650 days ≈ 1.8 billion messages, about 180 TB. They want it done in a week, with original dates, folders and read flags, while their people keep working. What would you do?
Primitive: OAuth2, OIDC and Distributed Token Authentication, for the delegated grant to the old provider.
Step 3.6: Build or Buy Each Part?
The problem: we now run MTAs, a search tier, an anti-phishing pipeline, an eDiscovery engine and an import service, with a team of a few dozen engineers. Each could be bought or run as a managed service instead. What would you do?
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Thousands of tenants | Tenant map at RCPT TO; tenancy checked on every call; DKIM key per domain; per-tenant quotas and reputation; cells, dedicated for the largest; a usage warehouse | Configuration; spare capacity per cell |
| 3.2 | Retention, eDiscovery, holds | Purge as the only real delete; a holds table every destroying path reads; Object Lock legal hold as a second layer; an Athena-searched corpus; audited | Longer storage; a third copy of the text |
| 3.3 | Phishing | Impersonation and lookalike checks; DMARC for own domains; banners at display time; click-time link checks; sandboxing; clawback on distinct reports | False positives; click friction |
| 3.4 | Residency and keys | Home + standby per jurisdiction; a secondary MX that runs every check and relays home continuously; owner epochs and a fencing lease on our own clocks; gated failover from the standby; per-tenant SSE-KMS with Bucket Keys; bodies repacked per tenant from the spool; revocation answered with 451 at RCPT TO; replica keys revoked separately | A second copy of everything |
| 3.5 | Ten years of mail | IMAP or archive import with a ledger; the parser in import mode; its own 25-task parser fleet and +80,000 write units for the job | A second write path; about $5.2K of writes per migration |
| 3.6 | Build or buy | Decided part by part | Contracts or teams |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Content never crosses the jurisdiction line; only the tenant map (domains, tenant IDs, regions, owner epochs) and aggregate billing counts are global. Inbound mail finds the standby on its own through the second MX record, and the standby hands it back home within the same jurisdiction.
Synthesizing vector architecture diagram...
Each cell adds a tenant map lookup, a phishing stage, a second stream consumer (the eDiscovery corpus), a holds layer that every destroying path reads, per-tenant keys, an audit pipeline and an import path.
The audit pipeline. Every admin action, sign-in, eDiscovery search, export and access to a mailbox by anyone other than its owner is an event (about 300M a day, 500 bytes each: 150 GB a day). API servers pack up to 500 events (about 250 KB) into each Amazon Data Firehose record, because Firehose bills every record in 5 KB increments, and 500 is also the most events its deaggregation splits out of one record. Firehose partitions by tenant and writes compressed files to an audit bucket in the tenant's home region under S3 Object Lock in compliance mode for the tenant's contract period: nobody, including us, can delete them before it ends. Dynamic partitioning allows 500 active partitions per stream by default, so the US's 15,000 tenants go to 30 streams and the EU's 10,000 to 20, chosen by tenant hash. The EU's audit traffic (40% of 150 GB a day, about 0.7 MB/s) fits even the small default Direct PUT throughput some regions start with (1 MiB/s per stream in Frankfurt). Delivery is at least once, so each event carries an event_id and readers deduplicate by it.
Tracing a phishing message
Synthesizing vector architecture diagram...
Authentication passed because the sender really owns acrne.example. The lookalike check, not DMARC, is what caught it.
Tracing a legal hold preventing deletion (numbered steps)
- Legal places
h_2291on Dana's and Raj's mailboxes. The admin API writes the hold and an audit event. - A job finds every message in scope and starts S3 Batch Operations to set Object Lock legal hold on the body-pack and attachment versions they use.
- Dana deletes the "Falcon" thread and empties her trash. The messages become
Hidden; she no longer sees them. - The nightly purge finds Dana's hidden messages past retention, reads the
holdstable, findsh_2291covers Dana and the dates, and skips them. - Compaction finds Dana's pack is 60% dead but locked: it copies the live messages of other mailboxes into a new pack and leaves the old version in place.
- A bug in a new purge release tries to delete one of the locked pack versions by version ID. S3 refuses; the purge job raises a P1 alarm. Even a delete without a version ID would only have added a delete marker.
- Months later, legal releases
h_2291. A job removes the Object Lock holds on versions no other hold references, and the next purge run deletes what retention allows.
Tracing an EU tenant's message, with the home region down
Synthesizing vector architecture diagram...
The sender never saw an error. The standby kept trying to relay the message home; once it became the owner, about 12 minutes in (R3.6), it parsed the queue itself.
R3.6 Numbers and Cost
Tenants and volume per jurisdiction (assumptions from the scope raise: 100 messages per seat a day, 30% of attachment bytes repeated within a tenant)
| US | EU | Total | |
|---|---|---|---|
| Seats | 6M | 4M | 10M |
| Tenants | 15,000 | 10,000 | 25,000 |
| Messages | 600M/day: 6,944/s, 20,833/s at peak | 400M/day: 4,630/s, 13,889/s at peak | 1B/day |
| Cells (about 2M seats each, 200M messages a day) | 3 | 2 | 5, each with a standby |
| Per cell, rounded per AZ | hot search nodes 96 TB ÷ 5.6 = 17.1 → 6 per AZ = 18; UltraWarm 402 TB ÷ 19.8 (90% of 22 TB) = 20.3 → 7 per AZ = 21; inbound MTAs 4,861/s ÷ 400 carried by two AZs → 7 per AZ = 21; outbound MTAs 2,083/s ÷ 250 → 5 per AZ = 15; parsers 139 vCPUs → 12 per AZ = 36; scanners 146 vCPUs → 13 per AZ = 39 | same per cell | 90 hot and 105 warm search nodes; 105 inbound and 75 outbound MTAs; 180 parser and 195 scanner tasks |
| Stored after 3 years, home | 60% of 83.3 PB = 50.0 PB | 33.3 PB | 83.3 PB |
| Standby copy | 50.0 PB | 33.3 PB | 83.3 PB |
KMS requests for tenant keys (per message: tenant body-pack writes 0.04 (1B × 20 KB ÷ 512 KB packs = 39M a day), new attachments 0.21, indexer reads 1.0, user reads 0.6, scanner reads 0.21, old-mail and attachment downloads 0.08, replication reads 0.25: 2.39 at home; plus 0.25 replica writes in the standby)
| Region | Without Bucket Keys | Default quota (shared, symmetric) |
|---|---|---|
| us-east-1 | 600M × 2.39 = 1.43B/day = 16,597/s; 49,792/s at a 3× peak | 100,000/s: 50% at peak |
| eu-central-1 | 400M × 2.39 = 956M/day = 11,065/s; 33,194/s at peak | 20,000/s: exceeded by 66% |
| Cost | (2.39 + 0.25) × 10⁹ × 30.4 = 80.3B × $0.03/10,000 ≈ $241K a month |
We budget Bucket Keys at a 90% reduction for tenant keys (thousands of keys, each reused less than one busy key would be): Frankfurt's tenant-key peak falls to about 3,300 a second and their cost to about $24K a month. The spool's single per-cell key adds about 2.9 requests per message (spool writes, the parsers', repacker's and outbound MTAs' reads, replication); one heavily reused key gets close to the 99% figure, about 400 a second at Frankfurt's peak and $2.6K a month. Together, about 3,700 a second: 19% of Frankfurt's quota. A quota increase alone would leave us paying about $241K and one throttling incident away from refusing mail. Tenants who turn Bucket Keys off for per-object visibility are budgeted against the quota individually.
eDiscovery corpus. Per message about 7 KB of text (4 KB of body, plus 0.3 attachments × about 10 KB of extracted text), about 2.3 KB compressed (an assumption: 3× compression): 2.3 TB a day, 2.52 PB after 3 years. Stored in S3 Intelligent-Tiering (the files are large, so its monitoring fee is negligible): about $0.0074 per GB-month in the US and $0.00845 in Frankfurt for a 10/20/70 mix of its frequent, infrequent and archive-instant tiers. One full-text search over a 5,000-seat tenant's 3 years: 5,000 × 100 × 1,095 = 548M messages × 2.3 KB = 1.26 TB scanned × $5/TB ≈ $6.30.
Failover timing (dependent steps add):
| Step | Time |
|---|---|
| Detection: the standby's canaries (every minute) fail twice | 2 min |
| Alarm, page, engineer online | 1 min |
| Confirm it's the region (canaries, several services, AWS Health) | up to 5 min |
Stop the standby's acks and wait out the fencing lease (the last ack_until, up to 60 s ahead, plus the 30 s margin) | up to 1.5 min |
| Move tenant ownership in the tenant map, then flip the ARC routing control, from the standby | seconds |
| DNS TTL (60 s) and client reconnects with jitter (up to 60 s) | 2 min |
| Total | about 12 minutes without writes for web, IMAP and API clients |
Inbound mail doesn't wait for any of this: senders move to the second MX on their own, and the standby spools it. RPO: metadata, the replication delay, usually a few seconds (global tables replicate asynchronously); the spool, 15 minutes for 99.99% of objects under S3 Replication Time Control; packs and attachments replicate without RTC, but anything from the last 7 days can be rebuilt from the replicated spool. Mail whose spool object hadn't replicated when a region was lost is delayed until that region returns; it would only be lost if the region's S3 data were destroyed.
Monthly cost at the 3-year point (list prices, decimal GB, 730 hours. US home at us-east-1 prices; EU home at Frankfurt prices: S3 Standard $0.0245 / $0.0235 / $0.0225 per GB by tier, Standard-IA $0.0135, Glacier IR $0.005, DynamoDB about 22% above us-east-1, OpenSearch about 20% above; us-west-2 and eu-west-1 standbys at their own prices, which for the lines used here equal us-east-1's except DynamoDB in Ireland, about 13% higher. "Round 2 lines" means Round 2's per-line cost times the jurisdiction's share of the 1B messages.)
| Line | US (60%) | EU (40%) |
|---|---|---|
| S3 storage, home | Standard 1,880 TB ≈ $40.0K; Standard-IA 1,926 TB ≈ $24.1K; Glacier IR 46,320 TB × $4 ≈ $185.3K: $249.4K | Standard 1,254 TB ≈ $28.8K; Standard-IA 1,284 TB × $13.5 ≈ $17.3K; Glacier IR 30,880 TB × $5 ≈ $154.4K: $200.5K |
| S3 requests, home | Round 2's $294.8K per 1B/day, plus tenant repacking: 1.19B packs of 512 KB instead of 914M ($1.4K more PUTs, $8.3K more transitions) and the repacker's ranged reads of the spool, 30.4B × $0.0004/1,000 ≈ $12.2K: $316.7K; × 0.6: $190.0K | 0.4 × $316.7K × about 1.08: $136.8K |
| DynamoDB, home | 0.6 × $786.5K ≈ $471.9K, plus the repacker's pointer switch (1 write unit per message: 11,574/s ÷ 0.7 × $0.00081 × 730 ≈ $9.8K per 1B/day) × 0.6 ≈ $5.9K: $477.8K | 0.4 × $786.5K × 1.22 ≈ $383.8K + $4.8K: $388.6K |
| OpenSearch, a domain per cell | hot 54 × $2,886 ≈ $155.9K; UltraWarm 63 × $1,956 ≈ $123.3K; warm storage 0.6 × $48.2K ≈ $28.9K; masters $3K: $311.1K | the same for 2 cells (36 hot, 42 warm) ≈ $207.4K × 1.2: $248.9K |
| Compute, Valkey, cross-AZ, load balancers, WAF, logs (Round 2 lines, $293.4K) | 0.6 × ≈ $176.0K, plus per-cell rounding (12 more MTAs of each kind, 6 parsers, 12 scanners across the 5 cells, ≈ $4.6K, split by share) ≈ $2.8K: $178.8K | 0.4 × × about 1.1 ≈ $129.1K + $2.0K: $131.1K |
| Internet egress | 1,095 TB tiered: $58.6K | 730 TB tiered: $40.3K |
| DynamoDB replicas in the standby (writes and storage, no backups) | writes: replicated items are billed as ordinary replicated writes, 1 unit per KB (check: we assume the 2× transaction charge applies only in the region that ran the transaction; if replicas pay 2× too, this line roughly doubles its write part). Per message that's 13.6 units for the transactional items (27.2 ÷ 2), minus 0.3 for condition checks, which aren't replicated, plus 1.5 ID-map items and 1 for the replica's own unread index: 15.8 of Round 2's 29.7, 53%. Per 1B/day: $257.7K × 0.53 ≈ $136.6K + the repacker's $9.8K + storage $185.2K = $331.6K; × 0.6: $199.0K | 0.4 × $331.6K × 1.13: $149.9K |
| S3 replication and the standby copy: spool with RTC into Standard, packs and attachments into Glacier IR | per 1B/day: replica PUTs (10.14B × $0.005 + 1.19B × $0.02 + 6.38B × $0.02 per 1,000) ≈ $202.1K; transfer 6,089 TB × $0.02/GB ≈ $121.8K; RTC on the spool 3,800 TB × $0.015/GB ≈ $57K; replica storage ≈ $348.7K; total $729.6K × 0.6: $437.8K | × 0.4: $291.8K |
| Standby compute (MTAs for the second MX, relaying all the time, and the warm slice of sending IPs; a warm 20%) | $35.8K | $26.2K |
| KMS: 40,000 keys we hold (20,000 tenants and their replicas; customer-held keys are billed to the customers) $40K; requests with Bucket Keys $26.7K; field-encryption data keys $2K | 60% of $68.7K: $41.2K | $27.5K |
| eDiscovery: corpus in Intelligent-Tiering (US 1,512 TB × $0.0074 ≈ $11.2K; EU 1,008 TB × $0.00845 ≈ $8.5K), Athena ≈ $15K (an assumed 2,000 searches a month averaging 1.5 TB) | $20.2K | $14.5K |
| Anti-phishing: sandbox (1B × 0.21 new attachments × 0.5% risky from first-time senders = 1.05M detonations a day × 60 s = 729 slots ≈ $53K), click-time checks ≈ $5K, models ≈ $12K | $42.0K | $28.0K |
| Audit, tenant admin and control plane | $9.6K | $6.4K |
| Region pair total | ≈ $2.25M | ≈ $1.69M |
| Business email total | ≈ $3.94M/month |
About $0.39 per seat a month: $0.38 in the US and $0.42 in the EU, where storage, DynamoDB and search all cost more and no CDN helps. Two numbers to argue about: the standby (replica tables, the second copy of 83 PB and its replication, standby compute) costs about $1.14M a month, 29% of the bill, the price of surviving a region within each jurisdiction; and compared with Round 2's $2.39M for the same billion messages, residency's price uplift, cells and the new features (keys, eDiscovery, anti-phishing, audit) add another $0.42M. A migration like step 3.5's adds its own $5.2K of provisioned writes for the 100 hours it runs.
Retention vs cost. Each extra year a tenant keeps its mail adds about 27.5 PB of S3 (75.3 KB per message × 1B × 365, all of it in Glacier IR by then) in each of home and standby, and about 544 TB of metadata:
| Per extra year kept | Math | Monthly, once that year is stored |
|---|---|---|
| S3 at home | US 16,491 TB × $4 ≈ $66.0K; EU 10,994 TB × $5 ≈ $55.0K | $121.0K |
| S3 in the standby | 27,485 TB × $4 | $109.9K |
| DynamoDB at home, storage and point-in-time recovery | US 326,400 GB × $0.30 ≈ $97.9K; EU 217,600 GB × $0.367 ≈ $79.9K | $177.8K |
| DynamoDB replicas | US 326,400 × $0.10 ≈ $32.6K; EU 217,600 × $0.113 ≈ $24.6K | $57.2K |
| eDiscovery corpus | 840 TB: US ≈ $3.7K, EU ≈ $2.8K | $6.5K |
| Total | ≈ $472K a month per extra year |
| Retention | Added monthly cost at steady state, vs 3 years |
|---|---|
| 3 years (our sizing) | – |
| 7 years | 4 × $472K ≈ +$1.89M |
| 10 years | 7 × $472K ≈ +$3.31M |
The users' search index stops at 3 years (older mail is found through eDiscovery or a slower archive search), so it isn't in these rows.
R3.7 Trade-Offs
Shared cells vs a cell per tenant
| Shared cells (chosen for most tenants) | A dedicated cell (the largest, or paid isolation) | |
|---|---|---|
| Blast radius | A cell's worth of tenants | One tenant |
| Utilization | Good: many tenants smooth each other's peaks | Poor: sized for one tenant's peak |
| Noisy neighbors | Quotas and budgets per tenant | None |
| Cost | Shared overhead | A full stack's fixed cost, passed on in price |
| Operations | 5 cells | Grows with every dedicated tenant |
Retention vs cost. The table in R3.6: every extra year is about $472K a month once stored, and a hold can keep a mailbox's mail long past its retention. We default to 3 years for ordinary mail and let tenants choose more, priced per seat, because "keep everything forever" is a cost that only grows.
Anti-phishing strictness vs friction
| Setting | What users see | Risk |
|---|---|---|
| Warn only (banners) | Every external message tagged; lookalikes flagged | People learn to ignore banners |
| Quarantine lookalikes and impersonation of protected users (our default) | A few real messages held for admin review | A partner with a new domain waits for a release |
| Quarantine all first-time external senders | Many real messages held | Business slows down; admins drown in reviews |
Closing the loop. Round 1 answered "how do we accept mail without losing it?" with 250 only after durable storage and a queue in front of the parser. Round 2 kept that and changed how bytes are stored (packs, hashes), how views stay exact (conditional counters, threads, an index fed from the stream) and how we treat the hostile internet in both directions. Round 3 added companies, courts, attackers and borders, and still every piece hangs off the same spine: mail is accepted only once it's durable, every change to a mailbox is one conditional transaction with the next change number, and everything else (search, eDiscovery, push, audit) follows that ordered record. The holds, the keys and the standby all protect that record rather than replace it.
R3.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| A home region outage | The standby's canaries fail; senders' connections to MX 10 fail | Senders move to MX 20 by themselves; the standby spools, queues and keeps trying to relay home. Home stops mailbox writes once the fencing lease runs out. After confirmation and the lease wait, the standby moves tenant ownership and then flips the ARC routing control, about 12 minutes; its parsers drain the queued mail; a reconciliation job (with a UIDVALIDITY bump on any folder whose UIDs collided) runs after failback. |
| A KMS outage in a region | GenerateDataKey and Decrypt fail | Reads that need decryption fail closed for affected tenants: no fallback to an unencrypted path. Writes to S3 fail, so MTAs answer 451 and senders retry for days: mail waits outside, not lost. Bucket Keys already cached may let some operations continue for a short while. |
| A retention bug tries to delete held mail | A purge release deletes by version ID objects under legal hold | Three layers: the purge job reads the holds table itself; Object Lock refuses the delete; any delete without a version ID only adds a marker. The refusal is a P1 alarm: it means a path ignored the holds table. |
| One tenant's IP reputation incident (a compromised account sends 400,000 phishing messages) | Complaints and deferrals rise for that tenant's traffic | The tenant's breaker trips on distinct complaining recipients and pauses its outbound; its traffic moves to the quarantine pool so other tenants' pools stay clean; its admin resets the account; our deliverability team files delisting requests if a blocklist picked up the IPs. |
| A customer revokes its key | S3 returns access errors for that tenant | Intended. We check the key state in both regions (primary and replica), delete the tenant's search indices and cached data keys, and tell its admins. If the key is re-enabled, we rebuild the indices from S3. |
| An import pushes a cell's writes | DynamoDB throttling in the tenant's cell during a migration | The import has its own queue, its own 25-task parser fleet and its own +80,000 provisioned write units, so live mail keeps its latency; if throttles appear anyway, the import workers slow down (they read the throttle metrics), never the live parsers. |
| A misused eDiscovery export | An admin exports a colleague's mailbox without a matter | Only the ediscovery-manager role can search or export; every action is in the tenant's Object Lock audit log, which the tenant's own security team reads. |
R3.9 Runbook and Incident Response
Golden signals, per cell and per tenant OPS 8 · REL 6
| Signal | Alarm | Severity | First action |
|---|---|---|---|
SMTP accept rate, and our 4xx answers | Accepts down 30% vs the same hour last week, or our 451s above 1% | P1 | Spool writes failing? KMS? A cell's queue? |
| Parse-queue age | Oldest message over 60 s for 5 min | P2 | Scale parsers; look for a poison message pattern in the DLQ |
Parse errors and MALFORMED rate | 2× normal | P3 | A new attack pattern, or a parser release? |
| Delivery success by destination | Deferrals over 5% at a large provider for 15 min | P2 | Which pool? Lower its rate; check the provider's postmaster pages |
| Bounce and complaint rates, per tenant | Hard bounces over 5% or distinct-recipient complaints over 0.1% in an hour | P2 | The tenant's breaker; contact its admin |
| Search latency and index lag | P95 over 300 ms, or stream iterator age over 60 s | P3 | OpenSearch health; indexer errors |
| Counter corrections | Any rise in the reconciler's corrections | P3 | A write path forgot the transaction: find it |
| KMS throttles | Any sustained ThrottlingException | P2 | Bucket Keys off for a large tenant? Quota? |
| Replication lag | Spool replication over 15 min, or global table lag over 5 s | P2 | The failover RPO is growing |
| Deletes refused under legal hold | Any | P1 | Pause purge and compaction; find the path that ignored holds |
Reputation incident for one tenant SEC 10
- Confirm: complaints from distinct recipients, deferrals and spam-trap hits concentrated on one tenant.
- Pause the tenant's outbound (command 6); its queued deliveries wait, nothing is dropped.
- Move the tenant to the quarantine pool before resuming, at a low rate.
- With the tenant's admin: find and reset the compromised accounts; review their sent mail.
- Check blocklists for the pool's IPs; request delisting once the source is stopped.
- Resume gradually while complaints stay under the threshold.
Backlog drain REL 7
- Check the queue's depth and age (commands 2 and 3).
- Scale the parsers (command 4): the cell runs 36 parser tasks, and we add 100. Each 4-vCPU task handles about 200 messages a second, so 100 extra tasks drain 20,000 a second beyond the arrival rate: a backlog of 3.6M messages clears in about 3 minutes.
- Watch DynamoDB throttling while draining; the write quota, not the parsers, is often the real ceiling.
- Redrive the DLQ only after the cause is fixed (command 5).
Regional failover REL 13
- Confirm it's the region: the standby's canaries, several services and AWS Health agree.
- Check replica lag (command 10) and spool replication lag to estimate what may be delayed.
- From the standby region, stop the lease acks and wait until the last
ack_untilplus 30 s has passed. Then move each affected tenant's owner with a new epoch (command 11), and only then flip the ARC routing control (command 12). ARC's cluster has endpoints in five regions; try them in turn until one answers. - Let the standby's parsers drain the spooled mail; confirm deliveries and sends succeed.
- Fail back tenant by tenant at quiet hours after replication catches up; run the reconciliation.
Go deeper: CLI playbook
Plain commands an on-call engineer runs one at a time. Replace names, times and IDs with real ones.
text# 1. Alarms firing for a cell aws cloudwatch describe-alarms --region eu-central-1 --state-value ALARM --alarm-name-prefix mail-c1- # 2. Age of the oldest message in the parse queue, last hour aws cloudwatch get-metric-statistics --region eu-central-1 --namespace AWS/SQS --metric-name ApproximateAgeOfOldestMessage --dimensions Name=QueueName,Value=parse-queue-c1 --start-time 2026-09-28T08:00:00Z --end-time 2026-09-28T09:00:00Z --period 60 --statistics Maximum # 3. Depth of the parse queue aws sqs get-queue-attributes --region eu-central-1 --queue-url https://sqs.eu-central-1.amazonaws.com/111122223333/parse-queue-c1 --attribute-names ApproximateNumberOfMessages ApproximateNumberOfMessagesNotVisible # 4. Scale the parsers aws ecs update-service --region eu-central-1 --cluster mail-c1 --service parser --desired-count 136 # 5. Redrive the dead-letter queue after the fix aws sqs start-message-move-task --region eu-central-1 --source-arn arn:aws:sqs:eu-central-1:111122223333:parse-dlq-c1 --destination-arn arn:aws:sqs:eu-central-1:111122223333:parse-queue-c1 # 6. Pause one tenant's outbound mail aws dynamodb update-item --region eu-central-1 --table-name tenant-policy --key '{"tenant_id":{"S":"t_globex"}}' --update-expression "SET outbound_paused = :t, paused_reason = :r" --expression-attribute-values '{":t":{"BOOL":true},":r":{"S":"complaint spike"}}' # 7. Pause purge and compaction aws ssm put-parameter --region eu-central-1 --name /mail/purge/paused --value true --type String --overwrite # 8. Is this pack version under legal hold? aws s3api get-object-legal-hold --region eu-central-1 --bucket mail-packs-euc1 --key t/t_acme/packs/2026/09/28/p_7f3a --version-id 3HL4kqtJlcpXroDTDmJ.rmSpXd3dIbrHY # 9. A tenant key's state, in both regions aws kms describe-key --region eu-central-1 --key-id arn:aws:kms:eu-central-1:444455556666:key/mrk-1a2b3c4d5e6f47a8b9c0d1e2f3a4b5c6 --query "KeyMetadata.KeyState" aws kms describe-key --region eu-west-1 --key-id arn:aws:kms:eu-west-1:444455556666:key/mrk-1a2b3c4d5e6f47a8b9c0d1e2f3a4b5c6 --query "KeyMetadata.KeyState" # 10. Replica status of the messages table, seen from the standby aws dynamodb describe-table --region eu-west-1 --table-name messages-c1 --query "Table.Replicas" # 11. First, move a tenant's ownership to the standby with a new epoch, run in the standby aws dynamodb update-item --region eu-west-1 --table-name tenant-map --key '{"tenant_id":{"S":"t_acme"}}' --update-expression "SET owner_region = :r, owner_epoch = owner_epoch + :one" --condition-expression "owner_region = :old" --expression-attribute-values '{":r":{"S":"eu-west-1"},":one":{"N":"1"},":old":{"S":"eu-central-1"}}' # 12. Then flip the EU pair's routing control to the standby (ARC data plane) aws route53-recovery-cluster update-routing-control-state --region eu-west-1 --endpoint-url https://<cluster-endpoint>/v1 --routing-control-arn arn:aws:route53-recovery-control::111122223333:controlpanel/0123456789abcdef/routingcontrol/abcdef0123456789 --routing-control-state On
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | A standby per jurisdiction; inbound continuity through the second MX; failover run from the standby and gated by a person; owner epochs; the spool replicated with RTC; cells as bulkheads REL 9 · REL 10 · REL 13 |
| Security | Tenancy checked on every call; retention classes per group of mail (finance 7 years, the rest 3); per-tenant keys the customer can revoke, replicas included; holds and an Object Lock audit log nobody can delete; anti-phishing with click-time checks and sandboxing SEC 3 · SEC 4 · SEC 7 · SEC 8 |
| Performance Efficiency | Tenants served from their own jurisdiction; eDiscovery moved off the interactive index onto a batch engine built for it PERF 1 · PERF 3 |
| Cost Optimization | About $3.94M a month; the standby ($1.14M), residency and every extra year of retention ($472K a month) shown line by line; Bucket Keys save about $217K a month on tenant keys and a quota breach; build or buy decided per part COST 5 · COST 8 · COST 11 |
| Operational Excellence | Golden signals per cell and tenant; runbooks for reputation incidents, backlogs and failover; legal and residency requirements written into the design OPS 1 · OPS 10 |
| Sustainability | Data kept in its own jurisdiction; retention by policy instead of forever; cold copies in Glacier IR SUS 1 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Routes by domain to a tenant, checks tenancy on every call, and uses cells to bound the blast radius, with a separate path for cross-tenant reporting.
- Treats legal holds as a layer that every destroying path reads at the moment it acts, and adds an independent second layer (Object Lock), knowing its delete-marker and lifecycle behavior.
- Separates eDiscovery (tenant-wide, batch, audited) from user search (per mailbox, interactive).
- Defends against targeted phishing with checks that don't depend on authentication failing.
- Designs residency as home and standby within a jurisdiction, uses the second MX for inbound continuity (with the standby relaying home all the time and filtering as hard as home), fences the old home with a lease on its own clocks, and runs failover from the standby behind a human gate.
- Explains exactly what customer keys promise, including replica keys and derived copies, and does the KMS quota arithmetic.
- Imports through a dedicated, idempotent, budgeted path.
Follow-up questions
-
"A customer revokes its key at noon. What exactly stops working?" Answer: once both the primary and the replica key are disabled, S3 can't decrypt that tenant's packs and attachments: reads fail and new mail for that tenant gets
451at our MTAs, so senders keep it and retry. Our 5-minute cache of the tenant's field data key expires, so subjects and senders in the metadata become unreadable. We delete its search indices and eDiscovery query caches. The 7-day spool, under our cell key, expires on schedule unless the tenant has a dedicated cell. Re-enabling the key restores everything except the indices, which we rebuild. -
"Why not keep the EU standby in the US for cheaper storage?" Answer: the contract says backups stay in the EU, and a standby copy is a backup. Glacier IR in Ireland costs the same $0.004 per GB-month as us-west-2, so residency doesn't even cost us on that line; the EU premium is in Frankfurt's home prices.
-
"Can a legal hold be placed on mail that arrives after the hold?" Answer: yes, and it must be. That's why the hold is a scope (mailboxes and dates) in its own table, read by every destroying path when it acts, rather than a flag stamped on existing messages. A job also extends Object Lock to new packs and attachments that held messages use.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
"A tenant_id column is isolation" | One missing filter leaks data; one tenant's load hits everyone. |
| "Copy held mail to a vault" | A snapshot misses later mail and doubles storage. |
| "S3 refuses deletes under Object Lock" | Only deletes of a specific version; a delete without a version ID adds a marker. |
| "Spam filters stop phishing" | Targeted phishing is authenticated and low-volume. |
| "Active-active for every tenant" | Last writer wins and local transactions let two regions both deliver. |
| "Disable the key and the customer is safe" | The replica key in the standby has its own state and must be disabled too. |
| "CloudFront for EU attachments" | Edge caches can't be kept inside the EU. |
| "Import by re-sending over SMTP" | New dates, spam checks, notifications and bounces. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: own domain, clients, attachment size, spam, retention | Restate Round 1 in 60 seconds | Restate Round 2 in 60 seconds |
| 5–15 min | Requirements, the SMTP contract, the REST API, IMAP's UIDs and modseqs | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.6: MX and durable accept → metadata vs bytes → async parse → SPF, DKIM, DMARC → inbound checks → 4xx vs 5xx | Steps 2.1–2.6: attachments by hash and a safe collector → exact counters → threading → search from the stream → hostile MIME → deliverability | Steps 3.1–3.6: tenants and cells → retention, holds, eDiscovery → anti-phishing → residency and keys → import → build or buy |
| 40–50 min | Numbers, cost, SES vs our own MTA | Storage tiers and object counts, index sizing, write units, fleets per AZ, budgets, cost | Per-jurisdiction sizing, KMS arithmetic, standby cost, retention per year, failover timing |
| 50–60 min | Failures and pillar check | Failures, gotchas, pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening any round: "A mailbox is a small, hot index over a large, cold pile of immutable messages, so I store the raw message durably before answering
250, parse it asynchronously, and commit each mailbox change as one conditional transaction with the next change number; views, search and sync all follow that record." - When scale arrives: "I'll run our own MTAs with cheap checks first, store attachments once by hash with a collector that re-checks before deleting, keep counters in the same transactions, feed search from the change stream, parse hostile mail as a stream with limits, and treat every
4xxfrom a receiver as a request to slow down."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "When is a message safe?" (REL 4) | When it's in durable storage and queued; only then do we answer 250, and a failed write answers 451 so the sender retries. | 1–2 | Step 1.1, R2.5 |
| "What if the parser crashes on a message?" (REL 5) | The queue redelivers it, the parser is idempotent, limits stop bombs, and a DLQ catches the rest; the raw bytes are never at risk. | 1–2 | Steps 1.3, 2.5 | |
| "What if a region fails?" (REL 13) | Senders fall back to the standby's MX on their own; clients move after a gated flip from the standby, in about 12 minutes. | 3 | Step 3.4 | |
| Performance | "How is the inbox fast with years of mail?" (PERF 3) | A 1 KB item per message sorted by folder and time, counters in the same transaction, and a cache checked against modseq. | 1–2 | Steps 1.2, 2.2 |
| "How does search stay under 300 ms?" (PERF 1) | One routed shard per monthly index, 2 KB documents, recent months hot and the rest warm. | 2 | Step 2.4 | |
| Security | "Is mail protected between servers?" (SEC 9) | STARTTLS in both directions, an MTA-STS policy so senders insist on TLS to us, and DKIM signatures so a receiver can detect changes; SPF and DMARC prove the sender. | 1–3 | Step 1.4, R2.5, 3.1 |
| "How do you stop CEO fraud?" (SEC 4) | Impersonation and lookalike checks against the tenant's directory and domains, banners, click-time link checks and clawback. | 3 | Step 3.3 | |
| "Can customers hold their own keys?" (SEC 8) | Yes: SSE-KMS per tenant with Bucket Keys; revocation must cover the replica key, and we delete derived copies. | 3 | Step 3.4 | |
| Cost | "What does it cost?" (COST 5) | About $3.7K, $2.39M and $3.94M a month; keeping data dominates, and object counts matter as much as bytes. | 1–3 | R1.7, R2.6, R3.6 |
| "SES or your own mail servers?" (COST 11) | SES for one company; at a billion a day, SES would cost about $5.6M a month, so we run MTAs and pay for the team instead. | 1–2 | R1.8, R2.7 | |
| Operations | "How do you know mail is healthy?" (OPS 8) | Accept rate, queue age, parse errors, deferrals and bounces by destination, complaints by tenant, index lag and counter corrections. | 2–3 | R3.9 |
| Sustainability | "Is this wasteful?" (SUS 4) | Attachments stored once per tenant, cold mail in cold classes, small index documents, and retention by policy. | 2–3 | R2.6, R3.7 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Accepting mail | MX, 250 after durable storage, a managed receiver | Own MTAs, cheap checks first, group commit, 451 on failure | Tenant routing at RCPT TO, a second MX in the same jurisdiction |
| Storage | Metadata apart from bytes, bytes first | Packs, attachments by hash, a safe collector, tiering and object counts | Dedup and keys per tenant, a standby copy, retention per year priced |
| Views | Folder queries, counters in the transaction | Exact counters, threads, search from the stream, change tokens and push | eDiscovery apart from user search |
| Hostile input | Authentication verdicts, quarantine, no backscatter | Streaming parser with limits, isolated scanning | Targeted phishing, click-time checks, sandboxing, clawback |
| Sending | SPF, DKIM, DMARC through SES; 4xx vs 5xx | Retry schedule, greylisting, shared token buckets, IP pools, warm-up, providers' rules | Per-domain DKIM, per-tenant reputation and breakers |
| Deletion | Soft delete | Collector with doom, re-check and generations | Purge as the only real delete; holds as a layer; Object Lock as a second |
| Evolving under new scope | Builds from one server step by step | Opens with "what breaks", fixes storage before views | Adds tenants, law, attackers and borders without changing the spine |