Design a Metrics Monitoring and Alerting System
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | An ops team monitors 500 servers and a few services | A company-wide platform: every service emits metrics | Monitoring sold as a product to many companies, worldwide |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Volume | ~50K series every 10 s: 5,000 samples/s | 100M active series: 10M samples/s sustained, 30M/s peak | ~1B active series in 3 regions: ~100M samples/s |
| Retention | 30 days | 30 days raw, 180 days at 5 min, 5 years at 1 hour | Per-tenant plans |
| Footprint | 1 region, 3 AZs | 1 region, 3 AZs; survives an AZ loss | 3 regions of cells, each with an alerting standby region |
| Targets | Alerting 99.9%; dashboards < 1 s; a page within 60 s of a rule firing | Alerting and ingest 99.95%; recent queries P95 < 50 ms, 30-day P95 < 500 ms | 99.99% per region; alerts keep firing in a region outage; one tenant can't slow another |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is Metrics Monitoring?
You Already Know One: a Car's Dashboard
A car's dashboard has two kinds of things on it. Gauges show numbers that change over time: speed, engine temperature, fuel. Warning lights come on when something is wrong: low oil pressure, a brake fault. You glance at the gauges now and then. The warning lights must come on when, and only when, something really needs you.
A metrics monitoring system does the same for software. Every server and service reports numbers every few seconds: CPU use, requests per second, error counts, request latency. The system stores those numbers, draws them as graphs (the gauges), and runs alert rules that notify a person when something is wrong (the warning lights).
| Car | Metrics system | Example |
|---|---|---|
| A gauge | A time series: one metric for one thing, over time | CPU use of host web-17 |
| The needle's position right now | A sample (also called a point): one timestamp and one value | (10:00:10, 0.42) |
| The whole dashboard | A dashboard: graphs over a time range | The checkout service's error rate, last 24 hours |
| A warning light | An alert rule that becomes a notification | "5xx errors above 5% for 5 minutes" pages the on-call engineer |
A series is named by a metric name plus labels (key-value pairs): http_requests_total{service="checkout", code="500", instance="web-17"} is one series. Change any label value and it is a different series.
What Makes It Hard
- The numbers never stop. Millions of samples arrive every second, forever. Almost all are written once and never changed.
- Queries span seconds to years. The on-call engineer wants the last five minutes, in detail. The capacity planner wants two years, in outline. One storage layout can't serve both cheaply.
- The number of series explodes. One careless label, such as a user ID, turns one series into millions. Storage cost and memory depend on how many distinct series exist, not only on how many samples arrive.
- False alarms are as bad as none. A team paged ten times a night for nothing stops reading pages, and then misses the real one.
- The monitor must outlive what it watches. An outage is exactly when people need their graphs and alerts. If our system fails with everything else, it fails when it matters most.
The Question the Whole Loop Answers
How do we store and query an endless stream of numbers cheaply, and raise an alarm people trust?
The answer gets sharper every round:
- Round 1: a time-series layout instead of a table of rows, a label index, scraping, alert rules that wait before firing, grouped notifications, and something outside the system that notices when the system itself dies.
- Round 2: scale: bit-level compression, a sharded and replicated in-memory tier fed by a stream, downsampled tiers for long ranges, defenses against series explosions and late data, and alert storms reduced to one page.
- Round 3: monitoring as a product: per-tenant limits and shuffle sharding, alerting that survives a region outage, alerts based on error budgets, controlled query costs, and a bill for every tenant.
Round 1 · Mid-level · "Monitor One Company's 500 Servers"
~35 min · SDE II (L5) · 1 region, 3 AZs · 500 servers, ~50K series · 5,000 samples/s · 30 days · alerting 99.9%
R1.1 Establish Design Scope
The interviewer says: "Our ops team runs about 500 servers and a handful of services. Today they find out about problems from customers. Design a system that collects metrics, shows them on dashboards, and pages the on-call engineer when something breaks." Before drawing anything, we ask questions and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Which metrics? | Host metrics (CPU, memory, disk, network) and service metrics (requests, errors, latency). | Two sources: an agent on every host, and the services' own counters. About 50K series in total (R1.7). |
| How often do we collect? | Every 10 seconds. | Each series gets a new sample every 10 s. With 50K series that's 5,000 samples a second. Short outages (under 10 s) can hide between samples. |
| How long do we keep data? | 30 days. | About 26 GB in total (R1.7). One machine's disk holds it; no tiers yet. |
| What queries? | Graphs by host and by service, over the last hour to the last month. | We must find series by their labels fast (step 1.2), and read a time range of one series without touching others (step 1.1). |
| Who gets alerts, and how? | The on-call engineer by pager; the team by email. | An alert router with two receivers, and a plan for when the pager service itself is down (R1.9). |
| Do we pull metrics from the servers, or do they push to us? | You tell us. | A real design choice (step 1.3). |
Out of scope for this round:
- Years of history. 30 days is enough; long-range tiers come in Round 2.
- Defending against series explosions. 500 servers under one team's control won't invent millions of series by accident. Round 2 brings that back.
- Many teams and many tenants. One ops team owns everything.
In a multi-round loop, the interviewer brings parts of the out-of-scope list back later. Write it where you can see it.
R1.2 Functional Requirements, Derived Step by Step
We read the problem one phrase at a time and turn each phrase into something the system does:
| Phrase from the problem | What the system does |
|---|---|
| "Collects metrics" | Every 10 s, get the current value of every series from every host and service |
| (implied) | Store each sample durably and in time order |
| "Shows them on dashboards" | Query a time range for series matching labels, at a chosen resolution (the step), and draw graphs |
| "Pages the on-call engineer" | Evaluate alert rules on a schedule; when a rule is true long enough, create an alert |
| (implied) | Send each alert to the right people, once, grouped with related alerts |
| "When something breaks" | Including when our own monitoring breaks |
Not yet: multi-year retention, cardinality defenses, alert inhibition across teams.
R1.3 Non-Functional Requirements: the Questions
We state each quality in words first, in the order we'd defend it. The numbers come in R1.7.
- Alerting is the most important path. Dashboards can be slow for a minute; a missed page can't be undone. Alerting must survive one server or one AZ failing.
- Alerts must be fresh. Once a rule's condition has held long enough, the page should arrive within about a minute.
- The workload is write-heavy and append-only. 5,000 samples a second arrive all day, and nobody edits them. Reads are rare in comparison, and almost always "one series, a time range".
- Reads are by time range and labels. "CPU for every host in
service=checkoutover the last 6 hours" is the typical query. - Monitoring must not share fate with what it watches. If the monitored servers or their network fail, we must still see it and page someone.
R1.4 The API
Everything in this round is internal. There are three contracts.
1. A sample. The logical data model every other part of the system speaks (shown as JSON for readability):
json{ "name": "node_cpu_seconds_total", "labels": { "instance": "web-17:9100", "job": "node", "mode": "user", "cpu": "3" }, "timestamp_ms": 1790510410000, "value": 81234.57 }
A counter like this one only goes up (seconds of CPU time used since boot), so a graph shows its rate per second, not its raw value. A gauge (memory in use, queue depth) is shown as it is.
With pull collection (step 1.3), this is what a host's agent serves when we ask it, in the plain-text format Prometheus-style agents use:
text# TYPE node_cpu_seconds_total counter node_cpu_seconds_total{cpu="3",mode="user"} 81234.57 node_cpu_seconds_total{cpu="3",mode="idle"} 1204521.02 # TYPE node_memory_MemAvailable_bytes gauge node_memory_MemAvailable_bytes 3.1254528e+09
The collector adds instance and job labels and the timestamp when it reads the page.
2. The query endpoint. A range query takes an expression, a start, an end and a step (the gap between the points the graph gets back):
httpGET /api/v1/query_range?query=sum+by+(instance)(rate(node_cpu_seconds_total{job="node",mode!="idle"}[5m]))&start=1790488800&end=1790510400&step=60 HTTP/1.1 Host: metrics.internal.example.com
httpHTTP/1.1 200 OK Content-Type: application/json { "status": "success", "data": { "resultType": "matrix", "result": [ { "metric": { "instance": "web-17:9100" }, "values": [ [1790488800, "2.91"], [1790488860, "3.04"], [1790488920, "2.87"] ] } ] } }
The query is written in PromQL, the Prometheus query language (a query language, not program code). rate(x[5m]) means "the per-second increase of counter x, averaged over the last 5 minutes". Six hours at a 60-second step returns 360 points per series, which is about what a graph a few hundred pixels wide can show.
3. An alert rule (YAML, in the format Prometheus-style rule evaluators read):
yamlgroups: - name: checkout interval: 15s rules: - alert: CheckoutHighErrorRatio expr: > sum(rate(http_requests_total{service="checkout", code=~"5.."}[5m])) / sum(rate(http_requests_total{service="checkout"}[5m])) > 0.05 for: 5m labels: severity: page team: payments annotations: summary: "Checkout 5xx ratio above 5% for 5 minutes" runbook: "https://wiki.example.com/runbooks/checkout-errors"
interval is how often the rule is evaluated. for is how long the condition must stay true before the alert fires (step 1.4). labels decide where it goes (step 1.5).
Recap
- Three contracts: a sample (name, labels, timestamp, value), a range query with a step, and an alert rule with a
forduration. - About 50K series every 10 seconds, kept 30 days, in one region.
- Alerting first, fresh within a minute, and independent of what it watches.
Let's build it, starting with the simplest version that works.
R1.5 Design Evolution: From a Table of Rows to Alerts People Trust
Every step follows the same pattern: a problem, your turn to think, the answer, and what the answer costs us. The cost is always the next problem.
Step 1.0: The Baseline
A relational database with one table. Every sample is one row: (metric_name, labels, timestamp, value), with an index on (metric_name, timestamp). A script on each server inserts its numbers every 10 seconds. Another script runs alert queries every minute and emails someone if a query returns rows.
Synthesizing vector architecture diagram...
Every sample is a row, and every question is a query against one big table.
What's good about it: it works on day one, and everyone knows SQL. What's wrong: 5,000 rows a second is 432 million rows a day. Each row repeats its name and labels (tens of bytes), and each index entry adds more. A row per sample costs roughly 50 to 100 bytes on disk once overheads and indexes are counted, against 16 bytes for the timestamp and value alone. And a graph of one host's CPU over a day has to find 8,640 rows scattered among 432 million.
Step 1.1: Storage and Range Queries Are Slow
The problem: after a week the table has 3 billion rows. The index is bigger than memory, so inserts slow down, and a one-day graph of one host takes seconds because its 8,640 rows are spread across the whole table. Storage grows by tens of gigabytes a day. What would you do?
Primitive: Write-Ahead Log & LSM-Trees
Synthesizing vector architecture diagram...
Writes are appends to the WAL and to open chunks in memory; every two hours the head becomes an immutable block, and old data leaves by dropping whole blocks.
Step 1.2: Find Every Series With service=checkout
The problem: the dashboard asks for http_requests_total{service="checkout", code=~"5.."}. We have 50,000 series. Which ones match?
What would you do?
Primitive: Trie Data Structure & Inverted Index
Step 1.3: Pull or Push?
The problem: 500 servers and a few services must get their numbers to us every 10 seconds. Either our collector asks each of them (pull, also called scraping), or each of them sends to us (push). What would you do? Which one, for this scope, and why?
Step 1.4: Alerts Fire and Clear Every Minute
The problem: a rule says "page if CPU > 85%". A busy host's CPU swings between 84% and 87% all afternoon. The on-call engineer gets 30 pages, each followed by a "resolved" a minute later. By evening they mute their phone. What would you do?
Synthesizing vector architecture diagram...
An alert has to earn its way to Firing, and has to fall well below the fire threshold to leave it.
Step 1.5: Five Alerts for One Outage Wake the On-Call Five Times
The problem: the checkout database slows down. Within a minute, five rules fire: checkout latency, checkout errors, payment latency, database CPU, and connection-pool exhaustion. The on-call engineer's phone rings five times for one problem. Worse: both of our collector copies (R1.6) evaluate the same rules, so every alert arrives twice. What would you do?
Synthesizing vector architecture diagram...
Five alerts from two evaluators become one notification.
Step 1.6: The Monitoring Server Died and Nobody Noticed
The problem: on Saturday the monitoring server's disk fills and it stops. No alerts fire, because the thing that fires alerts is down. On Monday a customer reports that checkout has been failing since Sunday. What would you do?
Synthesizing vector architecture diagram...
The alarm fires on silence, so it catches every way the pipeline can die: the evaluator, Alertmanager, the network, or the notifier.
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | One row per sample in a SQL table | Everything below |
| 1.1 | Slow storage and range reads | Time-series layout: series IDs, compressed chunks, WAL, 2-hour blocks | A specialized database |
| 1.2 | Find series by labels | Inverted index of posting lists, intersected | Index memory grows with series count |
| 1.3 | Pull or push | Pull with service discovery; a push gateway for batch jobs | Discovery to maintain; one scraper's reach |
| 1.4 | Flapping alerts | Scheduled rules, for durations, hysteresis, ratios | Slower alerts |
| 1.5 | Many pages for one outage | Alertmanager: dedupe, group, route | Routing config; 30 s wait |
| 1.6 | The monitor died silently | Watchdog alert to an outside dead man's switch | A second, independent system |
R1.6 Architecture v1
Now the concepts get names.
Synthesizing vector architecture diagram...
Two identical Prometheus servers scrape everything and send every alert to all three Alertmanagers, which share state by gossip and send one notification. The watchdog path sits outside the stack.
Why these pieces:
- Two identical Prometheus servers in different AZs, each scraping every target and evaluating every rule. This is the standard way to make Prometheus highly available: there is no replication between them, just two full copies. If one dies, the other keeps collecting and alerting. Each copy may have small gaps the other doesn't, which is fine for monitoring.
- Three Alertmanagers, one per AZ, forming a cluster. Each Prometheus sends every alert to all of them (not through a load balancer). They share silences and a notification log (which notifications were sent) by gossip. Before sending, each instance waits according to its position in the cluster (position ×
--cluster.peer-timeout, 15 s by default), so the first one sends and the others see it in the shared log and stay quiet. If gossip breaks, they may each send: a duplicate page, never a missing one. - Grafana draws dashboards from Prometheus 1, falling back to Prometheus 2. Dashboards are kept as files in version control, so the Grafana server itself holds nothing we can't rebuild.
- EC2 service discovery gives Prometheus the list of instances to scrape (every instance tagged
monitoring=on), so new servers are picked up without editing a file, and a missing one shows asup == 0. - The watchdog (step 1.6) runs in Lambda, CloudWatch and SNS: AWS services that share no servers, disks or configuration with our stack.
Tracing the three flows
1. A sample from scrape to disk
- At 10:00:10, Prometheus 1 fetches
http://web-17:9100/metrics(about 100 series from this host, a few KB of text). - It parses each line, adds
job="node"andinstance="web-17:9100", and stamps every sample with the scrape time. - Each label set is looked up in the head's index. A known one gives its series ID; a new one gets a new ID and new posting-list entries.
- The samples are appended to the WAL, then to each series' open chunk in memory. Within milliseconds they can be queried.
- At the next 2-hour boundary, the head is cut into a block on disk and the old WAL segments are removed.
2. A dashboard query
- Grafana asks for
rate(node_cpu_seconds_total{instance="web-17:9100", mode!="idle"}[5m])over 24 hours at a 60 s step. - The index intersects the posting lists: 7 non-idle modes × the host's CPU cores, since
node_cpu_seconds_totalhas acpulabel: 14 series on a 2-vCPU host (summed per host later). - Prometheus reads those series' chunks from 12 two-hour blocks and the head: 14 × 8,640 ≈ 121,000 samples, decoded in well under a second.
- It computes the rate at each of the 1,440 steps and returns the points.
3. An alert from breach to page
Synthesizing vector architecture diagram...
The page arrives about 7 minutes 40 seconds after the error started: about 2 minutes for the 5-minute rate window to carry the ratio past 5% (5 ÷ 12 × 5 min ≈ 2 min 5 s), the 5-minute for, and the 30-second group wait.
R1.7 Numbers
Targets
| Quality | Target | Why this number |
|---|---|---|
| Alerting availability | 99.9% of minutes with a working evaluate-and-notify path (43.8 min a month) | Two collectors and three Alertmanagers in three AZs |
| Alert freshness | Page within 60 s after a rule's for has elapsed | Worked out below |
| Dashboards | < 1 s for any graph over up to 30 days | One server, 50K series |
| Retention | 30 days | From R1.1 |
Series and samples
| Item | Math | Result |
|---|---|---|
| Host series | 500 servers × ~80 series each (CPU by mode, memory, each disk and network interface), after dropping the agent's collectors we don't use | 40,000 |
| Service series | ~10 services × ~1,000 (requests by route and status code, latency histogram buckets) | 10,000 |
| Total series | ~50,000 | |
| Samples per second | 50,000 ÷ 10 s | 5,000/s |
| Samples per day | 5,000 × 86,400 | 432 million |
| Scrapes per second | ~600 targets ÷ 10 s | 60/s |
The 80 series per host is an assumption: a host agent with every collector turned on emits several hundred series per host (CPU time alone is cores × 8 modes). Choosing what to collect is the first cardinality decision, even at this size.
Bytes
| Item | Math | Result |
|---|---|---|
| Raw sample | 8-byte timestamp + 8-byte float | 16 B |
| Compressed sample | 1.37 B in the Gorilla paper's workload (Round 2); Prometheus's own documentation says 1 to 2 bytes per sample on average | We plan with 2 B |
| Per day | 432M × 2 B | 0.86 GB |
| 30 days | 0.86 GB × 30 | ~26 GB |
| Disk per Prometheus | 26 GB + WAL + index + room to grow | 100 GB gp3 |
| Label index | 50,000 series × ~200 B (label strings once, posting entries) | ~10 MB |
| Memory for the head | 50,000 × ~8.3 KB per in-memory series (a conservative planning figure from Grafana Mimir's capacity guide: 2.5 GB per 300,000 series) | ~0.4 GB; an 8 GiB m7g.large is ample |
Alert freshness, step by step. These steps happen one after another, so the worst cases add up:
| Step | Worst case |
|---|---|
| The breach reaches a scrape | 10 s (scrape interval) |
| The next rule evaluation | 15 s (evaluation interval) |
for duration | the rule's own choice (5 min above); excluded from the target |
Alertmanager group_wait for a new group | 30 s |
| Delivery to the pager service | ~5 s (assumed) |
Total after for | ≈ 60 s |
A group that is already firing adds new alerts at the next group_interval (up to 5 minutes), which is why each team gets its own group.
This budget starts when the rule's expression becomes true, which is not the moment the problem starts. An expression over rate(...[5m]) lags the event: a jump from ~0% to 12% errors lifts the 5-minute ratio past 5% only after about (5 ÷ 12) × 5 min ≈ 2 min (the trace in R1.6). Shorter rate windows react faster but are noisier.
Rough monthly cost, three options (us-east-1 list prices; check the AWS Pricing Calculator before quoting)
| Line | Self-run Prometheus (chosen) | Amazon Managed Service for Prometheus + Amazon Managed Grafana | CloudWatch custom metrics |
|---|---|---|---|
| Collectors | 2 × m7g.large × $0.0816/h × 730 h = $119 | 2 small scrapers in agent mode (no local storage), ≈ $50 | CloudWatch agent on each host, no extra servers |
| Storage | 2 × 100 GB gp3 × $0.08 = $16 | ~26 GB × $0.03/GB-month ≈ $1 | included |
| Samples or metrics | – | 13.14B samples/month (5,000 × 2.628M s): first 2B at $0.90 per 10M = $180, next 11.14B at $0.35 per 10M = $390 | 50,000 metrics: first 10,000 at $0.30 = $3,000, next 40,000 at $0.10 = $4,000 |
| Queries and rules | – | $0.10 per billion samples processed; our rules and dashboards: ~$25 (estimate) | alarms $0.10 each per month and up |
| Alerting | 3 × t4g.small Alertmanagers ≈ $37 | included in the service | SNS |
| Dashboards | 1 × t4g.medium Grafana ≈ $25 | 5 editors × $9 + 20 viewers × $5 = $145 | CloudWatch dashboards |
| Total | ≈ $200 | ≈ $790 | ≈ $7,000+ |
Three lessons:
- CloudWatch prices per metric, and to CloudWatch every series is a metric. It's excellent for AWS services' own metrics and a few hundred custom ones, and very expensive for 50,000 series. The price of a monitoring system follows its series count.
- Managed Prometheus costs about $600 a month more here, in return for no servers to patch, no disks to fill (the cause of step 1.6's outage) and no upgrades. For a small ops team, that's a few engineer-hours a month, and it can easily be worth it. Its minimum rule evaluation interval is 30 s (a service quota), which adds 15 s to our freshness budget.
- We choose self-run for this round because it's cheap, it teaches every moving part, and the team already runs servers. We'd revisit it the first time the monitoring stack causes an incident.
R1.8 Trade-Offs
Pull vs push, for this scope
| Pull (chosen) | Push | |
|---|---|---|
| Down detection | up == 0 for free | Must alert on missing data |
| Inventory | Service discovery says what should exist | Only who talked recently |
| Batch jobs | Need a push gateway | Natural |
| Firewalled hosts | Collector must reach them | Outbound only |
Self-run vs managed
| Self-run Prometheus (chosen) | Managed Prometheus + Grafana | |
|---|---|---|
| Monthly cost here | ≈ $200 | ≈ $790 |
| Operations | We patch, size disks, upgrade | AWS runs storage, querying, rules, Alertmanager |
| Limits | Whatever our servers hold | Service quotas (for example 50M active series and ~1.67M samples/s per workspace by default, both adjustable) |
| When it wins | Small, stable, skilled team | No appetite to run stateful servers |
Static thresholds vs rates and ratios
| Static count ("> 100 errors/min") | Rate or ratio ("> 5% of requests fail") | |
|---|---|---|
| At peak traffic | Pages when nothing is wrong | Correct |
| At night | Silent while half the requests fail | Correct |
| Setting it | Easy, but must be re-tuned as traffic grows | Stable as traffic grows |
R1.9 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| A scrape target disappears | A host crashes; its scrape fails. | Prometheus records up{instance="web-17:9100"} = 0. A rule up == 0 with for: 2m pages. If the target was removed on purpose (the instance was terminated), service discovery drops it and nothing fires. If a whole job vanishes from discovery, absent(up{job="checkout"}) catches it, because "no series at all" is different from "a series at zero". |
| Prometheus restarts | The process restarts after an upgrade or crash. | On start it replays the WAL to rebuild the head: seconds at 50K series. The other copy keeps scraping and alerting meanwhile; this copy has a gap of a minute or so. |
| The disk fills | Writes fail; the server stops ingesting. | Alerts on the monitoring servers' own disk use fire days before (a normal rule, since the other copy is healthy). If both fail, the watchdog's silence pages through CloudWatch within 5 minutes. |
| The pager service is down | Alertmanager's notifications to it fail and retry. | Every page also goes to email (continue: true, step 1.5). A rule on alertmanager_notifications_failed_total is routed to email only, never to the failing integration. |
| An AZ fails | One Prometheus and one Alertmanager disappear. | The other Prometheus and two Alertmanagers keep working. Grafana falls back to Prometheus 2. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Two independent collectors and a three-node Alertmanager cluster in three AZs; a dead man's switch outside the stack; every page also by email. REL 6 · REL 11 |
| Performance Efficiency | A storage layout matched to the data: append-only compressed chunks read by time range, and an inverted index for labels. PERF 3 |
| Operational Excellence | This system is the operations pillar's main tool: it's how every other team knows their workload's health. So it gets the same care: alert rules and dashboards live in version control and are reviewed like code; every paging rule has a runbook link; alerts are on symptoms and ratios; and the stack watches itself. OPS 4 · OPS 8 · OPS 10 |
| Cost Optimization | Three options priced; CloudWatch's per-metric price named as the trap for high series counts; Graviton instances. COST 5 |
| Security | Light this round: collectors reach targets only on metrics ports, allowed by security groups; Grafana behind single sign-on. SEC 5 |
| Sustainability | Light this round: collect only the series we use. SUS 4 |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Explains why a table of rows is the wrong shape, and what a time-series layout does instead (series IDs, time-ordered compressed chunks, WAL, blocks).
- Finds series with an inverted index and intersects posting lists.
- Compares pull and push on down detection, not just on "scale".
- Uses
for, hysteresis and ratios to stop flapping, and knows the cost is slower alerts. - Groups and deduplicates alerts before they reach a person.
- Asks "who watches the watcher?" without being prompted.
Follow-up questions
-
"Your two Prometheus copies disagree slightly. Is that a problem?" Answer: no. They scrape at slightly different moments, so values differ in the last digits, and one may have a gap the other doesn't. Dashboards read one copy (with fallback); alerts come from both and are deduplicated by labels in Alertmanager. If exact agreement mattered (billing, say), metrics would be the wrong tool.
-
"A batch job runs for 20 seconds every night. How do you know if it failed?" Answer: it pushes its result (
last_success_timestamp, duration, records processed) to the push gateway before it exits, and we alert ontime() - last_success_timestamp > 26 * 3600: "no success in 26 hours". That catches both a failed run and a run that never started, which an "errors > 0" rule would miss. -
"Why does the watchdog go through CloudWatch? Couldn't the second Prometheus watch the first?" Answer: it can, and we do that too. But both copies share the same Alertmanager cluster, the same pager integration and often the same configuration mistake. The dead man's switch tests the whole path end to end, from somewhere that shares none of it.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Store metrics in a normal SQL table with indexes" | Random index writes and repeated labels; range reads touch rows all over the table. |
| "Push is always better" | Push loses free down detection and the list of what should exist. |
| "Alert on every threshold crossing" | Values near the threshold flap; use for and hysteresis. |
| "A static error count threshold" | Wrong at both peak and night; use ratios. |
| "The monitoring server can monitor itself" | A dead system reports nothing. Use a dead man's switch outside it. |
| "Put our 50K series in CloudWatch custom metrics" | About $7,000 a month, because it prices per series. |
Round 2 · Senior · "10M Points a Second and 100M Series"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 100M active series · 10M samples/s sustained, 30M/s peak · 5-year history · ingest and alerting 99.95%, survives an AZ loss
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We monitor 500 servers and ten services: about 50,000 series sampled every 10 seconds, 5,000 samples a second, kept 30 days. Two identical Prometheus servers in different AZs scrape every target, found through EC2 service discovery, and store samples in a time-series layout: each series gets an ID, its samples go into compressed, time-ordered chunks, new data lands in a WAL and an in-memory head, and every two hours the head becomes an immutable block. An inverted index of posting lists finds series by label. Rules run every 15 seconds with
fordurations and hysteresis, on ratios rather than counts. Both servers send alerts to a three-node Alertmanager cluster that deduplicates, groups by team and service, waits 30 seconds, and routes pages to the pager and email. A Watchdog alert feeds a CloudWatch dead man's switch outside the stack. About $200 a month. Open costs: everything lives on one server's disk and memory, there's nothing older than 30 days, and nothing stops a team from creating a million series."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: two full copies of a single-node time-series database, one alert router cluster, and a watcher outside.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Slow storage and range reads | Series IDs, compressed chunks, WAL, 2-hour blocks | A specialized database |
| 1.2 | Find series by labels | Inverted index, intersected posting lists | Memory grows with series count |
| 1.3 | Pull or push | Pull with service discovery | One scraper's reach |
| 1.4 | Flapping alerts | for, hysteresis, ratios | Slower alerts |
| 1.5 | Many pages for one outage | Alertmanager dedupe, group, route | 30 s wait |
| 1.6 | Silent monitor | Dead man's switch outside the stack | A second system |
Open costs: one server's memory and disk hold everything; no long-range history; no defense against series explosions or late data; an outage still produces many alerts across teams.
R2.1 The Scope Raise
Interviewer: "It worked, and now every team wants it. It becomes the company's monitoring platform: every service, container and batch job. We expect 10 million samples a second, bursting to 30 million. Capacity planners want five years of history, and a five-year dashboard must still load in about a second. Last month a team added
user_idas a label on one metric and our test cluster fell over. Mobile apps and batch jobs send data late. In our last outage, 2,000 alerts fired. And the platform must survive losing an AZ."
A scope raise is not the end of scoping. We ask back, and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| At what interval, so how many series? | Still 10 seconds for almost everything. | 10M samples/s ÷ one sample per 10 s = 100 million active series. One machine can't hold that in memory; we shard (step 2.2), and every byte per sample now matters (step 2.1). |
| What does 30 million mean? | Catch-up bursts: after a network problem, agents resend what they buffered. | The same series, three times the samples for a while. A buffer between the edge and storage absorbs it (step 2.2). |
| Five years at what detail? | Detail for a month. After that, trends: five-minute points for half a year, hourly for five years. | Downsampled tiers with the right aggregates, and queries that pick a tier by range (step 2.3). |
| Who can add labels? | Every team, through their own code. We can't review every change. | Limits enforced by the platform at ingest, per team, and alerts when a team's series count jumps (step 2.4). |
| How late is "late"? | Seconds for most; mobile and batch data can be hours old. | A short out-of-order window in memory, and a separate backfill path for older data (step 2.5). |
| What were the 2,000 alerts? | One AZ had network trouble. Every service in it alerted, each with several rules, all paging. | Grouping, inhibition and symptom-based paging (step 2.6). |
| What must survive an AZ loss? | Ingest, alerting and recent queries, with no acknowledged data lost. | The in-memory tier must be replicated or rebuildable (step 2.2). Recent queries: P95 < 50 ms; 30-day: P95 < 500 ms. |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Series | ~50K | 100M active |
| Samples/s | 5,000 | 10M sustained, 30M peak |
| Collection | One scraper pulls everything | Agents scrape locally and push to the center |
| Retention | 30 days raw | 30 days raw; 180 days at 5 min; 5 years at 1 hour |
| Tenancy | One ops team | Hundreds of teams, each with limits |
| Late data | Rare | Seconds to hours late |
| Survives | A server | An AZ, with no acknowledged sample lost |
| Availability | 99.9% alerting | 99.95% ingest and alerting |
| Queries | < 1 s | Recent P95 < 50 ms; 30-day P95 < 500 ms |
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | What breaks at the new scope |
|---|---|
| One Prometheus holds everything | 100M series need about 830 GB of memory with a real engine (R2.6), and one process can't append 10M samples a second. |
| One scraper for all targets | Millions of containers, many short-lived, across many networks. One scraper can't reach them all every 10 s, and a scrape that runs late shifts every timestamp. |
| Raw data kept 30 days, nothing longer | Five years of raw 10-second data is 2.2 PB even compressed, and a five-year graph would decode about 16 million points per series. |
| Trusting labels | One user_id label can add tens of millions of series in an hour and exhaust memory on every node. |
| "Late samples are rejected" | Mobile and batch data is lost, silently. |
| Alerts routed per team | A shared cause (an AZ) makes every team's alerts fire at once: 2,000 pages. |
| Two full copies | Two copies of 100M series doubles an already large memory bill, and neither copy is sharded. |
We fix them in this order: bytes per sample (2.1), sharding and replication (2.2), long ranges (2.3), cardinality (2.4), late data (2.5) and alert storms (2.6).
R2.3 New Requirements and API Additions
1. Batch ingest. Agents on every host (and in every Kubernetes pod network) scrape their local targets and push batches to the platform, using the Prometheus remote write protocol: a protobuf message, compressed with Snappy, sent by HTTP POST. It carries many series, each with its labels once and many samples.
httpPOST /api/v1/push HTTP/1.1 Host: ingest.metrics.internal.example.com Authorization: Bearer <team token> Content-Type: application/x-protobuf Content-Encoding: snappy X-Prometheus-Remote-Write-Version: 0.1.0
httpHTTP/1.1 200 OK
A 200 means durably accepted: the batch is in the stream (step 2.2). A 429 means the team is over its rate limit (retry later); a 400 means the batch was rejected (bad labels, timestamps too far in the future), and the body says why.
2. A query with a resolution hint. Same PromQL API as Round 1, plus a hint:
httpGET /api/v1/query_range?query=sum(rate(http_requests_total{service="checkout"}[1h]))&start=1632744000&end=1790510400&step=86400&resolution=auto HTTP/1.1
resolution is auto (the default: pick the coarsest tier that still gives at least one point per step), or raw, 5m or 1h to force a tier. The response says which tier answered, in a resolution_used field, so a graph can label itself "hourly data".
3. Per-team limits, set by the platform team and readable by each team:
json{ "team": "payments", "max_active_series": 2000000, "max_series_per_metric": 200000, "ingestion_rate_samples_per_s": 200000, "ingestion_burst_samples": 2000000, "max_label_names_per_series": 30, "max_label_value_length": 1024, "drop_labels": ["user_id", "session_id", "request_id", "email"], "out_of_order_window_s": 600 }
4. Silences and inhibition. A silence mutes matching alerts for a fixed time (planned maintenance). It is created through the Alertmanager API:
json{ "matchers": [ { "name": "cluster", "value": "payments-db-2", "isRegex": false, "isEqual": true } ], "startsAt": "2026-09-27T22:00:00Z", "endsAt": "2026-09-28T00:00:00Z", "createdBy": "maria@example.com", "comment": "Planned failover drill, CHG-4821" }
An inhibition rule mutes alerts while another alert is firing (step 2.6):
yamlinhibit_rules: - source_matchers: ['alertname="AvailabilityZoneImpaired"'] target_matchers: ['severity="cause"'] equal: ['availability_zone']
R2.4 Design Evolution: Scaling the Store, Taming the Alerts
Step 2.1: Raw Storage Is 16 Bytes Per Sample, 10 Million Times a Second
The problem: a raw sample is a 64-bit timestamp and a 64-bit float: 16 bytes. At 10M samples a second that's 160 MB/s, 13.8 TB a day. Just the last two hours, which we want in memory for fast queries and alerting, would be 1.15 TB. What would you do?
Step 2.2: One Node Can't Take 10 Million Samples a Second
The problem: 100M series, 10M samples a second, bursts to 30M. One process can't append that, and one machine can't hold the head in memory. If we spread series over many machines, what happens to the last two hours of data on a machine that dies? In Round 1 that data lived only in its memory and local WAL. What would you do?
Synthesizing vector architecture diagram...
Two consumer groups read one replicated stream, so the head exists twice and the stream holds every acknowledged sample for 24 hours.
Primitive: Message Queues vs Event Streams
Synthesizing vector architecture diagram...
A lost ingester costs a replica, never data: its twin answers while the replacement replays the stream from the last checkpoint.
Step 2.3: A Five-Year Graph Scans Billions of Samples
The problem: a capacity planner opens "requests per second, all services, last 5 years". At 10 s per sample that's 5 × 365 × 8,640 ≈ 15.8M samples per series, times thousands of series. And keeping five years of raw data would be 1.18 TB × 1,826 days ≈ 2.2 PB.
What would you do?
Synthesizing vector architecture diagram...
Every range reads the coarsest tier that still gives one point per step.
Step 2.4: Someone Put user_id in a Label
The problem: a team adds user_id to http_requests_total "to debug one customer". The metric had 30 series per service (routes × status codes). With 2 million active users, it now has up to 60 million. Within an hour, ingester memory climbs past its limit on every node.
What would you do?
Synthesizing vector architecture diagram...
Cheap checks reject bad batches before the stream; the exact series limit is enforced where the series live.
Step 2.5: Late and Out-of-Order Samples Are Rejected
The problem: a phone app buffers metrics while offline and sends them 40 minutes later. A nightly batch job reports its counters when it finishes, stamped with the times they were measured. A host reconnects after a network blip and resends the last 3 minutes. Our heads reject every sample older than the newest one they already have for that series, because a Gorilla chunk can only be appended to. What would you do?
Step 2.6: One Outage Produced 2,000 Pages
The problem: AZ b's network degrades for 20 minutes. 150 services have instances there. Each has 10 to 15 alert rules (latency, errors, saturation, pool exhaustion), and most fire: about 2,000 alerts. They are grouped by team and service, so about 150 groups page 60 on-call engineers, while the one person who can act (the network on-call) is buried under the same flood. What would you do?
Synthesizing vector architecture diagram...
Symptoms page, causes explain, and the shared cause pages the one team that can fix it.
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | 16 bytes per sample | Gorilla: delta of deltas, XOR; ~1.37 B/sample on the paper's workload | CPU; append-only chunks |
| 2.2 | One node can't take it | Hash-sharded ingesters behind a Kinesis stream; two in-memory replicas; checkpoint after upload | Two copies of the head; a stream |
| 2.3 | Five-year graphs | 5-minute and 1-hour rollups of min/max/sum/count; tier chosen by step; histograms for percentiles | Rollup jobs; approximations |
| 2.4 | user_id in a label | Gateway label rules and rate limits; per-team series limits enforced at ingesters; cardinality alerts | Dropped new series for offenders |
| 2.5 | Late data | 10-minute out-of-order window; backfill path to blocks | Late data visible within an hour |
| 2.6 | 2,000 pages | Symptom paging, inhibition, grouping by cause, silences | Tuning; risk of hiding a problem |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
Writes flow left to right into a replicated stream; two ingester groups hold the recent hours; blocks in S3 hold everything older; one read path serves dashboards and rule evaluation alike.
Why these pieces:
- Agents at the edge. Each host (or Kubernetes node) runs an agent that scrapes its local targets every 10 s and pushes batches to us. This keeps what pull was good at (the agent knows its targets and records
up) while no central scraper has to reach millions of endpoints. The agent keeps a local buffer on disk and retries with backoff, so a short outage on our side delays data rather than losing it. - An internal Network Load Balancer spreads long-lived HTTPS connections from agents over the gateways. Teams in other accounts reach it through PrivateLink.
- Ingest gateways are stateless: authenticate the team, apply label rules and rate limits, hash, route by age, and write to Kinesis. They scale with CPU.
- Kinesis Data Streams is the durable buffer and write-ahead log (step 2.2). Provisioned with 600 shards (R2.6).
- Ingesters, 25 per group, two groups (A in AZ a, B in AZ b), each holding every series' last few hours, plus local disk for building blocks.
- S3 holds all blocks: raw, 5-minute and 1-hour. S3 is designed for 99.999999999% (11 nines) object durability; that's a property of the stored blocks, not a promise about our whole pipeline.
- The compactor merges the two replicas' blocks into one, merges 2-hour blocks into 24-hour blocks (fewer, larger objects), builds rollups, merges backfilled blocks, and deletes data past each tier's retention.
- Store gateways serve queries on blocks in S3. They keep a small index header of each block on local disk and fetch only the chunks a query needs, with a chunk cache.
- The query frontend splits long queries into one-day pieces run in parallel, caches the results of past days (keyed with a cache generation that backfill bumps), and picks the tier (step 2.3).
- Rule evaluators run every team's rules through the same queriers, sharded so each rule group is evaluated by exactly one evaluator at a time, and send alerts to a three-node Alertmanager cluster (one per AZ).
- ElastiCache for Valkey holds query results and chunks.
The data model
A block is a directory of objects in S3 under blocks/<team>/<block-id>/:
| Object | Contents |
|---|---|
meta.json | Time range, resolution (raw, 5m, 1h), source (ingester A/B, compactor, backfill), stats |
index | Symbol table (every label string once), series with their label references and chunk offsets, posting lists |
chunks/000001 … | Compressed chunks, concatenated; one file up to 512 MB |
Keeping blocks per team makes retention, deletion and cost attribution a matter of prefixes. Round 3 turns teams into paying tenants.
The ingester's lease and checkpoint table (DynamoDB, managed by the Kinesis Client Library):
| Attribute | Example | Notes |
|---|---|---|
leaseKey (partition key) | shardId-000000000168 | One item per shard per consumer group (each group has its own table) |
leaseOwner | ingester-a-7 | Who reads this shard now |
checkpoint | a Kinesis sequence number | Everything before it is in an uploaded block |
leaseCounter | 918273 | Incremented by the owner as a heartbeat; a worker that sees it stop changing may take the lease |
The lease is decided by a counter the owner keeps incrementing, not by comparing clocks, which fits DynamoDB: its condition expressions have no server-side "now", so the library compares counters it has seen, on its own schedule.
Trace 1: ingest
Synthesizing vector architecture diagram...
The agent's 200 means the stream has the samples; each ingester group builds its own copy of the head from the same records.
Trace 2: a 30-day query
- Grafana asks for
sum by (service) (rate(http_requests_total{team="payments"}[1h]))over 30 days at a 15-minute step. - The query frontend picks the 5-minute tier (a step of at least 5 minutes but under an hour), splits the range into 30 one-day pieces, and finds 29 of them in the results cache. A cached day can go stale when the backfill path (step 2.5) adds late samples to it, so every backfilled block bumps a per-team cache generation for the days it covers; the generation is part of the cache key, and affected days are recomputed.
- Only today's piece runs: a querier asks the store gateways for today's older hours and both ingester groups for the last few hours, merges the replicas, and computes.
- The frontend stitches 29 cached days and 1 fresh one and returns 2,880 points per service.
Trace 3: an alert storm reduced to pages
- 14:02: packet loss in AZ b. Latency and error rules start going pending across 150 services.
- 14:04:
AvailabilityZoneImpaired{availability_zone="b"}(for: 2m) fires and pages the network on-call. - 14:05 onward: ~1,970
severity="cause"alerts fire and are inhibited: recorded, not sent. - ~30 symptom alerts fire; Alertmanager groups them by
alertnameandavailability_zone, and each affected team gets one page naming AZ b. - 14:24: the network recovers; the source alert resolves, the inhibition lifts, and everything still firing is sent normally. Nothing is.
R2.6 Numbers and Cost
Targets
| Quality | Target | Why this number |
|---|---|---|
| Ingest and alerting availability | 99.95% (21.9 min a month) | Survives an AZ; one level above Round 1 |
| Acknowledged data | 0 samples lost after 200, through any single node or AZ failure | The stream holds 24 h in 3 AZs |
| Recent queries (last 3 h) | P95 < 50 ms | From the scope raise |
| 30-day queries | P95 < 500 ms | From the scope raise |
| Alert freshness | Page within 90 s after for elapses | 10 s scrape + ~5 s ingest + 30 s query offset (R2.8) + 15 s evaluation interval + 30 s group wait |
How we measure the first line matters. Kinesis Data Streams' own SLA is 99.9% a month; an SLA is a credit commitment, not a measurement, but it's a warning not to claim more for a path that runs through one stream. Agents buffer and retry, so a short stream problem delays samples instead of losing them; we count ingest as available when samples land within 5 minutes of being sent.
We don't promise "five nines" (99.999%, which a common textbook design claims) for this round: one region can't honestly offer it, and we haven't built anything that survives a region. Round 3 raises the target with the design to back it.
Rates and bytes
| Item | Math | Result |
|---|---|---|
| Active series | 10M samples/s × 10 s | 100M |
| Peak raw ingress | 30M × 64 B (a common textbook estimate of a sample's in-memory size with its label references) | 1.92 GB/s = 15.36 Gbps |
| Sustained raw | 10M × 64 B | 640 MB/s = 5.12 Gbps |
| On the stream | We plan 16 B per sample after batching and compression (labels once per series per batch); an assumption to measure | 160 MB/s sustained, 480 MB/s peak |
| Compressed in blocks | 10M × 86,400 × 1.37 B | 1.18 TB/day |
| Last 2 hours | 10M × 7,200 × 1.37 B | 98.6 GB per replica |
| Label index | 100M × ~200 B | 20 GB per replica |
Memory: the data floor vs a real engine. The two lines above say the recent data is about 119 GB per replica. That's the floor: compressed bytes plus the index. A real engine spends far more per series: an open chunk per series, hash maps, per-series bookkeeping, label strings, query buffers. Grafana Mimir's capacity guide says to plan 2.5 GB of memory and 1 CPU core per 300,000 in-memory series (minimums, with 50% extra recommended):
| Item | Math | Result |
|---|---|---|
| Memory per replica | 100M ÷ 300K × 2.5 GB | 833 GB (≈ 8.3 KB per series) |
| CPU per replica | 100M ÷ 300K × 1 core | 333 cores |
| With headroom | memory × 1.5; CPU × 1.2 | 1,250 GB; 400 cores |
| Ingesters per group | 400 cores ÷ 16 vCPU = 25 × m7g.4xlarge (16 vCPU, 64 GiB) | 25 nodes, 1,600 GiB: memory fits too |
| Two groups | 2 × 25 | 50 ingesters |
| Series per ingester | 100M ÷ 25 | 4M |
A common textbook design sizes the hot tier from the data floor (three r6i.4xlarge, unreplicated). Planned from the engine's real per-series cost, the head needs about seven times the memory the floor suggests. This is the single most important number in the round, and it scales with series, not samples.
Stream
| Item | Math | Result |
|---|---|---|
| Shards for peak writes | 480 MB/s ÷ 1 MB/s per shard | 480 minimum |
| Provisioned | 480 ÷ 0.8 (stay under 80% at peak) | 600 shards |
| Records | 1,000 samples per record (~16 KB); peak 30,000 records/s ÷ 600 | 50 records/s per shard (limit 1,000) |
| Reads | 2 groups × up to 0.8 MB/s per shard at peak | 1.6 MB/s of the 2 MB/s a shard allows for all its consumers together |
| Shards per ingester | 600 ÷ 25 | 24 |
| Replay after a loss | ≤ 3 h × 0.27 MB/s average per shard ≈ 2.9 GB per shard. It reads at ~1.73 MB/s (what's left of 2 MB/s beside the other group), but new writes keep arriving at 0.27 MB/s, so it gains 1.47 MB/s: 2,880 MB ÷ 1.47 MB/s | ~33 minutes, if the ingester's CPU keeps up |
The replay estimate is why we keep two replicas: 33 minutes without recent data for 4% of all series would break the alerting target.
Gateways
| Item | Math | Result |
|---|---|---|
| CPU | Mimir's guide: 1 core per 25,000 samples/s for the equivalent component | 400 cores sustained; 1,200 at peak |
| Fleet | 40 × c7g.4xlarge (16 vCPU) = 640 cores | Handles 16M samples/s |
| Above 16M/s | Auto scaling adds gateways on CPU; until they're up (minutes), gateways return 429 and agents retry from their local buffers | Peak costs delay, not data |
We don't keep 1,200 cores idle for a burst that comes from agents catching up: agents buffer, so a few minutes of delay is the cheaper answer.
Query latency budgets. A recent query fans out to all 50 ingesters (both groups) and waits for all of them. The steps run in order, so they add up; the fan-out is parallel, so it costs the slowest of the 50 calls:
| Step | Budget |
|---|---|
| Frontend and results-cache check | 3 ms |
| Querier parses and plans | 2 ms |
| Fan-out to 50 ingesters, in parallel: the slowest reply | 25 ms |
| Merge replicas and compute | 10 ms |
| Network, both ways | 5 ms |
| Total | 45 ms |
The catch is the fan-out's tail. For the slowest of 50 calls to be under 25 ms 95% of the time, each call must be under 25 ms with probability 0.95^(1/50) ≈ 0.999: each ingester's P99.9, not its P95, must fit. That's a real requirement on ingester garbage collection and noisy neighbors, and a reason Round 3 shrinks each tenant's fan-out.
A 30-day query runs its one-day pieces in parallel too: with 29 of 30 days cached (trace 2), usually only one piece runs, taking about 20 ms for index lookups plus chunk fetches (a few ms from the cache; tens to a couple of hundred ms from S3 on a miss) and decoding. We budget 400 ms for the slowest piece and 50 ms for merging: under 500 ms, and cold-cache queries across all 30 days are what we load-test.
Storage tiers, recomputed (rollup values assumed at ~2 B each, since they compress worse than raw)
| Tier | Per series per day | × 100M series | Kept | Stored |
|---|---|---|---|---|
| Raw | 8,640 × 1.37 B = 11.8 KB | 1.18 TB/day | 30 days | 35.5 TB |
| 5-minute | 1,152 values × 2 B = 2.3 KB | 230 GB/day | 180 days | 41.5 TB |
| 1-hour | 96 values × 2 B = 192 B | 19.2 GB/day | 5 years (1,826 days) | 35.1 TB (7.0 TB per year) |
| Total at year 5 | 112 TB |
(Rounding 1.18 TB × 30 gives the 35.4 TB often quoted; unrounded it's 35.5.) The 5-minute tier is the biggest: half a year of four aggregates. Storage here is written per day and kept for a fixed number of days, so each tier's size is simply writes per day × days kept. Before the compactor merges them, both replicas' raw blocks exist for a day or so, adding about 1.2 TB.
Rough monthly cost (us-east-1 list prices, on-demand; check the AWS Pricing Calculator before quoting)
| Line | Math | ≈ Monthly |
|---|---|---|
| Ingesters | 50 × m7g.4xlarge × $0.6528/h × 730 h, + 5 TB gp3 for block building | $24,200 |
| Ingest gateways | 40 × c7g.4xlarge × $0.58/h × 730 h | $16,900 |
| Kinesis endpoints | Gateways and ingesters reach Kinesis through interface VPC endpoints, billed per GB in either direction: 420 TB written + 840 TB read = 1.26 PB; first 1 PB at $0.01/GB, the rest at $0.006 | $11,600 |
| Kinesis Data Streams | 600 shards × $0.015/h × 730 h = $6,570; 26.3B PUT payload units (25 KB each) × $0.014/M = $370 | $6,900 |
| Query path | 50 queriers (m7g.xlarge, 200 vCPU in all) × $0.1632/h × 730 h $5,960; 6 store gateways (m7g.2xlarge) + 2 TB gp3 $1,590; 3 frontends (m7g.xlarge) $360 | $7,900 |
| NLB | ~576 GB/hour processed ≈ 576 capacity units × $0.006/h × 730 h + hourly | $2,540 |
| S3 | Year 1: raw 35.5 + 5-min 41.5 + 1-hour 7.0 = 84 TB: 50 TB × $0.023 + 34 TB × $0.022 ≈ $1,900; requests ≈ $500 (estimate) | $2,400 |
| ElastiCache for Valkey | 6 × cache.r7g.xlarge at ~$0.35/h | $1,530 |
| Rule evaluators, compactor, Alertmanager | 6 × m7g.xlarge $715; 2 × m7g.2xlarge + 2 TB gp3 $640; 3 × m7g.large $180 | $1,540 |
| DynamoDB, CloudWatch | Lease tables; meta-monitoring | $350 |
| Total | ≈ $75,900 |
Synthesizing vector architecture diagram...
Memory for 100M series is the biggest line; storage, which people worry about first, is among the smallest.
That's about $760 per million active series per month. Four lessons:
- Series drive the bill. The ingesters are a third of it, and they're sized by series count. Every limit in step 2.4 is also a cost control.
- Storage is cheap once it's compressed and tiered: 84 TB for about $2,400 a month.
- The network path matters. Interface endpoints cost $11,600 a month here; putting the same traffic through a NAT gateway would cost $0.045/GB, about $57,000. The 16-bytes-per-sample assumption drives both lines, so we measure it first.
- Managed would cost far more at this size. Amazon Managed Service for Prometheus at 26.28 trillion samples a month: the first 2B at $0.90 per 10M, the next 250B at $0.35 per 10M and the rest at $0.16 per 10M (the published tiers we checked) come to about $425,000 a month before queries, and its default quotas (1.67M samples/s per workspace) would need raising. At Round 1's size it was the easy choice; here building saves several times its cost in engineers.
R2.7 Trade-Offs
How to replicate the hot tier
| 1 in-memory copy + the stream | 2 copies + the stream (chosen) | 3 copies, written synchronously (Mimir classic) | |
|---|---|---|---|
| Data lost on a node loss | None (the stream has it) | None | None (2 of 3 hold every sample) |
| Recent data during recovery | Unavailable for that node's series for ~30 min | Served by the twin | Served by the other two |
| Memory | 1× | 2× | 3× |
| Write path | Gateway → stream | Gateway → stream | Gateway → 3 ingesters, waits for 2 |
| Burst absorption | The stream | The stream | Ingesters must keep up, or writes fail |
We pay 2× memory because 30 minutes of blind alerting on 4% of series is not acceptable, and we skip the third copy because the stream already guarantees durability.
How many aggregates to keep in rollups
| 1 aggregate (average only) | 4: min, max, sum, count (chosen) | |
|---|---|---|
| Storage | 4× smaller | Baseline |
| Spikes in long graphs | Averaged away | Visible in max |
| Correct averages across series | No: an average of averages is wrong when counts differ | Yes: sum ÷ count |
Kinesis vs a Kafka cluster (Amazon MSK) for the stream
| Kinesis provisioned (chosen) | Amazon MSK | |
|---|---|---|
| Operations | No brokers; resharding by API | Brokers, partitions, storage to size |
| Cost at our size | Shards + endpoint traffic (~$18,500) | Brokers and storage; often cheaper at high throughput, after sizing |
| Consumer limits | 2 MB/s and 5 reads/s per shard, shared by consumers | Set by broker capacity |
| Retention | 24 h default, up to 365 days at extra cost | Configurable |
Build on open source vs managed
| Our own stack on Mimir/Thanos/Cortex-style components (chosen) | Managed Prometheus + Managed Grafana | |
|---|---|---|
| Cost here | ≈ $76K/month | ≈ $425K/month in ingestion alone |
| Team | A platform team to run it | Far less to run |
| Control | Our limits, tiers, rollups and alert routing | The service's quotas and features |
R2.8 Failure Modes
| Trigger | What you'd see | How the design responds | Drill |
|---|---|---|---|
| Stream backlog and memory pressure | Ingester consumer lag grows; memory climbs during catch-up. | Each ingester reads with a bounded in-flight budget and stops reading above 80% memory, so it lags instead of crashing; the stream holds 24 h. An alarm on the stream's iterator age (> 60 s) pages; the fix is more ingesters (fewer shards each). | – |
| An ingester dies | One of 50 nodes gone. | Its twin serves queries and rules; a replacement replays from the checkpoint in ~30 min (step 2.2). | – |
| A cardinality spike | One team's active series double in an hour. | Gateway label rules and the per-team local series limits cap it (step 2.4); the platform on-call gets the cardinality alert with the metric and label. | – |
| The compactor falls behind | Thousands of small 2-hour blocks, both replicas unmerged, rollups missing for recent days. | Ingestion doesn't slow: ingesters upload blocks without waiting for compaction. What degrades is reads: long queries open many more blocks, and long ranges fall back to raw data. We alarm on "oldest uncompacted block age" > 6 h and scale compactors (they split work by team and time range). See the note below. | TSDB write stall |
| Clock skew produces false alerts | A host's clock runs 15 s behind; its latest samples look old, so a rule over the last minute sees a dip or missing data. | Rules evaluate with a 30 s query offset (Prometheus's query_offset for rule groups): at 14:00:00 they evaluate as of 13:59:30, so samples a few seconds late have landed. Agents sync time with NTP (Amazon Time Sync Service on EC2). The gateway measures arrival − timestamp per host and alerts on hosts more than 30 s off. | – |
| An AZ is lost | Gateways, ingester group A (or B), one Alertmanager gone. | Kinesis and S3 are unaffected. The other ingester group serves everything; group A is rebuilt in another AZ from the stream. The remaining 27 gateways handle ~10.8M samples/s while auto scaling replaces the rest. | – |
Go deeper: why an LSM engine stalls writes when compaction falls behind, and why we don't. In an LSM-tree engine (RocksDB, Cassandra), writes go to a WAL and an in-memory table; full tables are flushed to small sorted files at level 0, and compaction merges them into larger levels. Every level-0 file must be checked on reads, so if compaction can't keep up with the flush rate, level-0 files pile up. To keep reads bounded and memory from running out, the engine deliberately slows, then stops, foreground writes when level-0 files pass a threshold: a write stall. Compaction itself rewrites data many times over (write amplification), so under a burst it competes with flushes for disk bandwidth. Our design avoids the stall by not making flushes wait for compaction: an ingester's 2-hour block goes straight to S3 as a finished, immutable object, and the compactor works on S3 later. The price is paid on reads instead, which is why we alarm on compaction lag. Why not a B-tree engine instead? A B-tree gives predictable reads with no background merging, but it updates pages in place: 10M samples a second spread over 100M series would be random writes all over the tree, far beyond what disks sustain, where an append-only layout writes sequentially. Time-series engines accept background merging to get sequential writes (step 1.1).
R2.9 Production Gotchas
1. One central scraper for everything
- Symptom: scrapes run late, timestamps drift, targets flap to
up == 0because a scrape timed out, not because the target died. - Cause: one process connecting to millions of endpoints every 10 seconds.
- Fix: agents scrape locally and push (R2.5); the central platform never connects to targets.
2. Querying raw data for months
- Symptom: a 90-day dashboard times out and slows every other query.
- Cause: 777,600 raw samples per series for 90 days, decoded for a graph 800 pixels wide.
- Fix: tiers chosen by step (step 2.3); a per-query limit on samples read.
3. No hysteresis or for
- Symptom: pages that resolve a minute later, dozens a night.
- Cause: firing on every evaluation where the condition is true.
- Fix:
for, separate clear thresholds, symptom-based paging (steps 1.4, 2.6).
4. Unbounded label values
- Symptom: memory grows steadily; one day every ingester runs out.
- Cause: IDs, emails, URLs with query strings, or error messages as label values.
- Fix: platform-wide forbidden labels, value-length limits, per-team series limits, cardinality alerts (step 2.4). Put per-request detail in logs or traces, not metrics.
5. Averaging percentiles
- Symptom: the "p99" on the yearly graph is far lower than any incident review remembers.
- Cause: averaging per-host or per-hour p99 values.
- Fix: record histograms, sum their buckets, compute the percentile last (step 2.3).
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | The stream as a replicated write-ahead log; two in-memory replicas in two AZs; checkpoint only after upload; backpressure by lagging, not crashing; per-team limits so one team can't break others. REL 4 · REL 5 · REL 10 · REL 11 |
| Performance Efficiency | Gorilla chunks; tiers chosen by query step; results caching of past days, invalidated per day by a backfill cache generation; fan-out tail budgets. PERF 3 · PERF 5 |
| Security | Per-team tokens at the gateway; least-privilege roles per component (below); TLS from agents; forbidden labels keep personal data such as emails out of metrics. SEC 3 · SEC 7 · SEC 9 |
| Cost Optimization | Series identified as the cost driver; ≈ $76K/month derived; interface endpoints instead of a NAT gateway; managed option priced. COST 5 · COST 6 · COST 8 |
| Operational Excellence | Alarms with first actions (below); symptom-based paging; inhibition; silences with authors and end times. OPS 8 · OPS 10 |
| Sustainability | Light this round: downsampling and retention per tier; Graviton instances; gateways scale with load. SUS 4 · SUS 5 |
Security in detail. Each team's token maps to a team ID at the gateway; the team ID is never taken from a request header. Gateways may only put records to the stream and write to the late-data prefix. Ingesters may read the stream, write to the lease table and write blocks under blocks/; they can't delete. Only the compactor may delete blocks. Store gateways and queriers are read-only on S3. Blocks are encrypted with SSE-KMS, and all traffic uses TLS.
Operations in detail.
| Alarm | Threshold | Severity | First action |
|---|---|---|---|
| Stream iterator age (consumer lag) | > 60 s for 5 min | P1 | Which ingester group? Memory stops? Add ingesters |
| Samples discarded, by reason | > 1% of a team's samples | P2 (team) | Show the team the metric and label |
| Active series, per team | 2× in 24 h | P2 (platform) | Find the new label; apply a drop rule |
| Rule evaluation missed iterations | > 0 for 10 min | P1 | Evaluator overloaded or queries slow? |
| Notifications failed | > 0 for 5 min | P1 (by email) | Integration down? Check the fallback path |
| Oldest uncompacted block | > 6 h | P3 | Scale compactors |
| Watchdog heartbeat missing | 5 min | P1 (CloudWatch → SNS) | The whole alert path is suspect |
R2.11 Round 2 Rubric and Follow-Ups
What a strong senior (L6) answer adds over L5
- Encodes a few samples with Gorilla by hand, and cites 1.37 bytes as one paper's workload, not a law.
- Knows that series, not samples, drive memory and cost, and sizes memory from the engine's per-series cost, not from compressed bytes.
- Keeps the hot tier durable and available through node and AZ loss, and explains what a node loss does to the last two hours.
- Recomputes rollup savings with the real interval and the number of aggregates.
- Defends against cardinality at the gateway and where the series live.
- Handles late data without corrupting chunks, and turns an alert storm into a few pages.
Follow-up questions
-
"Why not just add a third ingester group instead of the stream?" Answer: three groups written synchronously (a quorum of two) is a valid design and removes the stream bill, but then a burst must be absorbed by the ingesters themselves, and a slow ingester slows writes. With the stream, the gateway's
200depends only on Kinesis, and ingesters can lag, restart or be replaced without affecting agents. We'd reconsider if the stream's cost (about $18,500 a month with its endpoints) grew faster than its value. -
"A team says its alerts didn't fire during last night's incident. Where do you look?" Answer: in order: was the data there (discarded-samples counter for their team, consumer lag); did the rule evaluate (missed iterations, evaluation errors, and whether a label change made its selector match nothing, which looks like "all fine"); did it fire but get muted (inhibition or a silence); was it sent (notification failures). "No data" and "no problem" look the same to a threshold rule, which is why paging rules on important services also carry an
absent()companion. -
"The 16 bytes per sample on the stream turns out to be 40. What changes?" Answer: stream bytes grow 2.5×: 1,200 MB/s at peak needs 1,500 shards, which adds about $9,900 of shard-hours, and the endpoint traffic grows to about 3.15 PB (about $22,900 instead of $11,600, since volume beyond the first PB is cheaper). Ingesters, S3 and memory don't change: they depend on series and on the compressed format. We'd first try larger batches (labels sent once per series per batch) before paying for it.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Gorilla always gets 1.37 bytes per sample" | That's the paper's workload; Prometheus says 1 to 2 bytes. |
| "1-minute rollups cut storage 60×" | Only for 1-second data with one aggregate. At 10 s with four aggregates it's 1.5×. |
| "The hot tier doesn't need replication; it's only two hours" | Those two hours are what alerting reads. A lost node blinds its series. |
| "Size ingester memory from compressed bytes" | Engines spend kilobytes per series; plan from the engine's per-series cost. |
| "Limit samples per second to stop cardinality explosions" | New series, not samples, exhaust memory. |
| "Average the p99s" | Percentiles don't average; sum histogram buckets. |
| "Raise thresholds to stop alert storms" | Page on symptoms, inhibit causes, group by the shared cause. |
Round 3 · Architect · "Monitoring as a Product: Tenants, Regions, SLOs"
~45 min · Principal (L7) · 3 home regions of cells + 3 alerting standby regions · ~1B active series · ~100M samples/s · 99.99% alerting per region, through a region outage · one tenant can't slow another
R3.0 Where We Left Off
What the candidate says in the first 60 seconds of Round 3, and everything you need if you start here.
Round 2 in 60 seconds. "We run the company's monitoring platform: 100 million active series, 10 million samples a second, bursting to 30 million, in one region. Agents scrape locally and push with remote write. Stateless gateways enforce each team's label rules and rate limits, hash each series, and write to a 600-shard Kinesis stream, which is our replicated write-ahead log; a 200 means the stream has the data. Two groups of 25 ingesters, in two AZs, each read every shard and hold the last few hours in Gorilla-compressed memory, about 1.37 bytes a sample on the paper's workload but about 8 KB of memory per series in a real engine, which makes memory the biggest line of the bill. Every two hours they upload blocks to S3 and only then checkpoint. A compactor merges replicas and builds 5-minute and 1-hour rollups of min, max, sum and count; queries pick a tier by step. Per-team series limits are enforced at the ingesters; late data within 10 minutes goes into out-of-order chunks, older data through a backfill path. Symptom-based paging, inhibition and grouping turned 2,000 alerts into a handful of pages. About $76K a month, $760 per million active series. Open costs: one region, limits built for friendly teams, no way to bill anyone, and alerts based on thresholds that page too much or too late."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: a stream in the middle, two copies of the recent hours, blocks in S3, one read path for dashboards and rules.
Rounds 1–2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Slow storage | Series IDs, compressed chunks, WAL, blocks | A specialized database |
| 1.2 | Find series by label | Inverted index | Memory per series |
| 1.3 | Pull or push | Pull, then agents that push | Discovery; agents to run |
| 1.4 | Flapping | for, hysteresis, ratios | Slower alerts |
| 1.5 | Many pages | Dedupe, group, route | 30 s wait |
| 1.6 | Silent monitor | Dead man's switch outside | A second system |
| 2.1 | 16 B a sample | Gorilla | CPU; append-only chunks |
| 2.2 | One node too small | Stream + sharded, replicated ingesters | 2× memory; a stream |
| 2.3 | Five-year graphs | 5-min and 1-h rollups, 4 aggregates | Rollup jobs |
| 2.4 | user_id label | Gateway rules; series limits at ingesters | Dropped new series |
| 2.5 | Late data | 10-min window; backfill | Delayed visibility |
| 2.6 | 2,000 pages | Symptoms, inhibition, grouping | Tuning effort |
Open costs: a single region, so a region outage silences every alert; limits and query capacity shared by everyone; no idea what any team costs; threshold alerts that page on noise and miss slow burns.
R3.1 The Scope Raise
Interviewer: "The platform is good enough that we're going to sell it. Other companies will send us their metrics and rely on our alerts. Tenants range from a startup with a few hundred series to an enterprise with 10 million. One customer's giant query must never slow another's dashboard. Customers told us plainly: the moment they need alerts most is when a cloud region is having a bad day, so our alerting has to keep working through a region outage, including one of ours. They also want alerts based on SLOs, not just thresholds. And finance wants to know what each tenant costs us, so we can price it."
We ask back, and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How much in total, and how uneven? | Plan for about 1 billion active series and 100M samples/s. Thousands of tenants; the biggest has 10M series today and might grow 10×. | We can't run one giant cluster; we run cells of Round 2's size and place tenants in them, with shuffle sharding inside each cell (step 3.1). |
| Where are the tenants? | By volume: 50% North America, 30% Europe, 20% Asia-Pacific. Customers pick a home region, and want their data to stay in that part of the world. | Three home regions: us-east-1, eu-west-1, ap-southeast-1. Any standby must be in the same geography (step 3.2). |
| In a region outage, what must keep working? | Alerts, above all. Dashboards can be degraded for a while. | A second region per home region that runs alerting only, on the series alert rules need (step 3.2). |
| What's wrong with threshold alerts? | Customers get paged for brief blips, and don't get paged when errors are slightly high for days. | SLO burn-rate alerts as a product feature (step 3.3). |
| What do heavy queries look like? | Some tenants graph 90 days across millions of series, from dashboards that refresh every 30 s. | A query frontend with per-tenant queues, limits and caching, and recording rules for heavy dashboards (step 3.4). |
| How do we bill? | By active series and samples ingested, plus query usage and retention above the plan. It must be accurate, and explainable to the customer. | A metering pipeline per tenant, idempotent and reconciled (step 3.5). |
| Should we build this, or run it on something managed? | Leadership asks you to justify either answer. | A build-versus-buy answer with numbers (step 3.6). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Customers | Internal teams | Paying tenants, 100 to 10M series each |
| Active series | 100M | ~1B |
| Samples/s | 10M | ~100M, in 3 home regions (50/30/20M) |
| Footprint | 1 region, 3 AZs | 3 home regions of cells + 3 alerting standby regions |
| Isolation | Per-team limits | Per-tenant limits; cells; shuffle sharding; fair query queues |
| Alerting | Thresholds, symptom paging | + SLO burn-rate alerts; survives a region outage |
| Availability | 99.95% ingest and alerting | 99.99% alerting per region, through a region outage |
| Money | Internal cost | Per-tenant metering and billing; cost per million series |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| One cluster in one region | 1B series is ten Round 2 clusters. One cluster that big is one blast radius, and a region outage silences every tenant's alerts. |
| Every team's series spread over every ingester | One tenant's explosion or bad query touches every node, so it touches every tenant. |
| One shared query queue | A tenant's 90-day scan of millions of series queues ahead of everyone's dashboards. |
| Alerting in the same region as the data | During our region outage, nobody gets paged, which is exactly when customers need it. |
| Threshold rules | Page on blips; miss slow burns. Customers want "are we keeping our promise?", not "is CPU high?". |
| No usage records | We can't price, bill or explain a bill, or see which tenant is losing us money. |
R3.3 New Requirements and API Additions
1. Tenant identity on every call. Every write and query carries an API key. The gateway and the query frontend map it to a tenant ID; a tenant ID sent by the client in a header is ignored. Everything after that point (stream records, blocks in S3, cache keys, usage records) carries the tenant ID.
httpPOST /api/v1/push HTTP/1.1 Host: ingest.us-east-1.metrics.example.com Authorization: Bearer mk_live_7Hq... Content-Type: application/x-protobuf Content-Encoding: snappy X-Prometheus-Remote-Write-Version: 0.1.0
2. Per-tenant limits, set by the tenant's plan and overridable by us:
json{ "tenant_id": "t-4821", "plan": "business", "home_region": "us-east-1", "cell": "use1-cell-3", "max_active_series": 5000000, "ingestion_rate_samples_per_s": 500000, "ingestion_burst_samples": 5000000, "ingester_shard_size": 5, "querier_shard_size": 4, "max_samples_per_query": 50000000, "max_series_per_query": 100000, "max_query_length_days_raw": 32, "query_timeout_s": 120, "max_outstanding_queries": 100, "retention_days": { "raw": 30, "5m": 180, "1h": 1825 } }
3. An SLO definition (YAML), from which we generate recording and alerting rules (step 3.3):
yamlslo: name: checkout-availability tenant_id: t-4821 objective: 0.999 window: 30d sli: total: 'http_requests_total{service="checkout"}' errors: 'http_requests_total{service="checkout", code=~"5.."}' alerting: page: - { long: 1h, short: 5m, burn_rate: 14.4 } - { long: 6h, short: 30m, burn_rate: 6 } ticket: - { long: 3d, short: 6h, burn_rate: 1 } notify: { page: oncall-pager, ticket: team-tickets }
4. Usage export, one record per tenant, meter and hour:
json{ "tenant_id": "t-4821", "hour": "2026-09-27T14:00:00Z", "active_series_avg": 3814220, "samples_ingested": 1373119200, "samples_discarded": 0, "bytes_stored_gb": 2118.4, "query_samples_processed": 918442000123, "source_version": "meter-v7" }
R3.4 Design Evolution: Running Monitoring for Others
Step 3.1: One Tenant's Cardinality and Queries Hurt Everyone
The problem: tenant t-9 deploys a bug that doubles its series every hour, and runs a dashboard that scans 90 days of 2 million series every 30 seconds. In Round 2's design, its series live on every ingester and its queries run on every querier, so every tenant's dashboards slow down and some alerts evaluate late. What would you do?
Primitive: Circuit Breaker, Bulkhead & Fault Tolerance Patterns · Drill: One customer, one shard, one outage (the cross-tenant report question is answered in step 3.5)
Synthesizing vector architecture diagram...
t-9 can exhaust q07, q12, q30 and q41. t-4821 shares only q12, so three of its four queriers are untouched.
Step 3.2: Alerting Must Survive Our Own Region Outage
The problem: us-east-1 has a bad afternoon. Our cells there are down or unreachable. Our customers' own services in us-east-1 are failing too, and this is exactly the moment their on-call engineers need pages. But their alert rules, their data and their Alertmanagers are all in us-east-1. What would you do?
Primitive: Cloud Disaster Recovery & Multi-Region Active-Active
Synthesizing vector architecture diagram...
Two independent paths feed one Alertmanager cluster that spans the pair; either region alone can still page.
Step 3.3: Threshold Alerts Page Too Often, or Too Late
The problem: tenant t-4821 promises its customers that 99.9% of checkout requests succeed. Its rule "page if the 5xx ratio is above 5% for 5 minutes" pages for a 6-minute blip that used 1% of the month's allowance, and stays silent through three days at a 0.3% error ratio, which is three times the rate the promise allows and would use up the month's allowance in ten days. What would you do?
Step 3.4: Some Queries Scan Months Across Millions of Series
The problem: a tenant's executive dashboard has 20 panels, each over 90 days of 2 million series, refreshing every 30 seconds for 40 viewers. Each panel reads billions of samples. The queriers assigned to that tenant are pegged, and its own on-call can't load a graph during an incident. What would you do?
Synthesizing vector architecture diagram...
Most of a long dashboard comes from cache; what's left waits in the tenant's own queue and runs only on its own queriers.
Step 3.5: What Does Each Tenant Cost Us?
The problem: finance wants per-tenant revenue and cost. Today we know the whole bill (≈ $76K per cell) but not who uses it. One suggestion is to split the bill evenly across tenants. What would you do?
Synthesizing vector architecture diagram...
Each meter is written once per tenant per hour, so replays overwrite rather than double-count, and two independent meters check each other.
Step 3.6: Should We Build This at All?
The problem: leadership asks: AWS has CloudWatch, Amazon Managed Service for Prometheus and Amazon Managed Grafana; vendors sell hosted monitoring. Why are we running 10 cells ourselves? What would you do?
Further reading: the trade-off analysis playbook for how to present a build-versus-buy decision.
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Noisy tenants | Per-tenant limits; cells; shuffle-sharded ingesters and queriers | Placement logic; spare capacity |
| 3.2 | Region outage silences alerts | Standby region per home region, fed by agents directly; pair-spanning Alertmanager; per-tenant heartbeats | A standby cell; rules evaluated twice |
| 3.3 | Pages too often or too late | Multi-window, multi-burn-rate SLO alerts | Teaching SLOs; recording rules |
| 3.4 | Months across millions of series | Split, cache, fair per-tenant queues, limits, recording rules | Live newest 10 min; failed heavy queries |
| 3.5 | Cost per tenant | Idempotent hourly meters, reconciled, exported to S3 | A billing-grade pipeline |
| 3.6 | Build or buy | Cost per million series at each scale | Running it forever |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Each geography has a home region of cells and a standby region for alerting; configuration is written once and read locally everywhere; usage flows to one warehouse. The Europe and Asia-Pacific pairs have the same Alertmanager layout as North America (not drawn).
Why these pieces:
- Cells are Round 2's stack, unchanged, sized to 100M series. Growing means adding cells, never making one bigger, so our largest failure stays the size we have tested.
- Cell routers are thin: they read the tenant → cell table and forward. They hold no state that can't be reloaded.
- Standby cells hold only alert-referenced series for 8 hours, with no dashboards and no long-term storage.
- DynamoDB global tables carry tenants, limits, rules and SLOs, written in one control region (us-east-1) and read locally everywhere.
- Route 53 health checks with failover records send the customer-facing alert API (silences, acknowledgments) to the standby region when the home region is unhealthy. The records and health checks are set up in advance, because Route 53's data plane (answering queries and running health checks) keeps working when changing records might not.
- The outside watcher runs in a region that isn't part of the pair it watches, on CloudWatch and SNS.
Trace 1: a tenant query throttled
Synthesizing vector architecture diagram...
The noisy tenant waits in its own queue and on its own queriers; the other tenant's single query is taken on the next round-robin turn.
Trace 2: a region outage with alerts still firing
Synthesizing vector architecture diagram...
The standby never needed a failover: it was receiving and evaluating all along. Dashboards for us-east-1 tenants are down until 15:30; pages are not.
Trace 3: an SLO burn alert
- 09:00: tenant t-4821 deploys a bad config; 3% of checkout requests fail (burn rate 30).
- Recording rules update the 5-minute and 1-hour error ratios every minute.
- 09:29: the 1-hour ratio reaches
0.03 × 29 ÷ 60 ≈ 1.45%, above 1.44%, and the 5-minute ratio is 3%. The fast-burn page fires; about 2% of the month's budget is spent. - 09:40: the deploy is rolled back. The 5-minute ratio falls below 1.44% by 09:45, and the alert resolves, although the 1-hour ratio is still high.
- The page's annotation shows the budget left for the month: 40 minutes at 3% errors spent
0.03 × 40 ÷ 43.2 ≈ 2.8%of it, so 97.2% remains.
R3.6 Numbers and Cost
Targets
| Quality | Target | How we back it |
|---|---|---|
| Alerting availability, per region | 99.99% (4.4 min a month) | Two independent paths per geography (home and standby), each with its own stream; either can page |
| Alert delivery in a region outage | Pages continue with no failover step | The standby is fed by agents and evaluates continuously |
| Ingest | 99.99% of samples land within 5 minutes; none lost within the agents' buffer window | Stream in 3 AZs; agents buffer and retry |
| Noisy tenant | Other tenants' query P95 unchanged within 10% during a tenant's overload | Shuffle-sharded queriers, fair queues, limits |
| Billing | Gateway accepted = appended + discarded + backfill-routed, within 1% per tenant-hour | Reconciliation (step 3.5) |
Why two paths: each path depends on one region's stream, and Kinesis's SLA is 99.9% a month. Two paths that fail independently are unavailable together far more rarely than either alone. They aren't perfectly independent (the same code, the same agent, the same customer), which is why we still test region evacuation (R3.9) instead of trusting the multiplication.
Per region
| Region | Samples/s | Active series | Cells | Ingesters | Queriers (m7g.xlarge) | S3 after year 1 |
|---|---|---|---|---|---|---|
| us-east-1 | 50M | 500M | 5 | 250 | 250 | 5 × 84 TB = 420 TB |
| eu-west-1 | 30M | 300M | 3 | 150 | 150 | 252 TB |
| ap-southeast-1 | 20M | 200M | 2 | 100 | 100 | 168 TB |
| Home total | 100M | 1B | 10 | 500 | 500 | 840 TB |
| Standby (us-west-2, eu-central-1, ap-northeast-1) | 10M (10%) | 100M | 3 small | 50 | a few | none |
The per-cell numbers are Round 2's. Per-tenant retention plans change the storage column: a tenant on 13 months of 5-minute data costs more than one on 180 days, and the compactor's per-tenant bytes meter bills it. Blocks are stored per tenant, so a plan change is a retention setting on a prefix.
Standby sizing. The subset is 10% of samples (10M/s) and 10% of series (100M), about one Round 2 cell's worth spread over three regions: 50 ingesters (two groups), gateways for 10M samples/s, 600 shards in total, rule evaluators and Alertmanagers.
Forwarding vs dual sending (the alternative in step 3.2)
| Pair | Alert-subset bytes per month | Inter-Region price (list; check the corridor) | ≈ Monthly if forwarded |
|---|---|---|---|
| us-east-1 → us-west-2 | 5M samples/s × 16 B = 80 MB/s × 2.628M s = 210 TB | $0.02/GB | $4,200 |
| eu-west-1 → eu-central-1 | 48 MB/s → 126 TB | $0.02/GB | $2,500 |
| ap-southeast-1 → ap-northeast-1 | 32 MB/s → 84 TB | ~$0.09/GB | $7,600 |
| Total | 420 TB | $14,300 |
Asia-Pacific carries a fifth of the data and more than half of the transfer bill. With dual sending, none of this is charged to us.
Rough monthly cost (us-east-1 list prices used everywhere, which understates the Asia-Pacific and Frankfurt regions, where instances cost more; check the calculator)
| Line | Math | ≈ Monthly |
|---|---|---|
| Home cells | 10 × Round 2's compute and network (≈ $73,100 per cell: $75,900 without S3 and meta-monitoring); an upper bound, since endpoint traffic above 1 PB per region is cheaper | $731,000 |
| S3, year 1 | 10 × ≈ $2,400 | $24,000 |
| Standby cells | 50 ingesters $23,800; gateways for 10M/s $16,900; 600 shards $6,900; interface endpoints for 1.26 PB across three regions (each under 1 PB, so $0.01/GB) $12,600; NLBs $2,500; rules and Alertmanagers $1,500 | $64,200 |
| Control plane | Cell routers, global tables, usage pipeline, Athena, Route 53 health checks | $10,000 (estimate) |
| Meta-monitoring | Canaries, outside watchers, CloudWatch in third regions | $3,000 |
| Total | ≈ $832,000 |
That's about $830 per million active series per month, against $760 in Round 2. The extra $70 buys a standby for alerting, per-tenant isolation and billing. The home cells are nine-tenths of the bill, and in each cell the ingesters (memory for series) are the biggest line, so the lever that matters most is still cardinality: every tenant limit, forbidden label and cardinality alert is also a cost control. Steady fleets like ingesters are also the right place for Savings Plans, which lower their price in return for a commitment.
R3.7 Trade-Offs
Shuffle-shard size
| Small (3–4 workers per tenant) | Large (most of the cell) | |
|---|---|---|
| Blast radius of a noisy tenant | Its few workers | Everyone |
| A tenant's own capacity | Limited to its workers; big tenants need a bigger size | The whole cell |
| Recent-query fan-out | 6 ingester calls | 50 |
| Our rule | max(3, series ÷ 1M) ingesters per group; 4 queriers; dedicated cell above ~20M series | – |
Alerting in the same region vs a standby region
| Same region, 3 AZs | Standby region fed by agents (chosen) | Standby fed by forwarding | |
|---|---|---|---|
| Survives our region outage | No | Yes, with warm history | Only if the home gateways are up |
| Extra cost | – | ≈ $64K/month standby | Standby + ≈ $14K/month transfer |
| Customer impact | – | ~10% more bytes sent; a second destination in the agent | None |
| Data location | One region | Same geography | Same geography |
Burn-rate windows
| Fast windows only (1 h / 5 min) | Fast + slow (chosen) | Thresholds only | |
|---|---|---|---|
| Full outage | Condition true within a minute; page in about 2–3 minutes | Same | Depends on for |
| Slow burn (0.3% for days) | Missed | Ticket within a day | Missed or noisy |
| Pages for blips | Few | Few | Many |
Build vs buy (COST 11): step 3.6. At 1B series, building saves about $3.4M a month in list prices against managed ingestion alone, and it's the product we sell; at 50K series, managed would have been the sensible default.
Closing the loop. Round 1 asked "how do we store numbers and raise an alarm?" and answered with a time-series layout, for and hysteresis, grouping, and a dead man's switch. Round 2 asked "how do we do it for 100M series?" and answered with compression, a replicated stream-fed head, tiers and cardinality limits. Round 3 asked "how do we do it for strangers, through our own outages, and get paid?" and answered with cells, shuffle sharding, an alerting standby fed straight from the agents, burn-rate alerts and metering. The dead man's switch from Round 1 never went away: it became a canary tenant per cell and a heartbeat per tenant, watched from outside.
R3.8 Failure Modes
| Trigger | What you'd see | How the design responds | Drill |
|---|---|---|---|
| A home region outage | us-east-1 cells unreachable; canary tenants there go silent. REL 13 | Standby cells keep receiving alert series from agents and keep evaluating; the pair's Alertmanagers in us-west-2 send. Route 53 failover sends the alert API (silences, acknowledgments) to us-west-2. Dashboards for those tenants are down; we post status. On recovery, agents flush their buffers; samples older than 10 minutes go through backfill. | – |
| A runaway tenant | One tenant's series and samples triple in an hour; its ingesters' memory rises. | Its local series limits drop new series; its rate limit returns 429; only its 3–10 ingesters per group carry it. We notify the tenant with the offending metric and label; if needed, we add a drop rule for that label in its limits. | – |
| The query frontend is overloaded | Frontend CPU high; every tenant's queries slow at the front door. | Frontends are stateless and scale out; each rejects new queries with 429 above a per-instance concurrency cap, before parsing or splitting them (the expensive part). Per-tenant outstanding limits stop one tenant from filling the queues. Rule evaluation uses its own query path, so alerting isn't starved by dashboards. | – |
| A metering bug | Gateway accepted and ingester appended + discarded + backfill-routed disagree by 4% for a group of tenants. | The reconciliation alarm fires before invoicing; the affected hours are held. Because each hour's value is set, not added, we fix the bug and recompute those hours from the source counters, overwriting the items. | – |
| A bad rule pack reaches all tenants | A change to our built-in default rules (for example, a mistyped selector that matches everything) floods pages across tenants. | Built-in rule packs deploy through AppConfig with one environment per tenant cohort (the canary tenants, then 1% of tenants, then 10%, then everyone); we deploy to each environment in turn and stop automatically on a rise in alert volume. (AppConfig's own gradual strategies pace how many of the rule evaluators fetching the config get it, not which tenants, hence the environments.) Rollback redeploys the last good version (41) to each environment. Every pack is also checked before release by evaluating it against the last 7 days of canary data ("how often would this have fired?"). OPS 6 | – |
R3.9 Runbook
Signals for the monitor itself, per cell OPS 8 · REL 6
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Ingest lag: stream iterator age | > 60 s for 5 min | P1 | Which ingester group? Memory stops? Add ingesters or shards |
| Stream backlog: write throttling on the stream | > 0.1% of puts for 5 min | P2 | Hot shards? Raise the shard count |
| Dropped samples, by tenant and reason | > 1% of a tenant's samples | P2 | Notify the tenant; check for a new label |
| Cardinality growth | A tenant's series 2× in 24 h; a cell above 80% of its series capacity | P2 | Drop rule or limit; plan a tenant move |
| Rule evaluation lag | Missed iterations > 0 for 10 min; evaluation time > 50% of interval | P1 | Evaluators overloaded? A tenant's rule too heavy? |
| Notification failures | Failed sends > 0 for 5 min, per integration | P1 (routed by email) | Integration down? Check fallback |
| Canary alert delay | End-to-end > 3 min, or missing | P1 (outside watcher) | The cell's alert path is suspect |
| Meter mismatch | > 1% gateway vs ingester, per tenant-hour | P2 | Hold invoices for those hours |
Procedure: a tenant's cardinality spike
- Confirm in the tenant's cardinality report which metric and label grew (top-10 labels by new series in the last hour).
- If the tenant's own limits are containing it (new series dropped, other tenants' ingesters normal), notify the tenant and stop there.
- If its ingesters are above 80% memory, add a drop rule for the label to the tenant's limits (command below). Gateways pick up the change within a minute; its series stop growing.
- Watch the tenant's ingesters' memory fall as those series go stale and the next head cut drops them.
- Tell the tenant what was dropped and how to add the label back safely (as a log field or trace attribute).
Procedure: stream backlog
- Check iterator age per shard group: all shards (ingesters too slow) or a few (hot shards).
- All shards: are ingesters at their memory stop? Add ingesters to the group; each takes over whole shard ranges and replays them.
- A few shards: a tenant sending far above its rate through few series? Lower its rate limit. Otherwise raise the shard count (below); one call can at most double it, and the number of calls per day is limited, so raise it well above need.
Commands (AWS CLI; IDs and names are placeholders; check them before running)
textaws cloudwatch get-metric-statistics --region us-east-1 --namespace AWS/Kinesis --metric-name GetRecords.IteratorAgeMilliseconds --dimensions Name=StreamName,Value=metrics-use1-cell3 --start-time 2026-09-27T10:00:00Z --end-time 2026-09-27T10:30:00Z --period 60 --statistics Maximum aws kinesis describe-stream-summary --region us-east-1 --stream-name metrics-use1-cell3 aws kinesis update-shard-count --region us-east-1 --stream-name metrics-use1-cell3 --target-shard-count 900 --scaling-type UNIFORM_SCALING aws dynamodb update-item --region us-east-1 --table-name tenant-limits --key '{"tenant_id":{"S":"t-9"}}' --update-expression "SET drop_labels = list_append(drop_labels, :l)" --expression-attribute-values '{":l":{"L":[{"S":"user_id"}]}}' aws route53 get-health-check-status --health-check-id 1a2b3c4d-5e6f-7a8b-9c0d-1e2f3a4b5c6d aws appconfig start-deployment --application-id a1b2c3d --environment-id e4f5g6h --configuration-profile-id p7q8r9s --configuration-version 41 --deployment-strategy-id AppConfig.AllAtOnce
The first shows how far behind the slowest consumer of a cell's stream is. The fourth adds a drop rule in the control region, from where the global table carries it to every region. The last is the rule-pack rollback: it redeploys the last good version (41) to one cohort's environment at once, using AppConfig's predefined all-at-once strategy; we run it for each cohort's environment.
Incident flow OPS 10
Synthesizing vector architecture diagram...
The first question is the blast radius; cells and shuffle sharding exist so the answer is usually "one tenant" or "one cell".
Game days. Monthly: stop a cell's ingester group A and time the replay; send a canary tenant's cardinality to 3× and watch its limits hold. Quarterly: cut a home region off from its standby and from agents in a staging pair, and measure the canary alerts' delay. REL 12 · OPS 11
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | Cells as the unit of failure; shuffle sharding; an alerting standby per geography fed independently by agents; config written in one region; canary tenants and outside watchers; game days. REL 10 · REL 12 · REL 13 |
| Performance Efficiency | Smaller fan-out per tenant; results caching of past intervals, invalidated by backfill generations; fair queues; recording rules for heavy dashboards. PERF 3 · PERF 5 |
| Security | Tenant identity from API keys, never from headers; tenant ID in every record, key and cache entry; per-tenant S3 prefixes; blocks encrypted with SSE-KMS; least-privilege roles per component, and only the compactor may delete. SEC 2 · SEC 3 · SEC 8 |
| Cost Optimization | Cost per million series as the product's key number (≈ $830); metering per tenant; dual sending instead of paying inter-Region transfer; build vs buy by scale. COST 1 · COST 3 · COST 8 · COST 11 |
| Operational Excellence | Alerting first in every priority call; rule packs rolled out one cohort environment at a time with automatic stops; alarms with first actions; incident flow by blast radius; COEs. OPS 1 · OPS 6 · OPS 10 · OPS 11 |
| Sustainability | Data stays in the tenant's chosen geography; only 10% of series are duplicated for alerting; rollups and per-tenant retention; recording rules replace repeated heavy queries. SUS 1 · SUS 2 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Isolates tenants with limits, cells and shuffle sharding, and can do the overlap math.
- Designs alerting to survive the provider's own region outage, with warm history and no failover step, and prices the alternatives, including the higher Asia-Pacific transfer rates.
- Turns alerts into a product feature: SLOs, error budgets and multi-window burn rates, with the numbers.
- Protects shared query capacity with fair queues, limits, correct cache keys and recording rules.
- Meters usage idempotently and reconciles it, and knows cost per million series and what drives it.
- Answers build-versus-buy differently at each scale, with numbers.
Follow-up questions
-
"The standby fires an alert, but the home region says everything is fine. Who's right?" Answer: usually the one that sees the data sooner. If the home region is healthy, both evaluate the same rules on the same series, and small timing differences produce the same alert a few seconds apart, deduplicated by Alertmanager. A real disagreement means the subsets differ: a rule changed and the agent config hasn't refreshed, or the agent's second destination is failing. We alarm on "standby and home alert sets differ for more than 5 minutes" per tenant, and treat it as a bug in the subset, not as a vote.
-
"A customer disputes last month's bill. How do you explain it?" Answer: from the usage table: active series per hour, samples ingested and discarded, bytes stored, query samples, each with its source. We show the hours where usage jumped and, from the cardinality report, the metric and label that caused it. Because every hour's value was set once and reconciled against a second meter, we can show the same numbers they see on their usage dashboard.
-
"A tenant wants 13 months of raw data for an audit. What does it cost, and what changes?" Answer: storage only, mostly: raw is about 11.8 KB per series per day, so a 5M-series tenant adds
5M × 11.8 KB × 365 ≈ 21.5 TBa year at about $0.022/GB-month in S3 Standard, about $470 a month, or less if blocks older than 30 days move to a cheaper storage class they rarely read. Queries over raw data beyond 32 days still hit the query-length limit; for audits they'd use exports, not dashboards.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "One global limit for all tenants" | Fits no one; limits come from each tenant's plan. |
| "A cluster per tenant" | Thousands of idle clusters; use cells and shuffle sharding. |
| "Alerting in three AZs survives a region outage" | AZs protect against an AZ. |
| "Fail over DNS to a cold region for alerting" | No history: for timers and windows start empty. |
| "Forward alert data from the home region" | It stops when the home region fails, and costs inter-Region transfer. |
| "Cache query results by query text" | Without the tenant ID and tier in the key, tenants see each other's numbers. |
| "Add usage counters on every event" | Replays double-count; set a value per tenant-hour. |
| "Silences can expire with DynamoDB TTL" | TTL deletes eventually, within days. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping questions | Restate the Round 1 design in 60 seconds | Restate the Round 2 design in 60 seconds |
| 5–15 min | Requirements + API | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Design steps 1.0–1.6 | Design steps 2.1–2.6 | Design steps 3.1–3.6 |
| 40–50 min | Numbers (series × interval, bytes, alert delay) + trade-offs | Memory from series, tiers, cost, trade-offs | Cells, standby, cost per million series, trade-offs |
| 50–60 min | Failures + pillar check | Failures + pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening a round: "Before I design, let me ask a few scoping questions: how many series, at what interval, kept how long, and who gets paged?"
- When the scope is raised: "Here's what breaks in the current design, and here's the order I'll fix it in."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "How do you know your monitoring is working?" (REL 6) | An always-firing Watchdog alert feeds a dead man's switch outside the stack; in Round 3, canary tenants per cell are watched from another region. | 1, 3 | Steps 1.6, 3.2 |
| "What happens to the last two hours when a node dies?" (REL 11) | Nothing is lost: the stream holds 24 hours in three AZs, the twin replica serves queries, and a replacement replays from the last checkpoint. | 2 | Step 2.2 | |
| "How does one team or tenant not break everyone?" (REL 10) | Per-tenant limits, cells, and shuffle-sharded ingesters and queriers: a full overlap is 1 in 230,300 for queriers. | 2–3 | Steps 2.4, 3.1 | |
| "What if your region goes down?" (REL 13) | A standby region in the same geography, fed directly by agents, evaluates the alert rules all the time, so pages continue without a failover. | 3 | Step 3.2 | |
| Performance | "How do you store 10M samples a second?" (PERF 3) | Gorilla-compressed, append-only chunks per series, about 1.4 bytes a sample, with an inverted index over labels. | 1–2 | Steps 1.1, 2.1 |
| "How does a five-year graph load in a second?" (PERF 3) | Queries read 1-hour rollups of min, max, sum and count: 90× fewer values than raw 10-second data. | 2 | Step 2.3 | |
| Security | "How do you keep tenants apart?" (SEC 2, SEC 3) | Tenant identity comes from the API key, and the tenant ID is part of every record, block prefix and cache key. | 3 | R3.3, step 3.4 |
| "Can personal data end up in metrics?" (SEC 7) | Forbidden labels (user IDs, emails) are dropped at the gateway; per-request detail belongs in logs or traces. | 2 | Step 2.4 | |
| Cost | "What drives your bill?" (COST 1) | Active series: they set ingester memory, the biggest line, about $760–830 per million series a month. | 2–3 | R2.6, R3.6 |
| "Why not CloudWatch or managed Prometheus?" (COST 11) | Managed wins at 50K series; at 100M+ building costs several times less, and in Round 3 monitoring is the product. | 1–3 | R1.7, step 3.6 | |
| "Where does data transfer bite?" (COST 8) | Stream traffic through endpoints, and inter-Region transfer if we forwarded alert data, most of it from Asia-Pacific; agents send twice instead. | 2–3 | R2.6, R3.6 | |
| Operations | "How do you stop alert fatigue?" (OPS 10) | for and hysteresis, grouping, symptom paging with inhibited causes, and SLO burn-rate alerts. | 1–3 | Steps 1.4, 1.5, 2.6, 3.3 |
| "How do you change alert rules safely for thousands of tenants?" (OPS 6) | Rule packs go out one cohort environment at a time, canaries first, backtested on a week of data, with an automatic stop and a rollback that redeploys the last good version to each environment. | 3 | R3.8, R3.9 | |
| Sustainability | "How do you avoid storing what nobody reads?" (SUS 4) | Collect only used series, downsample, expire tiers, and replace repeated heavy queries with recording rules. | 1–3 | Steps 2.3, 3.4 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Storage | Time-series layout: series IDs, chunks, WAL, blocks; inverted index. | Gorilla by hand; tiers recomputed with the real interval and aggregates. | Per-tenant retention and storage metering. |
| Scale and durability | Two full copies and an Alertmanager cluster. | Stream as write-ahead log; two replicas; memory sized per series from the engine's real cost. | Cells; shuffle sharding; carve-outs for giant tenants. |
| Cardinality | Chooses which series to collect. | Limits at the gateway and at the ingesters; cardinality alerts. | Per-tenant limits from plans; cardinality as the cost lever. |
| Alert quality | for, hysteresis, ratios, grouping. | Symptom paging, inhibition, silences, late-data and clock-skew handling. | SLO burn-rate alerts with multiple windows. |
| Watching the watcher | Dead man's switch outside the stack. | Alarms on lag, drops and failed notifications. | Standby region fed independently; canary tenants; outside watchers. |
| Cost | Three options priced; CloudWatch's per-metric trap. | ≈ $760 per million series; memory is the bill. | ≈ $830 per million series; metering; build vs buy by scale. |
| Evolving under new scope | Builds from a table of rows, one problem at a time. | Opens with "what breaks" and rechecks the inherited numbers. | Changes the operating model: tenants, regions, billing, and who watches whom. |