Design a Scalable Distributed Notification System
1. Problem Statement & Scope Clarification
System Mission
Design a globally distributed, highly available, multi-channel notification platform (similar to Twilio, OneSignal, and Amazon SNS/SES) capable of delivering billions of notifications daily across Mobile Push (Apple APNs, Google FCM), SMS, Transactional Email, and Webhooks, with strict priority isolation, user preference compliance (quiet hours, category opt-outs), intelligent deduplication, and rate-limiting.
Functional Requirements
- Multi-Channel Dispatch Engine: Unified API supporting Mobile Push (iOS APNs HTTP/2, Android FCM v1), SMS (Amazon SNS / direct carrier routes), Email (Amazon SES), and In-App WebSockets.
- Strict Priority Queuing: Critical transactional messages (OTP, MFA, fraud alerts) must bypass marketing batches and achieve sub-second delivery.
- Template Rendering & Localization: Dynamically render multi-lingual templates with personalized user variables fetched from profile stores.
- Intelligent Deduplication: Suppress duplicate notifications for the same event trigger within a configurable sliding TTL window (e.g., preventing duplicate billing alert pushes).
- User Preference & Regulatory Compliance: Enforce per-user channel preferences, quiet hour dampening across local user timezones, and statutory opt-out compliance (TCPA, GDPR, CAN-SPAM).
- Device Token Lifecycle Management: Automatically handle device token invalidations (e.g., APNs
410 Gone/ FCMUNREGISTERED) without crashing worker pipelines.
Non-Functional Requirements (SLAs & SLOs)
- High Availability: uptime SLA for critical OTP and transactional notification paths.
- Delivery Latency:
- Critical Priority (OTP / MFA / Security Alerts): Delivery P95 , P99 .
- High Priority (Order confirmations, billing alerts): Delivery P99 .
- Low Priority (Marketing newsletters, daily digests): Delivery within 15 minutes of batch trigger.
- Scale: Ingest peak of ; deliver over .
- Durability: Zero acknowledged message loss ().
2. Capacity & Scale Estimation (Back-of-the-Envelope Math)
Traffic & Ingest Profile
- Total Daily Notification Volume: (2 Billion/day).
- Channel Distribution:
- Mobile Push (APNs / FCM): ().
- Transactional & Marketing Email (SES): ().
- SMS & WhatsApp (SNS / Twilio): ().
- Average Ingest QPS:
- Peak Ingest QPS ( multiplier during flash sales):
Network Bandwidth & Payload Sizing
- Average Ingestion Payload Size: (Metadata headers + template variables + user identifiers).
- Peak Ingest Network Bandwidth:
- APNs HTTP/2 Multiplexed Connection Sizing:
An HTTP/2 persistent connection to APNs handles across multiplexed streams:
Hosted across a fleet of 20 EKS Push Worker pods, each maintaining 6 persistent TLS connections to
api.push.apple.com.
Memory Sizing (Redis 24-Hour Sliding Deduplication Cache)
To prevent duplicate notifications from upstream retries, the Ingestion Service calculates a SHA-256 fingerprint for every notification:
- The deduplication window is per request (
deduplication_window_seconds, default , maximum for once-a-day alerts such as billing reminders). Size for the worst case where every producer asks for the 24-hour maximum: 2 Billion keys resident at once. - Each Redis key entry consumes: (SHA-256 stored as raw binary, not the 64-char hex string) + (TTL) + (dictEntry overhead) .
- Sized for the worst case on an Amazon ElastiCache Redis Cluster in cluster mode: shards
cache.r6g.2xlarge( each , giving headroom for fragmentation), one replica per shard across 3 AZs. Threecache.r6g.xlargenodes ( total) would overflow the worst case before fragmentation is even counted.
Historical Audit Storage (Amazon S3 Lake)
- All dispatched notifications stream into Amazon S3 Analytics Lake in compressed Parquet format via Kinesis Data Firehose:
3. High-Level Architecture & AWS Component Mapping
Synthesizing vector architecture diagram...
Follow a notification request from the top. Services send it through the NLB to the "Ingestion & Validation Tier" panel, where Redis drops duplicates seen in the last 24 hours and DynamoDB checks the user's preferences and opt-outs. In the "Physically Isolated Amazon SQS Queues" panel, it is routed by priority and channel: OTP and security codes go to their own FIFO queue, and push, email and SMS each get a separate queue with its own DLQ after 3 failed attempts. Each queue has a dedicated worker fleet in the "Channel-Specific Dispatch Fleets" panel that calls the provider (SNS, APNs, FCM, SES). Separate queues are the point: a million-message marketing campaign can fill the email queue, but a login code still arrives within seconds.
Data Flow Walkthrough
- Unified Ingestion:
- Internal microservices submit notification requests via gRPC or HTTPS to the Ingestion Fleet via an internal NLB.
- The Ingestion Service evaluates the Redis 24h Deduplication Ring via
SET NX EX. If the idempotency fingerprint exists, the request is immediately acknowledged as duplicate (DROPPED_DUPLICATE) without touching queues.
- User Preference & Quiet Hours Gating:
- The service fetches recipient channel opt-ins and local timezone quiet hours from DynamoDB (
NotificationUserDataTable). - If the user is currently in their quiet hours window ( local time) and priority is not
CRITICAL, the message is deferred: it is written to the scheduler table withdeliver_at = next 09:00 localand the producer receivesDEFERRED_QUIET_HOURS. Messages are dropped only when the producer explicitly setsdrop_if_deferred = true(e.g. a flash-sale push that is worthless by morning).
- The service fetches recipient channel opt-ins and local timezone quiet hours from DynamoDB (
- Physical Priority Queue Routing:
- Requests are routed to physically isolated Amazon SQS queues. High-priority OTP codes enter
SQS_Critical.fifo, ensuring zero Head-of-Line blocking from bulk marketing pushes.
- Requests are routed to physically isolated Amazon SQS queues. High-priority OTP codes enter
- Channel Dispatch & External Handshake:
- Dedicated EKS worker fleets pull from their respective SQS queues in batches of 10.
- Push workers multiplex payloads over persistent HTTP/2 connections to Apple APNs and Google FCM.
- Email workers dispatch via Amazon SES using dedicated warm IP pools.
- SMS workers dispatch via Amazon SNS with fallback carrier routing.
- Token Hygiene & DLQ Redrive:
- If APNs returns
HTTP 410 Gone, workers mark the device tokenINVALIDin DynamoDB. - Poison pill messages exceeding 3 retries route to channel DLQs for automated analysis.
- If APNs returns
Concrete Step-by-Step Request Walkthrough: Tracing Notification Lifecycle
| Step # | Event / Action | Component State | Distributed Transition | Output / Response |
|---|---|---|---|---|
| 1 | Auth service generates 2FA code | OTP generated in memory | Dispatches POST /v1/notifications with Priority: CRITICAL | Request terminates at NLB; forwards to Ingest ECS task |
| 2 | Sliding deduplication check | Redis Cluster connection pool | SET dedupe:<sha256(usr_99+SMS+tpl_security_otp_v2+key1)> 1 NX EX 120 | Redis returns OK (First submission); proceeding |
| 3 | Preference & quiet hours check | Fast point lookup in DynamoDB | Ingest verifies is_opted_in = true; priority CRITICAL bypasses quiet hours | Delivery validated in |
| 4 | Routing to isolated SQS FIFO | Dedicated queue SQS_Critical.fifo | SendMessage enqueues message with message deduplication ID | Message committed to queue in ; HTTP 202 returned to caller |
| 5 | Dedicated OTP worker pickup | EKS worker pod running Go runtime | Long-polling ReceiveMessage immediately retrieves OTP | Worker unpacks payload; skips bulk queue latency |
| 6 | Dispatch to carrier / APNs | Persistent carrier connection pool | Direct SMPP / HTTP/2 push emitted to Apple APNs / Carrier | Delivered to device in total |
| 7 | Delivery confirmation & receipt | Client device receives notification | APNs/FCM only confirm acceptance by the gateway; the app's notification service extension fires a lightweight beacon (POST /v1/notifications/{id}/events) on receipt and on tap | Notification state marked DELIVERED (then OPENED) in DynamoDB |
4. API Interface Design & Wire Protocol
1. Unified Notification Ingest Protocol (notification.proto)
protobufsyntax = "proto3"; package hispeeddesign.notification.v1; service NotificationService { rpc SendNotification (SendNotificationRequest) returns (SendNotificationResponse); rpc SendBatchNotification (SendBatchNotificationRequest) returns (SendBatchNotificationResponse); rpc GetNotificationStatus (GetNotificationStatusRequest) returns (GetNotificationStatusResponse); } enum PriorityLevel { PRIORITY_UNSPECIFIED = 0; PRIORITY_CRITICAL = 1; // Bypasses quiet hours, delivers via FIFO queue (OTP/MFA/Fraud) PRIORITY_HIGH = 2; // Transactional receipts, billing notices PRIORITY_LOW = 3; // Marketing promotions, weekly digests } enum ChannelType { CHANNEL_PUSH = 0; CHANNEL_SMS = 1; CHANNEL_EMAIL = 2; CHANNEL_IN_APP = 3; } message SendNotificationRequest { string user_id = 1; string idempotency_key = 2; // Client-supplied deduplication key PriorityLevel priority = 3; repeated ChannelType preferred_channels = 4; string template_id = 5; map<string, string> template_variables = 6; int32 deduplication_window_seconds = 7; // Default: 300s, max 86400s bool drop_if_deferred = 8; // Default false: quiet-hours traffic is rescheduled, not dropped } message SendNotificationResponse { string notification_id = 1; enum Status { QUEUED = 0; DROPPED_DUPLICATE = 1; DROPPED_OPT_OUT = 2; DROPPED_QUIET_HOURS = 3; // Only when the producer set drop_if_deferred = true DEFERRED_QUIET_HOURS = 4; // Default for non-critical traffic in quiet hours } Status status = 2; int64 timestamp_ms = 3; } message GetNotificationStatusRequest { string notification_id = 1; } message GetNotificationStatusResponse { string notification_id = 1; string user_id = 2; ChannelType channel = 3; string delivery_status = 4; // QUEUED, SENT, DELIVERED, FAILED int64 delivered_at_ms = 5; string failure_reason = 6; }
2. RESTful Ingestion Endpoint
httpPOST /v1/notifications Host: api.notification.aws.internal Authorization: Bearer <service_jwt_token> X-Idempotency-Key: 8a9b2c3d-4e5f-6a7b-8c9d-0e1f2a3b4c5d Content-Type: application/json { "user_id": "usr_alex_99", "priority": "CRITICAL", "preferred_channels": ["CHANNEL_SMS", "CHANNEL_PUSH"], "template_id": "tpl_security_otp_v2", "template_variables": { "otp_code": "849201", "expires_in_minutes": "5" }, "deduplication_window_seconds": 120 }
Response: 202 Accepted
json{ "status": "QUEUED", "data": { "notification_id": "ntf_718293847561029384", "user_id": "usr_alex_99", "assigned_queue": "SQS_Critical.fifo", "timestamp_epoch_ms": 1718000000120 } }
5. Data Models & Storage Architecture
DynamoDB Single-Table Schema (NotificationUserDataTable)
- Partition Key (
PK): Entity identifier. - Sort Key (
SK): Specific entity instance or sequence cursor. - Global Secondary Index 1 (
GSI1): Notification history index (GSI1-PK: USER#<user_id>,GSI1-SK: SENT#<timestamp>).
| Entity Type | PK (Partition Key) | SK (Sort Key) | GSI1-PK | GSI1-SK | Attributes & Types |
|---|---|---|---|---|---|
| User Preferences | USER#<user_id> | PREFERENCES | - | - | opt_in_marketing (BOOL), quiet_hours_start (STR), quiet_hours_end (STR), timezone (STR) |
| Device Push Token | USER#<user_id> | DEVICE#<token_hash> | - | - | device_token (STR), platform (IOS/ANDROID), updated_at (NUM), status (ACTIVE/INVALID) |
| Notification Audit | NOTIF#<notif_id> | METADATA | USER#<user_id> | SENT#<timestamp> | channel, priority, delivery_status, attempts, error_code |
| Message Template | TEMPLATE#<tpl_id> | LANG#<lang_code> | - | - | subject (STR), body_template (STR), required_vars (LIST) |
| Deferred Delivery | SCHED#<yyyy-mm-ddThh> (UTC hour bucket) | AT#<deliver_at_ms>#<notif_id> | - | - | payload_s3_key, priority, channel; a scheduler Lambda runs every minute, queries the current hour bucket, and enqueues rows whose deliver_at has passed |
Device Token Registration (Where Push Tokens Come From)
Push tokens are minted by the mobile OS, not by this system, so the app must hand them over on every launch (tokens rotate on reinstall, OS restore, and periodically on iOS):
httpPOST /v1/devices Authorization: Bearer <user_jwt> { "platform": "IOS", "device_token": "a1b2c3...", "app_version": "5.2.0", "locale": "de-DE", "timezone": "Europe/Berlin" } Response: 200 OK { "device_id": "DEVICE#7f3a...", "status": "ACTIVE" }
The Ingestion Service upserts USER#<user_id> / DEVICE#<sha256(token)> (the hash keeps the sort key fixed-length and avoids storing the raw token in an index), refreshes updated_at, and the same call also updates timezone on the PREFERENCES item so quiet-hours arithmetic follows the device.
Per-User Frequency Capping
Deduplication stops the same event twice; frequency capping stops too many different marketing events. A Redis counter per user and category enforces caps such as "at most 3 marketing pushes per day":
textINCR cap:usr_99:marketing:2026-09-20 -> 4 EXPIRE cap:usr_99:marketing:2026-09-20 86400
If the counter exceeds the cap, LOW-priority requests return DROPPED_FREQUENCY_CAP; CRITICAL and HIGH traffic is never capped.
Redis Deduplication Ring
- Key:
dedupe:<sha256_hash> - Value:
"1" - Command:
SET dedupe:<hash> 1 NX EX <window_seconds> - Behavior: If the key already exists, Redis returns
nil, and the Ingestion Service immediately drops the request withDROPPED_DUPLICATE.
Unlock Complete Architecture & Production Runbooks
You have explored the free architectural preview (~48%). Spend 1 Coin to unlock the remaining 6 production deep-dive sections for a full 24 hours.