AWS Cloud Services Architecture Matrix Cheat Sheet
Instant decision matrix mapping distributed system requirements to native AWS managed cloud services.
Instant decision matrix mapping distributed system requirements to native AWS managed cloud services.
1. AWS Core Services Architecture Map
Synthesizing vector architecture diagram...
Follow one request from the top. At the edge, Route 53 picks the region, CloudFront serves anything cacheable, and WAF drops malicious traffic, so only real work reaches the gateway. API Gateway or the ALB hands it to one of the compute styles in the "Compute Tier" panel: Fargate for long-running services, Lambda for event handlers, EC2 for full control. In the "Storage & Database Tier" panel, synchronous work reads and writes Aurora (transactions), DynamoDB (key lookups), S3 (files) or Redis (hot reads). Slow work is handed to the "Streaming & Asynchronous Processing" panel instead, SQS for tasks and Kinesis/MSK for event streams, so the request can return immediately, and messages that keep failing go to the DLQ instead of blocking the queue.
2. Compute Tier Selection Cheat Sheet
| Cloud Service | Execution Model | Cold Start | Max Execution | Best Suited For | Anti-Pattern |
|---|---|---|---|---|---|
| AWS Lambda | Event-driven serverless functions | 100ms - 2s | 15 Minutes | Spike-driven APIs, Webhooks, S3/SQS triggers | Stateful WebSocket servers, long-running ML training |
| ECS Fargate | Serverless container orchestration | 30s - 2m | Unlimited | Standard REST/GraphQL APIs, background workers | Instant burst traffic from 0 to 10k RPS in < 5s |
| EKS (Kubernetes) | Managed Kubernetes control plane | Node spin-up | Unlimited | Complex multi-tenant microservices, service meshes | Small engineering teams with low operational bandwidth |
| EC2 Spot Fleet | Ephemeral virtual compute instances | Node boot | Terminated with 2m notice | Batch processing, video transcoding, CI/CD runners | Critical databases, non-fault-tolerant master nodes |
3. Database Selection Cheat Sheet
| Database Service | Engine Paradigm | Latency Profile | Max Capacity | Primary Architectural Strength |
|---|---|---|---|---|
| Amazon Aurora | Relational (PostgreSQL / MySQL compatible) | 2ms - 15ms | 128 TiB per cluster | ACID compliance, complex JOINs, distributed storage auto-repair |
| Amazon DynamoDB | Distributed Key-Value & Document | 1ms - 5ms (Sub-ms with DAX) | Virtually unlimited | Predictable p99 latency at any throughput scale |
| Amazon MemoryDB | Redis-compatible durable in-memory | Sub-millisecond reads & writes | Terabytes in-memory | Transactional log replication with microsecond reads |
| Amazon OpenSearch | Distributed Lucene inverted index | 10ms - 50ms | Petabytes | Full-text search, fuzzy queries, log analytics, vector search |
| Amazon Timestream | Append-only Time-Series database | 5ms - 30ms | Exabytes (S3 tiered) | IoT telemetry, server metrics, time-window aggregations |
4. Messaging & Event Bus Decision Matrix
Synthesizing vector architecture diagram...
Start at the top and answer two yes/no questions. First: must messages be processed in strict order? If yes, ask whether several independent consumers need the same messages: if they do, use Kinesis or MSK, where each consumer reads the ordered log at its own position; if only one consumer does, use SQS FIFO. If order does not matter, ask whether many services should each get a copy: if yes, use SNS or EventBridge to fan out; if not, SQS Standard is the simplest queue with nearly unlimited throughput. Ordering costs throughput (FIFO queues and shards have limits), so only choose it when the business logic really needs it.
5. Multi-Region Disaster Recovery Patterns
| DR Pattern | RTO (Recovery Time Objective) | RPO (Recovery Point Objective) | Cost Index | Description |
|---|---|---|---|---|
| Backup & Restore | 4 to 24 hours | Hours | 💲 (Minimal) | Snapshots copied to secondary region; deployed on disaster. |
| Pilot Light | 10 to 30 minutes | Minutes | 💲💲 (Low) | Database replication alive; core compute launched on failover. |
| Warm Standby | 2 to 5 minutes | Seconds | 💲💲💲 (Medium) | Scaled-down production replica running 24/7 in secondary region. |
| Multi-Region Active-Active | Near Zero (< 10s) | Near Zero (sub-second) | 💲💲💲💲💲 (Highest) | Traffic served simultaneously from 2+ regions globally. |
6. Related Resources
- Complete reference manual: AWS Reference Architecture Matrix
- Distributed components: Distributed Cache, Message Queue, Object Storage