List the key functional requirements for the system (Ask the AI for hints if stuck)...
Product catalog with search: Users browse and search millions of products by keyword, category, price range, and rating. Elasticsearch powers full-text search with faceted filtering.
Persistent shopping cart: Users add, update, and remove items. The cart survives browser closures and device switches. Server-side persistence is essential because roughly 70% of carts are abandoned and many purchases happen when users return hours or days later.
Checkout with payment processing: Validates the cart, reserves inventory, charges Stripe/PayPal, and creates the order. PCI DSS compliance via tokenization: the backend never touches raw card data.
Order tracking: Users view order history and receive status updates (confirmed, shipped, delivered) via email or push.
List the key non-functional requirements (performance, scalability, reliability, etc.)...
99.9%+ checkout availability: Checkout downtime directly equals lost revenue. Different paths get different SLA targets based on business impact.
Sub-200ms browsing latency: Every 100ms of latency reduces conversion by roughly 1%. Aggressive caching and CDN usage required.
10-100x peak scaling: Flash sales generate traffic spikes that dwarf the baseline. The system must absorb these without overselling.
Estimate the scale of the system. Consider daily active users, read/write ratio, storage requirements, bandwidth, and any relevant QPS calculations...
5M DAU, 50K concurrent at peak. 100M page views/day (1,150/sec average, 11,500/sec flash sales). Checkout: 17 TPS steady, 500-1,000 TPS flash sales. Search: 350K queries/min.
Products: 50GB metadata (10M products at 5KB), 25TB images in S3. Orders: 36GB/year (100K orders/day at 1KB). Search index: 20GB in Elasticsearch.
9.5 GB/s total page traffic. CDN absorbs 95%+. Origin handles 475 MB/s normally. Checkout is only 5 MB/s bandwidth but heavy on DB and payment gateway load.
Define the APIs expected from the system. This is your chance to analyze and define the read and write paths so that you can come up with the high-level design...
GET /products?q=keyword&category=electronics&min_price=50&max_price=200&sort=relevance&page=1&limit=20
GET /products/:id
The search endpoint returns an inventory_status field ("in_stock", "low_stock", "out_of_stock") rather than exact counts. Exact counts change rapidly and would be stale by the time users see search results. Bucketed status tolerates staleness gracefully.
GET /cart
POST /cart/items { product_id, quantity }
PUT /cart/items/:id { quantity }
DELETE /cart/items/:id
POST /checkout
Headers: Idempotency-Key: <client-generated UUID>
Body: { cart_id, shipping_address_id, payment_method_id }
GET /orders
GET /orders/:id
The payment_method_id is a Stripe token, not raw card data. The frontend collects card details via Stripe Elements and exchanges them for a token directly with Stripe servers. The backend never sees card numbers.
Describe the overall system architecture. Identify the main components needed to solve the problem end-to-end. Use the diagramming tool to create a block diagram.
Two flows define the system: checkout (critical write path) and search (high-volume read path).
Checkout flow: Order Service validates cart, reserves inventory atomically, charges payment via gateway, creates order record, and sends confirmation.
UPDATE inventory SET reserved = reserved + qty WHERE product_id = X AND quantity - reserved >= qtySearch flow: query routed to Search Service, Elasticsearch processes with filters and ranking, results enriched from Redis cache.
Define the data model. Identify the main entities, their attributes, and relationships. Consider the choice of database type (SQL vs NoSQL) and justify your decision based on access patterns...
users (id, email UNIQUE, name, password_hash, created_at)
products (id, name, description, price, category_id, image_urls, created_at)
inventory (product_id PK, quantity, reserved, version)
orders (id, user_id, idempotency_key UNIQUE, total, status, created_at)
order_items (id, order_id, product_id, quantity, unit_price)
cart_items (user_id, product_id, quantity, added_at)
The inventory table is deliberately separate from products. Products are cached in Redis with a 5-minute TTL for fast browsing. Inventory is locked during checkout with row-level transactions. A checkout locking an inventory row should never block a product page read. Different access patterns need different tables.
The version column on inventory enables optimistic concurrency control. Both users read inventory (quantity=1, version=5). User A updates with WHERE version=5 and succeeds, incrementing to version=6. User B also specifies WHERE version=5, but the version is now 6, so zero rows are affected and the checkout returns "out of stock."
Product catalog changes propagate to Elasticsearch via CDC (Debezium reading the PostgreSQL WAL). The application writes only to PostgreSQL. CDC reliably streams changes to the search index. If Elasticsearch is temporarily down, CDC buffers and replays when it recovers.
products:{id} -> product JSON TTL: 300s
cart:{user_id} -> cart items TTL: 24h
session:{token} -> user context TTL: 30min
Cart data lives in both Redis (fast reads) and PostgreSQL (durable backup). Redis provides sub-millisecond cart operations during active browsing. PostgreSQL ensures cart survival across Redis failures.
Deep dive into 2-3 key components. Explain how they work, how they scale, discuss tradeoffs, capacity, and any relevant algorithms or data structures.
The GATE of this problem is inventory reservation: the mechanism that prevents overselling. Every other component can be eventually consistent, but inventory must be strongly consistent at checkout.
Inventory reservation: two concurrent checkouts attempt the last item. Row-level lock ensures only one succeeds. The second gets out-of-stock.
Two fields: quantity (total stock) and reserved (items held by in-progress checkouts). Available = quantity - reserved. If payment fails, only reserved decrements; quantity was never touched, so other users always saw correct availability.
Normal traffic: optimistic concurrency. UPDATE with WHERE version = X. Rarely conflicts when stock is sufficient. On conflict, retry with the new version.
Flash sales: pessimistic locking. SELECT FOR UPDATE serializes access. Slower per checkout but predictable under extreme contention where optimistic locking causes cascading retries.
Reservation TTL. Background job releases reservations older than 10 minutes. The Order Service refreshes the timestamp at each step as a heartbeat, preventing active checkouts from release.
Payment failure saga: inventory reserved, payment fails, compensating transaction releases the reservation, user notified of failure.
Checkout spans multiple services that cannot join a single distributed transaction. The saga uses compensating transactions: reserve inventory, charge payment, create order. If payment fails, release the reservation. If order creation fails after payment, refund and release. Every step is idempotent.
Worst case: payment succeeds at Stripe but the Order Service crashes before recording the order. Recovery: a job scans pending payments, queries Stripe by idempotency key, and creates the order or refunds. Stripe webhooks provide a backup notification path.
Product changes flow from PostgreSQL through CDC (Debezium) to Kafka to Elasticsearch (1-5s latency). Debezium tracks its WAL position and replays missed changes on ES recovery.
Offline Spark batch computes "users who bought X also bought Y" and stores results in Redis. Sub-millisecond serving. Daily staleness is acceptable since purchase correlations change slowly.
Every major decision involves a trade-off between consistency, availability, and complexity.
Strong consistency for money: inventory, payments, orders must be ACID-correct. Eventual consistency for everything else: product cache, search index, recommendations. The danger is mixing them up. Eventual consistency for inventory means overselling. Strong consistency for browsing means unnecessary database load.
PostgreSQL for orders and inventory (ACID). MongoDB could handle the product catalog (flexible schemas for clothing vs. electronics). The risk: no single transaction spans both databases. Cross-database data needs explicit reconciliation. Pragmatic approach: PostgreSQL for everything critical, Elasticsearch for flexible search.
Start monolithic, extract as pain points emerge. Search is the natural first extraction: different data store, independent scaling profile, no transactional coupling with orders.
Reservations hold inventory during checkout but abandoned ones waste inventory at peak demand. Queues waste nothing but make users wait. Most systems use reservations normally and switch to queues for ultra-high-contention events like sneaker drops.
Common Pitfall
Do not use eventual consistency for inventory during checkout. Candidates who propose async inventory updates will oversell. The checkout path requires synchronous, transactional inventory decrement. Everything else (search, cache, notifications) can be async, but the moment a customer commits to buying, the inventory check must be strongly consistent.
Each failure scenario maps to a specific architectural decision.
Stripe timeout: Retry with exponential backoff using the same idempotency key. After max retries, mark as "payment_pending."
Response lost after successful charge: Reconciliation job runs every 5 minutes, queries Stripe for charges matching pending orders. Creates the missing order or refunds if out of stock. Stripe webhooks provide a backup notification path.
Primary failure triggers replica promotion (15-30s). In-flight transactions have ambiguous outcomes. Clients retry with the same idempotency key; the server checks for existing orders first.
Carts recover from PostgreSQL. Cache misses fall through to DB. Request coalescing prevents stampede. Cache-warming job pre-loads top 10K products on recovery.
Flash sale handling: traffic spikes 10x, rate limiter protects the database, queue buffers checkout requests, inventory reservations with TTL prevent overselling, auto-scaler adds instances.
When 10,000 users hit checkout on 100 items: (1) Rate limiter protects downstream, (2) queue absorbs the burst, (3) inventory virtual pools split items across 10 rows reducing contention 10x, (4) graceful degradation disables recommendations and extends cache TTLs.