List the key functional requirements for the system (Ask the AI for hints if stuck)...
List the key non-functional requirements (performance, scalability, reliability, etc.)...
Estimate the scale of the system. Consider daily active users, read/write ratio, storage requirements, bandwidth, and any relevant QPS calculations...
Individual parking lots at peak traffic are known to produce 0.1 QPS "check out traffic", 0.01 QPS "check in traffic", and 0.01 QPS "purchase traffic".
100,000 parkings lots experiencing peak traffic simultaneously is highly unlikely but would yield a worst-case scenario load of 1k QPS check in traffic, 1k QPS "purchase" traffic, and 10k check out traffic.
Due to the nature of regional parking lots and timezones, we expect a more realistic ceiling of 1/4 that. These are the ceilings the business is comfortable designing for within the next 5 years of business:
250 QPS Payments
250 QPS Check In
2,500 QPS Check Out
QR Code requests scale with check ins and check outs; roughly 2750 QPS total.
Because this traffic is split across three regions, but not necessarily divided across those three regions evenly, projections calculated in a separate document estimate that the hottest region will still hit...
150 QPS Payments
150 QPS Check In
1,500 QPS Check Out
1,650 QPS QR Codes
These are what we design for ultimately.
Define the APIs expected from the system. This is your chance to analyze and define the read and write paths so that you can come up with the high-level design...
POST /Order
Reply is 202 with URL to check status (GET /Order/{id})
GET /Order/{id}
POST /LotActivity
Reply contains {activity:entry}
Logic
GET /QR/{orderId}
Describe the overall system architecture. Identify the main components needed to solve the problem end-to-end. Use the diagramming tool to create a block diagram.
To provide the fastest latency, the backend services will be deployed in three different regions. Traffic will be Geo DNS to the nearest deployment. The benefit is latency. While this divides traffic into thirds, it is not expected to be distributed evenly across those three. See capacity estimation.
Order service, lot activity service, and QR service are each and individual AWS lambda deployments.
Atomicity of payments + QRcode + DDB data writes is managed via saga design pattern and our reliance on the Payment Service's "at least once" event API, such that our saga will be called repeatedly by our PaymentService until we ack receipt of the event. Our handler for that event will be idempotent.
Define the data model. Identify the main entities, their attributes, and relationships. Consider the choice of database type (SQL vs NoSQL) and justify your decision based on access patterns...
Order
QR Code
User
LotActivity
Deep dive into 2-3 key components. Explain how they work, how they scale, discuss tradeoffs, capacity, and any relevant algorithms or data structures.
Order service, lot, activity, service, and QR service, are each served by their own individual AWS Lambda.
QR service has smallest workload per request (fetch code and return it). Anticipated latency ~50ms, little's law yields ~80 concurrent Lambdas at peak.
LotActivityService needs to read, think for single digit ms, then write. Anticipated ~100ms latency. little's law yields ~160 concurrent Lambdas at peak.
EntryListener is a DynamoDB data change listener ("Data Capture" design pattern) and will notice when an entry happens for the first time and writes a TTL onto the QRCode so that it disappears from database, per requirements. EntryListener is expected to run in batches and not expected to run concurrently (only 1). This data capture design pattern also implements a outbox pattern, where we ensure the TTL gets written whenever entry data changes, and we don't depend on the app logic owning the LotActivity mutation to properly handle the QRCode TTL.
Order Service handling POST needs to create a step functions "saga" and return a 202. Anticipated latency ~50ms, little's law yields ~80 concurrent Lambdas at peak.
Order Service handling GET needs to fetch and return an order status and return a 200. Anticipated latency ~50ms, little's law yields ~80 concurrent Lambdas at peak.
All these service endpoint Lambdas will contain only code for their endpoint, and not the entire service and GO will be used as the programming language, with minimal dependencies, so cold starts are expected to be minimal. Parking lots have gradual traffics growth characteristics and we don't expect cold start storms to blow out concurrency.
The StepFunctions Payment Saga will run lambdas in concurrency as well, to generate QR Code, get payment, and update order status along the way. It's expected to yield roughly ~160 concurrent Lambdas at peak also.
Total concurrent Lambdas at this point is ~560, and that's at the worst case traffic peak ceiling business doesn't expect to hit. AWS limit is 1000.