List the key functional requirements for the system (Ask the AI for hints if stuck)...
These describe what the system does.
List the key non-functional requirements (performance, scalability, reliability, etc.)...
These describe how well the system performs.
Define the APIs expected from the system. This is your chance to analyze and define the read and write paths so that you can come up with the high-level design...
For a data pipeline, the "API" is usually the entry point (Ingestion) or the exit point (Consumption). Let's design the Ingestion API for a producer sending event data.
Endpoint: POST /v1/events
Describe the overall system architecture. Identify the main components needed to solve the problem end-to-end. Use the diagramming tool to create a block diagram.
We will use the Lambda Architecture pattern. This divides the pipeline into three layers:
The Data Flow:
Deep dive into 2-3 key components. Explain how they work, how they scale, discuss tradeoffs, capacity, and any relevant algorithms or data structures.
Component: Apache Kafka (or AWS Kinesis).
user_id or event_id to ensure ordering within a specific partition.Component: Amazon S3 (or HDFS).
Component: Apache Spark.
Component: Apache Flink or Spark Streaming.
Component: Data Warehouse (Snowflake / BigQuery) or Presto/Trino.