List the key functional requirements for the system (Ask the AI for hints if stuck)...
List the key non-functional requirements (performance, scalability, reliability, etc.)...
Estimate the scale of the system. Consider daily active users, read/write ratio, storage requirements, bandwidth, and any relevant QPS calculations...
Define the APIs expected from the system. This is your chance to analyze and define the read and write paths so that you can come up with the high-level design...
POST /api/v1/upload
POST /api/v1/video/{id}/complete
DELETE /api/v1/RemoveVideo?id
GET /api/v1/search?q
GET /api/v1/metrics?q
GET /api/v1/video
GET manifest/playlist
PUT parts
CDN: It is used to host the frontend of the service. The response is fast because it is in edge locations. Most of the time it handles 80% of the reads since it can cache as well. Media is also served using CDN only
API Gateway: To route traffic based on the endpoint getting hit. Used for authentication and for rate limiting the requests. We will also make sure that based on the client the response is returned meaning phone users will get lightweight response compared to laptop and desktop users.
Kafka: It is used as a high throughput queue to store the requests that are received. With this even if there is a down time it is easy to recover as kafka will have the missed requests. User will receive 200 OK
Redis:To cache the requests and take care of concurrent redundant requests and reduce load on the db and response time. and to store sorted sets so that delivery is fast.
Blob Storage: Great way to store the media that is uploaded into the platform we can use services like S3 for this.
Mongo: Document DB used for this use case.
ElasticSearch: Used to make the searches faster since index is created.
OLAP Aggregations: Used to get metrics like likes, dislikes, view count etc.,
VideoService: Handles everything related to videos from upload to delete. After an upload is done a job lands on queue and is fan out to transcoding workers that produce multple renditions and a completion event.
SearchService: Hanldes the search functionality. Whenever a new upload happens the index is refreshed to keep it up to date.
MetricsService: Used to collect the metrics of the videos uploaded.
dashboards can lag briefly; counts converge
Define the data model. Identify the main entities, their attributes, and relationships. Consider the choice of database type (SQL vs NoSQL) and justify your decision based on access patterns...
For Search we will be using Elasticsearch with index on videoID
For storing the actual media we will use s3 or similar service.
For metadata of the uploads we can go ahead with Mongo
Videos:
ID
Title
Description
Duration
Storage_URL
Uploaded_At
Updated_AT
UserID
ChannelID
Visibility
It will be partitioned by video_id. For a viral video the metadata is servced by redis/CDN cache + read replicas.
We will be using Time-series events for raw view events and OLAP for aggregated view/retention data.
Views
Likes
Dislikes
Shares
Video is predictive. A frame references earlier frames. Codecs group frames into a Group of Pictures (GOP): one keyframe (I-frame, self-contained) followed by P/B frames that depend on the keyframe and each other. Inter-frame compression only works within a GOP — P/B frames can't be decoded without their preceding keyframe.
So if you cut a file at an arbitrary byte offset, you land mid-GOP. The worker's first P/B frame has no keyframe to reference → decode error, and the seam between segments is corrupt.
Where to cut: Split at GOP boundaries (keyframe-aligned): each worker gets one or more whole GOPs, starts exactly on a keyframe, and can encode its chunk independently and in parallel. That's why the transcoder first scans the file (index/EBML) to find keyframe positions, then cuts there.
Two caveats worth knowing: