List the key functional requirements for the system (Ask the AI for hints if stuck)...
List the key non-functional requirements (performance, scalability, reliability, etc.)...
Estimate the scale of the system. Consider daily active users, read/write ratio, storage requirements, bandwidth, and any relevant QPS calculations...
We will estimate QPS, network bandwidth and storage.
For QPS, assuming we have 100 million DAU, and on average, each DAU creates/joins 2 Zoom meetings per day, and view their upcoming Zoom meetings/meeting recordings 5 times. Also assuming peak QPS is twice the average QPS.
Peak write QPS = 100 million * 2 / 3600 / 24 * 2 = 4.6K
Peak read QPS = 100 million * 5 / 3600 / 24 * 2 = 11K
For network bandwidth, assuming 40% of our MAU are in meetings concurrently at peak, and each second, we send 5MB of video/audio data.
We will need 40 million * 5Mb = 200Tb/second of network bandwidth.
For storage, we need to store the meetings metadata, including meeting status, participants etc.
Every day we have 200 million new meetings. Assuming each meeting needs 5Kb of metadata storage, and 1Gb of video/audio storage, on each day, we need:
200 million * 5KB = 1TB of additional metadata storage,
200 million * 1GB = 200PB of additional object storage.
Any data that is more than 30 days old will be put into cold storage.
Define the APIs expected from the system. This is your chance to analyze and define the read and write paths so that you can come up with the high-level design...
I will create the JSON payload examples for the WebSocket signaling messages. This includes the room subscription, the SDP negotiation, the network path selection, and a user action.
Here are the sample payloads for your API design section.
1. Connecting & Subscribing to a Room (subscribe_room)
Sent by the client immediately after establishing the WebSocket connection to join a specific meeting room.
json
{
"event": "subscribe_room",
"timestamp": 1787000000000,
"data": {
"meeting_id": "zoom-meet-8839",
"user_id": "usr-alice-77",
"auth_token": "eyJhbGciOiJIUzI1NiIsIn..."
}
}
2. Exchanging Capabilities (send_sdp)
This payload contains the SDP Offer. It tells the server what codecs and resolutions Alice's device can handle.
json
{
"event": "send_sdp",
"timestamp": 1787000000150,
"data": {
"meeting_id": "zoom-meet-8839",
"sender_id": "usr-alice-77",
"sdp_type": "offer",
"sdp_payload": "v=0\no=alice 2890844526 2890842807 IN IP4 ://anywhere.com\ns=-\nt=0 0\nm=audio 49170 RTP/AVP 0\na=rtpmap:0 PCMU/8000\nm=video 51372 RTP/AVP 31\na=rtpmap:31 H264/90000"
}
}
3. Exchanging Network Paths (send_ice_candidate)
Alice sends her public IP address and port combinations so the media server knows where to direct the UDP traffic.
json
{
"event": "send_ice_candidate",
"timestamp": 1787000000220,
"data": {
"meeting_id": "zoom-meet-8839",
"sender_id": "usr-alice-77",
"candidate": {
"candidate": "candidate:842163049 1 udp 16777215 192.0.2.1 52345 typ srflx raddr 10.0.1.1 rport 52345",
"sdpMid": "video",
"sdpMLineIndex": 1
}
}
}
4. Mid-Meeting Action (room_event)
Sent instantly over the WebSocket when Alice clicks the "Mute" button during the call. The server broadcasts this exact payload to everyone else in the room.
json
{
"event": "room_event",
"timestamp": 1787000543000,
"data": {
"meeting_id": "zoom-meet-8839",
"sender_id": "usr-alice-77",
"action_type": "MUTE_AUDIO",
"is_enabled": true
}
}
For users to join meetings:
POST v1/join_meeting {
user_id: UUID,
meeting_id: UUID,
}
This should return a signaling URL + STUN/TURN credentials (or an SFU region).
For users to create meetings:
POST v1/create_meeting {
user_id: UUID,
invited_participants: List
created_at: Timestamp
}
For users to schedule a meeting for future:
POST v1/schedule_meeting {
user_id: UUID,
invited_participants: List
created_at: Timestamp,
meeting_title: String,
meeting_contents: String,
meeting_metadata: String
}
For users to record their meetings:
POST v1/record_meeting {
user_id: UUID,
meeting_id: UUID
}
For users to share their screen during meetings:
POST v1/share_screen {
user_id: UUID,
meeting_id: UUID
}
For users to end meetings:
POST v1/end_meeting {
user_id: UUID,
meeting_id: UUID
}
For users to view coming meetings:
GET v1/view_meetings {
user_id: UUID
}
For users to view existing chats in a meeting:
GET v1/view_chats {
user_id: UUID,
meeting_id: UUID
}
Describe the overall system architecture. Identify the main components needed to solve the problem end-to-end. Use the diagramming tool to create a block diagram.
First of all, all requests to the system go through external LB, edge API gateway and internal LB. The external LB distribute traffic to different API gateways. For API gateway, we enforce rate limiting and authentication. Then requests go to internal LB to be distributed to different internal services.
We will walk through the critical paths: Creating/scheduling meetings, starting and joining meetings.
For creating/scheduling meetings, we write the meeting metadata to redis, which synchronously gets propagated to the underlying database. It triggers an event on the kafka queue, which gets processed by the notification service, and sends an email/push notification to user's device.
For starting meetings, the meeting creator sends the request, which gets processed by the meeting service. Then the following steps happen:
NAT Traversal (STUN/TURN)
Most participants sit behind NAT/firewalls with private IPs, so they cannot accept incoming media connections directly. We deploy:
Clients get STUN/TURN credentials (short-lived, per-user) from join_meeting. The TURN pool is co-located with SFUs per region so relayed traffic stays regional.
Regional SFU selection
SFUs are deployed in pools per region. Meeting creation assigns the meeting to the SFU pool closest to the host (or majority of participants) — session:{meeting_id} in Redis stores the chosen region/SFU. Large meetings (100+ participants) use cascaded SFUs: a root SFU connects to leaf SFUs, each serving a subgroup, so no single SFU handles all media.
SFU failure & rehoming
SFUs are stateless (no persistent state; media is forwarded, not stored), so a failed SFU can be replaced without data loss. SFUs heartbeat to a controller; on failure, the controller marks the SFU unhealthy, reassigns the meeting to a healthy SFU in the same region, updates session:{meeting_id} in Redis, and pushes a signaling event telling all participants to re-negotiate with the new SFU. Participants experience a 1-3s audio/video blip.Join burst handling
A 100-person meeting can see all participants join within seconds, creating a signaling + SFU setup spike. Each SFU enforces admission control (max concurrent sessions per SFU); when a meeting's SFU is at capacity, the signaling node queues join requests and responds with a "retry after" backoff. This prevents a single popular meeting from overwhelming one media node.
Define the data model. Identify the main entities, their attributes, and relationships. Consider the choice of database type (SQL vs NoSQL) and justify your decision based on access patterns...
Here's all the data that we need to store:
For the metadata, we will use relational database, for several benefits:
The tradeoff is:
However, these are acceptable tradeoffs for our use case.
For meetings recording storage, we will use S3 object storage, with a CDN caching layer.
Here's some sample data models:
table users {
user_id: UUID,
user_name: String,
user_email: String,
user_registered_at: Timestamp,
premium_user: Boolean,
user_metadata: String
}
table meetings {
meeting_id: UUID,
meeting_scheduler: UUID,
meeting_title: String,
meeting_description: String,
meeting_status: String,
meeting_start_time: Timestamp,
meeting_expected_duration: String,
meeting_metadata: String
}
table participants {
meeting_id: UUID,
participant_role: String,
participant_user_id: UUID,
participant_joined_at: Timestamp,
participant_left_at: Timestamp,
participant_metadata: String
}
For querying users, meetings and participants metadata, we cache the results in redis layer with a short-lived TTL. It will be a write-through cache, so any writes to the cache are synchronously written to the database. While this increases write latency, it ensures consistency between cache and database, and data durability in case of cache crash.
If the number of meetings and users increase, we will need to shard the table and create read replicas.
For sharding, we will shard the user table by user_id, meetings table by meeting_id, and the participants table by meeting_id (so we can easily find all participants for a meeting). However, for participants table, sharding by meeting_id means we will have to do cross shard search when we try to find the meetings for a given user. This is an acceptable tradeoff though.
For read replicas, we will implement synchronous replication for all writes. While this increases a bit of write latency, it gives us strong consistency when reading meetings/users metadata.
Active session mapping (Redis, not SQL)
In-flight meetings need a routing lookup that lives only for the meeting's duration:
Key: session:{meeting_id} Value: { "sfu_id": "sfu-042", "signaling_node": "sig-7", "region": "us-east-1", "scheduled_end": 1710000000 } TTL: 24h (or scheduled duration)
Written on meeting creation, read on every join (the join_meeting API returns these connection coordinates), deleted on meeting end. Postgres is the source of truth for metadata; Redis is the source of truth for live routing. This split keeps hot read/write traffic off the relational DB and auto-cleans stale rooms via TTL.
Deep dive into 2-3 key components. Explain how they work, how they scale, discuss tradeoffs, capacity, and any relevant algorithms or data structures.
Here's all the data that we need to store:
For the metadata, we will use relational database, for several benefits:
The tradeoff is:
However, these are acceptable tradeoffs for our use case.
For meetings recording storage, we will use S3 object storage, with a CDN caching layer.
Here's some sample data models:
table users {
user_id: UUID,
user_name: String,
user_email: String,
user_registered_at: Timestamp,
premium_user: Boolean,
user_metadata: String
}
table meetings {
meeting_id: UUID,
meeting_scheduler: UUID,
meeting_title: String,
meeting_description: String,
meeting_status: String,
meeting_start_time: Timestamp,
meeting_expected_duration: String,
meeting_metadata: String
}
table participants {
meeting_id: UUID,
participant_role: String,
participant_user_id: UUID,
participant_joined_at: Timestamp,
participant_left_at: Timestamp,
participant_metadata: String
}
For querying users, meetings and participants metadata, we cache the results in redis layer with a short-lived TTL. It will be a write-through cache, so any writes to the cache are synchronously written to the database. While this increases write latency, it ensures consistency between cache and database, and data durability in case of cache crash.
If the number of meetings and users increase, we will need to shard the table and create read replicas.
For sharding, we will shard the user table by user_id, meetings table by meeting_id, and the participants table by meeting_id (so we can easily find all participants for a meeting). However, for participants table, sharding by meeting_id means we will have to do cross shard search when we try to find the meetings for a given user. This is an acceptable tradeoff though.
For read replicas, we will implement synchronous replication for all writes. While this increases a bit of write latency, it gives us strong consistency when reading meetings/users metadata.
Active session mapping (Redis, not SQL)
In-flight meetings need a routing lookup that lives only for the meeting's duration:
Key: session:{meeting_id} Value: { "sfu_id": "sfu-042", "signaling_node": "sig-7", "region": "us-east-1", "scheduled_end": 1710000000 } TTL: 24h (or scheduled duration)
Written on meeting creation, read on every join (the join_meeting API returns these connection coordinates), deleted on meeting end. Postgres is the source of truth for metadata; Redis is the source of truth for live routing. This split keeps hot read/write traffic off the relational DB and auto-cleans stale rooms via TTL.
For the actual flows, all requests to the system go through external LB, edge API gateway and internal LB. The external LB distribute traffic to different API gateways. For API gateway, we enforce rate limiting and authentication. Then requests go to internal LB to be distributed to different internal services.
We will walk through the critical paths: Creating/scheduling meetings, starting and joining meetings.
For creating/scheduling meetings, we write the meeting metadata to redis, which synchronously gets propagated to the underlying database. It triggers an event on the kafka queue, which gets processed by the notification service, and sends an email/push notification to user's device.
For starting meetings, the meeting creator sends the request, which gets processed by the meeting service. Then the following steps happen:
NAT Traversal (STUN/TURN)
Most participants sit behind NAT/firewalls with private IPs, so they cannot accept incoming media connections directly. We deploy:
Clients get STUN/TURN credentials (short-lived, per-user) from join_meeting. The TURN pool is co-located with SFUs per region so relayed traffic stays regional.
Regional SFU selection
SFUs are deployed in pools per region. Meeting creation assigns the meeting to the SFU pool closest to the host (or majority of participants) — session:{meeting_id} in Redis stores the chosen region/SFU. Large meetings (100+ participants) use cascaded SFUs: a root SFU connects to leaf SFUs, each serving a subgroup, so no single SFU handles all media.
SFU failure & rehoming
SFUs are stateless (no persistent state; media is forwarded, not stored), so a failed SFU can be replaced without data loss. SFUs heartbeat to a controller; on failure, the controller marks the SFU unhealthy, reassigns the meeting to a healthy SFU in the same region, updates session:{meeting_id} in Redis, and pushes a signaling event telling all participants to re-negotiate with the new SFU. Participants experience a 1-3s audio/video blip.Join burst handling
A 100-person meeting can see all participants join within seconds, creating a signaling + SFU setup spike. Each SFU enforces admission control (max concurrent sessions per SFU); when a meeting's SFU is at capacity, the signaling node queues join requests and responds with a "retry after" backoff. This prevents a single popular meeting from overwhelming one media node.
Idempotent join via participant_session_id
When a participant first joins,join_meetingreturns aparticipant_session_id(UUID) that the client stores. If the client's network blips and it reconnects, it sendsjoinagain with the same session id. The signaling server looks up the session:
The SFU also dedupes by session id, so a reconnect that races the signaling server can't double-count the participant. This keeps the participants table and the media fan-out consistent with "one participant = one session."
Simulcast & Last-N forwarding
Senders publish simulcast layers (360p/720p/1080p) in a single stream. The SFU forwards to each receiver based on their bandwidth (WebRTC sender-side bandwidth estimation drives layer switching):
Last-N caps the SFU's fan-out per viewer: for a 100-person call, each viewer receives at most 4 video streams instead of 99, while active-speaker detection keeps the forwarded set relevant. This bounds SFU egress bandwidth and protects weak connections from stalling.