List the key functional requirements for the system (Ask the AI for hints if stuck)...
List the key non-functional requirements (performance, scalability, reliability, etc.)...
Estimate the scale of the system. Consider daily active users, read/write ratio, storage requirements, bandwidth, and any relevant QPS calculations...
Define the APIs expected from the system. This is your chance to analyze and define the read and write paths so that you can come up with the high-level design...
POST /api/v1/scheduleCall
POST /api/v1/share
POST /api/v1/Join
POST /api/v1/End
GET /api/v1/ListMeetings
GET /api/v1/GetMeeting?id=
WS /api/v1/signal
peers don't send media to each other directly (mesh) and don't send it to the app servers. They send it to an SFU — a Selective Forwarding Unit that receives each participant's single upstream and forwards it to the others. It's what makes group calls scale: with mesh, N participants need N×(N-1) streams; with an SFU it's N uploads to one node.
create/join → signaling over WebSocket (SDP offer/answer) → ICE via STUN, TURN relay as fallback → media flows through the regional SFU.
Join burst — gateway + signaling LB absorb simultaneous joins at meeting start
Regional routing — SFU Selector pins the meeting to a nearby SFU pool
NAT traversal — direct path first, TURN relay if blocked
SFU failover — if a node dies, rehome the meeting to a healthy node
Reconnect — brief drop → ICE restart/resignal
Adaptive media — SFU degrades bitrate/resolution on poor network
Clean teardown — last/host leaves → release SFU, clear Redis session
Postgres
Users
user_id (PK), name, email, status, created_at
Meetings
meeting_id (PK), host_id (FK→Users), title, description, scheduled_time, is_recurring, passcode, created_at INDEX (host_id), INDEX (scheduled_time)
Participants
participant_id (PK), meeting_id (FK→Meetings), user_id (FK→Users), role (host/attendee), joined_at, left_at, status INDEX (meeting_id), INDEX (user_id)
Redis:
ActiveSession: meeting_id → { signaling_node_id, sfu_id, region, started_at }
Presence: meeting_id → set of user_ids currently connected
QoS Telemetry:
call_stats: timestamp, meeting_id, user_id, packet_loss, jitter, bitrate, rtt
1. SFU (Selective Forwarding Unit)
How it works: Each participant sends one upstream stream, encoded as simulcast — multiple layers (e.g., 1080p / 720p / 360p). The SFU subscribes each receiver to the best layer it can handle, based on the receiver's reported bandwidth. This is Last-N + simulcast: forward the top-4 active speakers, and pick the layer per subscriber.
Data structures: Per room, a routing table subscriber → {publisher, selected_layer}. Audio is prioritized (always forwarded); video layers downgrade first.
Adaptive behavior: If a subscriber's bandwidth drops, the SFU silently switches them to a lower simulcast layer — no renegotiation, no stall. That keeps the active speaker clear under bandwidth shifts.
Capacity: Bound each SFU by egress bandwidth, not CPU. E.g., a node with 10 Gbps egress at 1.5 Mbps/subscriber ≈ ~6,600 downstream subscriptions → cap at ~4,000 with headroom.
2. Signaling Service (WebSocket)
How it works: Stateful WebSocket cluster. Holds each participant's connection; relays SDP offer/answer and ICE candidates; broadcasts call-control events (mute, join, leave).
Stickiness (major): Route a user to the same signaling node for the room (room→node mapping in Redis), so a brief disconnect reconnects to the same node and resumes — no duplicate participant.
Reconnect: Client carries a rejoin token. On reconnect, the server matches the token → restores the participant instead of creating a new one (idempotent rejoin).
Signaling hiccup ≠ dropped call: Media flows peer↔SFU directly, so a signaling blip doesn't tear down the call — the client retries/resubscribes and media continues.
3. SFU Selector / Placement
How it works: On first join, the signaling node asks the Selector for a media node. The Selector uses least-loaded placement with region affinity — pick a healthy SFU in the user's region with capacity.
Admission control (the GATE): Each SFU reports load to the Selector (via a registry/heartbeat). If a node is at capacity, the Selector refuses placement and picks another. Redis holds the room → sfu_node pin so all joiners land on the same SFU.
Failover: Heartbeats miss → node marked unhealthy → rooms rehomed to a healthy node → peers get a renegotiation signal (ICE restart). Brief blip, call continues.
Relay containment (minor): STUN first, TURN fallback; multiple relay pools + autoscale; rate-limit abusive clients to the relay tier.