List the key functional requirements for the system (Ask the AI for hints if stuck)...
List the key non-functional requirements (performance, scalability, reliability, etc.)...
Estimate the scale of the system. Consider daily active users, read/write ratio, storage requirements, bandwidth, and any relevant QPS calculations...
POST /api/v1/group - groupName, Description, Visibility, createdAt
DELETE /api/v1/group
GET /api/v1/groups
PATCH /api/v1/group/{id}
GET /api/v1/search?q=
POST /api/v1/shareFile
POST /api/v1/Authenticate
GET /api/v1/tasks?assigneeId=&status=
POST /api/v1/tasks
PATCH /api/v1/tasks/{id}
WS sendMessage
CDN: It is used to host the frontend of the service. The response is fast because it is in edge locations. Most of the time it handles 80% of the reads since it can cache as well. Media is also served using CDN only
API Gateway: To route traffic based on the endpoint getting hit. Used for authentication and for rate limiting the requests.
Kafka: It is used as a high throughput queue to store the requests that are received. With this even if there is a down time it is easy to recover as kafka will have the missed requests. User will receive 200 OK
Redis:To cache the requests and take care of concurrent redundant requests and reduce load on the db and response time. and to store sorted sets so that delivery is fast.
Blob Storage: Great way to store the media that is uploaded into the platform we can use services like S3 for this.
Mongo: Document DB used for this use case.
Message Service: It is used to send and receive message. Uses Websockets to keep the messaging real time. When ever a message event occurs it sends that to Kafka topic which in turn used by notification service.
Notification Service: Used to send notification to subscribers.
Group Service: Handles all the Group related operations.
Connection Manager: Handles all the connections and this is where the persistent connection actually lives.
resume path: client reconnects with last_message_id → service replays missed messages from the log/DB.
For file storage we will use blob storage like s3.
We'll use postgres to store the group table
Group:
ID
Name
Description
Visibility
CreatedAt
UpdatedAt
Members
Users:
ID
Name
Phone No
Status
CreatedAt
UpdatedAt
Task:
ID
Name
Description
Assigned
CreatedAt
UpdatedAt
For messages we will use Mongo DB:
Message:
ID
To
From
Text
Attachments
CreatedAt
UpdatedAt
Automated daily snapshots of Postgres + MongoDB, with S3 versioning for blobs. Encryption at rest and in transit using AES-256 at rest and TLS in transit.
We'll be using redis for presence status and recent messages.
Multi-region failover for disaster recovery.
For searching we will be using elasticsearch
If a action fails we even retry it for a set amount of times lets say 3. If it fails more than that set limit it will go to the DLQ. We will be retrying using exponential backoff with jitter to avoid multiple retries at the same time. This will prevent any load on the respective service.
We will be having replicas to adhere to high availability so that when ever a service fails it's replica takes over. We can go even further by having regional failovers so that if a region goes down we can switch to the other region.
client disconnects, reconnects with last_message_id, node replays missed messages from the log/DB.
Kafka partition key = group_id, so a conversation's messages stay ordered even with retries and multiple consumers
connection map in Redis lets any node rebind the session; no sticky-socket dependency.