List functional requirements for the system (Ask the chat bot for hints if stuck.)...
List non-functional requirements for the system..
Estimate the scale of the system you are going to design...
Let's have 1 million DAU.
Every status has userId, a status value, and a timestamp.
UserId takes 8 bytes, status value takes 1 byte, and a timestamp takes 8 bytes. One status event takes 20 bytes.
Every user's status is changed 5 times daily.
It would give 100 million bytes every day of new statuses.
Let's keep this information for one year, and it takes 100 Gigabytes (two replications)
Let's use a key-value storage like AWS DynamoDB.
Define what APIs are expected from the system...
getUserStatus(userId,fromDate,toDate) returns the user status list for the given time range
updateUserStatus(userId, status, currentTime) update the user status
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
We would use a NOSQL solution like AWS DynamoDB.
It has concurrent write and read speeds.
The key is userId
The value is the timestamp and status.
UserId would be the partition key, and timestamp and status are created as the secondary index.
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
Loadbalancer provides DDoS protection and TLS termination, and forwards request to the right service nodes.
The User Presence Service handles the user status requests and it's a stateless service. AWS DynamoDB persists the user status information.
Most demanding users' statuses would be cached by the Cache.
The User Presence Service gets the user's status update from Kafka.
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
The user sends a request to get the user's status for the given time range to the LoadBalancer. The Load balancer distributes the load between the User Presence Service instances by using either round robin, weighted round robin, the lease connection number, or other strategies. The User Presence Service checks the cache if it has the user statuses; if so, it returns them to the client, otherwise it goes to the AWS DynamoDB to find user statuses by userId and the time ranges, after getting the user statuses, The User Presence Service puts the user's status to the cache.
The cache uses LRU, LFU or FIFO to store the user statuses efficiently.
The Kafka connects the external event source and when new user status update is consumed by the User Presence Service, it persists it to AWS DynamoDB.
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?