List functional requirements for the system (Ask the chat bot for hints if stuck.)...
List non-functional requirements for the system...
Estimate the scale of the system you are going to design...
Let's assume we have:
Based on these estimations, on average an active user issue 3 read requests per minute.
Total read requests rate = 10000 (peak users) * 3 (requests per minute per user) = 30000 read requests per minute.
Average write requests per minute - 5 * (1 (comment) + 2 (likes) + 1(reply) = 20
Assume each comment and its replies cost around 1KB. We need to store
5 (comments per minute) * 60 * 24 * 365 * 3 = 8GB data.
This data could be stored in a single database machine.
Define what APIs are expected from the system...
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
Topics Table
Comments Table
Users Table
Upvotes Table
Downvotes Table
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
The system should include the following items:
Client (Publisher + Receiver): The client plays two different roles in the system: comment publishers and comment receivers.
Servers (Dispatcher + Gateway Server): The servers play two roles in the system: the dispatcher broadcasts the comments to appropriate servers (and subsequently receivers), and the Gateway Server forwards them to the receivers.
Storage (Endpoint Store + comments store + in-memory subscription store): The system has three storage locations—the comment store stores comments, user information, etc. The Endpoint Store contains information on which gateway servers are subscribed to what topics. The In-memory subscription store resides in the Gateway Server and contains information on which receivers want to receive comments on which topics.
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
The request flow should be as follows:
There are two main request flows:
When a user logs in and starts to read a topic.
When a user publishes a comment on a topic:
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Gateway Server: In order to achieve real-time comments refresh, each gateway server maintains SSE connections to all receiving clients. Each server also contains an in-memory subscription store, so that it does not fan out comments to all receivers unnecessarily.
Endpoint Store: This storage is stand-alone so that dispatchers does not need to maintain a state, therefore, the publishing load can be flexibly distributed across all dispatchers.
Explain any trade offs you have made and why you made certain tech choices...
SSE is chosen over a pull-based model for communication between Gateway Server and the receiver as
SSE is chosen over web socket for the following reasons:
Try to discuss as many failure scenarios/bottlenecks as possible.
Examples of failure scenarios are as follows:
All components, except for publisher and receiver, should have duplication to distribute load and increase reliability.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?
In the future, more duplications and partitions should be introduced as the user scale up.