Estimate the scale of the system you are going to design...
Define what APIs are expected from the system...
Search endpoint:
GET /search
TODO: add more details about pagination
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
We can consider using a distributed NoSQL database for the tweet store like MongoDB. NoSQL databases are easy to scale and have high availability. They are also flexible, so we can change schema design as needed to accommodate different tweet content and attributes like hashtags.
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
The design can be broken into two main services -- one for handling search queries and another for indexing tweeted content.
High-level flow:
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
We can consider using ElasticSearch with full-text search to help index and search tweet content based on keywords, hashtags, and user IDs. We'll leverage MongoDB as the primary database for storing tweet information. We will have a separate indexing service to batch indexing. The service will be scheduled so we can group new tweets into larger indexing operations to reduce overhead and improve efficiency of index updates. We can adjust this indexing frequency based on the traffic load on the service or time intervals or both. We can optimize for indexing most recent tweets, since users probably want to view the most recent events.
Data indexing approach
Explain any trade offs you have made and why you made certain tech choices...
Redis vs ElasticSearch
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?