List functional requirements for the system (Ask interviewer if stuck)...
List non-functional requirements for the system...
Estimate the scale of the system you are going to design...
Define what APIs are expected from the system...
CRUD on /tags/:
Tag Search (/search?tags):
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
For the user table, we can use SQL tech, i.e postgres,
same for the others table,
I don't see a specific requirement on the table here
one important thing is normalization of data before storing it in the db,
lower_case
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design...
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
For the tag operation,
the client connects to the load balancer, the load balancer redirects the request to the API gateway, which requests to the correct service, the tagging service, the tagging service writes in the DB the content_link and the associated tags in the DB,
using change data capture and a log based message broker, the CDC is sent to flink, which puts the data into the inverted search index,
on the search path, the API gateway redirects to the data to the search service, which queries the search index, and returns the id of the link,
which are queries in the DB and returns to the client
we will make use of caches (i.e) redis to retrieve most common links
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
sharding and index
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?