Let's assume total customer size is 1M
And 10% of the customer size uses the system on a daily basis ==> roughly 10000 users/ day
In which lets say there 5 reads per 1 write
Data I will store:
Let's say all 10000 users create tinyURL ==> 10000 * 500= 5 MB/ day for DB1.
For DB2 100* 10000 is roughly 1MB/ day
that is DB1 will be 300MB/ month in peak times we can estimate like (300 MB * 1000/ month) roughly 300GB
and for DB2 1MB * 30 ==> 30MB/ month and in peak time it will 30 GB
Both are not large boundaries, so I want to use in-memory data base for the caching layer owing to its high availability and low latency which are very important for this system and then use NoSQL in the back ground to persist the database
The prefect database for this will be a persistent key-value store (Redis or memcached); because we need to account for peak traffics we need to have NoSQL in the background which will ensure seemless scalabilty while leveraging the benefits of memcached/Redis in the caching layer
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?