List functional requirements for the system (Ask the chat bot for hints if stuck.)...
regulate volume
List non-functional requirements for the system...
high availablity
scaliability
Estimate the scale of the system you are going to design...
1Million user
each user request 200 daily
about 50% active user daily
QPS = 1Million * 0.5 * 200 / 86400 > 1000
Define what APIs are expected from the system...
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
NOSQL like dynamodb to store bucket information
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
There are 2 strategys used for rate limiter to limite request: token bucket and leaky bucket. token bucket can work with suddenly increase requests but need to maintain token. leaky bucket cannot work with suddenly increase requests. We will use token bucket in this design
We maintain token in cache instead of in the request service because when there are lots of request request service do not need to maintain token which will avoid slow down the system
cache also have hot key issue, we can use a combination of userId and request type as the primary key to do partition and shard to make request evenly distribute. However, if an user send same request multiple time it will slow down the cache, then we can ask request service to maintain a token to mitigate the request to cache. request service timely synchronous it's token with cache.
each request service, cache and dynamDB have multiple replics to avoid single point failure
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?