List functional requirements for the system (Ask the chat bot for hints if stuck.)...
Need to build a web application
alphanumeric with 6-8 characters
assuming url's are stored for 1 year
each url size is 1KB
assuming always redirection happens when a short url is called
List non-functional requirements for the system...
Avoid collisions
scallable
Estimate the scale of the system you are going to design...
As the size of URL is 6 characters long, we can produce upto 62^6=5.6 billion unique urls.
writes 200rps assuming 1kb per write. So in total 200*1kb=200kb
in one day = 200*86400 = 17280000 17.2 GB / day
for 1 year 17.2*365 =6.27 TB / year as a bottle neck we consider 10TB/year
for reads, 20,000 rps
this means we need more reads capacity units than WCU's.
Since there is need for low latency and high scallability and since the reda requests are more we can use nosql database DynamoDB.
Define what APIs are expected from the system...
GenerateID
it takes the long url as input and generates a unique hash ID and stores it in out database. This will be called when we are creating the short url for a long url
Redirection
This API is a POST request which has short url as a parameter and it fetches the respective long url assigned to the short url be querying the dynamo db and redirects the call to the long url.
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
ID, long url, short url
these were the three columns in our dynamo db.
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
The client first gets the IP address for the givem domain naim through the DNS and then the call goes to one of the server allocated by the Load balancer. Then if the request is for creating a new short url then generateID api is called. Here the generate ID api will take the long url and then generates a short URL hash value only if the long url is not already present in the DB. I will use a incremental counter and do a base64 encoding to it so that there won't be any collisions. Once the hash value is there we put it the DB and return the same to the client. on the other hand if we want long url given the short url Rediction api is called which querie the Redis cache first and if available then a 302 request with the original url set as long url will be sent to the client for temporary redirection. if not found in cache then will look in DB and update the cache and returns 302.
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Explain any trade offs you have made and why you made certain tech choices...
I'm scalling the servers horizontally and using a load balancer to distributed the traffic evenly across all the servers. Since the rps is more. Also I'm using DynamoDB which is a NOSQL db because the there is no need of complex querying and also the data is large and also the read request are a lot.
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?