List functional requirements for the system (Ask the chat bot for hints if stuck.)...
1) When a long url is given as an input it's short alias needs to be returned.
2) Generate a short url
3) Store the mapping between short url and long url in a store.
List non-functional requirements for the system...
User growth
Latency
Storage
Concurrency
Estimate the scale of the system you are going to design...
The system which can handle 100 million hits atleast but which can scale as needed.
Define what APIs are expected from the system...
POST /api/tinyUrl
Host: bloomberg.com
Accept: application/json
Content-Type: application/json
Content-Length:
{
"longUrl": "https://bloomberg.com"
}
GET /api/tinyUrl?shortUrl=https://abcly
Host: bloomberg.com
Accept: application/json
Content-Type: application/json
Content-Length:
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
A map data store with the short url as the key and long url as the value
short url - string <-> long url - string
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
The request with a long url from a client goes to the server. Service looks up the database and retrieves if available. Else, returns an error.
For faster retrieval an LRU cache can be maintained before a DB call.
Collision detection - Given a long url ignore the case before applying a hash (to avoid collisions ?). To really check for anomalies - look up during the POST request in the map database. If there's a collision - needs to be resolved
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Database: A key-value like Redis suits the tiny url usecase very well because of it's faster lookups. The data stores can be sharded based on a hash for even better performance. Also, some of the key-value databases can be scaled out of the box.
Load balancer:
A load balancer can help to distribute the incoming traffic among the service instances. Can even begin with round robin to distribute. However, the strategy to balancing the load needs to change according to other factors like geographical location.
Cache:
A LRU cache which gets filled and evicted based on the url hits will help reduce the latencies. To serve globally, a regional cache based on a location can be useful too. The eviction policy is based on the least used url in a day.
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?