Datadabase size :
Short URL table will contain :
Therefore each record would require approx : 7+4+4 bytes = 15 bytes
78 billion record multiplied by 15bytes = 1097GB
{
"username" : String,
"email" : String,
"password" : String
}
{
"email" : String,
"password" : String
}
{
"token" : String
}
{
"longUrl" : String
}
{
"shortUrl" : String,
"expires_at" : date
}
{
"urls" : [
{
"shortUrl" : String,
"longUrl" : String,
"created_at" : date,
"expires_at" : date
]
}
}
Get URL Analytics
GET /api/urls/{short_code}/analyticsAuthorization: Bearer {token}{ "access_count": "int", "created_at": "date", "expires_at": "date"}
USER Table :
id - String - Primary Key
name - String - User's name
surname - String - User's surname
password - Hash
created_at - Date - Account creation date
URL
id - String - Primary key
long_url - String - full url
expires_at - date - Expiration date
access_count - int - number of access
user_id - string - reference key
SHORT_URL
id - String - Primary key
short_code - string - generate short url
url_id - string - external key
We need APIs for signup, login, URL converter. We need DB to cache most common URL redirect and to store longurl and short URL. We would need a load balancer to distribute the traffic and some monitoring and analytics in place. The URL convertor logic will use MD5
User submit registration or login details
Service validates credential and create or return user account and auth token
User send a request to create a short url
Service check user's creation limit
Service generate an unique short code if limit hasn't been reached
Service stores the mapping of the newly created url in database
User access to genearte shortUrl
Service look upo in the database
If found, lomg url is returned
Service update access count in databse
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
I gave priority to speed of retrieving data by suggetsing a chaching and DB storing solution.
They APIs should provide all it needs
We haven't considered different geolocation, and we haven't deep dived into any scaling options
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?