50 million users per month, assuming 1 million writes and 10kb per write
1*10=10GB per month
120GB per year
check_short_exsists(long_url)-> to prevent overwrites, returns none if empty else, the short string
create_short(long_url)-> creates short and returns the result
get_long(short_url)-> Given the short URL sends the long URL, enabling user to visit the page
generate_token(long_url) called by create_short API creates a short based on the hash
URL Table:
1)Primary Key (int)
2) Long URL (varchar)
3) Short URL (var char)
4) Created (DateTime)
5) Access times (int)
graph LR
A[User] -- Requests --> B((Load Balancer))
B -- Requests --> C{Rate Limiter}
C -- Check Rate Limit --> D{Add to Queue}
D -- Queue Size Check --> E{Deny Service}
E -- Deny Request --> F((Blocked Service))
D -- Allow Request --> G{User API}
G -- Process Request --> H[Database or Cache]
H -. Response .-> A
G -- Write Request --> J(Database Write)
J -- Replication --> K[Database]
G -- Read Request --> L((Cache))
L -- Cache Hit --> K
L -- Cache Miss --> M[Database]
N[Background Job] -- Periodically --> O((Database Cleanup))
O --> K
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
1) Using a sql db to allow for better scans to delete old shorts
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?