+ Takes in a URL (string), validates that it's a valid URL. Provides a consistent hashing function which will return the same string. The format of the URL will be of the form `https://shortn/
+ Users can access URLs and they will be redirected to a long URL.
+ Abuse vectors.
+ Authentication
+ Analytics
1 table, 840 bytes per row.
1 cache mapping with 10k entries. Each entry takes 828 bytes.
Assuming we will need to store 10 million websites, our table will be of size 8.4 GB. This fits in a single server but for resiliency, we will need to use more than one.
Our cache needs 8.2 MB, for resiliency, we will also need multiple servers.
2 endpoints:
+ /api/shortened_url [POST]
+ /
Our database needs a single table called UrlMappings:
UrlMappings:
MappingId: Int (4bytes)
OriginalUrl: string(200 characters) (800 bytes)
ShortenedUrl: string(7 characters) (28 bytes)
CreatedTime: datetime (4 bytes)
Total: 840bytes
2 major user journeys:
+ Users send a request to generate a URL.
The server will query the cache to see if it exists. If it doesn't exist it will query the server to see if the URL exists. If it doesn't exist:
The server will generate a random 6 letter code(capitals included). The server queries the DB until it finds the first code that isn't repeated. The server then inserts it into the DB and the cache.
+ Users send a request to query a certain shortened URL.
The server queries the cache. If the URL is stored in the cache, we redirect to it. If it's not in the cache, we query the database and redirect. If it's not in the database, we redirect to an error page.
We use a cache to address hotspots. The cache uses a LRU strategy.
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
For the Database, we can use something like MySQL, DynamoDB or Spanner.
For the cache, we can use something like redis or memcache.
For the database we can partition by MappingId, since the cache takes care of the hotspots.
For the servers we will use some load balancing and web servers with maybe flask, apache, etc...
No particular load balancing is needed here. It's purely stateless architecture.
Increased cost (cache) to reduce hotspots in DB.
Potential hotspots when users try to generate short urls for popular websites. In this case, the cache alleviates the load on the DB.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?