List functional requirements for the system (Ask the chat bot for hints if stuck.)...
List non-functional requirements for the system...
Estimate the scale of the system you are going to design...
long: 100 character = 100 byte
short:short domain + 10 characters -- 20 characters 20 byte
createAt: 10 bytes
updatedAt: 10 bytes
1 million url, 0.14GB
(1000 new url per day * 30 *12 * 3years= 1 million url in db)
shortUrl: 20 byte
count: each url is touched 1000 times -- 4 bytes
1 million url, 10E6 * 25 = 25 MB
target market small business, peek estimation:
load balancing, redis, cdn?
Define what APIs are expected from the system...
request body: {"url": "https://...", "expiration?: "}
response: {"shorturl": "https://short.url"}
response: 302/301 redirect, 302 is more suitable since its temporary redirect so the site doesnt cache the redirect
error response:
503 service unaviable, short url hasnt been updated to the db yet, ask client to retry in 1 min
410 gone, short url has expired, notify the client
Response:
{"shortUrl": "https://shorturl/...",
"longUrl": "https://...",
"cliclCount": 100,
"createdAt": "*******"
"expiredAt": "*******"
}
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
a sql db:
INDEX on shortUrl
url: [shortUrl, longUrl, expiratedAt, createdAt]
stat: [shortUrl, count]
the url is high read low write, stat is high write low read
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
SQL:
INDEX on shortUrl
url: [shortUrl, longUrl, expiratedAt, createdAt]
stat: [shortUrl, count]
the url is high read low write, stat is high write low read
choose sql over nosql because:
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
check my sequence diagram
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
URL shortening service:
1 url + timestamp -> hash -> base 62; then check for conflict. if conflict, hash again
2 INCR the count -> hash -> base 62 -> check for conflict. if conflict regenerate hash
option 2 is favored here since each count is unique, so its likely to create a unique url
1 check for conflict
2 if conflict, fall back to incr -> hash -> base 62, add the part to the end the alias
Load Balancer:
Server:
DB:
url table:
stat table:
Redis:
cache TTL:
Database:
Database TTL:
Explain any trade offs you have made and why you made certain tech choices...
Master slave replication vs peer to peer replication:
Master-slave: 1db handles all write and all handles read
peer-peer: all serves as write and read
peer-to-peer is preferable because:
AWS Lambda vs batch update count through redis:
Lambda:
Through redis:
Currently, redis batch count is perferable, but if more stat is requried, we can add the microservice here
Try to discuss as many failure scenarios/bottlenecks as possible.
webserver is off:
Redis failure:
DB failure:
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?
added lambda to give more stat
check for url security in case some site use our service for phishing