The total number of users = 500 million.
Total number of daily active users = 100 million
The average number of files stored by each user = 200
The average size of each file = 1 MB
Total number of active connections per minute = 1 million
Storage Estimations:
Total number of files = 500 million * 200 = 100 billion
Total storage required = 100 billion * 1 MB = 100 PB
Considering 1 server can handle 1000 requests concurrently, we would need 1 Million / 1000 = 1000 servers
User Authentication API:
2. File Upload API:
3. File Download API:
4. File Management API:
5. File Synchronization API:
6. Sharing and Collaboration API:
7. Version Control API:
8. File Search API:
For the tables required in this design, refer to the class diagram, the list of classes is not exhaustive but this is a good number of tables to start with.
Database Choice
Data Partitioning:
Regional or Geographical Partitioning:
Sharding Strategy:
Sharding Key Selection:
Replication:
Load Balancing:
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design...
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?