For the scope of our problem, we will only consider the following resources in our architecture,
For capacity estimation we need to consider that our internal components like Database Backup Component, VM-Disk Backup Component should be able to take backups and snapshots efficiently even when the number of database copies and vm-disks increase in number.
let us consider a system where we have 100 Virtual machines, 2 VM disks, that is 200 VM-disks total, 100 database servers in each region. So our systems should be able to handle replication and do backup for the above systems within reasonable times.
Replication
Backup
Backup Retention
Below are a few API's we would need for our system, these are not exhaustive lists but some of the most essential ones.
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
Below are the components which will be require to support the disaster recovery system.
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?