For application deployment, we can have a table called Deployment, with the properties like the following:
Since the deployments volume can be huge, we choose a Non-SQL database, like DynamoDB or Cassandra, for the better horizontal scalability
For cluster info, we need to in favour for consistency over availability in the situation of network partition. Therefore, we can choose a strong consistency metadata database, like Zookeeper or etcd.
Overall the systems conform Master-Slave architecture
Control plane (Master node)
Data plane (Worker node)
Explain how the request flows from end to end in your high level design. Also you could draw a sequence diagram using the diagramming tool to enhance your explanation...
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?