Capacity estimation
Database
DB schema
// Reservation
reserve(spec) -> Reservation
{
id, // output
cpu: { atleast: 4 }
memory: { at least: 16g }
count: 4,
accelerator: H100
...
}
// Release
release(id) -> HTTP status
We use an ACID database to persist server, VM information and reservations. This can be fronted with a cache, but when making reservations we do that in a database transaction to ensure strong consistency.
Since reservation may take up to minutes if new servers/VMs needs to be created, the services sends the request to a queue, where the request is picked up by the server pool is satisfy the request. When the request is satisfied, the result (e.g. new VM ids) are written back to the DB.
The monitoring service serves 2 purposes:
For metrics reporting, we can use a pull or push based model. We choose a push based model, buffered by an EventQueue, to support higher throughput of the metrics. The qps (assume each server/VM reports 1 / sec) is 600k / sec. These data can be sent to a TSDB for real-time visualization/alerting. However, for the reservation system we only care about static attributes (e.g. # of CPUs). We might want to analyze the dynamic data to allow over-allocation, for example.
To find and allocation VMs, we need to define the objective we are optimizing for. For example, it could be the lowest cost, or VM collocation on the same and nearby servers (bin-packing to save resources, vs. round robin allocation to spread out the load, best fit, first fit). Such objective can be configurable.
ReservationService is stateless and horizontally scalable. The DB and cache sizes are small enough to fit into a single cluster. We should enable replicas and hot standbys for high availability. The EventQueue needs partitioning (e.g. Kafka with 4 partitions so each partition handles 150k message/sec). If there is a sudden surge of events, the Kafka queue acts as a buffer.
When there is a reservation spike, we can address it in several ways: