~300K-500K servers, avg. reservation ~2hrs.
~20-50 active flags/server (sparse over a ~1000 vocab)
Peak alloc/release ≈ 5–10K ops/sec, bursty (10K concurrent users)
Inventory/health reads = 10x allocation volume
~100K total users/tenants in the org
gRPC: allocate / release / heartbeat sit on the low-latency hot path and heartbeat renewal is naturally a stream - gRPC fits;
rpc RequestServers(RequestServersReq) returns (RequestServersResp); message RequestServersReq { repeated string flags = 1; int32 count = 2; } message RequestServersResp { repeated Server servers = 1; }
REST (admin/read): adding flags and browsing inventory is opes-facing dashboard traffic.
Client -> allocate (gRPC) -> API Gateway -> Load Balancer -> Allocate APIs -> Index (roaring bitmaps)
Client -> Lease Mgr (TTL = reaper) -> Shared Store (etcd/FoundationDB - replicated Quorum) -> flag index (CAS reserve, return to pool)
Servers ->
Partitioned KV w/ CAS (etcd / FoundationDB)In-memory inverted index (Redis + roaring bitmaps)Postgres (durable catalog)