advanced

Design a GPU Cluster Scheduler for ML Training

Design a GPU cluster scheduler for large-scale ML training workloads like Meta AI Infra or Anthropics training platform. Support gang scheduling (all N pods start together or none), topology-aware placement across NVLink pairs and same-rack GPUs for AllReduce bandwidth, hierarchical fair-share quotas across orgs/teams/users, priority preemption with coordinated checkpointing, spot-instance handling with graceful drain, resource fragmentation mitigation, and integration with Kubernetes or Ray as the execution layer. Handle thousands of GPUs, queue times under 10 minutes at p50, and fault recovery without restart-from-scratch.

Loading the workspace...