Bending the Kubernetes scheduler
How we make the default Kubernetes scheduler pack GPU workloads into the cluster first and spill to overflow capacity only when it's full, using dynamic taints, the Descheduler, and Kueue.
Train in cluster, infer anywhere
How Krea shares GPUs between research and production with a custom solution that lets inference scale anywhere when training claims every GPU.