← All posts

#infrastructure

2 posts

Bending the Kubernetes scheduler

Bending the Kubernetes scheduler

How we make the default Kubernetes scheduler pack GPU workloads into the cluster first and spill to overflow capacity only when it's full, using dynamic taints, the Descheduler, and Kueue.

Train in cluster, infer anywhere

Train in cluster, infer anywhere

How Krea shares GPUs between research and production with a custom solution that lets inference scale anywhere when training claims every GPU.