Kubernetes 1.37 ("Garhwal") is out, and a lot of it is aimed at AI workloads: a stable Metrics API, rootless Kubelet in beta, and scheduler work built for GPU jobs. If you've been running GPU inference in containers and felt like the tooling was playing catch-up with the hardware, this release is where the gap starts closing.
I run inference on two NVIDIA DGX Sparks bridged over ConnectX-7, and the problems that setup keeps raising (seeing what the hardware is doing, isolating workloads, accounting for hardware the scheduler doesn't understand) are the ones 1.37 takes aim at. Here's what matters and why.
The Stable Metrics API Is a Bigger Deal Than It Sounds
For a long time, getting reliable resource metrics out of Kubernetes meant either trusting metrics-server and hoping it didn't drift, or bolting on a full Prometheus + Grafana stack before you could answer the question "what is this pod actually consuming right now." The Metrics API graduating to stable in 1.37 changes that baseline.
For AI/ML workloads this matters because inference pods don't behave like typical web services. A model serving endpoint can sit idle for minutes, then spike hard when a request batch hits. Standard autoscaling heuristics built on CPU and memory don't capture that shape well. A stable, well-defined Metrics API gives custom schedulers and HPA configurations a reliable foundation, and it makes custom metrics adapters (the kind you'd wire to GPU utilization or VRAM pressure) much more tractable to implement correctly.
DriftWatch has a small version of the same shape. The serving container sits at zero replicas until a request arrives, and retraining only runs when the drift check, a GitHub Actions workflow on a six-hour schedule, finds something. I sidestepped the autoscaling problem by putting the serving side on Azure Container Apps with scale-to-zero and accepting the cold start. On Kubernetes you'd be building that scaling signal yourself, and a stable Metrics API is what you'd want underneath it.
Why Rootless Kubelet in Beta Matters for AI Deployments
Rootless Kubelet is the kind of security hardening that gets skipped in the "we just need this to work" phase of standing up an AI cluster. The trouble is that GPU nodes are high-value targets. A compromised container on a DGX Spark with 128 GB of unified memory is sitting next to model weights, inference logs, and possibly API keys baked into environment variables.
Running Kubelet without root shrinks the blast radius: if a container escapes, it lands on a host where the Kubelet and container runtime aren't running as root. For enterprise deployments, especially RAG systems like DocQuery that sit on top of sensitive documents, I'd treat that as table stakes.
Rootless operation took a long time to get here because it has non-trivial interactions with volume mounts, device plugins (which is how GPU resources get exposed to containers), and some CNI configurations. Shipping it to beta in a release themed around AI/ML suggests the device plugin story has improved, but that's the piece I'd check before depending on it for GPU workloads.
The Scheduler Changes That Matter for GPU Jobs
AI/ML optimization in 1.37 is a set of changes rather than one headline feature, and two of them matter most for GPU work. Extended resource support in Dynamic Resource Allocation (DRA) went GA, so a pod can keep asking for GPUs with the familiar extended-resource syntax while a DRA driver handles the allocation underneath. And the Workload and PodGroup APIs for gang scheduling reached beta, along with workload-aware preemption.
Gang scheduling is the one I'd watch. If I moved my two-Spark setup onto Kubernetes, where one model runs tensor-parallel across both nodes, half a deployment would be useless: if one pod schedules and the other can't, the first sits there holding a node's worth of GPU memory and doing nothing. Gang scheduling places the whole group or none of it. That's the piece Kubernetes has been missing for multi-node inference, and it's a big reason people reached for Volcano.
I checked NVIDIA's release notes, and GPU Operator v26.7.0 is the release that added Kubernetes 1.37 support. If you run the operator, upgrade it alongside the cluster. It's the component that advertises your GPUs to the scheduler, and you don't want it a version behind the thing reading those advertisements.
Upgrade Deliberately
1.37 is worth taking seriously if you're running AI workloads in Kubernetes. That doesn't mean upgrading your production cluster tomorrow. Rootless Kubelet is still beta, so you don't want it on a node where a GPU workload failing to start costs you a real SLA. Try it in a non-critical environment first.
Start with the Metrics API changes. They're stable and backward compatible, and better observability costs almost nothing while paying off in every debugging session after. Then get the GPU Operator to v26.7.0 or later before you touch anything on the GPU nodes themselves.
I'll be watching rootless Kubelet as it moves toward stable. If your models process documents, code, or anything else you wouldn't want exfiltrated, the security argument is strong enough to run it in a dev environment now, so you learn its failure modes before production forces you to.
It's good to see Kubernetes catching up to GPU-heavy workloads. The scheduler was never designed with "this pod needs 128 GB of unified memory on a specific node topology" in mind, and 1.37 moves it closer. Just keep track of what's stable and what's still being sorted out.