Kubernetes 1.37 Starts the Long Goodbye for IPVS and kube-dns
The release brings scale-to-zero autoscaling and a tougher API server, but operators should read the deprecation section first.
The Kubernetes project shipped v1.37, named Garhwal, on 26 August. It contains 67 enhancements: 16 graduating to stable, 23 to beta, 27 new alpha features and one deprecation or removal. The release was led by Dipesh Rawat.
Release posts lead with features, and this one has several good ones. But for people who run clusters, the more consequential part sits near the bottom, where two long-standing components get end dates.
The deprecations
kube-proxy's IPVS mode is deprecated. Clusters running kube-proxy with mode: ipvs now log a warning at startup. The plan is for IPVS mode to be disabled by default by v1.40, still selectable through a feature gate, and removed entirely by v1.43. The project's reasoning is blunt: the kernel's IPVS API cannot implement Kubernetes Services on its own, so IPVS mode has always relied on iptables underneath anyway. IPVS mode was introduced back in v1.8 to fix iptables performance bottlenecks, and many teams adopted it for exactly that reason. Those teams now need a migration plan.
kube-dns is deprecated. CoreDNS has been the default since v1.13, and kube-dns lacks support for EndpointSlices and dual-stack Services. No new kube-dns packages are expected after v1.40. If an old cluster is still running it, this is the nudge.
Two smaller changes could also bite. Static Pods can no longer reference Secrets or ConfigMaps; that was always a bug, and the feature gate that let you keep relying on it has been removed. And the slow retirement of cgroup v1 continues: since v1.35 the kubelet refuses to start on cgroup v1 nodes unless failCgroupV1: false is set, and the project is clear that the override is temporary.
What is worth having
HorizontalPodAutoscaler scale to zero reaches beta and is enabled by default. With spec.minReplicas: 0, an HPA driven by object or external metrics, such as queue depth, can scale a workload down to nothing and back up when work arrives. It does not work with CPU or memory metrics, which need running Pods to report anything. For queue consumers and GPU batch jobs this is real money, and it removes one reason to add a separate event-driven autoscaler.
The API server protects etcd more firmly. The work on resilient watch cache initialisation is now stable and locked on. Instead of letting expensive list and watch requests pile onto etcd while the cache warms up, the API server delegates bounded requests and rejects the rest with HTTP 429. That is good for the control plane, but it moves the burden onto clients. Custom controllers and operators that do not respect Retry-After and back off properly will turn a cache warm-up into an error storm. Check your in-house operators.
SELinux volume mounts change behaviour. With SELinuxMount now stable and on by default, volumes whose CSI drivers opt in are mounted with a single SELinux context instead of being relabelled recursively. Pods with different SELinux labels sharing one volume on the same node, which used to work, can now fail to start. Setting seLinuxChangePolicy: Recursive on the Pod keeps the old behaviour, and it can still be disabled cluster-wide until v1.38. Clusters without SELinux are unaffected.
Elsewhere: the metrics.k8s.io API finally reaches stable after nearly nine years in beta; storage version migration moves in-tree and on by default; Pod certificates and ClusterTrustBundles are stable; gang scheduling for AI and HPC jobs reaches beta; KYAML, a stricter subset of YAML, is stable via kubectl get -o kyaml; and manifest-based admission policies, which keep working when etcd is unavailable, reach beta.
What to do
- Find out which kube-proxy mode you run, and start testing the alternative now rather than in the v1.40 cycle.
- Audit operators and controllers for 429 handling before upgrading control planes.
- If you use SELinux, look for volumes shared between Pods with different labels.
- If you run managed Kubernetes, check when your provider plans to offer 1.37. The deprecation clock started on 26 August whether or not your provider has caught up.
Kubernetes has become good at long, well-signposted retirements, and v1.37 is another example. The risk is not surprise; it is the habit of leaving warnings in logs until a release turns them into failures.
Sources