Kubernetes 1.36 Upgrade Blockers: containerd 2.0, cgroup v2 and IPVS Deprecation
Kubernetes 1.36 drops containerd 1.x, kubelet refuses cgroup v1 nodes since 1.35 and IPVS mode is on its way out. How to detect each blocker and fix it.
The three things most likely to block a Kubernetes 1.36 upgrade are all node-level. containerd 1.x is not supported from 1.36.0, so nodes need containerd 2.0 or later. The kubelet refuses to start on cgroup v1 nodes by default since 1.35. And kube-proxy IPVS mode is deprecated, with removal planned for 1.43. Control plane upgrades rarely trip on any of these. Node pools do.
That is why these blockers catch teams out. The control plane moves first, the dashboard goes green, and then the first node group roll fails because the AMI or node image underneath was built for a world of containerd 1.7 and cgroup v1. Here is what changed, how to find it and how to fix it.
What actually changes between 1.35, 1.36 and 1.37?
Here are the upstream dates and decisions in one place, taken from the Kubernetes blog and the EKS release notes:
| Change | Version | What it means |
|---|---|---|
failCgroupV1 defaults to true | 1.35 | Kubelet does not start on cgroup v1 nodes unless you opt out |
| Last release supporting containerd 1.x | 1.35 | 1.35 nodes can still run containerd 1.7 |
| containerd 1.x support dropped | 1.36.0 | Upgrade containerd to 2.0+ before or with the kubelet |
gitRepo volume type permanently disabled | 1.36 | Pods with gitRepo volumes are rejected by the kubelet |
Service externalIPs deprecated | 1.36 | Warnings now, removal planned for 1.43 |
KubeProxyIPVS feature gate added | 1.37 | IPVS still works, logs a deprecation warning |
| cgroup v1 opt-out removed | 1.38 (planned) | No more failCgroupV1: false escape hatch |
| IPVS off by default / removed | 1.40 / 1.43 (planned) | Move to nftables (or iptables) mode |
Kubernetes 1.37 shipped on August 26, 2026, so a lot of teams will plan 1.35 to 1.37 in one push. Treat 1.36 as the hard line. That is where an un-upgraded runtime stops being a warning and becomes a node that will not join.
Why does containerd 2.0 block the 1.36 upgrade?
The Kubernetes 1.34 cgroup driver post put it plainly: “The last Kubernetes release to offer this support will be the last released version of v1.35, and support will be dropped in v1.36.0.” The reason is that 1.36 relies on the kubelet asking the runtime for its cgroup driver through the RuntimeConfig CRI call, which only containerd 2.0+ (and CRI-O 1.28+) implement.
Bumping containerd from 1.7 to 2.x is not just a package update. Google’s GKE migration guide, which moved Linux nodes to containerd 2.0 at GKE 1.33, lists what breaks:
- Docker schema 1 images can no longer be pulled. Old images built years ago and pinned by tag are the usual culprit. Google’s own
startup-script:v1image is one of them. - The CRI v1alpha2 API is gone. Anything that talks to containerd’s CRI socket directly with an old client (some security agents, log shippers, custom tooling) fails.
registry.configs.tlsis removed, andregistry.mirrorsandregistry.configsare deprecated. If your node bootstrap writes a customconfig.tomlfor private registry mirrors or TLS, that file needs rewriting and testing.
This is exactly the kind of config that lives in a launch template nobody has opened in years.
Why is my kubelet failing on cgroup v1 nodes?
The Kubernetes blog’s October 6 post on the cgroup v2 shift spells it out: “Starting with Kubernetes v1.35, failCgroupV1 defaults to true, so the kubelet does not start on a cgroup v1 node by default.” You can set failCgroupV1: false in the kubelet configuration as a temporary override. That fallback still exists in 1.36 and 1.37, and it is scheduled for removal in 1.38.
cgroup v2 needs Linux kernel 5.8 or later and a distro that enables it by default: Ubuntu 21.10+, Debian 11+, RHEL 9+, Container-Optimized OS, Amazon Linux 2023 and Bottlerocket all qualify. Amazon Linux 2 does not.
The node OS is only half the job. Per the Kubernetes cgroup docs, applications that read cgroup files directly need versions that understand v2:
- Java: OpenJDK/HotSpot jdk8u372, 11.0.16, 15 and later
- Node.js: 20.3.0 or later. The 18 line does not reliably detect v2 memory limits and can size its heap from host memory, which ends in OOM kills.
- Go services using
uber-go/automaxprocs: v1.5.1 or later - Standalone cAdvisor: v0.43.0 or later, plus whatever monitoring and security agents read cgroupfs
The memory sizing problem is the nasty one, because it shows up as pods getting OOMKilled after the node migration, not as a failed upgrade.
What do EKS, AKS and GKE users need to check?
The managed providers each hit these changes through their node images, on slightly different schedules.
| Provider | containerd 2 | cgroup v2 | The usual blocker |
|---|---|---|---|
| EKS | Required before moving past 1.35 | AL2023 default; Bottlerocket v2 with failCgroupV1: false; Fargate still cgroup v1 | Amazon Linux 2 nodes. AWS stopped publishing EKS AL2 AMIs on Nov 26, 2025, and AL2 uses cgroup v1 |
| AKS | Ubuntu 24.04 and Azure Linux 3.0 images ship containerd 2.x; the 1.35 Ubuntu 22.04 image still lists containerd 1.7.29 | Ubuntu 22.04+ and Azure Linux are v2 | Pools pinned to Ubuntu2204 or still on Azure Linux 2.0, whose node images start being removed on Oct 31, 2026 |
| GKE | Linux nodes on containerd 2.0 since GKE 1.33 | cgroupv1 deprecated at 1.31, auto-migrated from 1.33, removed at 1.35 | Node pools explicitly pinned to cgroupv1 to dodge the auto-migration |
On EKS, the AL2 problem stacks with the version calendar. If you are also paying for old versions, read our breakdown of the EKS extended support cost first, because a 1.31 cluster heading for 1.36 has an AMI migration, a runtime migration and several API cleanups in front of it.
On AKS, the default Ubuntu OS SKU moves to Ubuntu 24.04 automatically when you upgrade to 1.35 or later. Pools that pinned the versioned Ubuntu2204 SKU do not, so those are the ones to look for.
Is kube-proxy IPVS mode going away?
Yes, slowly. The 1.37 sneak peek says clusters running kube-proxy in ipvs mode now log a deprecation warning, and KEP-5495 lays out the plan: a KubeProxyIPVS feature gate in 1.37, disabled by default by 1.40, removed by 1.43. EKS already flagged IPVS as deprecated in its 1.35 release notes and points customers to nftables mode.
The stated reason is maintenance. SIG Network no longer has maintainers who know the IPVS backend well, and nftables was designed to replace both IPVS and iptables. This is not a 1.36 blocker. It is a “do not build new clusters on it” signal, and a good item to fold into the same node work, since switching kube-proxy mode needs a careful rollout and testing of Services, NodePorts and any NetworkPolicy tooling that assumed IPVS.
How do you detect each blocker before upgrading?
Run these against every cluster, including the sandbox ones. Each check maps to one row of the blocker table below.
- Runtime version per node.
kubectl get nodes -o wideshows theCONTAINER-RUNTIMEcolumn. Anything readingcontainerd://1.xblocks 1.36. - The kubelet’s own warning. Start
kubectl proxyand querylocalhost:8001/api/v1/nodes/NODE/proxy/metrics, then grep forkubelet_cri_losing_support. A series withversion="1.36.0"means that node’s runtime is too old. - cgroup version. On the node (SSM, SSH or a debug pod with host access), run
stat -fc %T /sys/fs/cgroup/.cgroup2fsis v2.tmpfsis v1. TheOS-IMAGEcolumn from step 1 is a quick proxy: “Amazon Linux 2” means v1. - Who is relying on the opt-out. Through the same proxy, read
/api/v1/nodes/NODE/proxy/configzand checkkubeletconfig.failCgroupV1. Any node withfalseis living on borrowed time until 1.38. - kube-proxy mode.
kubectl -n kube-system get cm kube-proxy-config -o yamlon EKS, or thekube-proxyConfigMap on kubeadm clusters, and look formode: ipvs. - Schema 1 images. List every image running in the cluster and check its manifest with a tool like
crane manifest. A manifest with"schemaVersion": 1will not pull on containerd 2. - Removed APIs and volumes. Run a deprecated API scanner such as Pluto against live clusters and Helm charts, and search manifests for
gitRepo:volumes and ServiceexternalIPs.
The blocker table: what breaks, how to detect it, how to fix it
| Blocker | What breaks | How to detect | Fix |
|---|---|---|---|
| containerd 1.x | Nodes cannot run kubelet 1.36 | CONTAINER-RUNTIME column, kubelet_cri_losing_support | New node image with containerd 2.0+, rolled before or with 1.36 |
| Schema 1 images | ImagePullBackOff on containerd 2 nodes | Manifest scan for schemaVersion: 1 | Rebuild or re-push images, bump pinned tags |
| Custom containerd config | Mirrors, TLS or auth to private registries fail | Review bootstrap scripts for registry.configs | Rewrite for containerd 2 config_path host files and test |
| cgroup v1 nodes | Kubelet does not start (1.35+) | stat -fc %T /sys/fs/cgroup/, OS image | AL2 to AL2023 or Bottlerocket, Ubuntu 22.04 to 24.04, unpin GKE cgroupv1 |
| Old runtimes in apps | OOMKills after moving to cgroup v2 | Inventory Java, Node.js, Go automaxprocs versions | Upgrade runtimes or set heap limits explicitly |
failCgroupV1: false | Works today, breaks at 1.38 | configz check | Treat as a dated exception with an owner |
| kube-proxy IPVS | Deprecation warnings now, removal at 1.43 | kube-proxy ConfigMap mode | Plan nftables migration, new clusters off IPVS |
gitRepo volumes | Pods refused by kubelet on 1.36 | Manifest search | Init container or git-sync sidecar |
The bottom line
A Kubernetes 1.36 upgrade is mostly a node upgrade wearing a version number. If your nodes already run containerd 2 on a cgroup v2 OS, 1.36 is routine. If they run AL2, Ubuntu 22.04 with a pinned SKU, or a hand-built AMI with a custom containerd config, the version bump is the easy part and the OS migration is the real project. Do the detection pass now, while 1.35 still gives you the opt-outs, rather than at 1.38 when it does not.
If ingress is also on your list, our ingress-nginx to Gateway API migration checklist covers the other big change landing on most clusters this year.
Our Kubernetes 1.36 Upgrade Readiness Audit is a fixed-scope engagement: we run every check above across your clusters, produce the blocker table filled in for your node groups, add-ons and workloads, and hand back a sequenced upgrade plan with the node image changes spelled out. Book an upgrade readiness scoping call and we will tell you whether you are looking at a version bump or an OS migration.
Frequently Asked Questions
Does Kubernetes 1.36 support containerd 1.7?
No. Kubernetes SIG Node set 1.35 as the last release supporting containerd 1.x, and support is dropped in 1.36.0. Nodes need containerd 2.0 or later before, or at the same time as, the kubelet upgrade to 1.36. The kubelet_cri_losing_support metric with a version label of 1.36.0 tells you which nodes are still too old.
Why does the kubelet fail to start on my node after upgrading to 1.35?
Most likely the node still uses cgroup v1. Starting with Kubernetes 1.35, failCgroupV1 defaults to true, so the kubelet refuses to start on cgroup v1 hosts. You can set failCgroupV1: false in the kubelet config as a temporary bridge, but the fallback is scheduled for removal in 1.38. The real fix is a node OS that boots with cgroup v2.
Is kube-proxy IPVS mode removed in Kubernetes 1.37?
No, it is deprecated, not removed. Clusters running kube-proxy in IPVS mode log a deprecation warning, and 1.37 adds a KubeProxyIPVS feature gate. The upstream plan is to disable IPVS by default in 1.40 and remove it in 1.43. You have time, but new clusters should start on nftables or iptables mode.
Do EKS Fargate pods use cgroup v2?
Not according to AWS. The EKS release notes for Kubernetes 1.35 state that Fargate continues to use cgroup v1, while AL2023 uses cgroup v2 by default and Bottlerocket uses cgroup v2 but keeps failCgroupV1 set to false for compatibility. Fargate nodes are managed by AWS, so there is no node-level change for you to make, but check apps that read cgroup files directly.
How long does a Kubernetes 1.36 upgrade readiness check take?
For one cluster with Infrastructure as Code, the detection work in this post takes a day or two: node runtime and OS inventory, cgroup and kube-proxy mode checks, image manifest scans and an API deprecation scan. The fixes take longer, mostly because node OS migrations (AL2 to AL2023, Ubuntu 22.04 to 24.04) need workload testing, not just a version bump.
Complementary NomadX Services
Related Articles
Get Started for Free
We would be happy to speak with you and arrange a free consultation with our Kubernetes Expert in Dubai, UAE. 30-minute call, actionable results in days.
Every engagement is scoped by our principal architect, Adrian Vale: 20+ years in production engineering, 40+ professional certifications. Meet Adrian
Talk to an Expert