Get Current on Kubernetes - Before Your Cloud Provider Does It for You

A fixed-scope readiness audit plus upgrade execution for EKS, AKS and GKE. We clear the blockers, hop the versions, and hand back clusters that are off extended support and on a supported path.

Duration: 2-4 weeks Team: 1-2 Senior K8s Engineers + AI Agents

You might be experiencing...

Clusters sitting on versions that are billed at extended support rates, or about to be force-upgraded on the provider's schedule
Nobody is sure what breaks on Kubernetes 1.36: containerd 1.x nodes, cgroup v1 node images, kube-proxy in IPVS mode
ingress-nginx is archived and unpatched, but it fronts every production service and nobody wants to touch it
Upgrades keep slipping because each one turns into an open-ended project with no clear end date

Engagement Phases

Days 1-3

Readiness Audit

Inventory every cluster, version, upgrade policy, node image, container runtime, cgroup mode, kube-proxy mode, add-on version and ingress controller. Scan live clusters and Helm charts for removed and deprecated APIs.

Days 3-5

Upgrade Plan & Sign-off

Per cluster: in-place hops or blue/green to a fresh cluster, node OS migration path, add-on version pins for each hop, change windows and rollback plan. You approve the plan before we change anything.

Week 2

Blocker Remediation

Fix deprecated API usage, move node groups to containerd 2 and cgroup v2 images, rewrite custom containerd config, and plan kube-proxy away from IPVS. Ingress-nginx to Gateway API migration runs here when it is in scope.

Weeks 2-4

Upgrade Execution & Hand-back

One minor version at a time: control plane, add-ons, node groups, smoke tests. Then lock in the upgrade policy, add version checks to CI or your ops review, and hand over runbooks for the next upgrade.

Deliverables

Cluster version and blocker inventory across accounts and subscriptions
Deprecated and removed API report for live clusters and Helm charts
Per-cluster upgrade plan with in-place vs blue/green decision and rollback steps
Node image migration (for example AL2 to AL2023 or Bottlerocket, Ubuntu 22.04 to 24.04)
Upgraded clusters on a version in standard support
Optional: ingress-nginx to Gateway API migration with parallel data planes and DNS cutover
Upgrade runbook and a version calendar so the next upgrade is routine

Before & After

MetricBeforeAfter
ScopeOpen-ended upgrade projectFixed scope, agreed up front
Extended Support PremiumPaid on forgotten clustersClusters back in standard support
Upgrade TimingChosen by the cloud providerYour change window
Next UpgradeTribal knowledgeRunbook + version calendar

Tools We Use

Pluto kubent EKS Upgrade Insights ingress2gateway Terraform Karpenter

Frequently Asked Questions

What does the Kubernetes Upgrade Sprint cover?

A readiness audit and the upgrade itself. We inventory versions, node images, runtimes, add-ons and ingress controllers, fix deprecated APIs and node-level blockers, then upgrade EKS, AKS or GKE clusters one minor version at a time to a release in standard support. Ingress-nginx to Gateway API migration can be added to the same sprint.

Why is it fixed scope instead of time and materials?

Because the audit comes first. After the first few days we know the number of clusters, version hops, node image changes and ingress work involved, and we agree the scope and plan with you before execution starts. That keeps the sprint from turning into the open-ended upgrade project you were trying to avoid.

Can you get us off EKS extended support before the forced upgrade?

That is one of the most common reasons teams book the sprint. EKS bills versions in extended support at $0.60 per cluster-hour versus $0.10 in standard support, and auto-upgrades the control plane once extended support ends. We plan the hops so you land on a version that is still in standard support, then set the upgrade policy so it does not happen again.

Do you handle the Kubernetes 1.36 node blockers?

Yes. Kubernetes 1.36 drops containerd 1.x, the kubelet has refused cgroup v1 nodes by default since 1.35, and kube-proxy IPVS mode is deprecated. On managed clusters that usually means a node image change, so we migrate node groups to containerd 2 and cgroup v2 images and test your workloads on them before the version bump.

Do we upgrade in place or build a new cluster?

It depends on the number of hops and the state of the cluster. One or two minor versions on healthy node groups usually go in place. Three or more hops, an old node OS, or a lot of hand-built configuration often make a blue/green move to a fresh cluster faster and safer. The audit makes that call per cluster.

Will there be downtime?

The goal is none for workloads. Control plane upgrades on EKS, AKS and GKE keep the API available, and we drain nodes with PodDisruptionBudgets checked first. Ingress cutovers run as parallel data planes with a DNS switch, so rollback is a DNS change. Anything that needs a maintenance window is called out in the plan before we start.

Get Started for Free

We would be happy to speak with you and arrange a free consultation with our Kubernetes Expert in Dubai, UAE. 30-minute call, actionable results in days.

Every engagement is scoped by our principal architect, Adrian Vale: 20+ years in production engineering, 40+ professional certifications. Meet Adrian

Talk to an Expert