A bad version used to sit in production until someone noticed. A production container often has no shell, making debugging difficult. IKS AIR puts a canary with automatic rollback on every deploy, and a debug shell you attach without changing the app image.
Load tests can be automated. Completing the analysis and reflecting the findings into HPA and CPU/memory requests is the part that usually happens once at launch and then rarely again. IKS AIR fixes that with AIR Managed Autoscaling: it sizes each service continuously from real traffic instead of a guess.
At Intuit we run thousands of production services across hundreds of Kubernetes clusters. At that scale, developers were spending 30-40% of their time on YAML, deploys, and ops (the outer loop) instead of writing product code.
I’ve watched this happen more times than I’d like to admit: a new pod comes online, Kubernetes declares it “Ready,” and within seconds it’s drowning in production traffic it isn’t warmed up to handle yet. Latency spikes, you get a run of transient 5xx/504s, readiness starts flapping. This is a classic case of a “gate rush.” The fix is slow start at the load balancer or the mesh: trickle traffic in until the pod can take its full share.
Autoscaling is the biggest sustainability lever a platform team actually controls — that was my argument on stage. AI compute is exploding, someone has to pay the energy bill, and “better safe than sorry” over-provisioning is directly at odds with both.