A bad version used to sit in production until someone noticed. A production container often has no shell, making debugging difficult. IKS AIR puts a canary with automatic rollback on every deploy, and a debug shell you attach without changing the app image.
Load tests can be automated. Completing the analysis and reflecting the findings into HPA and CPU/memory requests is the part that usually happens once at launch and then rarely again. IKS AIR fixes that with AIR Managed Autoscaling: it sizes each service continuously from real traffic instead of a guess.
At Intuit we run thousands of production services across hundreds of Kubernetes clusters. At that scale, developers were spending 30-40% of their time on YAML, deploys, and ops (the outer loop) instead of writing product code.
I’ve watched this happen more times than I’d like to admit: a new pod comes online, Kubernetes declares it “Ready,” and within seconds it’s drowning in production traffic it isn’t warmed up to handle yet. Latency spikes, you get a run of transient 5xx/504s, readiness starts flapping. This is a classic case of a “gate rush.” The fix is slow start at the load balancer or the mesh: trickle traffic in until the pod can take its full share.
Ever walk into a meeting and feel an eerie sense of déjà vu? Same slide deck. Same “quick recap.” Same debate you were positive you already settled last week. I’ve started calling this the rerun: the meeting loops because the decision never got written down.
Autoscaling is the biggest sustainability lever a platform team actually controls — that was my argument on stage. AI compute is exploding, someone has to pay the energy bill, and “better safe than sorry” over-provisioning is directly at odds with both.