↓ Skip to main content

Operations, Done For You: Progressive Delivery and Debugging
Simplifying Kubernetes, Part 3

Simplifying Kubernetes - This article is part of a series.
Part 3: This Article

A bad version used to sit in production until someone noticed. A production container often has no shell, making debugging difficult. IKS AIR puts a canary with automatic rollback on every deploy, and a debug shell you attach without changing the app image.

Improved operations is the third piece of IKS AIR, after the application spec and AIR Managed Autoscaling. This post is how the platform builds progressive delivery and debugging in, so operations are done-for-you. We’ll end on where we’re going next: Agentic Operations.

I work on the platform engineering team at Intuit. This post is reposted with permission, with light edits, from the Intuit engineering blog, where it was originally published with my colleague Avni Sharma. Original content © Intuit Inc. The views expressed are my own and do not necessarily reflect the views of Intuit Inc.

Progressive delivery, out of the box
#

The IKS AIR paved path builds progressive delivery in by default as a canary, using Argo Rollouts. A new version gets a slice of live traffic first. Analysis decides whether that slice looks healthy. If it does, traffic keeps shifting. If it doesn’t, the rollout automatically goes back to the last stable revision without any user intervention.

The canary isn’t one set of percentages and wait times for the whole fleet. A REST API, a GraphQL gateway, and an AI agent do not all want the same gate, and a QAL environment does not want the same soak as production; some earlier environments or work-in-progress applications may skip the canary entirely so you can iterate. The platform still owns the Argo Rollouts object (step weights, pauses, how many samples, how long between them), and developers pick a profile for how the canary behaves.

Example of a conservative production canary: traffic ramps in steps with a soak and an analysis check at each one. Earlier environments typically run a shorter profile, and some skip the canary so you can iterate. The exact weights and pauses are platform defaults, not a contract every service lives under.

The analysis trait
#

The part you configure is an analysis trait on the component, the same idea as sizing in Part 2. Leave it off and AIR injects two checks: HTTP 5xx error rate, and a mesh-level GraphQL error rate. Both are easy to reason about and are what you want for the majority of cases.

A GraphQL service is why there are two checks by default. HTTP 200s can still carry application errors, so an HTTP error-rate template will happily promote a bad canary. The GraphQL check reads the mesh metrics that actually signal request failures, so it ships on by default next to the HTTP one instead of waiting for a team to discover they needed it.

But you can declare the trait with different list of analysisChecks when that pair is the wrong shape for the app. Beyond the two defaults you can gate on latency, CPU, memory, pod restarts, or either of two success-rate variants. Same rollback either way. Only the signal that trips it changes.

        - type: analysis
          properties:
            analysisChecks:
              # replaces the defaults: HTTP error-rate is no longer gating
              - name: mesh-graphql-error-rate
A new version deploys as a canary. The analysis trait selects which checks gate each step. Healthy checks let traffic continue; a failed gate automatically reverts to the prior stable version.

We’re also prototyping an AIOps path on Numalogic (the Apache 2.0 anomaly-detection library from Numaproj, a project Intuit originated and contributes to). It would score the canary against the service’s own historical golden signals instead of a fixed threshold. The appeal is that it needs no new concept: it arrives as one more check on the same trait, and rollback works the way it already does.

On the old kustomize path, all of this lived in each team’s Rollout-patch.yaml: the templates, and every threshold, sample count, and interval behind them. That was real power, and it was also a steady contributor to the 47 infrastructure commits a year we audited in Part 1: hand-maintained per service, and drifting from the paved path over time. Now AIR generates the same templates centrally. The trait is how you opt into a different set without needed to manage the complexity of the full set of Rollout-related CRs.

For the failure modes the gate measures, mean time to recovery stops being a function of how long it takes someone to respond to an alert and becomes a proactive recovery based on the automatic analysis of the canary, exposing only a fraction of production traffic to the issue.

Rollback by hand
#

The automatic path only fires during the canary, and only on the signals the analysis trait is watching. A version can still make it all the way to production and be wrong in a way none of those checks measure: a bad recommendation, a leak that hasn’t tripped yet, something a customer notices first. For that, the Developer Portal lets you roll an AIR service back to a prior revision without waiting for a metric to trip.

Debugging the black box
#

A common concern with abstractions is that they become a black box when things go wrong. How do you debug a service when you don’t have access to the underlying container or namespace? Taking away the namespace takes away the escape hatch, and an escape hatch is exactly what people want at 2 a.m. So the debugging story had to be part of the abstraction, not a follow-up to it.

Debug shell
#

To get around those hardened, shell-less containers, we provide a debug shell that launches a temporary, ephemeral container within the same pod. This debug container comes loaded with the vetted tools an engineer needs for interactive troubleshooting (a shell, a package manager, tcpdump, curl, language-specific utilities) without modifying the main application container or its image. When the debug session ends, the container goes away. It relies on ephemeral containers, which let you attach a debug container to a running pod without rebuilding or restarting it.

An engineer clicks Debug Shell; the platform attaches a temporary, ephemeral debug container to the pod alongside the hardened app container, without modifying the application container or its image.

Dumps, profiles, and log levels
#

We built one-click debugging workflows directly into our developer portal. For Java services, an engineer can trigger a thread dump or a heap dump on a specific pod with a single click. In production, a heap dump requires additional approval before it runs: a .hprof may contain production data, including customer data, so dumps are access-controlled, audited, retained under policy, and handled under the same data-protection controls as any other production data access. Automation also scrubs the dump as it is captured, and what remains is still treated as production data. Go services get goroutine profiles. Python services get stack traces. It’s the same click in the same portal either way.

Flipping the log level to DEBUG used to mean a restart. Now you can do it on a running Java or Python pod, and revert it from the same screen when you’re done. It also expires on its own, so a DEBUG flag someone set during an incident doesn’t quietly stay on for a month. You don’t rebuild the image and you don’t bounce the pod.

Individually these are small things. Together they’re the difference between filing a ticket with the platform team and just fixing your own problem. That is not where we started.

What we got wrong about debuggability
#

The 2022 plan treated debuggability as a dashboard problem. Give developers good enough telemetry, we figured, and they won’t miss the shell. That was the wrong call, and it took production to show us why.

Two things moved underneath us. Our application containers got hardened. That is unambiguously good, and we were asking for it elsewhere in the org ourselves. It also took away the package manager, then the shell itself. Meanwhile the Kubernetes primitive that makes an attachable debug container practical, ephemeral containers, didn’t reach stable until 1.25, in August 2022; before that it was alpha and behind a feature gate. We could have planned for that gap and didn’t. We assumed telemetry would cover it. So for a while the honest answer to “my pod is misbehaving and I can’t get a shell” was a ticket to our team, which is the exact thing the abstraction was supposed to eliminate.

The lesson we took wasn’t “abstract less.” It was that every capability you remove from a developer has to come back as something at least as good, and telemetry is not a substitute for interactive access when you’re staring at a hung JVM. That’s why the debug tooling above is a platform feature with its own images and audit trail, rather than a runbook telling people which kubectl commands they aren’t allowed to run.

What’s next: making operations agentic
#

And we’re not done. Everything we’ve described so far (autoscaling, progressive delivery, debugging) still relies on a human deciding when to act, even if the platform does the heavy lifting once they do. The next stage we’re exploring is closing that loop: an agent that decides how carefully to ship a change, and that can recognize a problem and resolve it, not just surface it.

A few directions we’re exploring:

  • Risk-scored canaries. We’d score the code change itself, then write the Argo Rollouts steps from that score: traffic weights and how long to soak at each step. A trivial change keeps a fast canary. A risky one gets smaller increments and longer pauses, so it picks up more signal before it owns production traffic. Today’s analysis trait already lets a service pick which checks gate the canary. This would keep the platform owning the Rollout config, but pick the pace from the change instead of a static profile.
  • Drift detection and remediation. We’d have an agent spot when a running service has drifted from its intended or approved spec (a manual hotfix, a stale config, an untracked change, an unsupported pattern) and reconcile it automatically, instead of waiting for an engineer to notice, or for it to surface as an incident down the road.
  • Agentic debuggability. We’d extend today’s one-click dumps and log-level flips into an agent that proactively pulls the right diagnostic artifacts at the right moment, correlates them against recent deploys and metrics, and proposes a root cause. Today an engineer starts that from scratch, after the fact.
  • Agentic rollback. We’d go beyond a fixed error-rate threshold during canary, with an agent that observes the service, reasons about why it looks unhealthy, and decides the right response, rather than just paging someone to take manual action.

None of this is shipped, but it’s the direction we’re excited about: agents taking on more of the judgment calls that today still land on an on-call engineer at 2 a.m.

The first Kubernetes generation at Intuit took the machines off developers’ plates. IKS AIR then took the YAML, then the sizing, and with this post the defaults for shipping a version and debugging a pod. Developers still own the app and the judgment calls that are specific to it. The platform owns the Kubernetes under it, which is how we change and improve that layer under thousands of services without a migration program every time the cluster moves.

I like working on this kind of problem: a platform that lets thousands of engineers move faster with less friction. If you’re curious what else we’re building, check out the Intuit engineering blog or the software engineering careers page.


Image Credits
#

A control room filled with lots of monitors image by Tom Donders

Todd Ekenstam
Author
Todd Ekenstam
Notes on K8s, Python, and Homelab stability.
Simplifying Kubernetes - This article is part of a series.
Part 3: This Article