At Intuit we run thousands of production services across hundreds of Kubernetes clusters. At that scale, developers were spending 30-40% of their time on YAML, deploys, and ops (the outer loop) instead of writing product code.
We couldn’t just tell people to get better at YAML. The same platform serves Intuit’s approximately 100 million customers worldwide, across products like TurboTax and QuickBooks. This was a platform problem, not a training problem.
I work on the platform engineering team at Intuit. This post is reposted with permission, with light edits, from the Intuit engineering blog, where it was originally published with my colleague Avni Sharma. Original content © Intuit Inc. The views expressed are my own and do not necessarily reflect the views of Intuit Inc.
The first 6x: managed namespaces#
Before Kubernetes, teams owned the machines: on-prem servers, or AWS accounts and EC2s they patched and scaled themselves. The first Kubernetes system at Intuit took that away and put them on multi-tenant clusters the platform team ran. We owned the nodes: scaling, AMIs, security patching. Security patching turned out to be the part teams valued most. They no longer had to do AMI rotations themselves; they could rely on the platform to do it, and that turned out to be the change they were happiest about.
The contract sat at the namespace. A team got a vended, managed namespace they could deploy into. They did not get a cluster, and they could not change the platform-managed parts (security policies and the rest), which role-based access control (RBAC) and Open Policy Agent (OPA) admission webhooks enforced. Compared to owning VMs, we measured roughly a 6x improvement in developer productivity, an internal estimate against our own pre-Kubernetes baseline rather than an audited benchmark. It worked, and we ran the majority of Intuit’s production compute on it.
Kustomize in every repo#
Inside that namespace, the app was still theirs to describe. Each service had a Kustomize project: a shared app-base with app-specific patches plus an overlay per environment. The platform team published centrally managed remote bases (horizontal pod autoscalers, canary rollouts, metrics), which represented Intuit policies and best practices (the paved path), and services pinned them to a version tag. A typical layout looked like this:
├── app-base
├── environments
│ ├── e2e-use2
│ ├── e2e-usw2
│ ├── prd-use2
│ ├── prd-usw2
│ ├── prf-usw2
│ ├── qal-usw2
│ ├── stg-usw2And an app-base/kustomization.yaml looked something like this: remote bases for the paved path pieces, then whatever else the team needed.
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
bases:
- <remote>/service-hpa-base?ref=v7.0.0
- <remote>/service-rollout-canary-base?ref=v7.0.0
- <remote>/metrics-service-base?ref=v7.0.0
resources:
- ConfigMap-envoy.yaml
- PrometheusRule.yaml
patchesStrategicMerge:
- Rollout-patch.yaml
- Hpa-patch.yaml
- Service-metrics-patch.yamlThat’s the syntax as it stood then. bases and patchesStrategicMerge have since been folded into resources and patches: even the tool describing the config drifts.
From a developer’s point of view, that was total flexibility. Inside their namespace they could patch anything, add anything, ship anything. That was also as far as the abstraction went. The cluster was ours. The application YAML was still theirs. We could not change Kubernetes out from under a pile of manifests in someone else’s repo.
The outer-loop trap#
That leftover YAML living in the dev teams’ repos began accumulating tech debt and drift from the paved path. Kubernetes 1.16 was the breaking point. Leaving app manifests in someone else’s Git repo had always been a trade-off against the platform team’s ability to move the cluster. At 1.16 that trade-off stopped paying, and it became the forcing function to abstract inside the namespace as well.
API deprecations#
Kubernetes 1.16 stopped serving the extensions/v1beta1, apps/v1beta1, and apps/v1beta2 versions of Deployment, DaemonSet, ReplicaSet, and StatefulSet. The paved path templates still used those apiVersions. Every service’s Kustomize project had to move to apps/v1 before we could upgrade a cluster, and it wasn’t a pure apiVersion string swap: apps/v1 makes spec.selector required and immutable, so manifests that had relied on it being defaulted from the pod template labels needed a real edit, and any Deployment whose selector had drifted from its labels needed a recreate.
The upgrade to Kubernetes 1.16 was painful for the platform team. We generated PRs against onboarded repos, sent a notice asking every team to merge, test, and roll the change through every environment, then ran an Intuit-wide program to track who had and who hadn’t. Clusters sat on the old Kubernetes version until that list went to zero. Bumping the remote base was the intended path (ship a new release tag, teams adopt it, patches get updated), but without the mandate, that mostly did not happen. That work wasn’t the app teams’ job: they weren’t Kubernetes experts, and they shouldn’t have to be. They wanted to write Java. They wanted the application abstracted the same way the cluster already was.
And 1.16 was not a one-off. Kubernetes ships three new versions a year (four, until the cadence changed with 1.22), and API deprecations come with them. 1.22 dropped the beta Ingress API; 1.25 dropped policy/v1beta1 PodDisruptionBudget, batch/v1beta1 CronJob, and PodSecurityPolicy, which for a multi-tenant platform meant re-implementing the whole security-policy layer on Pod Security Admission and OPA rather than just editing an apiVersion; 1.26 dropped autoscaling/v2beta2 HorizontalPodAutoscaler. Every one of those lived in YAML the app teams owned. Leave that YAML in their repos and the 1.16 “migration” playbook becomes a full-time job: program after program, just to clear deprecated APIs in time to upgrade the cluster before it fell below the minimum version supported by Amazon Elastic Kubernetes Service (EKS). The changes we could make behind the scenes, with no app-team involvement, were the easy ones: AMI rotations already worked that way. The YAML was the part we still didn’t own.
47 commits a year#
It wasn’t just the API deprecations. We audited Kustomize repos and found that the typical project carried 47 infrastructure commits a year. That’s an average across the repos we audited, not a fleet-wide census, but it held up everywhere we looked. These weren’t product-code changes. They were outer-loop YAML: HPA min/max and scale metrics, CPU and memory requests, Ingress annotations, rollout patches, remote-base version pins, and the overlays that follow a service through every environment. That doesn’t sound like much until you multiply it across thousands of services. And each of those commits still had to be reviewed, tested, and rolled through every environment.
A lot of the 47 were knobs teams turned after something broke. Many teams left the vended HPA and resource defaults untouched until an incident, then tuned them as part of the root-cause analysis. HPA is not simple, and it is easy to get wrong. Incidents where the app did not scale, or did not scale fast enough, were usually written up as “HPA didn’t work.” Almost always, HPA was working exactly as configured. The config was the problem.
App teams were already carrying that load on top of feature work, so a lot of worthwhile platform changes never made the list: newer load-balancer annotations, remote-base version bumps, best-practice improvements we’d already published. 47 is what shipped. Nobody counted what didn’t.
Where the time actually went#
Put 47 infrastructure commits a year on every service and you can start to see where some of the outer-loop time went. When we measured the development lifecycle, about 42% of engineering time was “inner loop”: coding, testing, debugging. Another 30-40% was that outer loop: deploying, operating, monitoring, and managing infrastructure. The remainder is the meetings, reviews, and on-call that neither loop captures. Industry research lands in a similar range: Bain & Company’s Beyond Code Generation, from their 2024 Technology Report, finds developers spend roughly half their time writing and testing code, meaning the other half goes somewhere else.
Roughly one in five support requests we fielded were Kubernetes questions, not application questions, and configuration complexity drove a lot of the paved-path policy drift across our fleet. We put a developer-time cost on those 47 commits and on the incidents that came from bad HPA, other Kubernetes misconfig, or shipping without progressive delivery. I won’t publish the number, but it was large enough to fund the next round of work.
The next 6x: take away the YAML#
The first 6x had taken away the cluster. The next 6x had to take away the YAML. We wanted application-level changes we could make centrally and roll out on a schedule the platform team controlled.
That’s why the 2022 writeup was called unlocking the next 6X in development velocity through application abstraction. Keep the managed namespace. Take an application spec instead of raw Kubernetes inside it. That roadmap is now a production runtime, IKS AIR, and this series is the story of what we actually built.
IKS AIR (Intuit Kubernetes Service, AI Runtime) runs on top of Kubernetes. A developer says what the application needs, and the platform writes the manifests, the traffic wiring, the autoscaling, and the safety nets.
That breaks into three pieces, and this series takes them one at a time. This post is about the application spec that replaced the YAML. Part 2 is how the platform sizes a service from its own traffic instead of asking a developer to guess. Part 3 is what it does when you ship a new version or need to debug a running pod, and where operations go next.
From 40 YAMLs to a single intent#
A new service used to mean 30-40 YAML files, or more. Deployments, Services, Ingresses, HPAs, and a pile of other custom resources. And Kustomize, for all its power, layers base patches and environment overrides in ways that are easy to get lost in. Often a file would be updated correctly in preprod but not in prod.
IKS AIR takes a single application spec instead. We used an OAM-like model of components and traits, on purpose, instead of exposing raw Kubernetes resources. We took the model, not the implementation: OAM’s own runtime machinery was more than we needed for a fleet where the trait catalog belongs to the platform. Components are what the app is: a webservice or a cronjob. Traits are how it behaves. A developer who declares a sizing trait does not need to know what it becomes underneath: an HPA, a pod-sizing baseline, and in some environments a VerticalPodAutoscaler (VPA). They describe the outcome. An in-cluster controller reconciles the spec into the concrete resources and owns them as its children. That’s the load-bearing part: because the generated objects are reconciled from the spec rather than rendered once at CI time, changing what a sizing trait expands to means shipping a new controller, with no PRs against anyone’s repo.
A spec looks something like this:
apiVersion: iks.intuit.com/v1beta1
kind: ExpressApplication
metadata:
name: my-api
spec:
components:
- type: webservice
name: my-api
image: <registry>/my-api:v1
traits:
- type: sizing
properties:
horizontal:
size: finetune
vertical:
size: finetune
environments:
preprod:
- name: qal
prod:
- name: prdThe platform turns that into the same manifests we’d been hardening for years in production.
With this, the 1.16 class of problem mostly goes away. Kubernetes version upgrades, observability agent swaps, other infra refreshes roll out centrally. Application teams don’t reconfigure or redeploy because the platform underneath them changed. That used to be weeks of migration work every upgrade cycle. Now the namespace is on the same footing as the managed cluster and AMI rotations: we can move the platform underneath a service without touching it.
Each service that moves to IKS AIR saves an estimated 25 engineering days a year in ongoing outer-loop work. Migration has taken about four days in the services moved so far, so it pays back inside the first two months and compounds for the life of the service. In practice, 47 infrastructure commits a year drops to roughly a dozen. The 25 days isn’t 35 YAML edits; it’s each of those changes carrying review, a test cycle, and a roll through every environment, plus the incident time the misconfigured ones generated.
The IKS AIR application spec is what made that possible. Part 2 covers IKS AIR Managed Autoscaling, and how size: finetune on the sizing trait drives the platform to set replica counts and pod sizing from historical utilization instead of asking a developer to guess.
If you’re curious what else we’re building, check out the Intuit engineering blog or the software engineering careers page.
Image Credits#
Cable network image by Taylor Vick

