↓ Skip to main content

AI-Powered Autoscaling: How the Platform Sizes a Service
Simplifying Kubernetes, Part 2

Simplifying Kubernetes - This article is part of a series.
Part 2: This Article

Load tests can be automated. Completing the analysis and reflecting the findings into HPA and CPU/memory requests is the part that usually happens once at launch and then rarely again. IKS AIR fixes that with AIR Managed Autoscaling: it sizes each service continuously from real traffic instead of a guess.

The ideal is a load test after every release, or at least every major release, because the performance profile moves over time as features expand and new dependencies get pulled in. Skip that loop and fleet CPU utilization ends up far lower than teams expect.

That’s a data problem more than a Kubernetes one, solvable with machine learning plus ordinary analysis of historical utilization.

In Part 1 we replaced those 30-40 YAML files with one spec. Set the sizing trait to finetune and you get AIR Managed Autoscaling: the platform sets both the replica count and the pod size from historical utilization and current load. Developers don’t set pod CPU, memory, or replica counts themselves.

AIR Managed Autoscaling has two main pieces. Recommenders compute the baseline from real traffic. IPA, the Intelligent Pod Autoscaler, is the in-cluster controller that owns both axes of scaling (horizontal replicas and vertical pod size) and keeps a human in the loop for baseline changes.

I work on the platform engineering team at Intuit. This post is reposted with permission, with light edits, from the Intuit engineering blog, where it was originally published with my colleague Avni Sharma. Original content © Intuit Inc. The views expressed are my own and do not necessarily reflect the views of Intuit Inc.

How it actually works
#

Two loops
#

The fast loop is ordinary HPA. Load jumps, one of the metrics the recommenders chose crosses its target, more replicas come up. That happens in seconds, in every environment, without anyone filing a ticket. But replicas don’t help until they’re Ready, so IPA factors startup time into how aggressively it scales out. Scale-down waits on a stabilization window so it doesn’t flap. If the service is pressing against maxReplicas, IPA can proactively raise that ceiling so HPA isn’t stuck, and alerts the platform team when it does.

The slow loop is the baseline. Recommenders run on a three-day cycle against about a week of production and performance-test metrics (TPS, CPU, memory and heap, GC, latency, errors) and propose one conservative floor that covers both axes: CPU and memory for the app container, plus the HPA baseline (persistent min and max replicas, and which metrics to scale on). Those are the changes you don’t want applied by surprise in production, so they go through developer approval. They show up in the same CI pipeline the team already uses and can have automated tests validate them.

Vertical size stays the same across environments, even when the builds differ, so a perf run is actually representative of prod. Horizontal min/max are per environment, because a test environment should not look like production. The scale signal is shared, though: every environment uses the same HPA metrics, just with different bounds.

HPA starts with two CPU utilization signals: pod-level CPU, which sums across every container in the pod including the sidecars, and CPU of the app container alone on only the pods that are Ready and serving traffic. The recommender can swap in a signal that better matches the traffic (busy-thread percentage for Java/Tomcat services, for instance). Developers who want a specific metric can set a constraint for it, but most services never need to; the recommender’s selection holds up across the fleet. Changing the metrics is a baseline change, not a runtime one.

HPA scales replicas in seconds, without anyone filing a ticket. The baseline (CPU, memory, min/max) is recomputed every three days from about a week of production and performance-test traffic, and needs approval in production.

Feeding the window
#

The week-long window only works if there’s data in it. Production traffic counts, as does traffic in the performance test environment. If a service has been quiet in prod for a week and nobody ran a load test either, the next recommendation is working from stale or empty history. That’s why we ask teams to keep feeding it realistic load tests in perf, especially before a new service goes to production, and whenever you need the baseline to reflect a pattern prod hasn’t seen yet.

If you’re expecting a real traffic event, run that profile in perf so it lands in the data window that IPA and the recommenders are analyzing. The recommender doesn’t score each run on its own; it looks at the week of production and perf traffic together and takes a conservative floor from whichever was harder.

Coordinating both axes
#

Both axes run together on purpose. An undersized pod will try to compensate by adding replicas. A service with too many replicas will hide that each pod is oversized. Drive them separately and you pay for the same capacity twice.

Upstream guidance is not to run HPA and VPA against the same resource metric, and for good reason. HPA on CPU divides usage by the CPU request. A VPA that raises the request drives measured utilization down even though real load hasn’t changed, and HPA reads that as slack and scales in. The two controllers end up chasing each other.

The recommenders don’t tune pod resource requests and HPA configuration in isolation. From that week of production and perf metrics they propose one baseline for the app container: CPU and memory, plus the HPA settings that go with it: which signals to scale on, and the persistent min and max replicas. The new baseline for both is applied at the same time, so CPU/memory requests and HPA move together, and the CPU request HPA divides by only moves when that baseline does.

Who can change what
#

IPA owns the HPA, the VPA, and the app-container CPU and memory. It enforces that at admission rather than by quietly reverting any changes: a hand-edit of those fields is rejected, and the object is unchanged.

When scaling of an app is managed by AIR Managed Autoscaling, leave CPU and memory empty on the Rollout and IPA fills them from the baseline. Put different values in git and the sync is rejected. If you do need to temporarily hand-manage sizing or scaling, like in the middle of an incident, IPA supports a pause annotation that releases the resource for manual changes.

AIR Managed Autoscaling takes about a week of production and performance-test metrics and drives both axes together: HPA on replica count, and an approved baseline on CPU and memory per pod.

Floors, and why faster is hotter
#

If a service needs a guaranteed floor (a hard SLA, a past incident) you can set constraints on the sizing trait: minimum CPU, memory, or replicas. Constraints win over the recommender. Most services never need them. The platform already keeps a conservative production floor on min replicas, and a baseline won’t reduce CPU or memory by more than about ten percent in one pass.

A slow-starting app with sharp traffic spikes is a different problem. New replicas can’t take traffic until they’re Ready, so the pods already running have to absorb the spike. The usual lever is to lower the HPA target, 40% instead of 60%, so scale-out starts sooner. That works, and it wastes headroom: you run the pods cold on purpose. The better fix is to make the app start faster. The sooner a new pod is Ready, the hotter you can run, 80% or 90%, because scale-out actually arrives in time. Startup time is what really sets your headroom.

Dealing with seasonal traffic
#

HPA scales on load it can already see, and the recommenders set baselines from historical metrics. Neither covers load you know is coming that hasn’t arrived yet. A tax deadline, Black Friday, a campaign that starts Tuesday at 9 a.m.: those are on the calendar, not in this week’s metrics. You want to pre-scale so the pods are Ready when the spike hits.

The Time-Based Autoscaler is for that. You give it a time window and a replica floor. When the window opens, HPA holds at least that many pods, even if they appear under-utilized. When it closes, HPA goes back to following load. A one-time event, like Tax Day or Black Friday, gets a start and end date. Repeating traffic, like weekly or monthly payroll, is defined with cron semantics. It handles both.

It still runs through HPA. Overlapping windows take the higher floor, and it can’t go above HPA’s max or below the min. A load test in the week before is still how you get the baseline right.

We didn’t get here on the first try
#

In our 2022 post we laid out a three-level maturity model: static configs for predictable load, dynamic right-sizing from real-time CPU, then fully autonomous sizing with no developer in the loop.

Level 1 and Level 2 shipped. Level 3 is where the plan changed shape, and where we found out how much headroom we’d been leaving on the table. Fleet-wide, average CPU utilization sat well below what the hardware could support. Not because teams were careless, but because the only feedback they had was pre-production load testing that rarely matched real traffic, so the rational move was to over-provision defensively. Multiply that across thousands of services and you get a platform running at a fraction of its real capacity.

Closing that gap wasn’t “wire up a controller to watch memory in real time.” We built AIR Managed Autoscaling around IPA, a Kubernetes custom resource with its own controller, so we could change how recommendations are computed without asking every service owner to touch their manifests again. Full autonomy turned out to be the wrong target for production. Immediate scale is automatic. Baseline changes still go through the team. “No developer intervention at all” became “no developer toil, but still a human in the loop for the decisions that matter.”

We rolled AIR Managed Autoscaling out to all existing production AIR services without downtime. Nobody had to reconfigure anything. Some apps were over-provisioned and wasting money. Others were under-provisioned and at operational risk. With AIR Managed Autoscaling now fleet-wide, we are reducing cost and improving operational excellence.

Next time: the last piece of the puzzle, how we build progressive delivery and debugging tools into the platform so teams get them without asking, and where that’s heading next: Agentic Operations.

If you’re curious what else we’re building, check out the Intuit engineering blog or the software engineering careers page.


Image Credits
#

Nighttime Kubernetes container yard image generated with Google Gemini from prompts written by the author. No real people, places, or Intuit systems are depicted.

Todd Ekenstam
Author
Todd Ekenstam
Notes on K8s, Python, and Homelab stability.
Simplifying Kubernetes - This article is part of a series.
Part 2: This Article