↓ Skip to main content

KubeCon EU 2024 Keynote: Innovating Responsibly
Sustainability and Autoscaling in Kubernetes

Autoscaling is the biggest sustainability lever a platform team actually controls — that was my argument on stage. AI compute is exploding, someone has to pay the energy bill, and “better safe than sorry” over-provisioning is directly at odds with both.

I work on the platform engineering team at Intuit. This is a writeup of a keynote panel I spoke on at KubeCon EU 2024 - jump to my segment at 9:56, or watch the full session below.

Key takeaways:

  • Sustainability starts with using cloud resources efficiently - autoscaling is one of the biggest levers a platform team has.
  • “Better safe than sorry” over-provisioning is a common trap, and it’s directly at odds with sustainability goals.
  • Correctly configuring autoscaling is a genuinely hard, data-heavy problem - one AI is well-suited to help solve.
  • At Intuit, we’re building an intelligent autoscaling recommendation system so teams get exactly the resources they need.
  • Cloud sustainability isn’t any one group’s job - it takes cloud providers, platform teams, and app teams working together.

The panel
#

“Keynote: Innovating Responsibly: How to Navigate Sustainability in the Era of Kubernetes” ran on Thursday, March 21, 2024 at KubeCon + CloudNativeCon Europe 2024 in Paris. See the session listing on the KubeCon schedule. From the session abstract:

With almost a decade of successful K8s adoption, the focus of cloud platforms now shifts towards maximizing resource utilization to drive sustained innovation for the future. In the face of pressing climate change concerns, sustainability emerges as a critical theme demanding attention. Despite its importance, the concept of building sustainable cloud workload remains ambiguous, leaving platform operators and users grappling with actionable steps. In this session, Aparna is joined by Todd Ekenstam and David Meder-Marouelli, industry experts and practitioners and together they will demystify platform efficiency and offer actionable insights on how to contribute to sustainability.

You’ll not want to miss this opportunity to leave with a practical to-do list for advancing innovation responsibly.

Four of us shared the stage:

The through-line for the panel was “responsible innovation” - balancing the rapid growth of cloud-native and AI workloads against their environmental cost, from energy-efficient ARM hardware to what one of my co-panelists called “code sobriety.” My part focused on the piece I deal with every day: what it actually means to run a large Kubernetes platform efficiently.

The cloud consumer’s responsibility
#

Most of us on stage aren’t building the hardware - we’re consuming cloud capacity built by someone else. As a cloud consumer, I think the most important lever for sustainability is using your resources as efficiently as possible - and a big part of that is effective autoscaling.

At Intuit, that’s not an abstract concern. We run 100% of our services on modern SaaS infrastructure, processing 65 billion machine learning predictions a day for 24,000+ financial institutions, moving $560 billion, and handling 3.6 billion requests during peak season at 99.999% availability. We’re building an AI-native development platform on top of Kubernetes and cloud-native software, and using AI and data analytics inside the platform itself.

Load like that isn’t constant. We depend on autoscaling to handle the heaviest bursts, and just as much on scaling back down the moment that capacity isn’t needed anymore.

The over-provisioning trap
#

Sizing and scaling workloads correctly is hard, and that difficulty breeds a “better safe than sorry” instinct. How do you really know you’ve sized a workload correctly for every condition it will ever see?

Faced with that doubt, teams often just throw resources at the problem - arbitrarily pre-scaling to a large number of pods, or a large number of nodes, “just in case.” It buys peace of mind, but it also means higher cost and higher resource consumption than the workload actually needs. That’s not sustainable, especially when better alternatives exist.

Autoscaling is hard to get right
#

Kubernetes and the ecosystem around it already give you the capability to scale workloads automatically and dynamically. The gap isn’t capability - it’s that configuring those systems correctly is genuinely hard for application developers to get right on their own.

Capacity planning is fundamentally a data problem, and that’s exactly the kind of problem AI is good at. We believe AI can have a real impact here, helping teams be far more efficient with their computing resources without requiring every developer to become an autoscaling expert.

Building intelligent autoscaling at Intuit
#

That’s what we’re building at Intuit: an intelligent autoscaling recommendation system that takes the guesswork out of sizing, reduces the burden on our developers, and helps ensure every workload has the resources it actually needs - no more, no less.

It’s a large engineering investment, and it isn’t easy. But I’d rather put that effort into innovation and optimization than simply buy more hydrocarbons.

Questions for the road ahead
#

We closed out with a few questions I don’t think have easy answers yet, but are worth every platform team asking themselves: Can automated or bot traffic and user traffic be given different qualities of service? How many autoscaling anti-patterns in your deployments can you find and address? And how do you develop trust in your autoscaling configuration in the first place?

Just like the Paris Climate Accord brought countries together around a shared problem, I think cloud sustainability becomes actionable when cloud providers, platform teams, and app teams work closely together on it - rather than treating it as any one group’s job alone.

Photos from the panel
#

Watch the video
#

For my co-panelists’ full perspectives - Adrienne on what cloud providers are doing, David on compute efficiency, and Aparna tying it together - watch the full keynote below. The CNCF also has a recap of the full Day 3 keynote block.

If you’re working on autoscaling, capacity planning, or platform efficiency and want to compare notes, reach out on LinkedIn.

Todd Ekenstam
Author
Todd Ekenstam
Notes on K8s, Python, and Homelab stability.