↓ Skip to main content

Kubernetes Slow Start for Cold Pods
Beating the Gate Rush

I’ve watched this happen more times than I’d like to admit: a new pod comes online, Kubernetes declares it “Ready,” and within seconds it’s drowning in production traffic it isn’t warmed up to handle yet. Latency spikes, you get a run of transient 5xx/504s, readiness starts flapping. This is a classic case of a “gate rush.” The fix is slow start at the load balancer or the mesh: trickle traffic in until the pod can take its full share.

It’s especially common with anything that needs a minute to get its legs under it - JVM class loading and JIT warmup, caches that start empty, connection pools that haven’t been established yet, TLS handshakes, model loading. None of that happens instantly, and readiness checks often don’t know or care. You can get the turnstile from the load balancer (ALB target groups, for instance) or from the service mesh (usually Envoy or Istio under the hood).

The core problem: readiness is binary, warm-up isn’t
#

Kubernetes readiness is a yes/no signal. Real workloads live in the gray zone in between - they can handle some traffic right after startup, just not their full share without timing out or fighting each other for CPU.

In practice that looks like cold instances pegging their threads or CPU, p95/p99 latency spiking, occasional 503s and 504s, and readiness flapping as the process gets overwhelmed and recovers and gets overwhelmed again.

Here’s the part that trips people up: just delaying readiness doesn’t fix this. Some apps only warm up under real load - caches don’t populate themselves, and code paths don’t get exercised, until actual requests start flowing through them.

What slow start actually does
#

Load balancer slow start (AWS ALB target groups, for example)
#

ALB slow start lets a newly healthy target take an increasing share of requests over a configured warm-up window. It’s a nice option because it doesn’t touch your app at all - the pod registers and passes health checks like normal, the load balancer just holds it back from full load until the window closes. You end up with fewer cold-start brownouts during rollouts and scale-ups, which matters most for apps that genuinely need real traffic to reach steady state.

Service mesh slow start (Envoy / Istio)
#

The mesh version works the same way in spirit: a new endpoint starts with a reduced load-balancing weight, and that weight ramps up over the slow-start window. Envoy calls this “slow start mode” - it affects upstream load-balancing weights specifically to keep freshly-started endpoints from getting hammered before they’re ready.

Istio wraps this in warmup config on a DestinationRule, which maps to Envoy’s slow-start window:

apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
  name: my-service
spec:
  host: my-service
  trafficPolicy:
    loadBalancer:
      simple: LEAST_REQUEST
      warmupDurationSecs: 60

Either way, the effect is the same: new instances don’t get slammed with full traffic the instant they’re marked ready, and you see fewer 5xxs during initialization.

Two things will bite you here if you don’t know to look for them. First, Istio’s warmup only takes effect under the ROUND_ROBIN or LEAST_REQUEST load-balancing policies - set it under any other policy and it silently does nothing, which is a surprisingly common way teams “turn on” slow start and see zero change. Second, the linear ramp described above is just the default - Envoy also has an aggression parameter that controls the curve. A value above 1 front-loads the caution (slow to start, fast to finish); below 1 does the opposite.

Why this belongs under operational excellence
#

Rollouts, node drains, and autoscaling events happen constantly in any system running at scale. Skip slow start and you’re implicitly betting that every new pod is production-grade the moment it starts - which is rarely true.

It also makes progressive delivery safer. Canary analysis only means something if you measure the canary after it’s warmed up; measure it too early and you’re chasing false positives, or worse, dismissing a real problem because the baseline itself was noisy. That’s why teams that take this seriously usually pair warm-up windows with rollout pacing and analysis delays.

And it makes capacity during traffic bursts more predictable. Without slow start, a burst forces you to scale out right when your newest pods are least able to help - you’re overloading exactly the instances you’re counting on. With slow start, you trade a short ramp-up delay for a lot fewer error spikes, and that’s usually the trade worth making.

One limitation worth being honest about: slow start only helps when pods come up progressively, with some warm pods already in rotation to absorb traffic while the new ones ramp. If your whole deployment rolls at once, or you scale out a big batch in one shot, every pod is cold at the same time and there’s no warm capacity for slow start to lean on. That’s exactly the situation BlaBlaCar ran into when they rolled out Istio’s warmup feature - they had to tune maxSurge/maxUnavailable alongside it so rollouts actually stayed progressive enough for warmup to do anything.

Slow start is also only half of this story. The same mismatch between “Kubernetes thinks this pod is gone” and “the load balancer has stopped sending it traffic” shows up in reverse when a pod terminates - which is what connection draining and PreStop hooks are for on the way out.

How to actually use slow start well
#

Slow start is a safety net, not a fix for broken readiness. If you can make readiness reflect true “ready for full load” - pre-warmed caches, initialized pools, readiness gated on warm-up - do that first. I’d think of it as a stack: get readiness right, pre-warm whatever you can ahead of time, and only lean on slow start for the warm-up that genuinely has to happen under live traffic.

Pick your window from measurement, not intuition. Run a realistic load test, roll pods while it’s running, and choose the smallest window that keeps latency and errors inside your SLO.

Watch out for windows that are too long, too - an overly generous slow start can quietly eat into your effective capacity during a burst. You scale out, but the new pods take so long to ramp up that they don’t actually help in time. This bites hardest when your scaling is reactive rather than predictive. AWS bounds ALB’s slow_start.duration_seconds to 30-900 seconds; whatever you’re configuring, give it a similar ceiling.

Make it observable once it’s live. You want to be able to answer whether new pods are actually seeing fewer requests early on, whether error and latency spikes during rollouts have gone down, and whether time-to-steady-state has gone up enough to notice - and whether that’s a trade you’re willing to keep making. Watching per-pod RPS during rollout, p95/p99 latency for new pods versus old, and 5xx/504 rates alongside readiness flaps will tell you most of what you need to know.

A checklist worth keeping around
#

If any of this sounds familiar - new pods that are “Ready” but still throw 5xx for the first 30-120 seconds, canaries that fail analysis early but pass on retry, scale-outs during a burst that don’t actually stop the bleeding - work through it in this order:

  1. Tighten readiness first (real dependency checks, warm-up gates where you can add them)
  2. Add pre-warming for anything deterministic (classes, caches, pools)
  3. Turn on slow start at the load balancer and/or mesh to meter real traffic during warm-up - and double check your mesh’s load-balancing policy actually supports it before you assume it’s working
  4. Make sure your rollouts stay progressive enough for slow start to matter (tune maxSurge/maxUnavailable) and that rollout tooling doesn’t advance canary analysis until warm-up has finished
  5. Put a ceiling on the slow-start window so it can’t become a capacity risk in its own right

The takeaway
#

Cold pods aren’t the same as warm pods, and slow start is what happens when you actually encode that into your traffic routing instead of hoping nobody notices. Reach for ALB slow start or mesh slow start - either way, you’re trading a boring ramp-up delay for a rollout that doesn’t page anyone at 2am.

(This post is AWS/Envoy-centric because that’s the stack I work in day to day. GCP’s Cloud Load Balancing has connection draining for pods on the way out, but no direct backend-service equivalent of ALB-style slow start on the way in - if you’re on GCP, the mesh-level approach via Istio/Envoy is where you’d look instead.)

Todd Ekenstam
Author
Todd Ekenstam
Notes on K8s, Python, and Homelab stability.