Running Kubernetes at scale creates a cost problem that is easy to overlook.
A service may receive heavy traffic for a few hours and almost no traffic for the rest of the day. A development environment may sit unused overnight. A batch processor may only need compute resources when a job arrives. Yet the Kubernetes cluster can continue keeping pods running because the deployment has a minimum replica count greater than zero.
For workloads with intermittent demand, those idle replicas can represent a meaningful amount of wasted compute.
Google Kubernetes Engine provides scale-to-zero capabilities that allow suitable workloads to reduce their running capacity when there is no work to process. When demand returns, the workload can scale back up.
The idea sounds simple, but scale-to-zero changes how you need to think about Kubernetes autoscaling. Scaling from 10 pods to 5 is mostly a capacity problem. Scaling from 1 pod to 0 introduces a startup problem because there is no running application instance when the next request or piece of work arrives.
That distinction matters.
What Scale-to-Zero Actually Means
Traditional Kubernetes autoscaling usually keeps at least one or more replicas running.
For example:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: worker-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: worker
minReplicas: 1
maxReplicas: 20Even when there is no work, one pod remains active.
Scale-to-zero changes the lower boundary:
Normal autoscaling
20 pods
|
|
5 pods
|
2 pods
|
1 pod <- minimumWith scale-to-zero:
20 pods
|
|
5 pods
|
1 pod
|
0 pods <- no active workloadThe difference between one and zero is important because a zero-replica workload has no application process available to handle work immediately.
That means scale-to-zero works best for workloads where the platform has a reliable way to determine when work has arrived and can tolerate the startup delay.
Why Idle Pods Cost More Than They Appear
Suppose a service normally runs two replicas:
2 replicas × 24 hours × 30 daysIf the service is only needed during a limited period, those replicas still consume resources outside that period.
This becomes more noticeable with development and non-production environments.
A team might have:
Development environments
Test environments
Preview deployments
Internal tools
Batch workers
Event processorsMany of these workloads do not need to be running continuously.
Keeping them at one or two replicas guarantees availability, but it also guarantees a baseline resource cost.
Scale-to-zero changes that trade-off.
When there is no useful work, the workload can stop consuming pod resources entirely.
Scale-to-Zero Is Not the Same as Stopping a Cluster
This distinction is important.
Scale-to-zero normally applies to the workload, not necessarily the entire GKE cluster.
For example:
GKE Cluster
│
├── API Service 3 pods
├── Background Worker 0 pods
├── Report Generator 0 pods
└── Internal Tool 1 podThe cluster can remain available while individual workloads reduce their replica count.
Cluster-level autoscaling can then work independently to adjust the underlying node capacity when nodes are no longer required.
Conceptually, the two layers look like this:
Application workload
|
v
Pod autoscaling
|
v
Required nodes
|
v
Cluster/node autoscalingThis separation is useful because Kubernetes has different scaling problems at different layers.
HPA or another workload-level mechanism decides how many application instances are required.
Node autoscaling decides how much underlying compute is needed to host those workloads.
The Biggest Problem: Waking Up From Zero
The difficult part of scale-to-zero is not shutting down pods.
It is knowing when to start them again.
Consider a worker deployment:
Queue
|
v
Worker Deployment
|
+-- Pod
+-- Pod
+-- PodWhen the queue is empty:
Queue: 0 messages
Workers: 0 podsNow a new message arrives.
Something needs to detect that message and cause the deployment to scale from:
0 -> 1Only then can the new worker start processing the message.
This creates a startup path:
New work
|
v
Metric / event changes
|
v
Autoscaling decision
|
v
Pod scheduled
|
v
Container starts
|
v
Application becomes Ready
|
v
Work is processedThe complete path can take significantly longer than simply sending work to an already-running pod.
Cold Start Becomes Part of the Architecture
With a minimum replica count of one, the application is already running.
A request arrives:
Request
|
v
Running Pod
|
v
ResponseWith scale-to-zero:
Request / Work
|
v
Scaling signal
|
v
Pod scheduling
|
v
Container startup
|
v
Application initialization
|
v
Ready
|
v
WorkThis is effectively a cold start.
The more work your application performs during startup, the longer this delay becomes.
For example, an application that needs to:
Load a large model
Establish multiple connections
Warm caches
Run database migrations
Initialize a complex runtime
Download configuration or dependencies
may not be a good candidate for aggressive scale-to-zero.
A small worker that starts in a few seconds is a much better fit.
Workloads That Benefit From Scale-to-Zero
Not every Kubernetes workload should use it.
Scale-to-zero is particularly interesting for workloads with naturally intermittent demand.
Batch Processing
A batch processor may run only when a job exists.
No jobs
|
v
0 workers
Jobs arrive
|
v
Workers start
|
v
Jobs processed
|
v
Workers scale downThere is little value in maintaining several idle workers when there is no batch workload.
Queue Consumers
Message-driven applications are another strong candidate.
The workload's demand is represented by pending work rather than HTTP traffic.
Useful signals might include:
Queue depth
Messages waiting
Jobs pending
Consumer lagWhen the queue is empty, the worker population can potentially approach zero.
Development Environments
Development and preview environments often remain unused for long periods.
For example:
Developer active
|
v
Pods running
Developer leaves
|
v
No useful workload
|
v
Pods scale downThis can reduce unnecessary resource consumption without requiring engineers to manually shut down environments.
Event Processing
Some event-driven workloads have highly uneven traffic.
A service might receive thousands of events during a short period and almost none afterward.
Maintaining a large idle deployment for the peak load is inefficient.
Scale-to-zero allows the workload to follow actual demand more closely.
Workloads That May Not Be Good Candidates
Some services need immediate availability.
For example, a latency-sensitive public API may not tolerate a cold start every time traffic returns after an idle period.
Imagine:
Traffic:
09:00 high
12:00 low
18:00 highIf the service reaches zero during the quiet period, the first requests at 18:00 may encounter startup latency.
For a customer-facing API, that may be unacceptable.
A better configuration might be:
Minimum replicas: 1
Maximum replicas: 20rather than:
Minimum replicas: 0
Maximum replicas: 20The correct decision depends on the application's latency requirements.
Scale-to-Zero and Event-Driven Workloads
Event-driven systems require special consideration because the event source must remain available even when the consumer does not.
For example:
Message Broker
|
v
Scaling Signal
|
v
Worker Deployment
|
v
ProcessingThe broker cannot disappear just because the worker has scaled to zero.
The system needs a durable source of truth for pending work.
This is one reason queues and event streams work well with scale-to-zero architectures.
The queue retains the work while the compute layer is temporarily absent.
Do Not Use CPU as the Only Signal for Zero
This is a subtle but important problem.
CPU-based HPA works well when pods are already running.
But if the deployment reaches:
0 podsthere is no application process consuming CPU.
Therefore, CPU cannot by itself tell the system that a new workload has arrived.
Consider:
Pods: 0
CPU: 0%
Queue: 5,000 messagesCPU says nothing is happening.
The queue says the opposite.
For scale-to-zero workloads, the scaling signal generally needs to represent external demand rather than activity inside an already-running pod.
That is why queue depth, request activity, event counts, or another externally observable signal can be more useful.
Startup Time Needs to Be Measured
Do not estimate cold-start behavior from development machines.
Measure the real deployment.
A useful test is:
1. Scale workload to zero
2. Generate new demand
3. Record when the scaling signal changes
4. Record when the pod is scheduled
5. Record when the container starts
6. Record when readiness succeeds
7. Record when the first workload completesThis gives you the actual startup path.
For example:
Signal detected 0.5 sec
Pod scheduled 2.0 sec
Container started 5.0 sec
Application ready 8.5 sec
First job completed 10.2 secThe exact numbers will vary by application and environment.
The important point is that scale-to-zero introduces measurable latency that should be treated as part of the system design.
Scale-to-Zero and Node Autoscaling
There is another optimization opportunity when workloads scale down.
Suppose a node is hosting the last few pods of a workload:
Node 1
├── Worker Pod
└── Worker Pod
Node 2
└── Worker PodAfter the workloads scale down:
Node 1
└── no workload
Node 2
└── no workloadNode-level autoscaling can potentially remove unnecessary infrastructure when capacity is no longer needed.
This is where workload scaling and cluster scaling can complement each other.
The complete optimization looks like:
Work decreases
|
v
Pods scale down
|
v
Nodes become underutilized
|
v
Node autoscaling removes capacity
|
v
Infrastructure cost decreasesHowever, the timing and configuration of the two layers matter.
If workloads scale up frequently, aggressively removing nodes can create unnecessary churn.
Scale-to-Zero Can Increase Operational Complexity
Reducing resource usage sounds simple, but the architecture becomes more dependent on correct scaling behavior.
You need to understand:
What triggers scale-up?
How long does startup take?
What happens to queued work?
What happens during a scaling failure?
How does the application signal readiness?
What happens if demand arrives during a deployment?These are not theoretical questions.
They directly affect reliability.
A worker that can safely restart and resume processing is generally easier to scale to zero than a stateful process that expects to remain alive continuously.
Common Mistakes
Setting Minimum Replicas to Zero Without Testing
Changing the minimum from:
minReplicas: 1to:
minReplicas: 0is not enough.
You need to test the entire wake-up path.
Ignoring Cold Starts
If users notice a 10-second delay after a quiet period, the cost savings may not justify the user experience.
Scaling on the Wrong Metric
A metric must tell the autoscaler that work actually exists.
CPU is often insufficient once the workload reaches zero.
Forgetting Pending Work
For queue consumers, work should remain safely stored while there are no active workers.
A scale-to-zero design should never depend on an in-memory queue inside the application process.
Making Scale-Down Too Aggressive
If a workload repeatedly goes:
0 -> 1 -> 0 -> 1 -> 0because demand arrives intermittently, the application can spend more time starting and stopping than doing useful work.
A small stabilization period or minimum replica count may be more appropriate.
Advantages and Disadvantages
Advantages
Lower idle resource consumption: Workloads that spend significant time inactive do not need to keep application pods running continuously.
Better fit for intermittent workloads: Batch jobs, event processors, queue consumers, and development environments can follow actual demand more closely.
Potential infrastructure savings: When workload capacity decreases enough, node-level autoscaling can also reduce unused cluster capacity.
Less manual environment management: Development and temporary environments can potentially reduce their footprint automatically instead of relying on developers to shut them down.
Disadvantages
Cold-start latency: Returning from zero requires scheduling and starting a pod before the application can process new work.
More complicated scaling design: The system needs a reliable signal that can trigger scale-up even when no application pods are running.
Potential startup failures: Image-pull problems, configuration errors, readiness failures, or dependency issues can prevent a workload from recovering quickly.
Poor fit for latency-sensitive services: Public APIs and other workloads that require immediate response may be better served by keeping at least one replica running.
More difficult troubleshooting: Engineers need to inspect both the workload and the mechanism responsible for waking it back up.
A Practical Decision Checklist
Before enabling scale-to-zero, evaluate the workload against a few questions.
Does the workload have long idle periods?
|
+-- No --> Keep a baseline replica count
|
+-- Yes
|
v
Can demand be detected while pods are at zero?
|
+-- No --> Redesign the scaling signal
|
+-- Yes
|
v
Is cold-start latency acceptable?
|
+-- No --> Keep at least one replica
|
+-- Yes
|
v
Can pending work survive without the pod?
|
+-- No --> Add durable state/queue
|
+-- Yes
|
v
Scale-to-zero may fitThis is a better way to evaluate the feature than simply asking whether the application can technically run with zero replicas.
When Scale-to-Zero Makes the Most Sense
The strongest candidates usually share three characteristics:
Demand is intermittent.
Work can wait briefly while capacity starts.
The workload has a reliable external demand signal.
A queue-based worker is a good example.
A latency-sensitive authentication API is usually a much weaker candidate.
The distinction is not about Kubernetes itself. It is about the business and technical requirements of the workload.
Summary
GKE scale-to-zero can reduce wasted compute for Kubernetes workloads that spend substantial periods doing nothing.
The biggest benefit comes from workloads such as batch processors, event consumers, background workers, development environments, and other services where demand is intermittent.
But scaling from one replica to zero is fundamentally different from ordinary autoscaling.
Once there are no pods, the application cannot observe CPU utilization or process incoming requests itself. The system needs another mechanism to detect demand and bring the workload back.
That makes three things particularly important:
Reliable demand signal
+
Acceptable cold-start time
+
Durable workload stateWhen those conditions are satisfied, scale-to-zero can become a useful part of a GKE cost and capacity strategy. When they are not, keeping a small baseline of running replicas is often the safer design.
The goal is not to run everything at zero.
The goal is to stop paying for capacity that the workload does not actually need, without turning the application's next request into a recovery operation.

Join the conversation! Your thoughts help the community grow.