Kubernetes autoscaling has traditionally focused on one question:
How many resources should a workload have while it is running?
For many applications, that is enough.
But event-driven workloads and intermittent services have a different requirement:
What if the workload should not consume compute resources at all when there is no work?
That is where scale-to-zero becomes useful.
Google Kubernetes Engine now provides native scale-to-zero capabilities for supported workloads, allowing Kubernetes applications to reduce their running capacity to zero when demand disappears and return to service when demand arrives.
The architectural difference is significant:
Traditional Autoscaling
High traffic -> 10 Pods
Medium traffic -> 5 Pods
Low traffic -> 1 Pod
No traffic -> 1 PodWith scale-to-zero:
High traffic -> 10 Pods
Medium traffic -> 5 Pods
Low traffic -> 1 Pod
No traffic -> 0 PodsThat final state can reduce idle compute consumption for workloads that do not need to remain continuously active.
Why Zero Matters
Suppose a service normally receives requests during business hours.
A conventional minimum replica configuration might look like:
spec:
replicas: 1Even when nobody is using the service, one pod remains scheduled.
Over a long period, that idle capacity can become meaningful.
For example:
24 hours
|
+--> 8 hours active
|
+--> 16 hours mostly idleIf the application can tolerate cold-start latency, keeping compute at zero during those 16 hours can improve infrastructure efficiency.
Scale-to-zero is therefore primarily about workload economics and utilization, not simply performance.
Scale-to-Zero Is Different From Horizontal Scaling
Horizontal Pod Autoscaling normally changes replica count based on a metric.
For example:
CPU > threshold
|
v
Increase replicas
CPU < threshold
|
v
Decrease replicasThe minimum replica count is often greater than zero.
Scale-to-zero adds another state:
0
|
1
|
2
|
3
|
...That means the system needs a mechanism to determine:
When should this workload start again if there are currently no pods?
This is the core challenge.
The Cold-Start Problem
Suppose an application has scaled to zero:
Deployment
|
v
0 PodsA request arrives.
There is no pod available to process it.
The platform needs to:
Receive demand
|
v
Create workload
|
v
Schedule pod
|
v
Start container
|
v
Application ready
|
v
Process requestThat takes time.
The user therefore experiences a cold start.
This is the central trade-off of scale-to-zero:
Lower idle cost
vs.
Higher startup latencyNot every workload can tolerate that trade-off.
Good Candidates for Scale-to-Zero
Scale-to-zero is particularly useful for workloads such as:
Internal tools
Development services
Batch workers
Scheduled workloads
Event-driven consumers
Low-traffic APIs
Preview environments
Tenant-specific services
Intermittent AI workloadsFor example:
Developer opens environment
|
v
Start workload
Developer finishes
|
v
Scale to zeroThat can be much more efficient than keeping every environment running continuously.
Poor Candidates for Scale-to-Zero
Some applications should remain available at all times.
Examples include:
High-traffic APIs
Latency-sensitive services
Critical control planes
Always-on messaging infrastructure
Real-time systems
Customer-facing services with strict latency SLOsIf a 5-second startup delay violates the application's SLO, scale-to-zero may be the wrong choice.
Scaling to Zero Requires a Trigger
A workload at zero replicas cannot measure its own CPU utilization.
There is no pod running.
Therefore:
CPU metric
|
X
No pod existsA scale-to-zero architecture needs an external or platform-level signal.
For example:
Incoming request
|
v
Queue depth
|
v
External metric
|
v
Scale workloadThe trigger must exist outside the workload itself.
This is one reason event-driven autoscaling and request-aware infrastructure are closely related to scale-to-zero designs.
Request-Driven Applications
Consider a simple API:
Client
|
v
Load Balancer
|
v
Service
|
v
PodsIf the pods scale to zero, the load balancer still needs a way to cause the workload to become available.
A scale-to-zero solution therefore needs to account for the complete request path:
Request
|
v
Ingress / Gateway
|
v
Activation Signal
|
v
Scheduler / Autoscaler
|
v
PodThe exact implementation depends on the GKE feature and workload type.
The architectural principle remains the same:
The activation mechanism must survive when the application itself is scaled to zero.
Event-Driven Workloads Are a Natural Fit
Queues make scale-to-zero easier to reason about.
Consider:
Producer
|
v
Queue
|
v
WorkerWhen there are no messages:
Queue depth = 0
Worker replicas = 0When messages arrive:
Queue depth = 500
|
v
Scale worker
|
v
5 Pods
|
v
Process messagesThis is an excellent use case because queue depth is an external signal that remains available even when the worker deployment has no running pods.
Scale-to-Zero and Background Workers
A background worker might normally be deployed like:
spec:
replicas: 2But if it only processes occasional jobs, those replicas may spend most of their time idle.
A scale-to-zero architecture can instead use:
No jobs
|
v
0 workers
Jobs arrive
|
v
Scale workers
Queue drains
|
v
0 workersThis can be particularly useful for bursty workloads.
Kubernetes Resource Requests Still Matter
Scale-to-zero does not eliminate the need for accurate resource configuration.
When pods start, Kubernetes still needs to schedule them.
For example:
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "1"
memory: "1Gi"If the requested resources are too high, the workload may take longer to schedule.
If they are too low, the application may become unstable under load.
Scale-to-zero changes replica lifecycle, not the fundamental resource management model.
Cold Start Optimization
If a workload uses scale-to-zero, startup performance becomes an application concern.
A slow startup might come from:
Large container image
Slow dependency initialization
Database migrations
Configuration loading
Secret retrieval
JIT compilation
Model loading
Network callsFor example, an ASP.NET Core service might spend significant time initializing:
Entity Framework
Dependency injection graph
Configuration providers
Caches
External clientsIf cold starts matter, optimize the startup path.
Avoid doing expensive work before the application can accept traffic unless that work is actually required.
Container Image Size Matters
Consider two images:
Image A = 150 MB
Image B = 1.5 GBThe larger image may take longer to pull, particularly when the required node does not already have the image cached.
For scale-to-zero workloads, startup latency is part of the user experience.
Smaller images can therefore have a direct operational benefit.
Use:
Multi-stage builds
Minimal runtime images
Dependency pruning
Layer optimizationwhere appropriate.
Database Connections Need Care
A workload that scales from zero to many replicas can create connection bursts.
For example:
0 Pods
|
v
Traffic spike
|
v
20 Pods
|
v
20 x database connectionsThe database may become the bottleneck before the application does.
For .NET applications, connection pooling helps, but developers should still understand:
Maximum connections
Pool size
Connection lifetime
Database capacity
Startup concurrencyAutoscaling the application without considering database capacity can simply move the bottleneck downstream.
Scale-to-Zero and External Services
Startup code often calls external systems.
For example:
Pod starts
|
+--> Identity service
+--> Configuration
+--> Secret manager
+--> Database
+--> Cache
+--> Third-party APIIf 50 pods start simultaneously, those systems may receive a sudden burst of traffic.
The application should therefore be designed for startup bursts.
Use appropriate:
Connection pooling
Retry policies
Backoff
Caching
Rate limiting
Circuit breakersDo not allow every new pod to perform unnecessary initialization calls simultaneously.
Graceful Scale-Down
Scaling to zero also requires correct shutdown behavior.
A pod should be given time to finish important work.
For a worker:
SIGTERM
|
v
Stop receiving work
|
v
Finish current message
|
v
Commit result
|
v
ExitIf the process exits immediately, the workload may lose work or cause duplicate processing.
Use:
Graceful shutdown
Message acknowledgement
Idempotency
Checkpointing
Termination grace periodswhere applicable.
Autoscaling Does Not Guarantee Cost Savings
Scale-to-zero can reduce compute consumption.
But total application cost includes more than pod runtime.
Consider:
Compute
+
Load Balancer
+
Storage
+
Database
+
Network
+
Logging
+
MonitoringIf compute is only a small portion of the workload's cost, scaling to zero may not significantly reduce the total bill.
Measure actual cost.
Scale-to-Zero and AI Inference
AI inference is an interesting use case.
Suppose a model-serving endpoint receives only occasional requests.
Keeping GPU resources running continuously can be expensive.
Scale-to-zero can theoretically provide:
No requests
|
v
0 compute
Request
|
v
Start inference workload
|
v
Model loads
|
v
InferenceBut AI workloads often have very large startup costs.
Model loading can take seconds or minutes.
GPU initialization can also be expensive.
Therefore, scale-to-zero for AI inference should be evaluated carefully.
The economics may be attractive, but the latency trade-off can be substantial.
AI Model Loading Is Different From API Startup
For a normal API:
Start process
|
v
ReadyFor an AI model server:
Start process
|
v
Initialize runtime
|
v
Initialize GPU
|
v
Load model
|
v
Load tokenizer
|
v
Warm caches
|
v
ReadyThat means a model-serving workload may need a different scaling strategy.
Possible approaches include:
Always-on minimum capacity
Warm instances
Partial scale-to-zero
Smaller startup model
Model caching
Asynchronous inferenceThe correct choice depends on latency requirements.
Scale-to-Zero and SLOs
Before enabling scale-to-zero, define the application's SLO.
For example:
Availability: 99.9%
Normal latency: <200 ms
Cold-start latency: <5 secondsThen measure whether the architecture satisfies the actual requirement.
If cold-start latency is:
20 secondsbut the application's SLO allows only:
2 secondsthe architecture is not appropriate without additional techniques.
Do not optimize for cost at the expense of a clearly defined service objective.
Observability Becomes More Important
A workload that frequently transitions between zero and running states needs visibility into those transitions.
Track:
Scale-up events
Scale-down events
Activation latency
Pod startup time
Image pull time
Request latency
Queue depth
Cold-start failures
Scheduling delaysA useful trace might look like:
Request received
|
+--> Scale-up triggered
|
+--> Pod scheduled
|
+--> Image pulled
|
+--> Container started
|
+--> Application ready
|
+--> Request processedThis allows teams to identify where cold-start latency actually comes from.
Common Mistakes
Using scale-to-zero for latency-sensitive APIs
Cold starts can violate application SLOs.
Ignoring the activation path
A zero-pod workload still needs an external mechanism to detect demand.
Using huge container images
Large images increase startup time.
Loading everything during application startup
Expensive initialization makes cold starts worse.
Ignoring database connection bursts
Rapid scale-up can overload downstream systems.
Forgetting graceful shutdown
Workers can lose or duplicate jobs if they terminate incorrectly.
Assuming zero pods means zero cost
Other cloud services can continue generating costs.
Scaling AI inference without measuring model startup
Model loading can dominate cold-start latency.
Using autoscaling without a representative load test
The right configuration depends on actual workload behavior.
Advantages and Disadvantages
Advantages
Lower idle compute consumption: Workloads can reach zero running replicas when there is no demand.
Better fit for bursty applications: Resources can follow intermittent workloads more closely.
Useful for development environments: Temporary services do not need to run continuously.
Good fit for event-driven workers: Queue depth can provide a natural activation signal.
Improved resource utilization: Infrastructure can be concentrated on workloads that currently need it.
Disadvantages
Cold-start latency: A request may need to wait for the workload to become ready.
More complex activation: Something outside the workload must detect demand.
Startup dependencies become important: Large images and slow initialization directly affect responsiveness.
Downstream services can experience bursts: Rapid scaling can create database or API connection spikes.
Not suitable for every workload: Strict low-latency applications may require warm capacity.
When Native Scale-to-Zero Makes Sense
Consider it when the workload has:
Low average utilization
Bursty traffic
Predictable activation signals
Tolerance for cold starts
Short or moderate startup time
Event-driven behaviorIt is less suitable when:
Traffic is continuous
Latency requirements are strict
Startup is very expensive
The service must always be immediately availableThe decision should be based on the workload's utilization curve rather than simply the desire to reduce replicas.
A Practical Adoption Strategy
Teams can introduce scale-to-zero gradually:
Identify low-utilization workloads.
Measure traffic patterns over time.
Determine whether the workload can tolerate cold starts.
Identify the external activation signal.
Measure current startup time.
Reduce container image size where practical.
Optimize application initialization.
Validate downstream database and service capacity during scale-up.
Implement graceful shutdown.
Add monitoring for activation and startup latency.
Run load tests with realistic burst patterns.
Compare actual infrastructure savings against the latency and operational trade-offs.
This makes the decision measurable rather than theoretical.
What Developers Should Take Away
Scale-to-zero changes the mental model of Kubernetes autoscaling.
Instead of:
How many pods should always be running?the question becomes:
Should this workload be running at all right now?That is valuable for intermittent workloads.
But zero replicas introduce a new state that must be handled correctly.
The application needs:
A reliable activation signal
Fast startup
Graceful shutdown
Resource limits
Downstream capacity
Strong observabilityWithout those pieces, scale-to-zero can simply replace idle compute costs with cold-start failures and operational complexity.
Summary
Native scale-to-zero for GKE workloads can make Kubernetes more efficient for applications that spend significant amounts of time idle.
The fundamental lifecycle becomes:
No demand
|
v
0 Pods
|
v
Demand arrives
|
v
Workload starts
|
v
Process traffic
|
v
Demand disappears
|
v
0 PodsThe main benefit is lower idle resource consumption.
The main cost is cold-start latency and the additional engineering required to make activation reliable.
For event-driven workers, development environments, intermittent services, and some bursty workloads, scale-to-zero can be a strong architectural fit.
For latency-sensitive applications and large AI model servers, keeping some capacity warm may still be the better choice.
The right implementation is therefore not simply "scale to zero whenever possible."
It is "scale to zero when the workload's utilization pattern, startup characteristics, and SLOs make zero capacity a safe and economically useful state."

Join the conversation! Your thoughts help the community grow.