Horizontal Pod Autoscaling has always been useful for Kubernetes workloads, but the standard CPU and memory metrics are often not enough for real production systems.
A service may be CPU-light while its request queue is growing. A worker may have plenty of memory but be falling behind on messages. An API can have normal resource usage while request latency or error rates are increasing.
Those are application signals, not infrastructure resource signals.
Google Kubernetes Engine now supports using PromQL queries with Horizontal Pod Autoscaling (HPA), giving Kubernetes workloads a more direct way to scale from Prometheus-style application and operational metrics.
The practical change is important: instead of building a separate custom metric pipeline for every Prometheus metric you want to use for scaling, teams can use PromQL to express the signal that should drive the autoscaler.
A simplified architecture looks like this:
Application
|
| Metrics
v
Prometheus / Managed Prometheus
|
| PromQL
v
GKE HPA
|
| Scaling decision
v
Deployment
|
+---- Pod
+---- Pod
+---- Pod
+---- ...This makes HPA much more useful for workloads where CPU and memory are poor indicators of actual demand.
What Is Horizontal Pod Autoscaling?
Kubernetes HPA automatically changes the number of replicas for a workload based on observed metrics.
A simple HPA might say:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-server
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-server
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70The basic idea is straightforward.
If the average CPU utilization goes above the target, Kubernetes can increase the number of pods.
When demand decreases, it can scale the workload back down.
The problem is that CPU utilization does not always represent application demand.
Consider an API that performs most of its work by waiting for external services.
It might look like this:
CPU: 25%
Memory: 40%
Requests: 4,000/sec
Latency: 850 msCPU-based HPA sees a healthy workload.
Users see a slow application.
This is where application-level metrics become useful.
Why PromQL Matters for Autoscaling
PromQL is the query language used by Prometheus.
It allows developers and platform teams to calculate useful signals from collected metrics.
For example:
rate(http_requests_total[2m])can represent the request rate over a time window.
Another query might calculate an error ratio:
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))A queue-oriented workload might have a metric such as:
sum(queue_messages_ready)The advantage is that the autoscaling signal can be based on what the application actually needs.
Instead of:
Scale when CPU > 70%you can design a policy closer to:
Scale when requests per pod exceed the service's target capacity.That is a much better fit for many workloads.
A Request-Rate Example
Suppose a service can comfortably process around 100 requests per second per pod.
You expose a request counter:
http_requests_totalPromQL can calculate the current request rate:
sum(rate(http_requests_total[2m]))If the service is receiving:
500 requests/secthen five pods may be appropriate.
The important part is that the autoscaler should generally use a per-pod target, rather than simply saying that 500 requests per second always requires five pods.
As workload behavior changes, the desired replica count should continue to reflect the relationship between observed demand and pod capacity.
Application Metrics Give You Better Scaling Signals
PromQL-based HPA becomes especially useful when the metric represents a bottleneck that actually determines capacity.
Examples include:
HTTP request rate
Queue depth
Messages processed per second
Active jobs
Request latency
Database connection pressure
Custom business workload metricsNot every metric is a good autoscaling signal.
For example, scaling directly on a business metric such as:
orders_created_totalmay not make sense unless the metric can be translated into workload pressure.
The question should always be:
Does this metric tell me that the current number of pods is insufficient?
If the answer is no, it probably should not drive HPA.
Using a Queue Metric
Background workers are one of the clearest examples.
Imagine a Kubernetes deployment consuming messages from a queue.
CPU might remain at 35% even when the queue is growing:
Queue depth: 10,000
CPU: 35%
Memory: 45%
Workers: 3CPU-based HPA may do nothing.
But queue depth provides a much more direct signal.
A PromQL query could aggregate the number of pending messages:
sum(queue_messages_ready)The autoscaling policy can then be designed around the amount of work waiting to be processed.
The exact target depends on processing time and workload characteristics.
For example, a team might decide that each worker should normally handle approximately 100 pending messages.
The target then represents workload capacity rather than a generic CPU percentage.
PromQL Lets You Transform Raw Metrics
One of the most useful parts of PromQL is that the metric exposed by the application does not have to be exactly the metric you want to scale on.
You can calculate rates, ratios, aggregations, and other derived signals.
Suppose every pod exposes:
requests_totalA raw counter is not usually the best autoscaling signal.
You can convert it into a rate:
sum(rate(requests_total[2m]))You can also aggregate across instances.
For a latency-based signal, you might calculate a percentile from a histogram:
histogram_quantile(
0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)The important engineering decision is not the PromQL syntax itself.
It is deciding whether the resulting metric is stable enough to drive scaling.
Avoid Scaling Directly on Noisy Metrics
Autoscaling reacts to changing measurements.
That creates a common problem.
Suppose request traffic looks like this:
10:00 100 req/s
10:01 180 req/s
10:02 90 req/s
10:03 220 req/s
10:04 110 req/sIf the scaling signal reacts immediately to every change, the deployment can repeatedly scale up and down.
This is called flapping.
A good PromQL query often uses a time window:
rate(http_requests_total[2m])instead of attempting to react to an instantaneous value.
Kubernetes HPA also has scaling behavior controls that can help prevent overly aggressive changes.
For example:
behavior:
scaleUp:
stabilizationWindowSeconds: 0
scaleDown:
stabilizationWindowSeconds: 300The exact settings should depend on how quickly the application starts and stops handling additional traffic.
Scale-Up and Scale-Down Are Different Problems
A service often needs aggressive scale-up and conservative scale-down.
Imagine a sudden traffic spike.
Waiting several minutes to add pods could cause latency to increase significantly.
Now consider a temporary traffic drop.
Immediately removing pods may be counterproductive if traffic is about to return.
That is why autoscaling policies should distinguish between:
Scale upand:
Scale downFor scale-up, fast reaction is usually more important.
For scale-down, stability and avoiding unnecessary churn are often more important.
PromQL Does Not Replace Good Application Metrics
Adding PromQL support does not eliminate the need for instrumentation.
If the application only exposes:
cpu
memorythen PromQL cannot magically discover the application's actual workload.
The application still needs meaningful metrics.
For an HTTP API, useful metrics might include:
http_requests_total
http_request_duration_seconds
http_requests_in_flightFor a worker:
jobs_processed_total
jobs_failed_total
jobs_in_progressFor a message consumer:
messages_received_total
messages_processed_total
messages_pendingThe quality of the autoscaler depends heavily on the quality of these signals.
HPA and Managed Service for Prometheus
Google Cloud's managed Prometheus capabilities can collect Prometheus-compatible metrics from GKE workloads.
That creates a useful path:
Application
|
v
Prometheus Metrics
|
v
Managed Prometheus
|
v
PromQL
|
v
HPAThe benefit for platform teams is that application metrics and autoscaling logic can use a common observability model.
You do not necessarily need to create a separate custom metrics system simply because a metric is coming from Prometheus.
This also makes existing Prometheus instrumentation more valuable.
A metric that was already useful for dashboards and alerting may now also become useful for scaling, provided it is suitable for that purpose.
A Practical Kubernetes Design
A production service might use three different categories of metrics.
Infrastructure Metrics
|
+---- CPU
+---- Memory
Application Metrics
|
+---- Request rate
+---- Error rate
+---- Latency
Workload Metrics
|
+---- Queue depth
+---- Jobs waiting
+---- Active tasksThe HPA should select the metric that best represents capacity pressure.
For an API, request rate may be better.
For a worker, queue depth may be better.
For a CPU-intensive service, CPU utilization may still be the right signal.
There is no requirement that every Kubernetes workload use the same autoscaling metric.
Common Mistakes
Using CPU Because It Is Easy
CPU is easy to configure, but easy does not mean correct.
If CPU is not correlated with workload pressure, HPA can make poor scaling decisions.
Choosing a Metric That Reacts Too Slowly
If the metric has a large aggregation window, the autoscaler may respond after users have already experienced degraded performance.
The measurement window should match the workload's behavior.
Scaling on Error Rate Without Thinking About the Cause
A rising error rate does not always mean that more pods will fix the problem.
The actual issue could be:
Database unavailable
External API failure
Bad deployment
Invalid configuration
Network failureAdding pods in those situations can increase load on an already failing dependency.
Ignoring Startup Time
Suppose a new pod takes 60 seconds to become ready.
An HPA policy that assumes additional capacity appears immediately will not behave as expected.
Scale-up thresholds and stabilization need to account for application startup time.
Creating a Circular Scaling Signal
Be careful when the metric itself depends on the number of pods.
For example, a metric that increases simply because more replicas exist can create confusing feedback.
The signal should represent demand or pressure, not merely replica count.
Advantages and Disadvantages
Advantages
Application-aware scaling: PromQL allows HPA to use signals that are much closer to actual workload demand than CPU or memory alone.
Reuses existing Prometheus metrics: Teams that already instrument applications with Prometheus-compatible metrics can potentially use those signals for autoscaling instead of creating an entirely separate measurement model.
Flexible metric calculations: PromQL can turn raw counters and histograms into rates, aggregations, ratios, and other derived signals that are more useful for scaling.
Better fit for heterogeneous workloads: API servers, queue consumers, workers, and data-processing services can each choose metrics that match their actual capacity constraints.
Closer integration between observability and operations: The same metrics used to understand application behavior can also influence scaling decisions, provided they are stable and appropriate for that purpose.
Disadvantages
More complex autoscaling configuration: CPU-based HPA is relatively easy to understand. PromQL-based scaling requires teams to understand both Kubernetes autoscaling and Prometheus query semantics.
Bad queries can produce bad scaling: A technically valid PromQL query can still be a poor autoscaling signal if it is noisy, delayed, or unrelated to actual capacity.
Observability becomes part of the scaling path: If the metric pipeline is unavailable or produces incorrect values, scaling behavior can be affected.
More difficult troubleshooting: When a workload scales unexpectedly, engineers may need to inspect the HPA configuration, PromQL query, metric collection, application instrumentation, and workload behavior together.
How to Choose a Good Autoscaling Metric
Before putting a PromQL expression into an HPA, ask a few practical questions.
Does the metric represent demand?
A good metric should increase when the workload needs more capacity.
Is the metric correlated with capacity?
If doubling the number of pods does not meaningfully improve the metric, it may not be a good scaling signal.
Is it stable?
Highly noisy metrics can cause unnecessary scaling activity.
Does it react quickly enough?
A metric that takes several minutes to reflect demand may be unsuitable for latency-sensitive workloads.
Can the team explain the target?
If nobody can clearly explain why the HPA target is 100 requests per pod instead of 50 or 500, the target probably needs more investigation.
Can scaling actually solve the problem?
This is the most important question.
If the bottleneck is a database with a fixed connection limit, adding more application pods may make the situation worse.
Autoscaling should increase useful capacity, not simply increase the number of processes.
When PromQL-Based HPA Makes Sense
PromQL-based HPA is particularly useful when the workload's capacity is better described by application or business-adjacent metrics than by CPU and memory.
Good candidates include:
High-throughput APIs
Message consumers
Queue workers
Event processors
Background job systems
Data processing workloads
Services with strong Prometheus instrumentationFor a simple CPU-bound service, standard resource-based HPA may still be sufficient.
The goal is not to replace CPU-based autoscaling everywhere.
The goal is to use a better signal when CPU is not the right representation of demand.
Summary
PromQL support for Horizontal Pod Autoscaling in GKE gives Kubernetes teams more control over what actually triggers scaling.
Instead of relying only on:
CPU utilization
Memory utilizationteams can build scaling policies around application-aware signals such as:
Request rate
Queue depth
Active work
Processing rate
LatencyThe real value comes from choosing the right metric, not simply from using PromQL.
A well-designed autoscaling policy connects three things:
Real workload demand
|
v
Meaningful application metric
|
v
Scaling decision
|
v
Additional capacityThat makes HPA more closely aligned with how the application actually behaves.
But PromQL-based autoscaling should be treated as an engineering control loop, not just another Kubernetes configuration option. Query quality, metric stability, startup time, dependency limits, and scale-down behavior all affect the result.
For teams already using Prometheus metrics on GKE, this capability makes those metrics more useful. The same observability data can help engineers understand the workload and, when chosen carefully, tell Kubernetes when more capacity is actually needed.

Join the conversation! Your thoughts help the community grow.