Kubernetes 1.37 HPA Scale to Zero: External Metrics Setup and Gotchas
Configure Kubernetes 1.37 HorizontalPodAutoscaler minReplicas 0 with external metrics, understand ScaledToZero, and avoid workloads that never wake up.
Kubernetes 1.37 promotes HorizontalPodAutoscaler scale-to-zero support to beta and enables it by default. A Deployment can now reach zero replicas during idle periods and wake when an object or external metric reports new work.
The best candidates are queue consumers, batch workers, and expensive CPU or GPU workers whose demand exists independently of the Pods. Request-driven web services need more care because a Kubernetes Service does not buffer traffic while no endpoint is ready.
Why CPU and Memory Cannot Wake Zero Pods
CPU and memory metrics come from running Pods. At zero replicas there is nothing to measure, so those signals cannot tell the HPA to create the first Pod.
Use a signal that exists while the workload is stopped, such as:
- messages waiting in a durable queue;
- Kafka consumer lag;
- pending jobs in a database;
- an object metric maintained by another Kubernetes resource;
- an external metric exposed through a metrics adapter.
Kubernetes 1.37 requires at least one object or external metric when minReplicas is zero. An HPA containing only resource metrics is rejected.
Expose an External Queue Metric
Assume Prometheus collects this metric:
queue_consumer_lag{namespace="default",name="worker_tasks"}Configure Prometheus Adapter to expose it through the External Metrics API. The exact Helm values depend on your installation, but the mapping follows this pattern:
rules:
external:
- seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
metricsQuery: 'sum(<<.Series>>{<<.LabelMatchers>>}) by (name)'
resources:
overrides:
namespace:
resource: namespaceVerify the signal before creating the HPA:
kubectl get --raw \
'/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks' \
| jq .If this request fails, fix the adapter, discovery rule, RBAC, or label selection first. Scale-to-zero is only as reliable as the metric that wakes the workload.
Configure the HPA
This HPA targets one replica for every 30 queued tasks, with a maximum of ten workers:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: queue-worker
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: queue-worker
minReplicas: 0
maxReplicas: 10
metrics:
- type: External
external:
metric:
name: queue_consumer_lag
selector:
matchLabels:
name: worker_tasks
target:
type: Value
value: "30"
behavior:
scaleDown:
stabilizationWindowSeconds: 300Start the Deployment with at least one replica. Manually scaling a Deployment to zero still means “paused”; the HPA will not claim ownership of a zero state it did not create.
Understand the ScaledToZero Condition
When the controller scales from one or more replicas to zero, it records ScaledToZero=True in HPA status. This tells later reconciliation loops that the HPA owns the zero state and should continue evaluating the wake-up metric.
Inspect it with:
kubectl describe hpa queue-workerAfter the workload wakes, the condition changes to ScaledToZero=False with reason NotScaledToZero. If the workload is at zero without the true condition, check whether an operator or deployment tool manually paused it.
Troubleshoot a Workload That Does Not Wake
Start with HPA conditions and events:
kubectl get hpa queue-worker -o yaml
kubectl describe hpa queue-worker
kubectl get deployment queue-worker -o jsonpath='{.spec.replicas}{"\n"}'ScalingActive=False with FailedGetExternalMetric points to the metrics pipeline. Confirm the series still exists at zero replicas and that its labels match the HPA selector.
Also verify that both kube-apiserver and kube-controller-manager support and enable HPAScaleToZero during a version-skewed upgrade. Before downgrading or disabling the feature, change affected HPAs to minReplicas: 1 and wake workloads currently at zero.
Production Guardrails
Measure cold-start time from metric creation to a ready worker. Include adapter polling, HPA synchronization, scheduling, image pulls, initialization, and readiness probes.
Use a durable queue with visibility timeouts and idempotent consumers. Alert on queue age, not only queue depth, so a small number of stuck tasks cannot hide behind a low count. Keep minimum replicas above zero when latency objectives cannot absorb a cold start.
Finally, test failure modes: unavailable metrics, a broken adapter, exhausted cluster capacity, slow image pulls, and a worker that fails readiness. Saving the last idle Pod is useful only when the system reliably returns from zero.
Before enabling this in production, work through the broader Kubernetes 1.37 upgrade guide and verify control-plane version skew.
Sources
Today I Fixed
Short real fixes from production — posted daily
Stay ahead of the curve
Get the latest DevOps, Kubernetes, AWS, and AI/ML guides delivered straight to your inbox. No spam — just practical engineering content.
Related Articles
AI-Driven Capacity Planning for Kubernetes Clusters (2026)
How to use AI and machine learning for Kubernetes capacity planning. Covers predictive autoscaling, cost optimization, tools like StormForge and Kubecost, and building custom ML models for resource forecasting.
AI-Powered Capacity Planning Across Multi-Cluster Kubernetes: Where This Is Heading in 2026
Capacity planning across dozens of Kubernetes clusters used to mean spreadsheets and quarterly guesswork. AI agents that correlate usage trends across clusters, predict when a cluster will run out of headroom, and recommend rebalancing are moving from research to real platform teams in 2026.
Build an AI Capacity Forecasting Tool with Prophet + Kubernetes Metrics
Reactive autoscaling fixes problems after they happen. Build a forecasting tool using Facebook's Prophet library on historical Prometheus metrics to predict capacity needs days ahead — before traffic spikes hit.