🎉 DevOps Interview Prep Bundle is live — 1000+ Q&A across 20 topicsGet it →
All Articles

Kubernetes 1.37 HPA Scale to Zero: External Metrics Setup and Gotchas

Configure Kubernetes 1.37 HorizontalPodAutoscaler minReplicas 0 with external metrics, understand ScaledToZero, and avoid workloads that never wake up.

DevOpsBoys4 min read
Share:Tweet

Kubernetes 1.37 promotes HorizontalPodAutoscaler scale-to-zero support to beta and enables it by default. A Deployment can now reach zero replicas during idle periods and wake when an object or external metric reports new work.

The best candidates are queue consumers, batch workers, and expensive CPU or GPU workers whose demand exists independently of the Pods. Request-driven web services need more care because a Kubernetes Service does not buffer traffic while no endpoint is ready.

Why CPU and Memory Cannot Wake Zero Pods

CPU and memory metrics come from running Pods. At zero replicas there is nothing to measure, so those signals cannot tell the HPA to create the first Pod.

Use a signal that exists while the workload is stopped, such as:

  • messages waiting in a durable queue;
  • Kafka consumer lag;
  • pending jobs in a database;
  • an object metric maintained by another Kubernetes resource;
  • an external metric exposed through a metrics adapter.

Kubernetes 1.37 requires at least one object or external metric when minReplicas is zero. An HPA containing only resource metrics is rejected.

Expose an External Queue Metric

Assume Prometheus collects this metric:

promql
queue_consumer_lag{namespace="default",name="worker_tasks"}

Configure Prometheus Adapter to expose it through the External Metrics API. The exact Helm values depend on your installation, but the mapping follows this pattern:

yaml
rules:
  external:
    - seriesQuery: '{__name__="queue_consumer_lag",name!=""}'
      metricsQuery: 'sum(<<.Series>>{<<.LabelMatchers>>}) by (name)'
      resources:
        overrides:
          namespace:
            resource: namespace

Verify the signal before creating the HPA:

bash
kubectl get --raw \
  '/apis/external.metrics.k8s.io/v1beta1/namespaces/default/queue_consumer_lag?labelSelector=name%3Dworker_tasks' \
  | jq .

If this request fails, fix the adapter, discovery rule, RBAC, or label selection first. Scale-to-zero is only as reliable as the metric that wakes the workload.

Configure the HPA

This HPA targets one replica for every 30 queued tasks, with a maximum of ten workers:

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: queue-worker
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: queue-worker
  minReplicas: 0
  maxReplicas: 10
  metrics:
    - type: External
      external:
        metric:
          name: queue_consumer_lag
          selector:
            matchLabels:
              name: worker_tasks
        target:
          type: Value
          value: "30"
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300

Start the Deployment with at least one replica. Manually scaling a Deployment to zero still means “paused”; the HPA will not claim ownership of a zero state it did not create.

Understand the ScaledToZero Condition

When the controller scales from one or more replicas to zero, it records ScaledToZero=True in HPA status. This tells later reconciliation loops that the HPA owns the zero state and should continue evaluating the wake-up metric.

Inspect it with:

bash
kubectl describe hpa queue-worker

After the workload wakes, the condition changes to ScaledToZero=False with reason NotScaledToZero. If the workload is at zero without the true condition, check whether an operator or deployment tool manually paused it.

Troubleshoot a Workload That Does Not Wake

Start with HPA conditions and events:

bash
kubectl get hpa queue-worker -o yaml
kubectl describe hpa queue-worker
kubectl get deployment queue-worker -o jsonpath='{.spec.replicas}{"\n"}'

ScalingActive=False with FailedGetExternalMetric points to the metrics pipeline. Confirm the series still exists at zero replicas and that its labels match the HPA selector.

Also verify that both kube-apiserver and kube-controller-manager support and enable HPAScaleToZero during a version-skewed upgrade. Before downgrading or disabling the feature, change affected HPAs to minReplicas: 1 and wake workloads currently at zero.

Production Guardrails

Measure cold-start time from metric creation to a ready worker. Include adapter polling, HPA synchronization, scheduling, image pulls, initialization, and readiness probes.

Use a durable queue with visibility timeouts and idempotent consumers. Alert on queue age, not only queue depth, so a small number of stuck tasks cannot hide behind a low count. Keep minimum replicas above zero when latency objectives cannot absorb a cold start.

Finally, test failure modes: unavailable metrics, a broken adapter, exhausted cluster capacity, slow image pulls, and a worker that fails readiness. Saving the last idle Pod is useful only when the system reliably returns from zero.

Before enabling this in production, work through the broader Kubernetes 1.37 upgrade guide and verify control-plane version skew.

Sources

🔧

Today I Fixed

Short real fixes from production — posted daily

Browse fixes
Newsletter

Stay ahead of the curve

Get the latest DevOps, Kubernetes, AWS, and AI/ML guides delivered straight to your inbox. No spam — just practical engineering content.

Related Articles

Comments