🎉 DevOps Interview Prep Bundle is live — 1000+ Q&A across 20 topicsGet it →
All Articles

Kubernetes 1.37 Native Histograms: Prometheus Migration Guide

Enable Kubernetes 1.37 native histograms safely, configure Prometheus 3, migrate queries and dashboards, and validate observability cost and accuracy.

DevOpsBoys5 min read
Share:Tweet

Kubernetes 1.37 enables native histograms by default, giving platform teams a practical path away from manually chosen classic histogram buckets. That sounds like a simple monitoring upgrade, but enabling ingestion without planning queries, dashboards, storage, and rollback can leave you with two representations of the same metric and a confusing bill.

This guide explains what changed, how to prepare Prometheus, and how to migrate without losing alert coverage.

What changed in Kubernetes 1.37?

Native histograms reached beta in Kubernetes 1.37 and are enabled by default. Kubernetes components can expose both classic histogram series and a native histogram representation. The native format stores the distribution in one histogram sample instead of generating a separate time series for every configured bucket.

Classic histograms remain available during the transition. That compatibility is useful because existing PromQL expressions, recording rules, dashboards, and remote-write consumers may still depend on _bucket, _sum, and _count series.

Native histograms do not automatically make every query faster or every monitoring system cheaper. Their benefit depends on the bucket schema, scrape frequency, metric cardinality, retention, and whether downstream systems preserve the native data type.

Prerequisites

Before changing production scraping, confirm the complete metrics path supports native histograms:

  • Kubernetes 1.37 components are exposing the new samples.
  • Prometheus can ingest native histograms. Prometheus 3.0 or newer is the safest baseline for the per-job settings used below.
  • Any agent, remote-write receiver, long-term store, and query layer supports the native histogram wire format.
  • Your dashboard and alert tooling understands native histogram PromQL results.
  • You have a storage and series-count baseline from before the migration.

If one component silently drops native samples, the scrape may look healthy while the data you expect is unavailable later in the pipeline.

Enable ingestion in Prometheus

Start with one Kubernetes scrape job or a canary Prometheus instance. A transition configuration can ingest native samples while retaining classic buckets:

yaml
scrape_configs:
  - job_name: kubernetes-apiservers
    scrape_native_histograms: true
    always_scrape_classic_histograms: true
    kubernetes_sd_configs:
      - role: endpoints

scrape_native_histograms: true tells Prometheus to accept native histograms. always_scrape_classic_histograms: true keeps classic histogram samples when both formats are present. The second option is valuable during migration, but it also means temporarily ingesting both representations.

Reload Prometheus and check its configuration and target pages. A successful target alone is not sufficient; query a known histogram and confirm native samples are stored.

Understand the PromQL difference

A classic histogram quantile typically aggregates bucket series by le:

promql
histogram_quantile(
  0.99,
  sum by (le, job) (rate(apiserver_request_duration_seconds_bucket[5m]))
)

With native histograms, aggregate the histogram samples themselves:

promql
histogram_quantile(
  0.99,
  sum by (job) (rate(apiserver_request_duration_seconds[5m]))
)

Do not replace all production rules at once. Create parallel recording rules, compare the classic and native results over representative traffic, and investigate meaningful differences. Native histograms can provide better resolution, so exact equality is not the goal. You are looking for consistent operational conclusions and stable alert behavior.

Functions such as histogram_avg, histogram_count, and histogram_sum can make native-histogram queries easier to read. Verify each function against the Prometheus version used by your query layer.

A safe migration sequence

1. Inventory histogram consumers

Search dashboards, alerts, recording rules, SLO tooling, and ad-hoc runbooks for _bucket, histogram_quantile, and affected Kubernetes metric names. Include remote teams that query your central metrics platform.

2. Establish a baseline

Record active series, samples ingested per second, storage growth, scrape duration, query latency, and remote-write volume. Without a baseline, you cannot prove whether the migration helped.

3. Run dual exposition

Enable native ingestion plus classic preservation for a limited job. Keep the experiment long enough to include peak traffic and at least one normal incident or load test.

4. Create parallel queries

Add native versions of critical latency panels and alerts under temporary names. Compare quantiles, rates, missing-data behavior, and label dimensions. Confirm your alert notification templates still receive the labels they expect.

5. Migrate recording rules first

Dashboards often consume recording rules. Updating the rule layer first gives you one controlled compatibility boundary and avoids duplicating complicated PromQL across many dashboards.

6. Remove classic ingestion carefully

Only disable always_scrape_classic_histograms after all critical consumers have moved and your retention window no longer requires immediate comparison. Keep a documented rollback configuration.

Validation queries and checks

Use these checks during rollout:

  • Prometheus targets remain healthy and scrape errors do not increase.
  • Native histogram samples appear for the intended Kubernetes jobs.
  • Critical quantile panels show data for the same label scope as before.
  • Alerts fire in a controlled test and resolve correctly.
  • Remote-write receivers preserve the new sample type.
  • Series count, ingestion rate, storage use, and query duration stay within expected limits.

Also inspect mixed-version clusters. During a Kubernetes upgrade, not every component may expose the same representation at the same time. Queries and recording rules must tolerate that transition.

Common migration mistakes

Assuming default enablement means Prometheus is ready

Kubernetes producing native histograms does not guarantee that Prometheus or the remote backend accepts them. Validate every hop.

Dropping classic buckets too early

Existing rules that reference _bucket will return no data once classic buckets disappear. Dual scraping buys time to move consumers safely.

Comparing quantiles as exact numbers

Classic and native histograms use different bucket behavior. Compare trends, error-budget decisions, and alert thresholds rather than demanding identical decimals.

Ignoring the temporary duplication cost

Dual exposition is a migration tool, not necessarily a permanent setting. Track the extra samples and storage while both formats are retained.

Rollback plan

Keep the previous Prometheus configuration in version control. If a receiver, dashboard, or alerting path fails, restore classic histogram ingestion first, reload Prometheus, and verify the established recording rules. Do not downgrade Kubernetes solely to roll back observability ingestion; the compatibility controls belong in the metrics pipeline.

  • Upgrade the Prometheus path before relying on Kubernetes native histograms.
  • Canary one job and retain classic buckets temporarily.
  • Measure ingestion, storage, query latency, and alert results.
  • Migrate recording rules, then dashboards and alerts.
  • Test remote write and long-term queries.
  • Document rollback and ownership.
  • Remove duplicate classic ingestion only after consumer migration is complete.

For the wider cluster change plan, use the Kubernetes 1.37 upgrade guide. For monitoring fundamentals and dashboard structure, see the Prometheus and Grafana monitoring guide.

Prepare for Kubernetes interviews

If you are preparing for Kubernetes operations and observability questions, this CKA preparation course on Udemy can help you build a structured study plan.

Affiliate link: DevOpsBoys may earn a commission at no extra cost to you.

Official sources

🔧

Today I Fixed

Short real fixes from production — posted daily

Browse fixes
Newsletter

Stay ahead of the curve

Get the latest DevOps, Kubernetes, AWS, and AI/ML guides delivered straight to your inbox. No spam — just practical engineering content.

Related Articles

Comments