Kubernetes 1.37 Native Histograms: Prometheus Migration Guide
Enable Kubernetes 1.37 native histograms safely, configure Prometheus 3, migrate queries and dashboards, and validate observability cost and accuracy.
Kubernetes 1.37 enables native histograms by default, giving platform teams a practical path away from manually chosen classic histogram buckets. That sounds like a simple monitoring upgrade, but enabling ingestion without planning queries, dashboards, storage, and rollback can leave you with two representations of the same metric and a confusing bill.
This guide explains what changed, how to prepare Prometheus, and how to migrate without losing alert coverage.
What changed in Kubernetes 1.37?
Native histograms reached beta in Kubernetes 1.37 and are enabled by default. Kubernetes components can expose both classic histogram series and a native histogram representation. The native format stores the distribution in one histogram sample instead of generating a separate time series for every configured bucket.
Classic histograms remain available during the transition. That compatibility is useful because existing PromQL expressions, recording rules, dashboards, and remote-write consumers may still depend on _bucket, _sum, and _count series.
Native histograms do not automatically make every query faster or every monitoring system cheaper. Their benefit depends on the bucket schema, scrape frequency, metric cardinality, retention, and whether downstream systems preserve the native data type.
Prerequisites
Before changing production scraping, confirm the complete metrics path supports native histograms:
- Kubernetes 1.37 components are exposing the new samples.
- Prometheus can ingest native histograms. Prometheus 3.0 or newer is the safest baseline for the per-job settings used below.
- Any agent, remote-write receiver, long-term store, and query layer supports the native histogram wire format.
- Your dashboard and alert tooling understands native histogram PromQL results.
- You have a storage and series-count baseline from before the migration.
If one component silently drops native samples, the scrape may look healthy while the data you expect is unavailable later in the pipeline.
Enable ingestion in Prometheus
Start with one Kubernetes scrape job or a canary Prometheus instance. A transition configuration can ingest native samples while retaining classic buckets:
scrape_configs:
- job_name: kubernetes-apiservers
scrape_native_histograms: true
always_scrape_classic_histograms: true
kubernetes_sd_configs:
- role: endpointsscrape_native_histograms: true tells Prometheus to accept native histograms. always_scrape_classic_histograms: true keeps classic histogram samples when both formats are present. The second option is valuable during migration, but it also means temporarily ingesting both representations.
Reload Prometheus and check its configuration and target pages. A successful target alone is not sufficient; query a known histogram and confirm native samples are stored.
Understand the PromQL difference
A classic histogram quantile typically aggregates bucket series by le:
histogram_quantile(
0.99,
sum by (le, job) (rate(apiserver_request_duration_seconds_bucket[5m]))
)With native histograms, aggregate the histogram samples themselves:
histogram_quantile(
0.99,
sum by (job) (rate(apiserver_request_duration_seconds[5m]))
)Do not replace all production rules at once. Create parallel recording rules, compare the classic and native results over representative traffic, and investigate meaningful differences. Native histograms can provide better resolution, so exact equality is not the goal. You are looking for consistent operational conclusions and stable alert behavior.
Functions such as histogram_avg, histogram_count, and histogram_sum can make native-histogram queries easier to read. Verify each function against the Prometheus version used by your query layer.
A safe migration sequence
1. Inventory histogram consumers
Search dashboards, alerts, recording rules, SLO tooling, and ad-hoc runbooks for _bucket, histogram_quantile, and affected Kubernetes metric names. Include remote teams that query your central metrics platform.
2. Establish a baseline
Record active series, samples ingested per second, storage growth, scrape duration, query latency, and remote-write volume. Without a baseline, you cannot prove whether the migration helped.
3. Run dual exposition
Enable native ingestion plus classic preservation for a limited job. Keep the experiment long enough to include peak traffic and at least one normal incident or load test.
4. Create parallel queries
Add native versions of critical latency panels and alerts under temporary names. Compare quantiles, rates, missing-data behavior, and label dimensions. Confirm your alert notification templates still receive the labels they expect.
5. Migrate recording rules first
Dashboards often consume recording rules. Updating the rule layer first gives you one controlled compatibility boundary and avoids duplicating complicated PromQL across many dashboards.
6. Remove classic ingestion carefully
Only disable always_scrape_classic_histograms after all critical consumers have moved and your retention window no longer requires immediate comparison. Keep a documented rollback configuration.
Validation queries and checks
Use these checks during rollout:
- Prometheus targets remain healthy and scrape errors do not increase.
- Native histogram samples appear for the intended Kubernetes jobs.
- Critical quantile panels show data for the same label scope as before.
- Alerts fire in a controlled test and resolve correctly.
- Remote-write receivers preserve the new sample type.
- Series count, ingestion rate, storage use, and query duration stay within expected limits.
Also inspect mixed-version clusters. During a Kubernetes upgrade, not every component may expose the same representation at the same time. Queries and recording rules must tolerate that transition.
Common migration mistakes
Assuming default enablement means Prometheus is ready
Kubernetes producing native histograms does not guarantee that Prometheus or the remote backend accepts them. Validate every hop.
Dropping classic buckets too early
Existing rules that reference _bucket will return no data once classic buckets disappear. Dual scraping buys time to move consumers safely.
Comparing quantiles as exact numbers
Classic and native histograms use different bucket behavior. Compare trends, error-budget decisions, and alert thresholds rather than demanding identical decimals.
Ignoring the temporary duplication cost
Dual exposition is a migration tool, not necessarily a permanent setting. Track the extra samples and storage while both formats are retained.
Rollback plan
Keep the previous Prometheus configuration in version control. If a receiver, dashboard, or alerting path fails, restore classic histogram ingestion first, reload Prometheus, and verify the established recording rules. Do not downgrade Kubernetes solely to roll back observability ingestion; the compatibility controls belong in the metrics pipeline.
Recommended rollout checklist
- Upgrade the Prometheus path before relying on Kubernetes native histograms.
- Canary one job and retain classic buckets temporarily.
- Measure ingestion, storage, query latency, and alert results.
- Migrate recording rules, then dashboards and alerts.
- Test remote write and long-term queries.
- Document rollback and ownership.
- Remove duplicate classic ingestion only after consumer migration is complete.
For the wider cluster change plan, use the Kubernetes 1.37 upgrade guide. For monitoring fundamentals and dashboard structure, see the Prometheus and Grafana monitoring guide.
Prepare for Kubernetes interviews
If you are preparing for Kubernetes operations and observability questions, this CKA preparation course on Udemy can help you build a structured study plan.
Affiliate link: DevOpsBoys may earn a commission at no extra cost to you.
Official sources
Today I Fixed
Short real fixes from production — posted daily
Stay ahead of the curve
Get the latest DevOps, Kubernetes, AWS, and AI/ML guides delivered straight to your inbox. No spam — just practical engineering content.
Related Articles
Build a Complete Kubernetes Monitoring Stack from Scratch (2026)
Step-by-step project walkthrough: set up Prometheus, Grafana, Loki, and AlertManager on Kubernetes using Helm. Real configs, real dashboards, production-ready.
Prometheus High Cardinality Causing OOM — How to Find and Fix It (2026)
Prometheus is crashing with OOMKilled or running out of memory. The culprit is almost always high cardinality metrics — labels with thousands of unique values. Here's how to find which metrics are killing your Prometheus and exactly how to fix it.
AI-Powered Kubernetes Anomaly Detection: Beyond Static Thresholds
Static alerts miss 40% of real incidents. Learn how AI and ML-based anomaly detection — using tools like Prometheus + ML, Dynatrace, and custom LLM runbooks — catches what thresholds can't.