🎉 DevOps Interview Prep Bundle is live — 1000+ Q&A across 20 topicsGet it →
All Articles

Kubernetes 1.37 Pod-Level Resource Managers: NUMA, CPU and Memory Guide

Test Kubernetes 1.37 pod-level resource managers with CPU, memory, and topology policies, then verify NUMA-aware allocation safely.

DevOpsBoys6 min read
Share:Tweet

Kubernetes traditionally makes CPU Manager, Memory Manager, and Topology Manager decisions around individual containers. Kubernetes 1.37 adds beta support for pod-level resource managers, allowing tightly coupled containers in a pod to receive coordinated CPU, memory, and NUMA placement.

The feature is valuable for latency-sensitive services, network functions, and accelerator-adjacent workloads, but it is disabled by default and has strict node-policy requirements. Treat it as a controlled performance feature, not a cluster-wide switch.

What pod-level management changes

A pod can contain an application container plus sidecars that share data paths or latency requirements. Container-scoped allocation can place those containers across different NUMA nodes, increasing cross-socket memory access and making performance less predictable.

With pod-level resource management, kubelet resource managers coordinate an allocation for the pod as a unit. The containers can draw from the pod-level resource budget, while Topology Manager evaluates their combined placement.

This behavior builds on pod-level resources; it does not remove container resource controls or make every workload NUMA-aware automatically.

Requirements and limitations

Kubernetes 1.37 marks the capability beta, but PodLevelResourceManagers remains disabled by default. Plan around these constraints:

  • Enable both PodLevelResources and PodLevelResourceManagers where required by your control-plane and node configuration.
  • The feature is for Linux nodes.
  • CPU Manager must use the static policy for pod-level CPU allocation.
  • Memory Manager must use the Static policy for pod-level memory allocation.
  • BestEffort memory-manager behavior is not supported for this feature.
  • Nodes need appropriate reserved CPU and memory configuration before static policies can operate safely.
  • Workloads must actually define pod-level resources to receive pod-level allocation behavior.

Check the exact Kubernetes 1.37 documentation and your distribution's feature-gate mechanism. Managed services may expose new gates later than upstream Kubernetes.

Prepare a test node pool

Use a dedicated node pool before enabling the feature in production. Mixing experiments with general-purpose workloads makes performance results difficult to interpret and raises the blast radius of kubelet configuration mistakes.

A conceptual kubelet configuration includes the feature gates and compatible manager policies:

yaml
featureGates:
  PodLevelResources: true
  PodLevelResourceManagers: true
cpuManagerPolicy: static
memoryManagerPolicy: Static
topologyManagerPolicy: single-numa-node
topologyManagerScope: pod
reservedSystemCPUs: "0-1"

The correct reserved CPUs, memory reservations, and topology policy depend on node hardware. Do not copy the example values directly to production. Inspect sockets, NUMA nodes, CPU siblings, huge pages, and operating-system reservations first.

After changing kubelet policies, follow the documented node drain and kubelet restart procedure for your environment. CPU and memory manager checkpoints make casual in-place policy changes risky.

Define pod-level resources

The pod specification can declare a shared resource budget at the pod level while containers retain their own configuration where needed. A simplified example looks like this:

yaml
apiVersion: v1
kind: Pod
metadata:
  name: numa-demo
spec:
  resources:
    requests:
      cpu: "4"
      memory: "4Gi"
    limits:
      cpu: "4"
      memory: "4Gi"
  containers:
    - name: application
      image: registry.k8s.io/pause:3.10
    - name: helper
      image: registry.k8s.io/pause:3.10

Use an application image with real CPU and memory behavior for performance testing. The pause image only illustrates the placement of the pod-level fields.

Guaranteed-quality workloads are the clearest starting point: keep requests equal to limits and use whole CPU values. This makes exclusive CPU assignment and placement results easier to reason about.

Choose the Topology Manager policy carefully

single-numa-node provides the strongest alignment requirement. The kubelet admits the pod only when the requested resources can be satisfied from one NUMA node under the active hint providers. This is useful for strict low-latency workloads but can reject pods on fragmented nodes.

Less restrictive policies may admit more workloads but provide weaker alignment. Choose according to the application SLO, not simply the desire to maximize scheduling density.

The scheduler can select a node that appears to have sufficient aggregate capacity while kubelet later rejects the pod because a valid topology hint is unavailable. Monitor admission failures and leave headroom in the specialized node pool.

Verification workflow

1. Confirm feature-gate and policy state

Inspect kubelet configuration on every canary node. A workload moving between inconsistently configured nodes will produce misleading results.

2. Check pod admission

Create one test pod and inspect events:

bash
kubectl describe pod numa-demo

Topology admission errors should be treated as capacity or policy signals, not fixed with repeated blind restarts.

3. Inspect allocated CPUs

On the node, compare the CPU Manager checkpoint with the CPUs visible to each container. Confirm the application's threads remain on the intended exclusive set.

4. Verify NUMA locality

Use node tools such as lscpu, numactl, and workload-specific telemetry to verify socket and memory locality. Kubernetes admission alone does not prove that the application's internal thread and memory behavior is optimal.

5. Query the PodResources API

The kubelet PodResources API and related metrics can expose allocated devices and CPU information to node-level monitoring agents. Use it to compare expected and actual placement across the canary pool.

6. Measure application outcomes

Track tail latency, throughput, CPU throttling, memory bandwidth, page faults, and cross-NUMA traffic before and after. A technically aligned allocation is only successful if it improves a workload outcome.

Safe rollout strategy

Label and taint the experimental nodes. Admit only workloads that explicitly tolerate the pool, then roll out in stages:

  1. Validate one synthetic workload.
  2. Move a non-critical production replica.
  3. Compare it with an unchanged control replica.
  4. Exercise rescheduling, node drain, and restart behavior.
  5. Test full-node and fragmented-capacity scenarios.
  6. Expand only after several representative load windows.

Add alerts for topology admission failures, unavailable exclusive CPUs, node pressure, and pods stuck outside Ready state.

Common failure modes

Pod remains pending or is rejected

The node may have enough total CPU and memory but no single NUMA node that satisfies the request. Inspect pod events, topology policy, current exclusive CPU allocation, and per-NUMA free memory.

Feature appears to do nothing

Confirm both feature gates, pod-level resource fields, and the static CPU and memory policies. Enabling only PodLevelResources is not the same as enabling pod-level resource managers.

Kubelet fails after policy changes

Manager checkpoint state may conflict with the new configuration. Drain the node and follow the official policy-change procedure instead of deleting checkpoint files on a live node without a recovery plan.

Performance gets worse

CPU pinning and NUMA alignment cannot repair application-level contention, undersized limits, noisy interrupts, or a sidecar doing unexpected work. Compare hardware counters and application profiles with your control group.

Production checklist

  • Use a dedicated, homogeneous Linux node pool.
  • Verify feature gates and manager policies on every node.
  • Reserve resources for the OS and Kubernetes daemons.
  • Start with Guaranteed pods and whole CPU requests.
  • Test NUMA admission under fragmented capacity.
  • Measure application SLOs, not only allocation state.
  • Document the drain, rollback, and checkpoint procedure.

Read the Kubernetes 1.37 upgrade guide before enabling new beta features across a cluster.

Prepare for Kubernetes interviews

For structured practice on scheduling, resources, and production operations, consider this CKA preparation course on Udemy.

Affiliate link: DevOpsBoys may earn a commission at no extra cost to you.

Official sources

🔧

Today I Fixed

Short real fixes from production — posted daily

Browse fixes
Newsletter

Stay ahead of the curve

Get the latest DevOps, Kubernetes, AWS, and AI/ML guides delivered straight to your inbox. No spam — just practical engineering content.

Related Articles

Comments