🎉 DevOps Interview Prep Bundle is live — 1000+ Q&A across 20 topicsGet it →
All Articles

Kubernetes 1.37 Storage Hardening: Secure emptyDir with noexec, nosuid and mode

Harden Kubernetes writable volumes using v1.37 bindMountOptions and emptyDir mode, with manifests, tests, and Alpha feature warnings.

DevOpsBoys4 min read
Share:Tweet

A read-only container root filesystem does not automatically make every writable path safe. A compromised process may still download a payload into an emptyDir volume, mark it executable, and run it. Kubernetes 1.37 introduces two Alpha features that address this gap directly: bind mount options and configurable emptyDir permission modes.

These features are not enabled by default. Test them in a non-production cluster and verify container runtime support before changing workload manifests.

The two feature gates

VolumeBindMountOptions adds bindMountOptions to a container's volumeMount. On Linux, it can enforce flags such as:

  • noexec: block direct execution of binaries from the mount;
  • nosuid: ignore set-user-ID and set-group-ID bits;
  • nodev: do not interpret device files on the mount.

EmptyDirVolumeMode adds mode to an emptyDir definition. It can create a shared temporary directory with permissions such as 0750 or the /tmp-style sticky-bit mode 01777.

Enable the gates on the API server and kubelet in a test cluster:

text
--feature-gates=VolumeBindMountOptions=true,EmptyDirVolumeMode=true

Managed Kubernetes services may not expose Alpha gates. Check your provider's supported feature policy before planning adoption.

Block execution from a writable volume

yaml
apiVersion: v1
kind: Pod
metadata:
  name: hardened-workspace
spec:
  os:
    name: linux
  containers:
    - name: app
      image: alpine:3.22
      command: ["sleep", "3600"]
      securityContext:
        readOnlyRootFilesystem: true
        allowPrivilegeEscalation: false
      volumeMounts:
        - name: workspace
          mountPath: /workspace
          bindMountOptions:
            - noexec
            - nosuid
            - nodev
  volumes:
    - name: workspace
      emptyDir: {}

Verify noexec from inside the container:

bash
kubectl exec hardened-workspace -- sh -c '
  printf "#!/bin/sh\necho should-not-run\n" > /workspace/test.sh
  chmod +x /workspace/test.sh
  /workspace/test.sh
'

The direct execution attempt should fail. An interpreter may still be able to read a script and execute its contents, so noexec is one layer, not a complete sandbox. Keep seccomp, capability removal, non-root execution, image controls, and network policy in the design.

Give shared emptyDir storage a sticky bit

Multi-container Pods often use emptyDir as a shared /tmp or build workspace. Mode 01777 allows each process to create files while preventing ordinary users from deleting files owned by another user.

yaml
apiVersion: v1
kind: Pod
metadata:
  name: shared-tmp
spec:
  containers:
    - name: producer
      image: alpine:3.22
      command: ["sleep", "3600"]
      volumeMounts:
        - name: tmp
          mountPath: /tmp/shared
    - name: observer
      image: alpine:3.22
      command: ["sleep", "3600"]
      volumeMounts:
        - name: tmp
          mountPath: /tmp/shared
  volumes:
    - name: tmp
      emptyDir:
        mode: 01777

Check the resulting permissions:

bash
kubectl exec shared-tmp -c producer -- stat -c '%a %A %n' /tmp/shared

For application-private scratch data, a narrower mode such as 0750 may be more appropriate. Match the owner and group design to the Pod security context.

Runtime and scheduling behavior

bindMountOptions depends on the container runtime supporting the CRI mount_options field and advertising that support through runtime features. Kubernetes uses node-declared features to keep these Pods away from incompatible nodes. If an incompatible Pod reaches a node anyway, kubelet rejects it instead of silently dropping the options.

emptyDir.mode does not require the same runtime capability. However, when the API server has the gate and kubelet does not, kubelet can fall back to the traditional behavior. Verify the effective mode on every node pool during rollout.

Other important limits:

  • both capabilities are Linux-focused; Windows does not implement these Unix flags and modes;
  • image volumes do not support bindMountOptions;
  • fsGroup processing can override group permissions from emptyDir.mode;
  • PersistentVolume mountOptions operate at the storage/CSI layer, while bindMountOptions affect the container bind mount.

Production rollout checklist

  1. Inventory workloads with writable emptyDir, PVC, Secret, ConfigMap, and projected mounts.
  2. Identify paths that must execute binaries; do not apply noexec blindly.
  3. Confirm API server, kubelet, scheduler, and runtime support in a test node pool.
  4. Add policy checks for approved mount options and permission modes.
  5. Test application startup, sidecars, init containers, package extraction, and upgrade paths.
  6. Roll out by namespace and watch Pod scheduling and kubelet events.
  7. Keep a rollback manifest while the features remain Alpha.

Want deeper hands-on Kubernetes administration practice? Explore this CKA preparation course on Udemy. Affiliate link: DevOpsBoys may earn a commission at no extra cost to you.

Sources

🔧

Today I Fixed

Short real fixes from production — posted daily

Browse fixes
Newsletter

Stay ahead of the curve

Get the latest DevOps, Kubernetes, AWS, and AI/ML guides delivered straight to your inbox. No spam — just practical engineering content.

Related Articles

Comments