🎉 DevOps Interview Prep Bundle is live — 1000+ Q&A across 20 topicsGet it →
All Fixes
Today I Fixed

Spring Boot Pod Randomly Restarting — Liveness Probe Timing Issue

kubectlMay 31, 202645 minutes to fixkubernetestroubleshooting

The Problem

Spring Boot application pod would restart randomly, roughly every 2-3 hours. No consistent pattern. Logs showed the app was handling requests fine right before the restart.

bash
kubectl get pod my-spring-app-xxx
# Restarts: 7

What Happened

Spring Boot does periodic GC (garbage collection). During a full GC pause, the JVM freezes for 2-3 seconds. The liveness probe has a 1-second timeout. During GC, probe times out, Kubernetes thinks the pod is dead and kills it.

bash
kubectl describe pod my-spring-app-xxx | grep -A10 "Liveness"
# Liveness: http-get http://:8080/actuator/health/liveness
# Delay=10s Timeout=1s Period=10s #Success=1 #Failure=3
#
# Warning Unhealthy: Liveness probe failed: context deadline exceeded

The 1-second timeout was causing false positives during GC pauses.

The Fix

Increased timeout and added a startup probe to handle slow JVM startup separately:

yaml
startupProbe:
  httpGet:
    path: /actuator/health/liveness
    port: 8080
  failureThreshold: 30    # 30 x 10s = 5 minutes for startup
  periodSeconds: 10
 
livenessProbe:
  httpGet:
    path: /actuator/health/liveness
    port: 8080
  initialDelaySeconds: 0   # startupProbe handles initial delay
  periodSeconds: 20
  timeoutSeconds: 10       # 10s timeout — enough headroom for GC
  failureThreshold: 3

No restarts in 5 days since the fix.

Root Cause

Default probe settings (1s timeout) are too aggressive for JVM applications. Spring Boot's Actuator liveness endpoint normally responds in under 10 ms but blocks during GC. The fix: always set timeoutSeconds: 5-10 for JVM apps, and use startupProbe to separate slow startup from slow runtime.