The Problem
Spring Boot application pod would restart randomly, roughly every 2-3 hours. No consistent pattern. Logs showed the app was handling requests fine right before the restart.
kubectl get pod my-spring-app-xxx
# Restarts: 7What Happened
Spring Boot does periodic GC (garbage collection). During a full GC pause, the JVM freezes for 2-3 seconds. The liveness probe has a 1-second timeout. During GC, probe times out, Kubernetes thinks the pod is dead and kills it.
kubectl describe pod my-spring-app-xxx | grep -A10 "Liveness"
# Liveness: http-get http://:8080/actuator/health/liveness
# Delay=10s Timeout=1s Period=10s #Success=1 #Failure=3
#
# Warning Unhealthy: Liveness probe failed: context deadline exceededThe 1-second timeout was causing false positives during GC pauses.
The Fix
Increased timeout and added a startup probe to handle slow JVM startup separately:
startupProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
failureThreshold: 30 # 30 x 10s = 5 minutes for startup
periodSeconds: 10
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 0 # startupProbe handles initial delay
periodSeconds: 20
timeoutSeconds: 10 # 10s timeout — enough headroom for GC
failureThreshold: 3No restarts in 5 days since the fix.
Root Cause
Default probe settings (1s timeout) are too aggressive for JVM applications. Spring Boot's Actuator liveness endpoint normally responds in under 10 ms but blocks during GC. The fix: always set timeoutSeconds: 5-10 for JVM apps, and use startupProbe to separate slow startup from slow runtime.