AWS CloudFront 502 Bad Gateway: Fix Origin TLS, DNS, and Connection Errors
Diagnose CloudFront 502 errors by checking origin DNS, certificate names, TLS chains, ports, and edge-function failures in the right order.
159 articles
Diagnose CloudFront 502 errors by checking origin DNS, certificate names, TLS chains, ports, and edge-function failures in the right order.
Fix GitHub Actions self-hosted runner registration failures caused by the 2026 minimum-version brownout and prepare for full enforcement.
Diagnose Cloudflare API error 403 by checking token scopes, account and zone resources, authentication headers, and resource-scoped roles.
HTTPRoute created but traffic returns 404 or its status shows Accepted=False? Diagnose parentRefs, listeners, hostnames, ReferenceGrant, backend ports, and controller status step by step.
Messages sitting in an SQS queue forever, never picked up by your consumer? Here is exactly how to diagnose visibility timeout misconfiguration, IAM permission gaps, dead-letter queue redirects, and consumer polling issues.
Ingress or LoadBalancer Service showing EXTERNAL-IP as <pending> forever? Here is exactly how to diagnose missing ingress controllers, cloud provider quota limits, and misconfigured service annotations blocking IP assignment.
terraform destroy failing with 'DependencyViolation' or hanging on a resource that refuses to delete? Here is exactly how to diagnose orphaned dependencies, out-of-band changes, and deletion protection blocking a clean destroy.
Multi-stage Docker builds re-running every layer from scratch even when nothing relevant changed? Here is exactly how to diagnose layer ordering, build context, and BuildKit cache export issues causing multi-stage cache misses.
HorizontalPodAutoscaler scales up fine but never scales back down, leaving you paying for pods you don't need? Here is exactly how to diagnose stabilization windows, metric server lag, and pod disruption budgets blocking scale-down.
helm rollback failing with 'another operation is in progress' or leaving the release in a broken state? Here is exactly how to diagnose stuck release secrets, failed hooks, and resource conflicts blocking a Helm rollback.
GitLab Runner showing as offline or stale even though the process is running? Here is exactly how to diagnose token issues, network connectivity, version mismatches, and concurrency exhaustion behind a stuck runner.
npm install failing with 'ERESOLVE unable to resolve dependency tree'? Here is exactly how to read the conflict, find the real peer dependency mismatch, and fix it without blindly reaching for --force.
terraform apply stuck with no output, no progress, no error? Here is exactly how to diagnose provider API rate limits, state lock deadlocks, dependent resource waits, and network issues causing Terraform to hang indefinitely.
CI failing with 'No space left on device' on GitHub Actions or self-hosted runners? Here is exactly how to diagnose what's eating disk — Docker layers, build artifacts, or the runner's own accumulation — and fix it for good.
kubectl port-forward starts fine but every connection gets refused or hangs? Here is exactly how to diagnose pod readiness, wrong target port, network policy, and CNI causes behind port-forward failures.
ArgoCD Application stuck in Progressing status forever, never reaching Healthy? Here is exactly how to diagnose stuck rollouts, missing health checks, hook failures, and resource hooks that never complete.
Getting 'network not found' or 'network ... declared as external, but could not be found' from Docker Compose? Here is exactly how to diagnose and fix external network references, stale networks, and project name mismatches.
sts:AssumeRole failing with AccessDenied even though the role exists and the policy looks right? Here is exactly how to diagnose trust policy, permission boundary, session policy, and external ID causes.
Getting 403 Forbidden from S3 even though you're sure the bucket policy is right? Here is exactly how to diagnose IAM policy, bucket policy, ACL, block-public-access, and KMS key permission causes.
Node stuck in NotReady status and pods getting evicted or stuck Pending? Here is exactly how to diagnose kubelet, container runtime, network plugin, and disk pressure causes — and fix each one.
Terraform failing with 'no available releases match the given constraints' or two modules requiring incompatible provider versions? Here is exactly how to diagnose and fix provider version conflicts without breaking your state.
Build a CLI tool using Claude API that automatically collects kubectl logs, events, describe output, and resource metrics from broken pods — then generates root cause analysis and step-by-step fix commands in plain English.
CrashLoopBackOff in Kubernetes? This guide covers every cause — bad command, OOMKill, readiness probe killing container, config missing, dependency not ready — with exact kubectl commands to find and fix each one.
RDS connection timeout in production? These 6 fixes cover security group rules, VPC subnet routing, parameter group connection limits, max_connections exceeded, SSL/TLS misconfig, and connection pool exhaustion — with exact AWS CLI commands.
Nginx Ingress returning 503 Service Unavailable? These 5 fixes cover no healthy backends, wrong service name, port mismatch, readiness probe failures, and connection refused — with exact kubectl commands.
GitHub Actions workflow not running on push, PR, or schedule? These 7 fixes cover branch filters, path filters, workflow_dispatch, YAML syntax, and permission issues with exact examples.
Existing AWS resources not in Terraform state? Use terraform import to bring them under management, fix state drift, and avoid accidental resource deletion — with exact commands for EC2, RDS, S3, VPC, and EKS.
HorizontalPodAutoscaler not scaling up or stuck at min replicas? These 6 fixes cover missing Metrics Server, wrong metric name, resource requests not set, and cooldown periods — with exact kubectl commands.
Docker builds ignoring cache and rebuilding from scratch every time? These 5 fixes — Dockerfile layer ordering, BuildKit cache mounts, .dockerignore issues, and multi-stage optimizations — cut build times by 60-80%.
Getting 'no basic auth credentials' or 'denied: Your authorization token has expired' when pushing to AWS ECR? Here are the exact commands to fix authentication for Docker, GitHub Actions, and Kubernetes.
ArgoCD sync failing with 'unknown field', 'strict decoding error', or 'field is immutable'? Here are the exact causes and kubectl/ArgoCD commands to fix each one without recreating resources.
StatefulSet pods stuck in Pending, Init, or CrashLoopBackOff? This guide covers the 6 most common causes — PVC binding, headless service missing, init container failures, pod identity issues — with exact kubectl commands to diagnose and fix each.
Fix for Kubernetes deployments stuck in rollout — covering insufficient cluster resources, failing readiness probes, PodDisruptionBudgets blocking drain, and max unavailable settings causing stalls.
How to resolve Kubernetes field manager conflict errors when using server-side apply — what causes them, how to identify the conflicting manager, and three different ways to fix it cleanly.
Step-by-step troubleshooting guide for PersistentVolumeClaims stuck in Pending state — covering StorageClass issues, provisioner problems, capacity limits, and access mode mismatches.
If your ECS task stops immediately with exit code 1 or exit code 137, this guide covers how to find the actual error, common causes like missing env vars, IAM permissions, and memory limits, and exact steps to fix each.
Terraform's null_resource is deprecated in favor of terraform_data. Learn what changed, how to migrate your existing null_resource blocks, and common patterns that now have better native solutions.
Fixing Helm errors like 'Error: chart not found', 'repository not found', and 'failed to fetch' when running helm install or helm upgrade. Covers outdated repo index, auth issues, and OCI registry problems.
If your Kubernetes pods are being evicted with 'The node was low on resource: ephemeral-storage' or 'DiskPressure' status, this guide walks you through finding the cause and fixing it permanently.
If your ArgoCD application is stuck showing Unknown sync status and refresh doesn't help, this guide walks you through the exact steps to diagnose and fix it — including RBAC issues, webhook problems, and cluster connectivity errors.
Getting state conflicts or duplicate resource errors with terraform import? Learn how to import existing AWS resources, fix config mismatches, and use moved blocks to rename state entries.
RDS CPU suddenly at 90%+? Learn how to use Performance Insights, pg_stat_statements, EXPLAIN ANALYZE, and connection pooling with RDS Proxy to find and fix slow queries fast.
Fix the ERR max number of clients reached error in Redis. Learn how to diagnose connection pool exhaustion, tune maxclients, fix connection leaks in Python and Node.js, and right-size your pool.
HPA scales on CPU but ignores your Prometheus or SQS custom metrics? Learn how the custom metrics adapter works, fix common errors, and use KEDA as a drop-in alternative.
Debug and fix Nginx Ingress 502/504 errors caused by 'upstream connect error or disconnect/reset before headers' in Kubernetes. Step-by-step with real commands.
Fix Grafana panels stuck on 'No data' or spinning forever. Covers datasource issues, time range mismatches, variable resolution failures, Prometheus scrape interval mismatches, and broken panel JSON.
PVC is Bound but your pod is stuck in ContainerCreating with a volume mount error? Here's how to diagnose and fix CSI driver issues on Kubernetes and EKS step by step.
ECS service ignoring your new task definition revision? Fix it with force-new-deployment, understand pinned ARNs vs LATEST, check the deployment circuit breaker, and verify which revision is actually running.
Kubernetes CronJob running the same job multiple times? Getting duplicate executions or jobs running concurrently when they shouldn't? Here are the fixes.
Building Docker images for both ARM64 and AMD64 fails with QEMU errors, manifest issues, or wrong architecture? Here are the exact fixes.
Compare the top AI-assisted Kubernetes debugging tools. K8sGPT, kubectl-ai, and k9s each help differently — here's when to use which and what they actually do.
Terraform applying to the wrong environment because workspace state is confused? Here's how to diagnose, fix, and prevent workspace state mismatches.
Fix GitHub composite action inputs that are empty or not passed correctly. Learn inputs syntax, env mapping, secrets handling, and debugging steps.
FluxCD source-controller or kustomize-controller getting OOMKilled or stuck in reconciliation loop? Here's how to diagnose and fix it.
ArgoCD PreSync or PostSync hooks failing silently? Here's how to find the real error, fix hook job issues, and stop your deployments from getting stuck.
Redis is hitting maxmemory and evicting keys you didn't expect to lose, or refusing writes entirely. Here's how to diagnose which eviction policy is wrong for your use case and fix it properly.
Your self-hosted runner shows 'Offline' in GitHub even though the service is running. Here's how to actually diagnose the cause — network, token expiry, or service crash — and fix it for good.
HashiCorp Vault restarts sealed and won't come back up, blocking every service that reads secrets from it. Here's how to diagnose unseal failures and fix the root cause, not just unseal-and-pray.
Consumer lag climbing and you don't know why? Here's a step-by-step way to find whether it's slow processing, rebalancing storms, or partition skew — and how to actually fix each cause.
Your GitHub Actions artifact upload is failing with 'upload artifact failed' or size limit errors. Here are the exact causes and fixes for the most common artifact upload failures.
You created a silence in AlertManager but alerts are still firing. Here are the 6 most common reasons silences fail and exactly how to fix each one.
kubectl drain hanging with no output? PodDisruptionBudget or DaemonSet pods are blocking it. Here's how to diagnose and fix it without nuking your cluster.
Your Kustomize overlay patch isn't modifying the base resource. Here's every reason patches silently fail and exactly how to debug and fix each one.
Your site hits an infinite redirect loop after adding TLS to Nginx Ingress. Here's every reason it happens and exactly how to fix each one.
Your pods are getting evicted with 'The node was low on resource: ephemeral-storage'. Here's why disk pressure evictions happen and exactly how to fix them.
Prometheus shows your target as DOWN in the Targets page. Here's every reason a scrape target goes down and exactly how to debug and fix each one.
Your Kubernetes pod can't access AWS services even though IRSA is configured. Here's every reason IRSA fails and exactly how to debug and fix each one.
Multi-stage Docker builds failing mid-build? Fix RUN cache misses, COPY path errors, wrong base image, build args not passing between stages, and secret leaking in intermediate layers.
Your EKS pods are Pending, nodes should be scaling up but aren't. Here's every reason EKS managed node groups fail to scale and exactly how to fix each one.
Your pod crashes immediately or has no shell. Here's how to use kubectl debug with ephemeral containers to inspect the container filesystem, processes, and environment without modifying the pod spec.
Your namespace has been Terminating for hours. kubectl delete won't help. Here's exactly why this happens and how to force-remove a stuck namespace safely.
Helm says the release is deployed but the app isn't responding. Here's why this happens and how to diagnose what's actually broken.
Your ALB shows targets as unhealthy and traffic isn't reaching your app. Here's every reason target health checks fail and exactly how to fix each one.
Your pods randomly restart with liveness probe failures, even when the app is healthy. Here's every reason this happens and how to tune probes correctly.
Your Docker build works perfectly on your machine but fails in GitHub Actions, GitLab CI, or Jenkins. Here's every reason this happens and exactly how to fix it.
You ran terraform apply and it deleted something it shouldn't have. Here's how to recover from accidental Terraform destroys before they become a disaster.
Getting 'no space left on device' during docker build? Here's exactly what's eating your disk and the commands to fix it fast without losing your images.
Node drain stuck or failing? Getting 'cannot delete Pods' or 'PodDisruptionBudget' errors? Here are all the fixes for kubectl drain problems.
Your pod worked fine before you updated resource limits, now it's in CrashLoopBackOff. Here's exactly why this happens and how to fix it without downtime.
Your EKS Cluster Autoscaler isn't scaling up, scale-down isn't working, or nodes spin up but stay empty. Here's every cause and the exact fix.
Your ArgoCD App of Apps pattern stopped syncing. Child apps aren't created, parent shows OutOfSync, or sync is stuck. Here are every cause and the exact fix.
Your pods suddenly can't talk to each other after applying a NetworkPolicy. DNS is failing, services are unreachable, or inter-namespace traffic is blocked. Here's every cause and the exact fix.
Your GitHub Actions workflow can't authenticate to AWS using OIDC. You're getting 'Not authorized to perform sts:AssumeRoleWithWebIdentity' or token errors. Here's every cause and the exact fix for each one.
Your ECS services can't find each other. Service Connect or Cloud Map DNS isn't resolving. Here's every cause — wrong namespace, missing IAM, wrong DNS config, VPC resolver issues — and exactly how to fix each one.
terraform init fails with S3 backend errors — access denied, bucket does not exist, state lock issues, wrong region. Here's every cause and the exact fix for each one.
kubectl says 'context not found', 'no configuration has been provided', or 'error loading config file'. Here's every cause — wrong KUBECONFIG path, missing context, corrupted config — and exactly how to fix each one.
Prometheus is crashing with OOMKilled or running out of memory. The culprit is almost always high cardinality metrics — labels with thousands of unique values. Here's how to find which metrics are killing your Prometheus and exactly how to fix it.
Docker daemon fails to start, or the systemd service crashes immediately. Here's how to diagnose dockerd startup failures from storage driver conflicts to corrupted config files.
CodeBuild exits with status 1, times out mid-build, or fails with cryptic phase errors. Here's how to diagnose DOWNLOAD_SOURCE, BUILD, and POST_BUILD failures with specific fixes.
ArgoCD Image Updater detects a new image tag but doesn't update the Application. Here's how to diagnose and fix annotation errors, registry auth issues, write-back problems, and sync failures.
Datadog dashboards show no data, hosts appear offline, or custom metrics aren't showing up. Here's how to systematically diagnose and fix Datadog agent issues on Kubernetes and VMs.
Your workflow runs are still slow even with actions/cache. Cache miss every time, cache key conflicts, wrong paths — here's how to diagnose and fix GitHub Actions caching.
Init container errors block your main containers from starting. Here's how to find the root cause fast — OOMKilled, permission errors, missing secrets, and more.
CloudFormation stack stuck in ROLLBACK_FAILED or UPDATE_ROLLBACK_FAILED state? Here's every cause and the exact steps to recover without losing your resources.
Your Prometheus alert should have fired 30 minutes ago but nothing happened. Here's every reason alerts silently fail — routing, inhibition, receivers, and rule syntax.
StatefulSet pods stuck in Init, Pending, or CrashLoopBackOff state? Here are every cause and the exact fix — PVC ordering, headless service, and init container issues.
GitLab Runner failing to register, showing offline, or not picking up jobs? Here's how to debug and fix runner registration and connection issues.
Node shows DiskPressure condition and pods are getting evicted? Here's how to find what's eating disk space and fix it permanently.
Vault Agent Injector not mounting secrets into your pod? Here's how to debug and fix Vault secret injection issues in Kubernetes step by step.
Getting 'admission webhook denied the request' or webhook timeout errors in Kubernetes? Here's how to debug and fix admission webhook issues step by step.
Kubernetes CronJob missed schedule or not triggering? Here's how to debug and fix CronJob scheduling issues — timezone problems, startingDeadlineSeconds, concurrencyPolicy, and more.
Getting connection timeout or upstream timed out errors through NGINX Ingress? Here's how to debug and fix timeout issues between NGINX and your backend services.
Terraform failing with version constraint conflicts between modules? Here's how to diagnose which module is causing the conflict and resolve it without breaking everything.
Secrets are unavailable and workflows fail when PRs come from forks. Here's why GitHub blocks fork access to secrets by default and the right way to handle it safely.
Pod can't resolve service names or external domains? DNS failures inside Kubernetes pods are caused by CoreDNS issues, ndots config, search domains, or network policies. Here's how to debug and fix each.
Docker builds taking 10+ minutes every time? Here's how to fix layer caching, use BuildKit properly, and cut build times by 80% with multi-stage builds and cache mounts.
Lambda function hitting the 15-minute limit or timing out before that? Here's how to find what's slow, increase timeout properly, and redesign for async patterns.
RDS instance hits 100% storage and your database goes read-only. Here's the immediate fix, how to prevent it with autoscaling, how to monitor free storage, and what's eating your disk.
Your GitHub Actions job times out after 6 hours or hits a custom timeout limit. Here's every cause — hung Docker builds, hanging tests, stuck deployments, missing timeout config — and the exact fix.
Your browser shows 'Access to fetch blocked by CORS policy' when loading from S3 or CloudFront. Here's every cause — missing CORS config, wrong AllowedOrigins, preflight failures — and the exact fix.
CloudFront returns 403 Forbidden but your S3 bucket or origin looks fine. Here's every cause — OAC misconfiguration, bucket policy missing, wrong origin domain, geo-restriction — and the exact fix.
Your pod starts but the ConfigMap or Secret isn't mounted as expected — missing files, stale values, wrong keys, or permission errors. Here's every cause and the exact fix.
Fix AWS ECR Docker push access denied and no basic auth credentials errors. Check IAM permissions, ECR login tokens, region, repository, and image URI.
Fix Jenkins Pipeline Groovy errors caused by bad interpolation, missing brackets, CPS restrictions, and declarative syntax with working examples.
EKS worker nodes stuck in NotReady or not appearing at all? Here are all the causes and step-by-step fixes for node bootstrap failures.
Getting Terraform lock file conflicts, provider version mismatches, or 'required provider not installed' errors? Here's every cause and how to fix it.
Fix Docker push permission denied or unauthorized errors in GitHub Actions. Check registry login, token scopes, package access, and workflow permissions.
EKS pods can't connect to RDS? Fix RDS connection timeouts from Kubernetes — covers security groups, VPC peering, subnet routing, and IAM auth issues.
Kubernetes readiness probe keeps failing? Learn the exact commands to diagnose misconfigured HTTP, TCP, and exec probes and fix them permanently.
Kubernetes Job stuck in Running, CronJob never triggering, or pods completing but Job shows Failed? Here are the real causes and fixes for each scenario.
Getting 'connection refused' between Docker Compose containers? Here are the 7 most common causes and how to fix each one, with working examples.
Pods stuck in Pending on EKS Fargate? Here are the 8 most common reasons Fargate pods won't schedule and exactly how to fix each one.
Your Service exists but traffic isn't reaching pods. Curl times out, 502s keep coming. Here's every reason a Kubernetes Service fails to route and how to fix each one.
Your PersistentVolumeClaim is stuck in Pending and your pod won't start. Here's every reason this happens and exactly how to fix it.
Your Prometheus /targets page shows red. Services are running but Prometheus can't scrape them. Here's every reason this happens — wrong port, NetworkPolicy blocks, ServiceMonitor label mismatch, auth — and exactly how to fix each one.
Your ALB returns 504 Gateway Timeout but the app seems fine. Here's every reason this happens — backend timeouts, keepalive mismatches, health check failures — and exactly how to fix each one.
Getting AccessDenied when Terraform tries to read or write your S3 backend? Here's every cause and the exact fix.
Your Kubernetes node is showing NotReady status. Here's every cause and the exact fix for each one.
Your ECS task starts and then immediately stops or keeps restarting. Here's every reason this happens and how to debug and fix it.
Pods suddenly showing 'Evicted' status in Kubernetes? Here's every reason nodes evict pods and exactly how to prevent it from happening again.
Getting 502 Bad Gateway from your Nginx Ingress Controller? Here's every cause and the exact fix for each one.
Getting 'Access Denied' or 'is not authorized to perform' errors in AWS? Here's how to diagnose and fix every IAM permission issue — EC2, EKS, Lambda, S3, and CLI.
Your helm upgrade ran successfully but nothing changed in the cluster. Here's every reason this happens and how to fix each one.
Comparing the three most popular CI/CD platforms head-to-head: features, pricing, speed, and when to pick each one in 2026.
Getting 'permission denied' when running kubectl exec? Here are all the real reasons it happens and exactly how to fix each one.
Getting 'Error: UPGRADE FAILED: another operation (install/upgrade/rollback) is in progress' in Helm? Here's exactly why it happens and how to fix it in under 2 minutes.
Fix AWS Application Load Balancer unhealthy targets. Covers health check misconfigurations, security group issues, target group problems, and EKS-specific ALB controller debugging.
Fix Terraform plan showing unexpected resource destruction. Covers state drift, provider upgrades, import mismatches, lifecycle rules, and safe recovery strategies.
Fix PodDisruptionBudget misconfigurations that block kubectl drain during cluster upgrades, node maintenance, and autoscaler operations. Real scenarios and step-by-step solutions.
Fix Kubernetes DNS resolution failures caused by CoreDNS misconfigurations, ndots issues, and pod DNS policies. Real troubleshooting scenarios with step-by-step solutions.
Fix the common Helm error 'has no deployed releases' that blocks upgrades. Step-by-step diagnosis and 4 proven solutions including history cleanup and force replacement.
Pods won't delete and stuck in Terminating? Here's how to diagnose finalizers, graceful shutdown issues, and force-delete stuck pods step by step.
VPA restarting your pods every time it adjusts resources? Here's how to stop the evictions using Kubernetes 1.35's In-Place Pod Resize feature.
Ingress-NGINX is officially being retired. Your ingress rules will stop working. Here's the step-by-step migration plan to Kubernetes Gateway API before it's too late.
ArgoCD app won't sync? Stuck in OutOfSync or Progressing state forever? Here's every cause and how to fix each one step by step.
GitHub Actions failing with 'no space left on device'? Here's how to free disk space on runners, optimize Docker builds, and handle large monorepos.
Your terraform plan looks clean but apply blows up? Here's how to fix provider conflicts, state drift, and dependency errors step by step.
Pods can't resolve hostnames? Getting NXDOMAIN or 'no such host' errors? Here's how to diagnose and fix CoreDNS issues in Kubernetes step by step.
ImagePullBackOff is one of the most common Kubernetes errors. This guide covers every root cause — wrong image names, missing auth, network issues, rate limits — with step-by-step debugging and fixes.
cert-manager Certificate stuck in a non-Ready state is a common Kubernetes TLS issue. This guide covers every root cause — DNS challenges, RBAC, rate limits, and issuer problems — with step-by-step fixes.
Pods stuck in Pending on EKS are caused by a handful of known issues — insufficient node capacity, taint mismatches, PVC problems, and more. Here's how to diagnose and fix each one.
Debug a failed GitLab CI pipeline step by step. Fix YAML errors, missing variables, runner problems, Docker-in-Docker failures, and job dependency issues.
Terraform state lock errors can block your entire team. Learn why they happen, how to safely unlock state, and how to prevent lock conflicts for good.
OOMKilled crashes killing your pods? Here's the real cause, how to diagnose it fast, and the exact steps to fix it without breaking production.
CrashLoopBackOff, OOMKilled, exit code 1, exit code 137 — Docker containers restart for specific, diagnosable reasons. Here is how to identify the exact cause and fix it in minutes.
Your pod says Pending and nothing is happening. Here's how to diagnose every possible reason — insufficient resources, taints, PVC issues, node selectors — and fix them fast.
Helm upgrade failing silently? Release stuck in pending state? This guide covers the 10 most common Helm errors DevOps engineers hit in production — with exact commands and fixes.
Your CI/CD pipeline failed and you don't know why. This complete debugging guide covers GitHub Actions, Jenkins, and ArgoCD failures with real error messages and step-by-step fixes.
The most complete Kubernetes troubleshooting guide for 2026. Learn how to diagnose and fix Pod crashes, ImagePullBackOff, OOMKilled, CrashLoopBackOff, networking issues, PVC problems, node NotReady, and more — with exact kubectl commands.