Deleting the pods made it worse INC-2371

Open2 versionsKubernetes · Easy · Fix · about 25 min ·Linux + Kubernetes

Lab machine

A private machine with its own Kubernetes cluster. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Teo Marin opened INC-2371 at 10:08SEV-2

10:02, a config change for depot-7. 10:05, the pods did not pick it up, so on-call deleted them. 10:06, every Waybill pod is crash-looping.

The old pods were still running with the config they started with. The new pods read the new config and die on it.

"I deleted them so they would reload. Now none of them come up." (on-call)

kubectl logs --previous shows what a crashed container said before it died. The fix belongs in manifests/, the source of truth, or the next apply puts the broken config back.

Your task

Find what the crashing containers object to, fix it in manifests/ keeping depot-7's settings, apply, and get both replicas Ready and stable.

On the machine

  • . ./env.sh then kubectl get pods, kubectl describe pod
  • kubectl logs deploy/waybill --previous
  • manifests/

Timeline

10:02Hazmat label template enabled for depot-7 in manifests/configmap.yaml, applied.
10:05"Pods weren't picking up the new config, so I deleted them."
10:06Every Waybill pod in CrashLoopBackOff. depot-7 scanners offline.
10:08INC-2371.

Done when

  1. Both Waybill replicas are Ready and stop restarting.
  2. /healthz through the Service shows depot-7's configuration loaded.
  3. The live ConfigMap and manifests/configmap.yaml agree.

Hints

Hint 1

The current container may be starting; ask for the previous one's logs.

Hint 2

The crash message names the file and the parser that rejected it.

Hint 3

JSON is strict. It allows no comma after the last property, and only straight double quotes.

Hint 4

Fix the manifest, apply it, and let the kubelet restart the containers (or restart the rollout).

Show the solution

Run `kubectl logs deploy/waybill --previous` to see the JSON parse error and where it is. Fix `manifests/configmap.yaml` (a trailing comma after the last field, or curly quotes pasted from chat around a value) and `kubectl apply -f manifests/`. The kubelet refreshes mounted ConfigMaps within about a minute, and the crashing containers recover on their next restart; `kubectl rollout restart deployment/waybill` makes it immediate.