Pending after the node migration INC-2524

Open2 versionsKubernetes · Easy · Fix · about 20 min ·Linux + Kubernetes

Lab machine

A private machine with its own Kubernetes cluster. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Mara Okafor opened INC-2524 at 07:40SEV-2

Last night's node migration renamed the pool labels. Every node is Ready, but the edge copy of Waybill has been Pending for forty minutes and the depot scanners are queueing offline.

waybill-edge serves the depot scanners and must run on the edge pool, next to the depot VPN gateways. The migration moved pool labels to the northstar.example/ prefix, and other teams have already switched to the new name.

"Migration went fine, every node is Ready. I renamed the pool labels to the northstar.example/ prefix like we agreed. Didn't touch any workloads." (Sasha)

A Pending pod is waiting for the scheduler to find it a node. The scheduler writes down why it could not.

Your task

Get both waybill-edge pods running on the edge pool again by fixing the workload in manifests/ and applying it. Leave the node labels as the migration left them.

On the machine

  • kubectl get pods, kubectl describe pod <name>
  • kubectl get nodes --show-labels
  • manifests/

Timeline

01:00Node migration CHG-4455 renames pool labels to northstar.example/pool.
07:00waybill-edge is redeployed. Both pods stay Pending.
07:25"Every node is Ready. Didn't touch any workloads." (Sasha)
07:40INC-2524 lands with you.

Done when

  1. Both waybill-edge pods are Running and Ready, and the Service answers.
  2. The pods can still only run on edge-pool nodes.
  3. No node carries the retired `pool` label.
  4. manifests/ matches the cluster.

Hints

Hint 1

`kubectl describe pod` ends with the scheduler's reason. Read it before changing anything.

Hint 2

Compare every placement rule in the pod spec (nodeSelector and nodeAffinity) with `kubectl get nodes --show-labels`.

Hint 3

Putting the old label back on nodes undoes the migration other teams already rely on.

Hint 4

Change the placement rule in `manifests/deployment.yaml` to the new label key, then apply it.

Show the solution

`kubectl describe pod` says no node matches the pod's node selector or affinity. The pods still ask for the retired `pool=edge`, either in `nodeSelector` or in a required `nodeAffinity` term. Change that rule in `manifests/deployment.yaml` to the new key, `northstar.example/pool: edge`, apply it, and watch the new pods schedule with `kubectl get pods -w`.