The rollout that never started INC-2530

Open2 versionsKubernetes · Medium · Fix · about 25 min ·Linux + Kubernetes

Lab machine

A private machine with its own Kubernetes cluster. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Dmitri Vos opened INC-2530 at 10:45SEV-3

Ivo applied Waybill 2.0.0 forty minutes ago. `kubectl rollout status` is still waiting, every pod is still 1.4.0, and no new pod has even been created.

2.0.0 carries the new depot-3 label format, and peak-season label printing starts on Monday. The namespace runs under a ResourceQuota, the depot's agreed share of the cluster.

"I bumped the memory request to 1Gi so it never OOMs again over peak. Should be harmless, the node has 8Gi." (Ivo)

"Please don't just raise the quota. That's the depot's agreed budget." (Mara)

When a Deployment waits without new pods appearing, the interesting part is one level down.

Your task

Ship Waybill 2.0.0 within the depot's quota. Size the memory request from the measurements, fix manifests/, and apply it.

On the machine

  • kubectl rollout status deployment/waybill, kubectl get rs
  • kubectl describe rs <newest>, kubectl describe resourcequota depot-budget
  • staging/memory-2.0.0.txt and staging/cpu-2.0.0.txt: what 2.0.0 really uses
  • manifests/

Timeline

10:05Ivo applies Waybill 2.0.0 with a 1Gi memory request.
10:20kubectl rollout status still waiting. All pods 1.4.0.
10:40"Please don't just raise the quota." (Mara)
10:45INC-2530 lands with you. The new label format ships with 2.0.0.

Done when

  1. All three Waybill pods run 2.0.0 and are Ready.
  2. Each pod requests at least the 140Mi 2.0.0 uses, and three pods plus one rollout pod fit the budget.
  3. The depot-budget ResourceQuota is unchanged.
  4. manifests/ matches the cluster.

Hints

Hint 1

The Deployment only waits. The ReplicaSet is the object that tried to create a pod and failed.

Hint 2

`kubectl describe resourcequota depot-budget` shows what is used and what is allowed, resource by resource.

Hint 3

A rolling update runs one extra pod for a while. Your requests must leave room for it.

Hint 4

`staging/` says how much memory and CPU 2.0.0 really needs. Size the request that does not fit from it.

Show the solution

`kubectl describe rs` on the new ReplicaSet shows FailedCreate: exceeded quota, and names the resource. Compare `kubectl describe resourcequota depot-budget` with the requests in `manifests/deployment.yaml`. Size the request that does not fit from `staging/`: memory just above the 140Mi peak (about 160Mi, limit 256Mi), or CPU around 50m to 100m, so that four pods fit the budget during a rollout. Apply it and watch `kubectl rollout status deployment/waybill` finish.