Killed at a limit nobody chose INC-2388

Open2 versionsKubernetes · Hard · Investigate · about 35 min ·Linux + Kubernetes

Lab machine

A private machine with its own Kubernetes cluster. Starting takes about 30 seconds. Sessions last up to 60 minutes.
Teo Marin opened INC-2388 at 12:20SEV-2

Docklight moved to Kubernetes last week. Its worker restarts every few seconds with exit code 137, and on-call wants to shrink the batch size the carrier contract fixes.

"Node has gigabytes free, must be a bug in the worker. Lower BATCH_MB to 8?" (on-call)

The carrier contract fixes batches at 64 MiB. The platform budget for this namespace is in capacity.md.

Your task

Find why the worker is killed and size its memory limit (and request) for a 64 MiB batch plus the runtime, within the 256 MiB budget. Keep BATCH_MB at 64, fix manifests/, apply, and watch batches complete.

On the machine

  • kubectl describe pod (Last State, Reason, Exit Code, Limits)
  • kubectl logs of the worker
  • kubectl describe limitrange
  • manifests/, capacity.md

Timeline

Last weekDocklight worker migrated to Kubernetes; limits copied from a chart default.
12:00Carrier batches go back to 64 MiB after a contract update.
12:01Worker pod: OOMKilled, restarting.
12:20"Raise BATCH_MB down to 8?"

Done when

  1. The worker processes 64 MiB batches continuously without being killed.
  2. It keeps a memory limit of 256 MiB or less, with a request no larger than the limit.
  3. manifests/ matches the cluster.

Hints

Hint 1

The kill is recorded in the pod's container status, not in the node's free memory.

Hint 2

Find where the pod's limit comes from: the manifest, or a namespace default (`kubectl describe limitrange`).

Hint 3

The batch size is a contract; the limit is yours to size.

Hint 4

Set an explicit memory limit (and a sensible request) in the manifest, within the budget, and apply it.

Show the solution

`kubectl describe pod` shows OOMKilled, exit 137, at a 64Mi limit. That limit is either the chart default copied into `manifests/worker.yaml`, or, when the manifest sets none, the namespace LimitRange default (`kubectl describe limitrange`). Set an explicit memory limit in `manifests/worker.yaml` that covers a 64 MiB batch plus the Node runtime (for example 160Mi, request 128Mi), apply it, and watch several batches complete with no restarts.