OpenProject intermittently unreachable (web pod OOM-killed)

Symptom

  • Users report that https://projects.kup.tirol is “sometimes” unavailable for a minute or two, then works again.

  • No alert fires, and the ArgoCD application shows Synced and Healthy the whole time.

  • kubectl -n kup-openproject get pods shows a non-zero restart count on openproject-web that grows by a few restarts per week.

  • kubectl -n kup-openproject describe pod <web-pod> reports Last State: Terminated, Reason: OOMKilled, Exit Code: 137.

Root cause

The web container runs out of memory because its limit sits below Puma’s normal working range.

OpenProject 17 runs Puma with several worker processes and settles at roughly 1.6 to 2.0 GiB RSS under everyday load. A 2 GiB limit therefore leaves no headroom at all. The container spends days pinned at its ceiling, the kernel reclaims page cache continuously, and any load spike pushes it over the edge.

Two independent effects then make the outage visible to users.

Each OOM kill costs roughly one to two minutes. The startup probe alone waits 60 seconds before its first check, so Rails needs that long before the pod can become ready again.

A rollout costs another one and a half to three minutes. The chart ships strategy.type: Recreate, which terminates the old pod before starting the new one. With replicas: 1 that is a hard outage on every image or chart bump, and Renovate merges those regularly.

The chart justifies Recreate with a writable volume. That justification does not apply on kup6s: persistence is disabled, attachments live in S3, and the web pod mounts only emptyDir volumes.

Diagnosis

Confirm the OOM kill and its timestamp:

kubectl --context kup6s -n kup-openproject get pods -l openproject/process=web
kubectl --context kup6s -n kup-openproject describe pod <web-pod> | grep -A6 "Last State"

Check how close the container runs to its limit:

kubectl --context kup6s -n kup-openproject top pod --containers | grep openproject-web
kubectl --context kup6s -n kup-openproject get deploy openproject-web \
  -o jsonpath='{.spec.template.spec.containers[0].resources}{"\n"}'

Distinguish a real heap problem from a page-cache artifact by comparing RSS against cache. A high container_memory_rss with low container_memory_cache means the Ruby heap genuinely needs the memory, so raising the limit is the correct answer rather than a workaround:

container_memory_rss{namespace="kup-openproject",pod=~"openproject-web.*",container="openproject"}
container_memory_cache{namespace="kup-openproject",pod=~"openproject-web.*",container="openproject"}

Plot the working set over several days. A flat line exactly at the limit for many hours before a restart confirms the diagnosis:

container_memory_working_set_bytes{namespace="kup-openproject",pod=~"openproject-web.*",container="openproject"}

Rule out CPU starvation before you change memory. Compare container_cpu_cfs_throttled_periods_total against container_cpu_cfs_periods_total; a ratio near zero means CPU is not involved.

Fix

Size the web container at the chart default and remove the single point of failure. Both settings live in dp-kup/internal/openproject/.

Set the limit in config.yaml:

resources:
  web:
    requests:
      memory: 1Gi
      cpu: 250m
    limits:
      memory: 4Gi
      cpu: "1"

Set the replica count and strategy in charts/constructs/openproject-helm.ts:

replicaCount: 2,
strategy: {
  type: 'RollingUpdate',
},

Regenerate the manifests and deploy:

cd dp-kup/internal/openproject
npm run build
git add config.yaml charts/ manifests/ && git commit && git push origin main

ArgoCD runs with automated sync and self-heal, so the change rolls out on its own.

Important

npm run build regenerates the hocuspocus-secret-auto-generated secret in manifests/openproject.k8s.yaml on every run. That rotation is accepted for normal version bumps, because those restart the web and hocuspocus deployments together and both pick up the new value.

Revert that one line when your change touches only one of the two deployments, as a web-only resource change does. Both pods read the value through a secretKeyRef without a checksum annotation, so a pod keeps the old value until it restarts. Restarting only the web pod leaves it out of step with hocuspocus and silently breaks collaborative editing.

Verify that both replicas run and that the limit applied:

kubectl --context kup6s -n kup-openproject get deploy openproject-web \
  -o jsonpath='replicas={.spec.replicas} available={.status.availableReplicas} strategy={.spec.strategy.type}{"\n"}'
curl -s -o /dev/null -w "%{http_code}\n" https://projects.kup.tirol/login

Prevention

With two replicas and RollingUpdate, the defaults resolve to maxUnavailable: 0 and maxSurge: 1. Both a rollout and a single OOM kill now stay invisible to users, because at least one pod always serves traffic.

Scheduling spreads the replicas across separate nodes, which also covers a single node failure.

No alert covers this class of failure yet. KubePodCrashLooping does not fire at a handful of restarts per week, so a container that dies every other day stays invisible until someone complains. Add alerts on container OOM kills, on restart rate, and on containers exceeding roughly 85 percent of their memory limit.

See also

Review resource limits whenever you override a chart default downward. The OpenProject chart defaults to 4Gi for the web container; the kup6s deployment had halved that without a recorded reason.