Troubleshoot Penpot¶
This guide covers failures seen on this deployment and how to resolve them.
The frontend pods are OOMKilled¶
Symptom: penpot-frontend pods restart immediately with exit code 137 and log nothing.
kubectl -n penpot get pod -l app.kubernetes.io/name=penpot-frontend \
-o jsonpath='{.items[*].status.containerStatuses[0].lastState.terminated.reason}'
Cause: the frontend image ships worker_processes auto with worker_connections 65535.
On the 20-core EX workers, nginx starts 20 workers and preallocates 65535 connection slots per worker, which costs about 560Mi per replica before it serves a request.
Kubernetes CPU limits do not reduce the worker count, because nginx reads the host CPU affinity rather than the cgroup quota.
Resolution: the deployment mounts its own nginx.conf with 2 workers and 4096 connections, which brings a replica down to about 7Mi.
Verify the mount survived an upgrade:
kubectl -n penpot exec deploy/penpot-frontend -- head -1 /etc/nginx/nginx.conf
The first line must be a comment from dp-infra/penpot/files/nginx.conf, not worker_processes auto;.
ArgoCD cannot create the namespace¶
Symptom: the application fails to sync with one of these messages.
admission webhook "namespaces.mutating.projectcapsule.dev" denied the request:
You do not have any Tenant assigned: please, reach out to the system administrators
admission webhook "namespaces.mutating.projectcapsule.dev" denied the request:
namespace is not owned by any tenant
Cause: Capsule intercepts every service account, because capsuleUserGroups contains system:serviceaccounts.
The ArgoCD application controller has no tenant, so Capsule denies both the creation and the update of any namespace.
Resolution: the namespace webhook exempts the controller by username in dp-infra/capsule/charts/constructs/operator.ts.
Verify the exemption is live:
kubectl get mutatingwebhookconfiguration capsule-dynamic-webhook -o yaml | grep -A2 exclude-argocd-controller
Creating the namespace by hand does not help. ArgoCD still patches it to add its tracking annotation, and Capsule denies that patch too.
The application never reaches Synced¶
Symptom: the three ExternalSecret resources stay OutOfSync and their metadata.generation climbs every few seconds.
kubectl -n penpot get externalsecret penpot-app-secrets-es -o jsonpath='{.metadata.generation}'
Cause: the manifest writes refreshInterval: 1h and the API server stores 1h0m0s.
ArgoCD sees a permanent difference and self-heal re-applies the object in a loop.
Resolution: write the normalized form 1h0m0s in the construct, rebuild, and push.
The Helm release never installs¶
Symptom: the HelmChart job fails with a schema error such as additional properties 'podDisruptionBudget' not allowed.
Cause: chart 1.10.0 ships a values.schema.json with additionalProperties: false for each component, and names the disruption budget key pdb.
Resolution: use pdb, rebuild, and validate the values before pushing:
helm template penpot penpot/penpot --version 1.10.0 -f values.yaml
Assets fail to load after a reload¶
Symptom: images upload without an error but return 404 after a page reload.
Check the backend log for the object storage:
kubectl -n penpot logs deploy/penpot-backend | grep -i s3
List the bucket to see whether the object arrived:
uv run --with boto3 python -c "
import boto3
s3 = boto3.client('s3', endpoint_url='https://hel1.your-objectstorage.com')
print([o['Key'] for o in s3.list_objects_v2(Bucket='assets-penpot-kup6s').get('Contents', [])][:5])
"
If the bucket stays empty, switch the backend to the filesystem backend as a fallback.
Set persistence.assets.enabled to true, storageBackend to fs, and backend.replicaCount to 1, because the volume is ReadWriteOnce.
Verification and invitation mails never arrive¶
Symptom: registration reports that a verification mail was sent, but nothing arrives. The backend log shows a transport failure rather than a rejection:
kubectl -n penpot logs deploy/penpot-backend | grep -A2 "MessagingException"
jakarta.mail.MessagingException: Exception reading response
Cause: Mailjet resets connections from the egress IPs of the dedicated EX workers. The TCP handshake completes and the connection is reset before the SMTP greeting, on every Mailjet port. Confirm it from the node that runs the backend:
kubectl run mjtest --rm --restart=Never --image=curlimages/curl:8.11.0 \
--overrides='{"spec":{"nodeName":"kup6s-ex-hel-2","tolerations":[{"operator":"Exists"}]}}' \
--command -- curl -sS -v --max-time 12 smtp://in-v3.mailjet.com:587
A healthy result is 220 in.mailjet.com ESMTP Mailjet.
Recv failure: Connection reset by peer means the block is still in place.
This is neither a credential problem nor a blocked outbound port.
Authentication never happens, and smtp.office365.com:587 answers normally from the same node.
Mailjet support ticket 4282793 tracks the unblock request for 157.180.107.49 and 157.180.107.250.
Activate an account without the verification mail¶
Registration already creates the profile, its default team, and the password. Only the flag that the verification link would set is missing:
kubectl -n penpot exec penpot-postgres-1 -c postgres -- \
psql -U postgres -d penpot -c \
"update profile set is_active = true where email = 'person@example.com';"
Check the state first:
kubectl -n penpot exec penpot-postgres-1 -c postgres -- \
psql -U postgres -d penpot -c "select email, is_active from profile;"
Exports fail or the exporter restarts¶
Symptom: PNG or PDF export fails, and the exporter pod shows exit code 137.
kubectl -n penpot describe pod -l app.kubernetes.io/name=penpot-exporter | grep -A3 "Last State"
Cause: headless Chromium exceeded the 2Gi memory limit on a large board.
Resolution: raise resources.exporter.limits.memory in config.yaml, rebuild, and push.
Database connections run out¶
Penpot’s connection pool is capped at 20 per backend replica, and the admin console opens its own pool.
Check the totals against max_connections, which is 100:
kubectl -n penpot exec penpot-postgres-1 -c postgres -- \
psql -U postgres -tAc "select count(*), state from pg_stat_activity group by state;"
Raise max_connections in postgres-construct.ts together with the PostgreSQL memory limit if the total approaches the ceiling.
Check the backups¶
kubectl -n penpot get backup.postgresql.cnpg.io
kubectl -n penpot get scheduledbackup.postgresql.cnpg.io
Write out backup.postgresql.cnpg.io in full.
A bare backup resolves to the Longhorn resource instead.
A backup that reports a cache miss means the S3 credentials secret is missing.