Traefik returns 404 for half of all requests (stale middleware cache)¶
Symptom¶
Users report that a site is “sometimes” broken and works again on reload.
Roughly half of all HTTPS requests return
404, the rest succeed. The ratio matches the Traefik replica count.The response body is
404 page not foundwithcontent-type: text/plain, which is Traefik’s own 404 rather than the application’s styled error page.The backend application logs almost no 404s, because the requests never reach it.
Many unrelated hosts are affected at once, across namespaces and teams.
Nothing alerts, and every pod reports
Ready.
Confirm the pattern with a request series:
for i in $(seq 1 40); do curl -s -o /dev/null -w "%{http_code} " https://<host>/; done
Root cause¶
One Traefik replica cannot resolve a middleware that the Ingress references, so it disables every router that uses it.
Traefik treats an unresolvable middleware as a fatal router error.
It does not skip the middleware and it does not fail open: it sets the router to disabled.
A disabled router matches nothing, so the request falls through to the default handler, which answers 404.
The middleware itself is present and healthy.
Only one replica’s kubernetescrd provider cache is missing it, which is why exactly one replica serves 404 while the others are fine.
The cache goes stale when a replica starts in a window where it cannot list the Middleware resources, typically during a node reboot when many pods restart at once. The replica builds its configuration without the middleware and caches that result.
The state never repairs itself. Traefik corrects its cache from watch events, and a Middleware object that nobody edits emits no events. On kup6s the CrowdSec bouncer middleware had been unchanged for 132 days, so the affected replica stayed broken until it was restarted.
Important
A single unresolvable middleware takes down every route that references it. The CrowdSec bouncer is attached to most public routers, so one stale replica degraded customer sites, the AAF instances, and the entire ops surface at the same time.
Diagnosis¶
Ask each Traefik pod what it thinks the routing table looks like. The API listens on port 8080 and answers per pod, which is what makes the replicas comparable:
kubectl --context <cluster> -n traefik get pods -o wide
Query one pod at a time from inside the cluster, substituting each pod IP:
kubectl --context <cluster> -n traefik run curlcheck --rm -i --restart=Never \
--image=curlimages/curl:latest --quiet -- \
curl -s "http://<pod-ip>:8080/api/http/routers?search=<router-name-fragment>&per_page=50"
A healthy replica reports "status":"enabled".
A stale replica reports the cause directly:
"middlewares": ["crowdsec-crowdsec-bouncer@kubernetescrd"],
"error": ["invalid middleware \"crowdsec-crowdsec-bouncer@kubernetescrd\" configuration: invalid middleware type or middleware does not exist"],
"status": "disabled"
Confirm that the middleware genuinely exists before blaming the configuration:
kubectl --context <cluster> get middleware -A
Count the damage across all routers on the suspect pod to see the real scope:
for pg in 1 2 3 4 5; do
curl -s "http://<pod-ip>:8080/api/http/routers?per_page=200&page=$pg"
done | grep -c 'invalid middleware'
Note
The routers API paginates and silently truncates. Page through it rather than trusting a single request, or you will underestimate how many hosts are affected.
Compare replicas directly when you want proof before acting.
Resolve the hostname to one pod at a time against the websecure entrypoint on port 8443:
curl -sk -o /dev/null -w "%{http_code}\n" \
--resolve <host>:8443:<pod-ip> https://<host>:8443/<path>
Testing port 8000 is misleading.
The HTTP to HTTPS redirect happens at the entrypoint, so a broken replica still answers 301 there.
Fix¶
Delete the replica with the stale cache. The replacement rebuilds its provider cache from a fresh list:
kubectl --context <cluster> -n traefik delete pod <stale-traefik-pod>
The remaining replicas keep serving traffic, so this causes no downtime.
Verify that the 404s are gone and that no router remains disabled:
for i in $(seq 1 40); do curl -s -o /dev/null -w "%{http_code} " https://<host>/; done
Prevention¶
Nothing in the stock monitoring catches this.
Every pod stays Ready and the deployment reports full availability, so no Kubernetes-level alert has anything to fire on.
At the time of the incident Traefik metrics were not scraped either, which left the ingress tier entirely unobserved.
The incident of 2026-09-14 ran for eight days and surfaced only because a user mentioned intermittent 404s.
Monitoring added on 2026-09-14¶
Traefik had exported Prometheus metrics on its metrics entrypoint all along, but nothing scraped them.
Enabling traefik in the kup6s monitoring config activates a PodMonitor, which is what the metrics below depend on.
A ServiceMonitor does not work here: the Traefik Service does not publish the metrics port, only the Pod does.
The alert TraefikReplicaRoutingDivergence watches for exactly this class of failure.
It compares the share of 404 responses each replica serves on the websecure entrypoint and fires when one replica exceeds the best one by more than 25 percentage points for 15 minutes.
Comparing replicas against each other, rather than against a fixed error rate, is what makes the rule trustworthy. Scanner traffic produces a steady drizzle of 404s on every replica alike, and that shared baseline cancels out in the difference. A middleware that someone deleted on purpose degrades all replicas equally, so the rule correctly stays silent.
Note
Router state itself cannot be measured.
Traefik exposes no router metric by default, and a disabled router does not exist at runtime, so it emits nothing even with addRoutersLabels enabled.
The per-pod entrypoint counters are the usable signal.
The rule carries a minimum-traffic guard of 0.1 requests per second per pod.
Ingress on kup6s sees only about 0.3 req/s per replica, and Prometheus keeps the series of deleted pods for several minutes after a rollout.
Without the guard, a dead series carrying two requests would register as a divergence of 1.0.
Still open¶
The threshold is derived from the measured baseline, not from incident data: the broken pod’s counters were deleted with it, and scraping only began afterwards. Watch it over the first weeks.
Nothing catches a fault that hits every replica at once. That needs synthetic probing from outside, which is tracked with the wider alerting work.
Whether an unresolvable middleware should disable a router at all remains a Traefik design question. Failing closed on a security middleware is defensible, and it is not configurable in any case, but it does convert one stale cache into a platform-wide outage.
See also
A middleware reference that is merely misspelled produces the same disabled router and the same 404, but on every replica at once rather than on one.
Always reference middlewares by their fully qualified name, including the provider suffix.