Symptom
A deploy script checks the edge on the loopback port and gets nothing back, while the public site answers 200:
$ curl -sS -o /dev/null -w 'code=%{http_code}\n' http://127.0.0.1:3000/healthz
curl: (52) Empty reply from server
code=000
curl exits with 52 and prints 000 as the status. nginx's access log shows status 444 and 0 bytes for the request:
"GET /healthz HTTP/1.1" 444 0 "-" "curl/8.21.0" "-"
The script decides the new release is unhealthy (or "can't read the version") and rolls back a perfectly good deploy.
When it happens
The production edge serves only its own host names, and catches everything else in a default server that closes the connection:
server {
listen 80;
server_name example.com;
include /etc/nginx/snippets/app.conf;
}
# Anything else (a hostname not configured above) gets nothing.
server {
listen 80 default_server;
return 444;
}
This setup is common and sensible, because it keeps scanners and spoofed Host headers away from the app. Behind a tunnel or a CDN (for example cloudflared forwarding plain HTTP with the public Host kept), real traffic always carries the public name. A check that calls http://127.0.0.1:PORT/... sends Host: 127.0.0.1:PORT instead, so nginx picks the default server.
The dev stack often has only one server block with no server_name, so the same check passes there and fails only in production.
Cause
nginx chooses the server block by the request's Host header. With no match, it uses the default_server for that listen address. return 444 is nginx's non-standard code meaning "close the connection without sending a response", so the client sees an empty reply rather than a 4xx. The health endpoint behind it is never reached, so you can't tell "app down" from "wrong host name".
Fix
Option 1 (preferred): ask for the public name, the way real traffic does. The edge stays as strict as it is:
EDGE=http://127.0.0.1:${LB_PORT:-3000}
EDGE_HOST=$(printf '%s' "$PUBLIC_URL" | sed -E 's#^[a-z]+://##; s#[/:].*##') # https://example.com -> example.com
edge() { curl -s --max-time 10 ${EDGE_HOST:+-H "Host: $EDGE_HOST"} "$EDGE$1"; }
edge /healthz # {"ok":true}
edge /openapi.json # read the running version from here
Every check that goes through the edge must use the helper: health, the version probe, smoke tests. In the reference incident, a separate monitoring script still graded the public URL by the loopback port and mailed a CRITICAL on a healthy deploy. That script needed the public URL passed in too.
Option 2: give health its own answer in the default server, and keep 444 for everything else:
server {
listen 80 default_server;
location = /healthz {
allow 127.0.0.1; allow 172.16.0.0/12; deny all;
return 200 '{"ok":true}';
}
location / { return 444; }
}
This only proves that nginx is up. To check the app as well, proxy_pass that location to the app's upstream instead of return 200. Note that the allowed source addresses depend on how the check reaches the container (published port, Docker network).
Verify
curl -sS -o /dev/null -w '%{http_code}\n' http://127.0.0.1:3000/healthz # 000, exit 52
curl -sS -o /dev/null -w '%{http_code}\n' -H 'Host: example.com' http://127.0.0.1:3000/healthz # 200
Reproduced with nginx 1.30.5 (nginx:1.30-alpine) and curl 8.21.0. With Option 2 in place, /healthz without a Host answered 200 and / still came back empty (000).
Before you change a check, read your production config's catch-all server block (grep -n -A3 default_server). Whether it returns 444, a 404 page or redirects decides what a bare-IP check will see.
Notes
- If TLS terminates at nginx, the default server is often
listen 443 ssl default_server; ssl_reject_handshake on;. A check by IP over https then fails in the TLS handshake instead, so use --resolve example.com:443:127.0.0.1, which sets SNI and Host together. - A check that can't tell "no answer" from "unhealthy" is dangerous in an automatic-rollback deploy. Log the curl exit code and the status, not only "health failed".
- Keep the 444 catch-all. The fix belongs in the checker, not in loosening the edge.
The full body — free, open to anyone, no key.
Source: Diagnosed and fixed in WITAN's own server-side deploy script on 2026-10-01, during the first deploy of v0.19.0 with it. The health and version checks asked the loopback port with no Host header. The production edge (nginx behind a Cloudflare Tunnel, with a default_server that returns 444 for unknown names) closed those requests, so a healthy v0.19.0 failed both checks. The fix (send the public Host, and pass the public URL to the ops check) was merged on 2026-10-01 and shipped in v0.19.1, and the checks were confirmed by hand on the production host (200, the right version). Both behaviours were reproduced on 2026-10-01 with nginx 1.30.5 (nginx:1.30-alpine) and curl 8.21.0. The TLS-handshake note follows the repo's non-tunnel template (ssl_reject_handshake on) and was not reproduced.