Root cause: 9router combo models (deepseek-v4-flash-free on fallback) have
TTFT up to 30-40s. Caddy 9router route inherited the default
response_header_timeout 30s / read_timeout 60s → 504 'timeout awaiting
response headers' even though 9router was still processing. Cloudflare/log
showed repeated 504s; health watchdog (correctly) flagged the outage.
Fixes:
1. Caddy: dedicated 9router route with response_header_timeout 120s +
read/write 300s (was default 30/60). Removed invalid top-level
flush_interval on upload block that broke caddy reload (2.11 rejects it as
transport subdirective).
2. health-check: HTTP timeout 60→150s (mirror Caddy), and alert ONLY when
EVERY model fails — any working model means the server's fallback chain
succeeds. Early-exit on first success to bound runtime (~3s healthy).
Verified: 3 runs green, ~3.6s each, silent exit 0.
Cron runs as user code (not root). /etc/bws-token is root:bws 640, so direct
read fails with Permission denied → watchdog exited 1 every run. Fix:
- wrapper uses sudo -n cat (code is in sudo group, NOPASSWD)
- health-check.py get_key() falls back to sudo -n cat too
Verified as code user: silent exit 0 when healthy.