Root cause: 9router combo models (deepseek-v4-flash-free on fallback) have
TTFT up to 30-40s. Caddy 9router route inherited the default
response_header_timeout 30s / read_timeout 60s → 504 'timeout awaiting
response headers' even though 9router was still processing. Cloudflare/log
showed repeated 504s; health watchdog (correctly) flagged the outage.
Fixes:
1. Caddy: dedicated 9router route with response_header_timeout 120s +
read/write 300s (was default 30/60). Removed invalid top-level
flush_interval on upload block that broke caddy reload (2.11 rejects it as
transport subdirective).
2. health-check: HTTP timeout 60→150s (mirror Caddy), and alert ONLY when
EVERY model fails — any working model means the server's fallback chain
succeeds. Early-exit on first success to bound runtime (~3s healthy).
Verified: 3 runs green, ~3.6s each, silent exit 0.
Cron runs as user code (not root). /etc/bws-token is root:bws 640, so direct
read fails with Permission denied → watchdog exited 1 every run. Fix:
- wrapper uses sudo -n cat (code is in sudo group, NOPASSWD)
- health-check.py get_key() falls back to sudo -n cat too
Verified as code user: silent exit 0 when healthy.
Real PR-Agent analytics logs wrap fields under 'record': {...}. The parser
now unwraps that before extracting command/pr_url/message/level, so
/api/analytics and /api/metrics show real data (verified with actual format
from production logs).
openai/auto/best-coding and openai/auto/claude-sonnet return
'No active credentials for provider: auto' on 9router (broken upstream
key). Primary openai/claude-opus-4-8 + fallbacks now all verified
working via litellm against 9router.asepharyana.my.id.