fix(gateway): align stuck-recovery threshold with batch timeout (300s→120s)

recovery-worker used a hardcoded STUCK_PROCESSING_AGE_MS=300_000 while
messagesCleanup's default and AI_ANALYSIS_PROCESSING_TIMEOUT_MS are both
120s. Rows stuck between 2 and 5 minutes were never reverted by the
recovery worker — they looked permanently stuck (and accumulated under
load) even though the batch budget had long passed. Now derives the
threshold from config so the two knobs can never drift again.
This commit is contained in:
asepharyana
2026-09-24 17:52:17 +07:00
parent deed7bdfb0
commit 045cdf1f75
@@ -26,8 +26,16 @@ import {
const logger = createChildLogger("ai-recovery"); const logger = createChildLogger("ai-recovery");
/** Revert messages stuck in `processing` for longer than this. */ /**
const STUCK_PROCESSING_AGE_MS = 300_000; * Revert messages stuck in `processing` for longer than this.
* Kept in lockstep with the batch processing timeout
* (AI_ANALYSIS_PROCESSING_TIMEOUT_MS, default 120s): a row sitting past the
* batch budget is a leak, not a legitimate slow batch. messagesCleanup's
* default was lowered 300s→120s in 2026-08-24; this constant was missed and
* stayed at 300s — messages looked stuck for up to 5 minutes before recovery
* touched them.
*/
const STUCK_PROCESSING_AGE_MS = config.AI_ANALYSIS_PROCESSING_TIMEOUT_MS;
/** /**
* Starts the periodic recovery worker. * Starts the periodic recovery worker.