85204ca6f03cf4093afd5b9ab8d29a8b149f6a87
Even with the prompt firewall (username vs content), the LLM can still occasionally mis-apply a content-level zero-tolerance flag (sara / conflict_instigation) to a message whose ONLY violation is the username (e.g. 'matikanetanyahu'). The auto-delete eligibility check only recognized exact offensive_username flags, so such false positives still deleted the message. Add a belt-and-suspenders guard in isNicknameOnlyViolation: if the flag set is entirely username-attributable (offensive_username/sara/conflict_instigation) AND the analysis text corroborates that the violation is username-only with clean message content, route to nickname-reset instead of message deletion. Adds 6 test cases covering the real matikanetanyahu scenario and the false-positive/negative boundaries.
Description
Bete Discord moderation watcher
29 MiB
Languages
TypeScript
95.8%
CSS
1.4%
Shell
1%
Nix
0.9%
PLpgSQL
0.5%
Other
0.4%