Commit Graph
23 Commits
Author SHA1 Message Date
MythEclipseandClaude Opus 4.8 817f1ce3df fix(ai-moderation): add Furina/Genshin character name exception to prevent false positive furry flags
Nama karakter game/anime populer seperti 'Furina' dari Genshin Impact
sering kena false positive sebagai 'sexual_deviation' karena kemiripan
fonetik dengan kata 'furry'. Menambahkan aturan eksplisit bahwa nama
karakter fiksi normal bukan referensi furry fetish, dgn pengecualian
jika konteks pesan secara eksplisit membahas aspek fetish/seksual.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 12:04:10 +07:00
MythEclipse 25c86dcc30 fix(ai-moderation): raise retry backoffs and cooldowns, make 429 retryable, reduce prompt false positives 2026-06-05 21:47:02 +07:00
MythEclipse d7a35e8377 refactor(ai-moderation): expand prompt rules for lyrics and literature
Update moderation guidelines to prevent false positives on song lyrics, poems, memes, and literary quotes, ensuring political or revolutionary content is not flagged as conflict instigation unless accompanied by explicit incitement.
2026-06-05 20:05:32 +07:00
MythEclipse 49ada183b2 feat(ai-moderation): implement auto-fallback for streaming and refine prompt rules
- Add automatic fallback to streaming mode if the provider rejects non-streaming requests with a 400 error
- Refactor `llmChat` to use an internal execution function to support retry logic with modified parameters
- Update moderation prompt to explicitly allow Japanese pop culture terms (e.g., "moe", "waifu", "wibu") to prevent false positive sexual deviation flags
2026-06-05 19:44:06 +07:00
MythEclipse 08c624fcf4 fix(ai-moderation): use generic sender name and force descriptive media analysis to prevent bad global cache poisoning 2026-06-05 18:20:43 +07:00
MythEclipse 2f3d7e1d61 feat(ai-moderation): introduce user reputation and channel culture context
Implements a context-aware moderation system by tracking user behavior
and channel-specific norms to improve AI decision-making accuracy.

- Adds `user_reputations` table to track trust scores, clean streaks,
  and infraction history.
- Adds `channel_cultures` table to store AI-generated summaries of
  channel-specific norms and slang.
- Implements `userReputationStore` to autonomously update user scores
  based on moderation outcomes (clean vs. flagged).
- Implements `cultureLearner` and `channelCultureStore` to manage
  evolving channel contexts.
- Enhances LLM prompts to inject user reputation (trust scores,
  history) and channel culture summaries, enabling "wisdom-based"
  moderation (e.g., giving benefit of the doubt to high-trust users).
- Integrates reputation and culture updates into the existing
  `aiAnalyzer` pipeline.
2026-06-05 18:04:57 +07:00
MythEclipse f057bf1f0b refactor(ai-moderation): implement distributed locking and content-based caching
Refactors the AI moderation pipeline to improve concurrency control and
cache efficiency by moving from user-centric to content-centric caching.

- Implements a distributed locking mechanism for media analysis using
  `acquireMediaAnalysisLock` to prevent redundant LLM vision calls across
  multiple pods.
- Transitions text moderation caching from `user_mod:userId:hash` to a
  purely content-based `text_mod:hash` approach to increase hit rates.
- Enhances `getPendingMessagesByConversation` with atomic transactions
  and `FOR UPDATE SKIP LOCKED` to safely transition messages from
  `pending` to `processing` state.
- Adds `processing` status to the `AIStatus` type and database schema to
  track active analysis lifecycles.
- Implements polling logic in `llmModerationClient.ts` to wait for
  in-progress media analyses.
2026-06-05 16:56:46 +07:00
MythEclipse 399919ded0 feat(ai-moderation): add anti-evasion rule for emoji spelling
- Instructs the AI to decode combinations of regional indicator emojis (e.g., 🇬 🇦 🇾) and custom letters spelling out words, rather than dismissing them as 'just a series of emojis'
- Added a specific few-shot example (Contoh 15) to demonstrate flagging this technique when used to spell banned words
2026-06-05 16:23:59 +07:00
MythEclipse 6b3f2cdecd feat(ai-moderation): tighten rules for anatomical vulgarity and BL mentions
- Added explicit zero-tolerance rule for anatomical/sexual vulgarity (e.g. titten, kontol), explicitly forbidding the AI from passing them off as 'casual conversation' or 'jokes'
- Expanded sexual_deviation rule to explicitly cover brief mentions of BL (Boys Love), yaoi, yuri, and LGBT topics, instructing the AI to flag them regardless of casual context
2026-06-05 16:16:21 +07:00
MythEclipse a0bf7dfdd9 chore(ai-moderation): harden LLM prompt and lexical scanner against evasion techniques and cross-lingual vulgarities 2026-06-05 15:28:04 +07:00
MythEclipse 9a02ac8d17 feat(ai-moderation): flag excessive religious jokes and satire as sara 2026-06-04 19:31:08 +07:00
MythEclipse b0278de51a feat(ai-moderation): add URL analysis rules to moderation prompt to prevent domain-based false positives 2026-06-04 16:35:14 +07:00
MythEclipse 1c46a8a084 feat(ai-moderation): add conflict instigation, offensive username, and discrimination moderation flags and prompt rules 2026-06-04 12:49:24 +07:00
MythEclipse c1de1279a5 There are no staged changes to commit. The staging area is empty — git status shows a clean working tree with nothing staged. 2026-06-02 23:07:11 +07:00
MythEclipse 1f3d1ac3f4 API Error: API returned an empty or malformed response (HTTP 200) — check for a proxy or gateway intercepting the request 2026-06-02 23:06:55 +07:00
MythEclipse 774472b6ba perf(discord-gateway): add tiktoken token counting, vision LRU cache, PromptMode few-shot splits, and Piscina maxThreads config 2026-06-02 22:50:53 +07:00
MythEclipse f163ead3cf feat(discord-gateway): add local badword pre-filter, accurate token estimation, Piscina maxThreads config, and PromptMode support 2026-06-02 22:50:23 +07:00
MythEclipse 6d2da7ce1d perf(discord-gateway): start AI moderation analysis immediately, running attachment uploads in parallel instead of blocking 2026-06-02 22:49:57 +07:00
MythEclipseandClaude Opus 4.8 62d7d19b53 fix: analysis output wajib deskriptif — bukan generic placeholder
Sebelumnya: 'Pesan hanya berisi attachment tanpa teks yang melanggar'
Sekarang: harus deskriptif berdasarkan tipe konten:
- Text only: '[user] membahas tentang <topik>. <konteks>.'
- Image only: 'Gambar berupa <jenis>. Terlihat <isi>.'
- Text+Image: '[user] mengirim <gambar> sambil membahas <topik>.'

Tambahkan contoh baik vs buruk di OUTPUT_INSTRUCTIONS sebagai format wajib.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:24:21 +07:00
MythEclipseandClaude Opus 4.8 a69de85564 fix: prompt dua mode — gambar+teks vs gambar saja
Sebelumnya aturan 'percaya teks terlebih dahulu' membuat model abaikan
deskripsi gambar saat teks kosong. Semua image-only message di-clean.

Fix:
- SYSTEM_RULES: pisah Mode 1 (teks+gambar) dan Mode 2 (hanya gambar)
- Mode 2: deskripsi gambar jadi bukti utama, WAJIB dibaca
- Gambar terminal/chat/editor kode/casual → clean
- Gambar dengan elemen judi NYATA (chip, roulette, odds) → flag
- 2 contoh few-shot baru: terminal clean, situs judi flag

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:15:53 +07:00
MythEclipseandClaude Opus 4.8 55ee8ade16 fix: vision model hanya deskripsi, tidak memutuskan moderasi
Root cause: vision model diminta untuk 'flag' dan 'menilai' gambar,
sehingga screenshot terminal/chat biasa diklaim sebagai 'situs perjudian'.

Fix:
- vision prompt: HANYA deskripsi objektif (objek, teks, layout, jenis gambar)
- larang tegas kata 'gambling', 'judi', 'pelanggaran', 'harus dihapus'
- tambah buildGeneralImageVisionPrompt di discord-gateway stickerPrompt
- MEDIA_INSTRUCTIONS: tegaskan batch LLM adalah hakim, vision hanya saksi mata
- deskripsi netral (terminal, chat, editor kode) tidak boleh jadi dasar flag

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:08:41 +07:00
MythEclipseandClaude Opus 4.8 f1ddca5eee fix: false positive gambling detection + route collision + WS events (#1)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 10:44:27 +07:00
MythEclipseandClaude Opus 4.8 c48a0c5e3b refactor: split monolith into 3 microservices (frontend, backend, discord-gateway)
- Extract services into services/{frontend,backend,discord-gateway}
- Create packages/shared/ for shared logger, errors, utils, types
- Setup Modular MVC pattern in backend (controller→service→repository)
- Setup event-driven architecture in discord-gateway with Redis pub/sub
- Move Docker files to infra/docker/ with per-service Dockerfiles
- Update docker-compose.yml to use Traefik-only routing (no port exposes)
- Update GitHub Actions deploy workflow for multi-service matrix build
- Fix all import paths and resolve type errors across all services
- All 3 services pass tsc --noEmit clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 21:44:29 +07:00