Commit Graph
31 Commits
Author SHA1 Message Date
MythEclipseandClaude b508a39ff7 fix(ai-moderation): prevent false positive SARA flag for IMPHNEN project name
- Add explicit system rule that IMPHNEN is the project's own name, not religion
- Rename 'Imphnemia 11:17' example to 'Kitabonia 11:17' to avoid name collision
- Ensures mentioning/promoting the project URL is not flagged as SARA

Co-Authored-By: Claude <noreply@anthropic.com>
2026-06-12 23:28:09 +07:00
MythEclipseandClaude fbc2184c6e feat(ai-moderation): add user profile self-learning system
Add user_profiles table, store, and background learner worker
that summarizes user communication style, topics, and personality.

- New user_profiles table (user_id PK, guild_id, profile_summary, last_analyzed_at)
- userProfileStore.ts — CRUD (get/update) following channelCultureStore pattern
- userProfileLearner.ts — background worker: queries 100 recent msgs per user,
  calls LLM for personality summary, updates every 12h
- Inject <user_profile> XML tag per-message in moderation prompt
- Start worker alongside cultureLearner in aiAnalyzer.ts
- Migration 0008 for user_profiles table

Co-Authored-By: Claude <noreply@anthropic.com>
2026-06-12 20:11:34 +07:00
MythEclipse 8b114ca278 refactor(ai-moderation): expand SARA and religious blasphemy detection rules
Update the moderation prompt to include high-priority detection categories for:
- Fake scripture/verse parodies
- Claims of divinity or false religious movements
- Misuse of theological terms as internet slang/memes
- Mockery of religious figures and rituals

This change ensures stricter enforcement of SARA (Suku, Agama, Ras, Antargolongan) policies by explicitly defining religious blasphemy and parody as high-severity violations.
2026-06-11 02:40:11 +07:00
MythEclipse a3ef3bf7f1 fix(ai-moderation): add rule to prevent false positive on QWERTY typos like 'ngodonf' 2026-06-06 17:31:23 +07:00
MythEclipse 71a8a9e6e1 fix(ai-moderation): patch structural bypass vulnerabilities in NLP pipeline
- Implement Pre-computation Normalization for Polyglot Obfuscation.

- Inject Ontological Graph for literal translation evasion (e.g., 'kostum hewan').

- Enforce Entropy-Triggered Routing to deny softmax fallback exploitation.

- Format discord-gateway codebase.
2026-06-06 15:38:36 +07:00
MythEclipseandClaude Opus 4.8 da885339f9 fix(ai-moderation): remove user history from prompt to eliminate confirmation bias loop
user_history (riwayat flag sebelumnya) dan clean_streak/total_infractions
dikirim ke LLM setiap kali menganalisis pesan — ini bikin self-fulfilling
prophecy: user yg pernah kena false positive jadi makin gampang dituduh
lagi, dan link Instagram pun dianggap sexual_deviation cuma karena
riwayat user.

Changes:
- Hapus getUserRecentInfractions dari text batch path
- Hapus getUserRecentInfractions dari media analysis path
- Hapus import getUserRecentInfractions yg gak dipakai
- Ubah instruksi prompt dari 'jadilah lebih tegas jika riwayat jelek'
  jadi 'setiap pesan dinilai berdasarkan isinya sendiri'

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 13:55:24 +07:00
MythEclipseandClaude Opus 4.8 edb55fb5bf fix(ai-moderation): add religious expression safe list to stop Astaghfirullah false positives
Common Indonesian religious expressions like 'Astaghfirullah', 'Astaga',
'Alhamdulillah', 'Subhanallah', dll were being flagged as vulgar_language
by the LLM. Added explicit rule that these are normal religious/cultural
expressions in Indonesia - not vulgar language - even in all-caps or
with repeated letters.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 13:52:31 +07:00
MythEclipseandClaude Opus 4.8 cfe7230a55 fix(ai-moderation): prevent LLM from hallucinating furry slang on common names and unknown words
- Adds explicit rule that Indonesian names/nicknames like 'Sapik' (Syafik),
  'Ayang', 'Dek', 'Bang', 'Mas', etc. are NOT furry or sexual_deviation references
- Adds rule prohibiting the LLM from inventing slang meanings for words
  it doesn't recognize - default to innocent until proven guilty
- Prevents false positive cascade where LLM confuses names with furry slang

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 13:22:35 +07:00
MythEclipseandClaude Opus 4.8 817f1ce3df fix(ai-moderation): add Furina/Genshin character name exception to prevent false positive furry flags
Nama karakter game/anime populer seperti 'Furina' dari Genshin Impact
sering kena false positive sebagai 'sexual_deviation' karena kemiripan
fonetik dengan kata 'furry'. Menambahkan aturan eksplisit bahwa nama
karakter fiksi normal bukan referensi furry fetish, dgn pengecualian
jika konteks pesan secara eksplisit membahas aspek fetish/seksual.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 12:04:10 +07:00
MythEclipse 25c86dcc30 fix(ai-moderation): raise retry backoffs and cooldowns, make 429 retryable, reduce prompt false positives 2026-06-05 21:47:02 +07:00
MythEclipse d7a35e8377 refactor(ai-moderation): expand prompt rules for lyrics and literature
Update moderation guidelines to prevent false positives on song lyrics, poems, memes, and literary quotes, ensuring political or revolutionary content is not flagged as conflict instigation unless accompanied by explicit incitement.
2026-06-05 20:05:32 +07:00
MythEclipse 49ada183b2 feat(ai-moderation): implement auto-fallback for streaming and refine prompt rules
- Add automatic fallback to streaming mode if the provider rejects non-streaming requests with a 400 error
- Refactor `llmChat` to use an internal execution function to support retry logic with modified parameters
- Update moderation prompt to explicitly allow Japanese pop culture terms (e.g., "moe", "waifu", "wibu") to prevent false positive sexual deviation flags
2026-06-05 19:44:06 +07:00
MythEclipse 08c624fcf4 fix(ai-moderation): use generic sender name and force descriptive media analysis to prevent bad global cache poisoning 2026-06-05 18:20:43 +07:00
MythEclipse 2f3d7e1d61 feat(ai-moderation): introduce user reputation and channel culture context
Implements a context-aware moderation system by tracking user behavior
and channel-specific norms to improve AI decision-making accuracy.

- Adds `user_reputations` table to track trust scores, clean streaks,
  and infraction history.
- Adds `channel_cultures` table to store AI-generated summaries of
  channel-specific norms and slang.
- Implements `userReputationStore` to autonomously update user scores
  based on moderation outcomes (clean vs. flagged).
- Implements `cultureLearner` and `channelCultureStore` to manage
  evolving channel contexts.
- Enhances LLM prompts to inject user reputation (trust scores,
  history) and channel culture summaries, enabling "wisdom-based"
  moderation (e.g., giving benefit of the doubt to high-trust users).
- Integrates reputation and culture updates into the existing
  `aiAnalyzer` pipeline.
2026-06-05 18:04:57 +07:00
MythEclipse f057bf1f0b refactor(ai-moderation): implement distributed locking and content-based caching
Refactors the AI moderation pipeline to improve concurrency control and
cache efficiency by moving from user-centric to content-centric caching.

- Implements a distributed locking mechanism for media analysis using
  `acquireMediaAnalysisLock` to prevent redundant LLM vision calls across
  multiple pods.
- Transitions text moderation caching from `user_mod:userId:hash` to a
  purely content-based `text_mod:hash` approach to increase hit rates.
- Enhances `getPendingMessagesByConversation` with atomic transactions
  and `FOR UPDATE SKIP LOCKED` to safely transition messages from
  `pending` to `processing` state.
- Adds `processing` status to the `AIStatus` type and database schema to
  track active analysis lifecycles.
- Implements polling logic in `llmModerationClient.ts` to wait for
  in-progress media analyses.
2026-06-05 16:56:46 +07:00
MythEclipse 399919ded0 feat(ai-moderation): add anti-evasion rule for emoji spelling
- Instructs the AI to decode combinations of regional indicator emojis (e.g., 🇬 🇦 🇾) and custom letters spelling out words, rather than dismissing them as 'just a series of emojis'
- Added a specific few-shot example (Contoh 15) to demonstrate flagging this technique when used to spell banned words
2026-06-05 16:23:59 +07:00
MythEclipse 6b3f2cdecd feat(ai-moderation): tighten rules for anatomical vulgarity and BL mentions
- Added explicit zero-tolerance rule for anatomical/sexual vulgarity (e.g. titten, kontol), explicitly forbidding the AI from passing them off as 'casual conversation' or 'jokes'
- Expanded sexual_deviation rule to explicitly cover brief mentions of BL (Boys Love), yaoi, yuri, and LGBT topics, instructing the AI to flag them regardless of casual context
2026-06-05 16:16:21 +07:00
MythEclipse a0bf7dfdd9 chore(ai-moderation): harden LLM prompt and lexical scanner against evasion techniques and cross-lingual vulgarities 2026-06-05 15:28:04 +07:00
MythEclipse 9a02ac8d17 feat(ai-moderation): flag excessive religious jokes and satire as sara 2026-06-04 19:31:08 +07:00
MythEclipse b0278de51a feat(ai-moderation): add URL analysis rules to moderation prompt to prevent domain-based false positives 2026-06-04 16:35:14 +07:00
MythEclipse 1c46a8a084 feat(ai-moderation): add conflict instigation, offensive username, and discrimination moderation flags and prompt rules 2026-06-04 12:49:24 +07:00
MythEclipse c1de1279a5 There are no staged changes to commit. The staging area is empty — git status shows a clean working tree with nothing staged. 2026-06-02 23:07:11 +07:00
MythEclipse 1f3d1ac3f4 API Error: API returned an empty or malformed response (HTTP 200) — check for a proxy or gateway intercepting the request 2026-06-02 23:06:55 +07:00
MythEclipse 774472b6ba perf(discord-gateway): add tiktoken token counting, vision LRU cache, PromptMode few-shot splits, and Piscina maxThreads config 2026-06-02 22:50:53 +07:00
MythEclipse f163ead3cf feat(discord-gateway): add local badword pre-filter, accurate token estimation, Piscina maxThreads config, and PromptMode support 2026-06-02 22:50:23 +07:00
MythEclipse 6d2da7ce1d perf(discord-gateway): start AI moderation analysis immediately, running attachment uploads in parallel instead of blocking 2026-06-02 22:49:57 +07:00
MythEclipseandClaude Opus 4.8 62d7d19b53 fix: analysis output wajib deskriptif — bukan generic placeholder
Sebelumnya: 'Pesan hanya berisi attachment tanpa teks yang melanggar'
Sekarang: harus deskriptif berdasarkan tipe konten:
- Text only: '[user] membahas tentang <topik>. <konteks>.'
- Image only: 'Gambar berupa <jenis>. Terlihat <isi>.'
- Text+Image: '[user] mengirim <gambar> sambil membahas <topik>.'

Tambahkan contoh baik vs buruk di OUTPUT_INSTRUCTIONS sebagai format wajib.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:24:21 +07:00
MythEclipseandClaude Opus 4.8 a69de85564 fix: prompt dua mode — gambar+teks vs gambar saja
Sebelumnya aturan 'percaya teks terlebih dahulu' membuat model abaikan
deskripsi gambar saat teks kosong. Semua image-only message di-clean.

Fix:
- SYSTEM_RULES: pisah Mode 1 (teks+gambar) dan Mode 2 (hanya gambar)
- Mode 2: deskripsi gambar jadi bukti utama, WAJIB dibaca
- Gambar terminal/chat/editor kode/casual → clean
- Gambar dengan elemen judi NYATA (chip, roulette, odds) → flag
- 2 contoh few-shot baru: terminal clean, situs judi flag

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:15:53 +07:00
MythEclipseandClaude Opus 4.8 55ee8ade16 fix: vision model hanya deskripsi, tidak memutuskan moderasi
Root cause: vision model diminta untuk 'flag' dan 'menilai' gambar,
sehingga screenshot terminal/chat biasa diklaim sebagai 'situs perjudian'.

Fix:
- vision prompt: HANYA deskripsi objektif (objek, teks, layout, jenis gambar)
- larang tegas kata 'gambling', 'judi', 'pelanggaran', 'harus dihapus'
- tambah buildGeneralImageVisionPrompt di discord-gateway stickerPrompt
- MEDIA_INSTRUCTIONS: tegaskan batch LLM adalah hakim, vision hanya saksi mata
- deskripsi netral (terminal, chat, editor kode) tidak boleh jadi dasar flag

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:08:41 +07:00
MythEclipseandClaude Opus 4.8 f1ddca5eee fix: false positive gambling detection + route collision + WS events (#1)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 10:44:27 +07:00
MythEclipseandClaude Opus 4.8 c48a0c5e3b refactor: split monolith into 3 microservices (frontend, backend, discord-gateway)
- Extract services into services/{frontend,backend,discord-gateway}
- Create packages/shared/ for shared logger, errors, utils, types
- Setup Modular MVC pattern in backend (controller→service→repository)
- Setup event-driven architecture in discord-gateway with Redis pub/sub
- Move Docker files to infra/docker/ with per-service Dockerfiles
- Update docker-compose.yml to use Traefik-only routing (no port exposes)
- Update GitHub Actions deploy workflow for multi-service matrix build
- Fix all import paths and resolve type errors across all services
- All 3 services pass tsc --noEmit clean

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 21:44:29 +07:00