Commit Graph
63 Commits
Author SHA1 Message Date
MythEclipseandClaude fbc2184c6e feat(ai-moderation): add user profile self-learning system
Add user_profiles table, store, and background learner worker
that summarizes user communication style, topics, and personality.

- New user_profiles table (user_id PK, guild_id, profile_summary, last_analyzed_at)
- userProfileStore.ts — CRUD (get/update) following channelCultureStore pattern
- userProfileLearner.ts — background worker: queries 100 recent msgs per user,
  calls LLM for personality summary, updates every 12h
- Inject <user_profile> XML tag per-message in moderation prompt
- Start worker alongside cultureLearner in aiAnalyzer.ts
- Migration 0008 for user_profiles table

Co-Authored-By: Claude <noreply@anthropic.com>
2026-06-12 20:11:34 +07:00
MythEclipseandClaude 6effb51b5d fix(ai-moderation): don't cache text-only results for messages with media
- hasMediaContent now also checks evidence.attachments from metadata
  (not just DB attachment records), catching the race where attachment
  DB rows aren't inserted yet when analysis runs.
- Cache-hit guard: treat cached entries as miss when the message has
  media evidence in metadata, so stale 24h-freezes are avoided.
- Cache-write guard: skip storing text-only analysis results for
  messages whose metadata shows attachments/stickers/embeds. This
  prevents a text-only 'clean' result (from failed vision) being
  frozen for 24h, blocking future re-analysis with full media context.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-06-12 18:45:33 +07:00
MythEclipseandClaude Opus 4.8 07032ab521 refactor: atomic, DRY, and logging improvements across codebase
- Split llmModerationClient.ts (2170 lines) into 5 focused sub-modules
- Split aiAnalyzer.ts (1282 lines) into 4 modular pipelines
- Split messages.db.ts (826 lines) into 5 domain-specific modules
- Moved shared schema to @bete/shared, eliminated backend duplication
- Added createChildLogger to all voice-recording and AI moderation modules
- Extracted tryCommandThenFallback, normalizeMediaState, DEFAULT_VOICE_STATUS
- Created shared pagination.ts utility, eliminated 5+ cursor-pagination duplications
- Created shared messageMapper.ts for row mapping
- Standardized backend error handling with asyncHandler
- Added frontend createLogger utility and useAsyncAction hook
- Added structured logging to frontend hooks, socket, and API client

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 19:46:08 +07:00
MythEclipseandClaude Opus 4.8 3614d32701 fix: resolve architecture disconnects and codebase weaknesses
- Standardize MessageRecord types — single source of truth from @bete/shared
- Clean up config: remove unused GUILD_ID/TEXT_GUILD_ID/TEXT_CHANNEL_ID, fix WEBSERVER_PORT default (3001), remove default admin password
- Move mascot_chat_messages table to Drizzle schema with proper migration
- Remove runtime DDL (CREATE TABLE IF NOT EXISTS) from mascot-chat repository
- Remove phantom analytics/ module from documentation
- Add better-sqlite3 dependency to root devDependencies
- Replace 'as any' casts with proper type assertions across AI moderation
- Add error logging to silent catch blocks in LLM client
- Apply Biome formatting and import organization

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 13:07:01 +07:00
MythEclipseandClaude Opus 4.8 4becf0d6f1 refactor: comprehensive codebase cleanup and architecture hardening
- Sprint 1 (Quick Wins): Remove dead analytics modules, fix 4 unresolved
  imports, replace 3 console.warn with logger, remove mock-crc import
- Sprint 2 (Architecture): Create MascotChatRepository, AnalysisRepository,
  3 Zod schemas (mascot-chat, analysis, voice), deduplicate error classes,
  move 3 SQL queries from routes to repository
- Sprint 3 (Complexity): Replace 7 any types with proper interfaces,
  extract 6 helpers from prepareMediaMessage (CC 85 -> ~15)
- Sprint 4 (Config): Remove 22 dead env vars from .env, add 30 missing
  vars to .env.example, standardize naming

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 10:16:04 +07:00
MythEclipse 483bc86236 refactor(ai-moderation): batch all media messages into single LLM call instead of one per message 2026-06-06 17:34:45 +07:00
MythEclipse 9287cb7319 fix(ai-moderation): guard against stale empty imageUrl rows from base64→URL migration 2026-06-06 17:18:41 +07:00
MythEclipse 71a8a9e6e1 fix(ai-moderation): patch structural bypass vulnerabilities in NLP pipeline
- Implement Pre-computation Normalization for Polyglot Obfuscation.

- Inject Ontological Graph for literal translation evasion (e.g., 'kostum hewan').

- Enforce Entropy-Triggered Routing to deny softmax fallback exploitation.

- Format discord-gateway codebase.
2026-06-06 15:38:36 +07:00
MythEclipse 5701b5f15f feat(ai-moderation): use cached sticker image URLs directly instead of re-uploading base64 data 2026-06-06 14:13:18 +07:00
MythEclipseandClaude Opus 4.8 da885339f9 fix(ai-moderation): remove user history from prompt to eliminate confirmation bias loop
user_history (riwayat flag sebelumnya) dan clean_streak/total_infractions
dikirim ke LLM setiap kali menganalisis pesan — ini bikin self-fulfilling
prophecy: user yg pernah kena false positive jadi makin gampang dituduh
lagi, dan link Instagram pun dianggap sexual_deviation cuma karena
riwayat user.

Changes:
- Hapus getUserRecentInfractions dari text batch path
- Hapus getUserRecentInfractions dari media analysis path
- Hapus import getUserRecentInfractions yg gak dipakai
- Ubah instruksi prompt dari 'jadilah lebih tegas jika riwayat jelek'
  jadi 'setiap pesan dinilai berdasarkan isinya sendiri'

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 13:55:24 +07:00
MythEclipseandClaude Opus 4.8 e6e44b30c7 fix(ai-moderation): remove indonesianTextNormalizer to stop hallucinated slang flags
The rule-based badword detector was injecting [normalized_text] and
[normalization_notes] tags into the LLM prompt that caused false
positive hallucinations - the LLM started associating innocent words
('sapik', 'furina') with furry/sexual_deviation due to misleading
context injected by the normalizer.

Removed:
- indonesianTextNormalizer.ts (full file deletion)
- formatModerationTextEvidenceForPrompt import/usage in llmModerationClient
- formatModerationTextEvidenceForPrompt import/usage in conversationContext
- stale re-exports in index.ts

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 13:27:24 +07:00
MythEclipseandClaude Opus 4.8 817f1ce3df fix(ai-moderation): add Furina/Genshin character name exception to prevent false positive furry flags
Nama karakter game/anime populer seperti 'Furina' dari Genshin Impact
sering kena false positive sebagai 'sexual_deviation' karena kemiripan
fonetik dengan kata 'furry'. Menambahkan aturan eksplisit bahwa nama
karakter fiksi normal bukan referensi furry fetish, dgn pengecualian
jika konteks pesan secara eksplisit membahas aspek fetish/seksual.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 12:04:10 +07:00
MythEclipse 25c86dcc30 fix(ai-moderation): raise retry backoffs and cooldowns, make 429 retryable, reduce prompt false positives 2026-06-05 21:47:02 +07:00
MythEclipse 2f3d7e1d61 feat(ai-moderation): introduce user reputation and channel culture context
Implements a context-aware moderation system by tracking user behavior
and channel-specific norms to improve AI decision-making accuracy.

- Adds `user_reputations` table to track trust scores, clean streaks,
  and infraction history.
- Adds `channel_cultures` table to store AI-generated summaries of
  channel-specific norms and slang.
- Implements `userReputationStore` to autonomously update user scores
  based on moderation outcomes (clean vs. flagged).
- Implements `cultureLearner` and `channelCultureStore` to manage
  evolving channel contexts.
- Enhances LLM prompts to inject user reputation (trust scores,
  history) and channel culture summaries, enabling "wisdom-based"
  moderation (e.g., giving benefit of the doubt to high-trust users).
- Integrates reputation and culture updates into the existing
  `aiAnalyzer` pipeline.
2026-06-05 18:04:57 +07:00
MythEclipse f057bf1f0b refactor(ai-moderation): implement distributed locking and content-based caching
Refactors the AI moderation pipeline to improve concurrency control and
cache efficiency by moving from user-centric to content-centric caching.

- Implements a distributed locking mechanism for media analysis using
  `acquireMediaAnalysisLock` to prevent redundant LLM vision calls across
  multiple pods.
- Transitions text moderation caching from `user_mod:userId:hash` to a
  purely content-based `text_mod:hash` approach to increase hit rates.
- Enhances `getPendingMessagesByConversation` with atomic transactions
  and `FOR UPDATE SKIP LOCKED` to safely transition messages from
  `pending` to `processing` state.
- Adds `processing` status to the `AIStatus` type and database schema to
  track active analysis lifecycles.
- Implements polling logic in `llmModerationClient.ts` to wait for
  in-progress media analyses.
2026-06-05 16:56:46 +07:00
MythEclipse d1d4510ce8 chore(format): auto-format files with biome 2026-06-04 18:52:56 +07:00
MythEclipse 946343ec8c fix(ai-moderation): skip caching error results to prevent false positives from transient failures 2026-06-04 17:42:23 +07:00
MythEclipse ce8a42f5fd feat(ai-moderation): eliminate LLM-based badword detection in favor of pure rule-based matching 2026-06-04 17:33:08 +07:00
MythEclipse b3d42118bd fix(ai-moderation): use cached status instead of hardcoding "clean" for cache hit results 2026-06-04 17:30:37 +07:00
MythEclipse 269a13d96a fix(ai-moderation): restrict vision cache pre-download check to stickers and custom emojis only 2026-06-04 17:20:52 +07:00
MythEclipse aaebe168cf fix(ai-moderation): fallback to discord_url when uploaded_url is null for attachment images 2026-06-04 17:06:57 +07:00
MythEclipse fc52fe5c1d feat(ai-moderation): cache per-user moderation results for text-only messages to skip redundant LLM calls 2026-06-04 16:54:06 +07:00
MythEclipse 65a65fdf39 feat(ai-moderation): fetch web content from URLs in text-only messages for accurate LLM analysis 2026-06-04 16:32:08 +07:00
MythEclipse 259b0babb5 feat(ai-moderation): parse moderation category from LLM fallback response instead of using static defaults 2026-06-04 13:49:35 +07:00
MythEclipse 3fc26592fc feat(ai-moderation): add simple text-only LLM fallback when normal analysis fails for small models 2026-06-04 13:40:25 +07:00
MythEclipse e55ee342b1 fix(ai-moderation): downgrade operational logs to debug and add missing debug logs to silent early-return paths 2026-06-04 13:13:55 +07:00
MythEclipse d568cab88f fix(ai-moderation): add defensive logging to silent failure paths in media analysis 2026-06-04 13:06:35 +07:00
MythEclipse c304d4ca7c fix(ai-moderation): implement exponential backoff for LLM and vision API retries 2026-06-04 13:02:02 +07:00
MythEclipse aad62bc66b fix(ai-moderation): strip fallback text from LLM analysis content 2026-06-03 20:59:10 +07:00
MythEclipse b7d4e66600 refactor: remove json_schema response format, simplify to json_object only 2026-06-03 00:40:42 +07:00
MythEclipse b98aba5748 Tidak ada perubahan staged. Staging area kosong — git diff --cached tidak menghasilkan output. Silakan git add file yang ingin di-commit terlebih dahulu. 2026-06-02 23:09:30 +07:00
MythEclipse 3f3a1184fe Tidak ada perubahan staged untuk dibuatkan commit message. Working tree bersih — git status menunjukkan tidak ada file yang di-stage. Silakan git add file yang ingin di-commit terlebih dahulu, lalu saya bisa bantu tulis commit messagenya. 2026-06-02 23:08:48 +07:00
MythEclipse 75983a390c There are no staged changes to commit — the working tree is clean (nothing to commit, working tree clean). Stage some files first with git add, then I can craft the commit message. 2026-06-02 23:07:21 +07:00
MythEclipse 130f45a859 The git add -A command needs your approval to stage the modified file. Once staged, I'll craft the commit message. 2026-06-02 23:04:16 +07:00
MythEclipse a227c883e3 There are no staged changes to commit. The staging area is empty — stage some files first with git add and I can help craft the message. 2026-06-02 23:00:42 +07:00
MythEclipse aac80e569f perf(discord-gateway): parallelize all media downloads into single Promise.all batch 2026-06-02 22:59:33 +07:00
MythEclipse c18954a263 The staging area is empty — there are no staged changes to summarize. Stage some files first with git add and I can help craft the message. 2026-06-02 22:58:41 +07:00
MythEclipse 9f0d0d06ad If you meant to summarize the most recent (already committed) changes on this branch (41 commits ahead of origin/master), let me know and I can help craft a message for a squash or a summary instead. 2026-06-02 22:57:03 +07:00
MythEclipse 8150c5a20c fix(discord-gateway): update tiktoken import to encoding_for_model and replace includeMediaInstructions boolean with mode enum 2026-06-02 22:52:05 +07:00
MythEclipse acb7e86f07 No staged changes found. The staging area is empty. 2026-06-02 22:51:47 +07:00
MythEclipse f163ead3cf feat(discord-gateway): add local badword pre-filter, accurate token estimation, Piscina maxThreads config, and PromptMode support 2026-06-02 22:50:23 +07:00
MythEclipse 6d2da7ce1d perf(discord-gateway): start AI moderation analysis immediately, running attachment uploads in parallel instead of blocking 2026-06-02 22:49:57 +07:00
MythEclipse bf6ec728c4 chore(discord-gateway): remove stale eslint-disable comment for sticker cache stats 2026-06-02 22:20:46 +07:00
MythEclipse 00851e76da There are **no staged changes** to work with — the working tree is clean and git diff --cached returned nothing. 2026-06-02 22:16:06 +07:00
MythEclipse 22b8e1a698 I can see the git command needs permission to run. Could you approve the git diff --cached command so I can see the staged changes and write the commit message? 2026-06-02 22:13:29 +07:00
MythEclipse 9b1f6b4a8d There are no staged changes to summarize. 2026-06-02 22:13:03 +07:00
MythEclipse 57d28a08af feat(ai-moderation): add per-batch timeout for text-only moderation calls 2026-06-02 22:12:32 +07:00
MythEclipse 4c5d76ea38 No staged changes detected. Would you like me to check unstaged changes instead? 2026-06-02 22:12:25 +07:00
MythEclipse 19028cf244 refactor: migrate discord-gateway from winston to @bete/shared/logger and utils
- Replace all local logger imports (../../shared/logger/logger.js) with @bete/shared/logger across 26 files
- Remove winston dependency, add pino to discord-gateway package.json
- Delete shared/logger/logger.ts (winston-based, 132 lines) and serialization.ts (109 lines)
- Replace local retryWithBackoff imports with @bete/shared/utils across 6 files
- Delete shared/utils/retry.ts (42 lines)
- Add CustomLogger type alias to @bete/shared/logger for backwards compatibility
- Remove logger param from all retryWithBackoff calls and uploadToTele interfaces
- Frontend: convert entity type files to re-exports from shared/api/client.ts
- Full monorepo typecheck clean (4/4 packages)

35 files changed, 42 insertions(+), 488 deletions(-)
2026-06-02 21:23:12 +07:00
MythEclipseandClaude Opus 4.8 363676b608 fix: remove dead OpenAI client instance and fix ChatCompletion type imports
- Delete unused OpenAI WAF bypass client block (60+ lines) from both monolith
  and microservice llmModerationClient.ts
- Replace OpenAI.Chat.Completions.ChatCompletion with type-only ChatCompletion
  import from openai/resources/chat/completions
- All LLM chat calls already go through centralized llmClient.ts helper

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:22:42 +07:00