perf(automod): compress prompts ~40% + semantic cache via AI_LLM_EMBEDDING_MODEL
Build & Deploy (Nix) / build-and-deploy (backend) (push) Successful in 3m4s
Build & Deploy (Nix) / build-and-deploy (discord-gateway) (push) Successful in 2m29s
Build & Deploy (Nix) / build-and-deploy (proxy) (push) Successful in 2m33s
Build & Deploy (Nix) / build-and-deploy (backend) (push) Successful in 3m4s
Build & Deploy (Nix) / build-and-deploy (discord-gateway) (push) Successful in 2m29s
Build & Deploy (Nix) / build-and-deploy (proxy) (push) Successful in 2m33s
Prompt overhaul (token-frugal, same quality): - rules.ts 28KB -> 10.3KB: every normative rule kept (safe lists, SARA 6 kategori, LGBT/Israel zero tolerance, anti-evasion, decision tree, evasi hierarchy, image rules) with duplicated phrasing removed - examples.ts 24.7KB -> 20KB: all 31 teaching examples kept; analysis strings shortened, redundant categories/policy_version dropped from example outputs (both optional in the response schema) - output.ts 13.8KB -> 6.8KB: compressed schema + personality + format rules; CRITICAL bans on generic analysis and reply-context requirement retained - system.ts: MEDIA_INSTRUCTIONS compressed, key rules kept Semantic moderation cache (AI_LLM_EMBEDDING_MODEL): - New embeddingClient.ts: OpenAI-compatible embeddings + cosine similarity; degrades gracefully when model/key unset - textCacheStore: stores embedding JSON per verdict, findSimilarTextModeration reuses near-duplicate verdicts (min 0.97 cosine, processing locks skipped) - moderationOrchestrator: after exact-hash miss, embed text-only targets and reuse stored verdict for near-duplicates -> skips expensive chat completion for spam variants; fresh verdicts written back with embedding - Config: AI_LLM_EMBEDDING_MODEL / MIN_SIMILARITY (0.97) / MAX_CANDIDATES (30) - Migration 0012: ADD COLUMN embedding to text_analysis_cache (idempotent) - .env.example documents the new vars
This commit is contained in:
@@ -77,6 +77,8 @@ AI_ANALYSIS_ENABLED=false # Enable AI content moderation (default:
|
||||
AI_LLM_BASE_URL=https://9router.asepharyana.my.id/v1 # LLM API base URL (default)
|
||||
AI_LLM_MODEL=text # LLM text model name (default: text)
|
||||
# AI_LLM_VISION_MODEL= # Vision model for image analysis (falls back to AI_LLM_MODEL)
|
||||
# AI_LLM_EMBEDDING_MODEL= # Embedding model for semantic moderation cache (optional; enables near-duplicate text reuse to save LLM calls)
|
||||
# AI_LLM_EMBEDDING_MIN_SIMILARITY=0.97 # Min cosine similarity to reuse a cached verdict (default: 0.97)
|
||||
AI_LLM_MAX_CONCURRENT=5 # Max concurrent LLM API calls (default: 5)
|
||||
AI_LLM_IMAGE_MAX_DIMENSION=1024 # Max image dimension in pixels before resize (default: 1024)
|
||||
AI_LLM_TEXT_BATCH_SIZE=20 # Max messages per text-only moderation batch (default: 20)
|
||||
|
||||
Reference in New Issue
Block a user