asepharyanaandClaude Code e351d74fa4 refactor(llm-api): implement clean architecture following scraper pattern
Split monolithic 1012-line main.rs into layered hexagonal architecture:
- Domain: entity types and LlmError enum
- Application: prompt building, sampler construction, tool call parsing
- Infrastructure: LlamaEngine wrapping llama-cpp-2 with isolated unsafe transmute
- Presentation: Axum handlers, middleware (auth), error chain, router
- Config: type-safe AppConfig with LazyLock
- Bootstrap: Application struct with build() + run()

Resolves build_sampler/build_sampler_params duplication.
Adds simple web chat UI at GET /.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-07-25 15:07:19 +07:00

llm-api

OpenAI-compatible LLM inference server using llama-cpp-2 (Rust).

Model: MiniCPM-V-4.6 Q4_K_M (505 MB)
Engine: llama.cpp via llama-cpp-2 crate
Domain: ai.asepharyana.my.id

API

GET /health

{"status": "ok", "model": "minicpm-v-4.6-q4_k_m"}

GET /v1/models

OpenAI-compatible model listing.

POST /v1/chat/completions

OpenAI-compatible chat completions.

curl https://ai.asepharyana.my.id/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "minicpm-v-4.6",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 100
  }'

Development

# Build
cargo build --release

# Run (with model path env var)
MODEL_PATH=/path/to/model.gguf ./target/release/llm-api

# Or use default path
./target/release/llm-api

Docker

docker compose -f docker-compose.yml up -d

Benchmark

Framework Model Size tok/s vs PyTorch
PyTorch BF16 2.48 GB 0.97 1.0x
llama.cpp Q4_K_M 505 MB 39.1 40.3x 🏆

Infrastructure

  • Traefik router: ai.asepharyana.my.idllm-api:8080
  • Network: app-shared-net
  • Docker Compose: see llm-api.yml
S
Description
No description provided
Readme
162 KiB
Languages
Rust 72.6%
HTML 21.8%
Shell 3.9%
Jinja 1.7%