- Unified synchronous generate() core with callback; both streaming and non-streaming paths run it via spawn_blocking (context: std Mutex). - build_prompt now passes tool definitions to the template (was dead) and embeds assistant tool-call history as XML matching the parser format; fixes double <tool_response> wrap and template set-scoping bug. - Tokenize with AddBos::Never (template owns <s>) to remove double BOS. - Streaming: preserve inter-word spaces (per-chunk trim removed), add [DONE] + usage chunk, emit error events, single-shot tool_calls delta. - Strict model validation (400 on unknown model); health/UI/README aligned to minicpm5-1b-fable5-v2-thinking; auth returns JSON errors; n_ctx/ n_batch/n_threads env-configurable. - Added 18 unit tests; cargo check/clippy/fmt clean. - scripts/smoke-test.sh for post-deploy verification on the VPS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
87 lines
2.3 KiB
Markdown
87 lines
2.3 KiB
Markdown
# llm-api
|
|
|
|
OpenAI-compatible LLM inference server using `llama-cpp-2` (Rust).
|
|
|
|
**Model:** MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking (Q8_0)
|
|
**Engine:** llama.cpp via `llama-cpp-2` crate
|
|
**Domain:** [ai.asepharyana.my.id](https://ai.asepharyana.my.id)
|
|
|
|
## API
|
|
|
|
### `GET /health`
|
|
```json
|
|
{"status": "ok", "model": "minicpm5-1b-fable5-v2-thinking"}
|
|
```
|
|
|
|
### `GET /v1/models`
|
|
OpenAI-compatible model listing.
|
|
|
|
### `POST /v1/chat/completions`
|
|
OpenAI-compatible chat completions.
|
|
|
|
The server serves a single model and rejects unknown model ids with `400`:
|
|
|
|
```bash
|
|
curl https://ai.asepharyana.my.id/v1/chat/completions \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "minicpm5-1b-fable5-v2-thinking",
|
|
"messages": [{"role": "user", "content": "Hello!"}],
|
|
"max_tokens": 100
|
|
}'
|
|
```
|
|
|
|
## Development
|
|
|
|
```bash
|
|
# Build
|
|
cargo build --release
|
|
|
|
# Run (with model path env var)
|
|
MODEL_PATH=/path/to/model.gguf ./target/release/llm-api
|
|
|
|
# Or use default path
|
|
./target/release/llm-api
|
|
```
|
|
|
|
### Environment variables
|
|
|
|
| Var | Default | Description |
|
|
|-----|---------|-------------|
|
|
| `MODEL_PATH` | `/models/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-Q8_0.gguf` | GGUF model file |
|
|
| `API_KEY` | *(empty = auth off)* | Bearer token required on `/v1/chat/completions` |
|
|
| `SERVER_PORT` | `4010` | Listen port |
|
|
| `RUST_LOG` | `info` | Log level |
|
|
| `N_CTX` / `N_BATCH` / `N_THREADS` | `8192` / `512` / `4` | llama.cpp context/batch/threads |
|
|
|
|
### Smoke test (setelah deploy)
|
|
|
|
```bash
|
|
./scripts/smoke-test.sh http://127.0.0.1:4010 # tanpa auth
|
|
./scripts/smoke-test.sh https://ai.asepharyana.my.id "$API_KEY"
|
|
```
|
|
|
|
## Deploy (Nix + systemd)
|
|
|
|
```bash
|
|
nix build .#default --impure --option sandbox false
|
|
# GitHub Actions: nix copy ssh://vps → systemctl restart llm-api
|
|
```
|
|
|
|
> **Legacy (2026-08-02):** Docker compose dihapus dari produksi. Deploy sekarang Nix+systemd.
|
|
|
|
## Benchmark
|
|
|
|
> *Historic* (MiniCPM-V-4.6). Kept for reference; numbers predate the current
|
|
> MiniCPM5-1B Thinking model.
|
|
|
|
| Framework | Model Size | tok/s | vs PyTorch |
|
|
|-----------|-----------|-------|------------|
|
|
| PyTorch BF16 | 2.48 GB | 0.97 | 1.0x |
|
|
| **llama.cpp Q4_K_M** | **505 MB** | **39.1** | **40.3x** 🏆 |
|
|
|
|
## Infrastructure
|
|
|
|
- Caddy reverse proxy: `ai.asepharyana.my.id` → `127.0.0.1:4010`
|
|
- systemd unit `llm-api`, deploy Nix via GitHub Actions
|