feat: render GGUF Jinja template via minijinja crate
- Replaced manual prompt building with GGUF's chat template rendered through minijinja (Rust Jinja2 engine) - Simplified template: removed multi-step tool detection, reasoning extraction (not needed at template level) - Added reasoning_content separation: clean_text() returns (reasoning, answer) tuple - Added reasoning_content field to ResponseMessage and SseDelta for OpenAI-compatible output - Embedded template at build time via include_str! from templates/chat_template.jinja - Built-in support for enable_thinking, tool_definitions, tool_calls, tool_response
This commit is contained in:
@@ -6,6 +6,7 @@ edition = "2021"
|
||||
[dependencies]
|
||||
# LLM inference
|
||||
llama-cpp-2 = "0.1"
|
||||
minijinja = "2"
|
||||
|
||||
# HTTP server
|
||||
axum = { version = "0.8", features = ["json"] }
|
||||
|
||||
Reference in New Issue
Block a user