OpenCode Free returns HTTP 400 for muse-spark-1.3-contributor-free when tool_choice is non-auto. Declare forceAutoToolChoiceModels quirk and normalize explicit tool_choice to auto.
Caveman/Ponytail injection now matches each target wire format instead of
assuming an OpenAI-shaped body:
- Chat arrays append a text block; Responses arrays append input_text and
create typed message items
- Claude inserts before the final cache-control block; Gemini preserves the
snake/camel systemInstruction wrapper
- Kiro updates systemPrompt and its mirrored first-user prefix atomically,
rolling back if the pair fails to converge
- Format label decides Claude/Gemini before the wire-shape sniff, since their
bodies also carry messages[]/contents[] and Anthropic rejects a "system"
role inside messages[]
- Delimiter-aware dedup makes injection exact-idempotent across retries, so
distinct prompts sharing a long prefix are no longer collapsed
- Every write is fail-open on frozen or proxied bodies
Saver order and X-9Router-Token-Saver: off behavior are unchanged.
Fixes#3202.
- add Headroom extras status + install endpoints
- show Headroom version + code/ml extras in Token Saver UI
- fix Windows interpreter selection to read from env with headroom-ai
Add four shared Caveman prompt fragments (no invented abbreviations,
preserve user language, no self-reference, no decoration) across all six
levels, and remove ULTRA contradictions around abbreviations/arrow
shorthand. Adds regression tests for the prompt rules.
Co-authored-by: Cursor <cursoragent@cursor.com>
Clients echo full message history each turn including reasoning_content,
which the Kimchi OpenAI gateway counts as input tokens. Multi-turn convos
balloon to 100k+ tokens and the model returns empty content.
KimchiExecutor.transformRequest now strips reasoning_content from assistant
messages when it exceeds an 8-char threshold, preserving the 1-char
placeholder injectReasoningContent sets and keeping content intact.
Co-authored-by: Cursor <cursoragent@cursor.com>