Commit Graph
25 Commits
Author SHA1 Message Date
asepharyana 9fd0fb0181 fix(anthropic): non-stream /v1/messages returned OpenAI shape; strip EOS sentinel from text deltas
BUG: Anthropic client (sourceFormat=CLAUDE) hitting a claude-target provider
got a raw OpenAI chat.completion body on non-streaming requests. The
needsTranslation(CLAUDE,CLAUDE) gate is false when target===source, so the
translator never ran; a claude-transport executor replying OpenAI JSON
(opencode/big-pickle) leaked choices[]/prompt_tokens to the client, which
Anthropic SDKs cannot parse (no content[] blocks, no type:"message").

FIX: shape-aware guard toClaudeMessageShape() in nonStreamingHandler — when
sourceFormat is CLAUDE, convert any OpenAI-shape body to a proper Claude
message (type, content blocks with thinking/text/tool_use, stop_reason via
finish mapping, usage input/output tokens). Claude-shaped bodies pass through.

Also: strip <|im_end|>/<|endoftext|>/<|eot_id|> EOS sentinels from Claude
text deltas in openai-to-claude and kiro-to-claude translators — the upstream
EOS token leaks into the final text_delta (observed 'OK<|im_end|>').

Tests: anthropic-nonstream-shape.test.js (6: eos strip + shape guard incl
tool_calls→tool_use, pass-through, finish mapping). 28/28 related tests green.
2026-09-23 15:24:15 +07:00
asepharyana 2bc9d67d26 merge: pull mhiqrambg/9router-mibp-version into master
Merge the MIBP fork (v1.0.14, synced to decolua v0.5.81) into our master
(0.5.86) at merge-base a8c9d380. Keep HEAD's infra policy (untracked
lockfile, mirror-configurable Dockerfile, decolua GHCR/DockerHub, README)
while absorbing the fork's engine features:

- feat(providers): freebuff provider + executor + OAuth + usage tracking
- feat(providers): cline free-tier models, Freebuff catalog sync
- feat(proxy-pools): pool egress geo probe, proxy-pool fitness + retry
- fix(usage): hide noAuth providers (devin-cli, mimo-free) from usage list
- test(harness): DATA_DIR isolation so tests never write the real DB
- fix(codebuddy-intl): probe token in connection test, OAuth by identity
- chore(guards): durable markers so fixes aren't silently dropped

Resolutions:
- registry/index.js regenerated deterministically (122 providers, alpha
  order). trae/windsurf/devin-cli stay hidden per HEAD security posture
  (no tool-calling / local-agent shell access) — not re-enabled.
- nonStreamingHandler: drop the generic unconditional unwrapDataEnvelope
  call; envelope unwrap stays scoped to clineEnvelope-quirk providers
  (unwrapClineEnvelope), fixing a latent mibp bug where non-opted-in
  providers ({success,data} bodies) were stripped.
- Drop fork-local Docker lockfile policy (package-lock.json, AGENTS.md,
  .npmrc verify scripts): this repo keeps package-lock untracked (nix
  build deploy). .npmrc (audit=false/fund=false) kept.
- Keep gitbook-pages workflow enabled (ours); mibp disabled it.
- Restore 13 upstream tests mibp deleted (they cover features we keep).

Verified: 2841 tests, 2687 pass, fail set byte-identical to HEAD (zero
new regressions); providers/alias/oauth baselines regenerated to merged
code and all green.
2026-09-23 11:44:58 +07:00
MUH. IQRAM BAHRING 0ac771bad8 merge: sync upstream v0.5.81 into MIBP fork
# Conflicts:
#	.gitignore
#	Dockerfile
#	open-sse/handlers/chatCore.js
#	open-sse/providers/registry/cline.js
#	open-sse/providers/registry/index.js
#	open-sse/services/usage.js
#	open-sse/utils/streamHandler.js
#	package.json
#	src/app/(dashboard)/dashboard/profile/page.js
#	src/app/(dashboard)/dashboard/providers/[id]/page.js
2026-09-19 11:35:34 +08:00
Aaron 822aa958d1 fix(opencode): cloak Responses requests that already have tools
Free-tier Zen models reject Responses requests with 403 FreeTierError
when client tools are present but the fingerprint quartet is missing.
Apply the fingerprint tools to every OpenCode request, canonicalise
case variants of the quartet (Bash->bash) without duplication, and
restore the caller's original spellings on the response side via a
request-local WeakMap threaded through the existing toolNameMap.
2026-09-19 10:15:45 +07:00
MUH. IQRAM BAHRING b0505067b2 feat(providers): add Cline free-tier models and sync Freebuff catalog
- Cline: 6 free models, cline-cli product headers, API-key auth,
  {data} envelope unwrap, workos: prefix handling
- Freebuff: muse-spark 1.3 → 1.2 (upstream withdrawal 2026-09-07),
  DeepSeek V4.1 Flash rename
2026-09-11 12:21:35 +08:00
LLL 248d7da01c revert(qoder): drop the Responses usage plumbing from shared code
The merged Qoder work also rewrote shared translator/handler code so that
/v1/responses clients got token usage on response.completed. That changed
behaviour for every provider, not just Qoder: proxies saw input tokens
rise by the 2000-token context buffer, and the plain token mapping was
replaced by one that always adds input_tokens_details.

A probe confirms the Qoder benefit does not depend on those edits: the
executor's coalescer already emits one include_usage-style finish chunk, so
a Claude client receives input_tokens and cache_read_input_tokens with
every shared file at its original state. Only the Responses path relies on
the shared translator, and that path has no Qoder-owned seam to put it in.

Reverts the shared files to their pre-PR state and drops the Responses
usage test. The Cline envelope unwrap in nonStreamingHandler.js, which
landed after the PR in the same file, is kept.
2026-09-10 23:13:06 +07:00
Nick Nyanjui 122f23eebc fix(cline,airforce): unwrap {success,data} envelope, add live catalog, and refresh airforce free models
Cline (api.cline.bot) wraps non-stream chat completions in
{"success":true,"data":{...choices...}}, which both the dashboard model-test
ping and the proxy non-stream path read at top level, producing "Provider
returned no completion choices for this model" (#3644). Unwrap the envelope
before usage extraction and response translation; the error envelope
({"success":false,...}) never matches and passes through untouched.

Scoped through `transport.quirks.clineEnvelope` so only cline/clinepass opt
in — no other provider's response body is ever rewritten.

Also adds a live Cline catalog: `fetchClineRawModels()` is shared between
`resolveClineModels()` (full catalog, including free-tier ids such as
z-ai/glm-5.3-flash) and `resolveClinepassModels()` (cline-pass/* only), wired
into /v1/models, the per-provider models route, and the combo selector's
model picker with the static catalog kept as fallback.

Refreshes the dead api-airforce free models (anthropic/claude-3.7-sonnet,
moonshot/kimi-k2.6, google/gemini-2.5-flash) with the live gpt-oss-120b,
gpt-oss-20b and kimi-k2.7-code, plus passthroughModels, forceStream and a
suggested-models filter.
2026-09-10 22:48:28 +07:00
LLL 1f10f9e5c4 fix(qoder): report usage to all clients and stop inlining large attachments
- Coalesce Qoder's empty finish-in-delta frame with the later choices:[] usage
  frame so OpenAI and Claude clients receive prompt_tokens, completion_tokens
  and cache-hit tokens (the dashboard already saw them)
- Upload inlined images through /api/v2/image/upload like qodercli, and stub
  oversized non-image files instead of stuffing 30MB+ data URIs into
  agent_chat_generation
- Emit response.completed -> response.usage for chat-native upstreams so
  /v1/responses clients (Codex CLI, sub2api) no longer log 0/0/0
- Keep Claude message_delta.usage working when usage arrives without choices[0]
- Escalate to the smallest advertised Qoder context tier (200K/400K/1M) when
  the estimated prompt no longer fits max_input_tokens
- Pass apiKey for PAT connections and list hidden enable:false catalog keys
  from /v1/models
2026-09-10 22:08:19 +07:00
nguyenha935 d06e0d26c6 fix(translator): preserve Responses Lite tools across Chat providers
Codex Responses Lite clients routed to a chat-native OpenAI-compatible
provider lost tool use in three places: non-streaming Chat responses
leaked the raw chat.completion envelope instead of Responses output
items, internal reasoning continuity fields leaked into the outbound
Chat body causing some upstreams to reject the request, and the
Responses to Chat request translator ignored additional_tools,
custom_tool_call, and custom_tool_call_output items entirely.

Also fixes apiType (chat vs responses) for openai-compatible nodes
being resolved from the immutable provider ID instead of the stored
node config, so editing a node's API Type had no runtime effect.
2026-08-05 13:27:25 +07:00
decoluaandCursor a625ea9fd8 refactor(log): unify request lifecycle logging with session-colored tags
Collapse scattered per-request console lines (request/routing/auth/pending/
usage/stream-usage/stream) into 3 correlated lines: request, transform,
done. Add stable per-session color tag so concurrent request lines are
easy to follow, surface thinking intent, always-on full error logging
for debug, re-enable warn level, and uppercase keyword labels. Also fix
usage overview cards wrapping (5 cards -> grid-cols-5).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-10 18:01:20 +07:00
Elio Bonfim Júnior dcf1927f22 feat(pxpipe): PXPIPE token saver — multimodal prompt compression (#2465)
Add pxpipe as an experimental fifth Token Saver: Claude-format request
bodies above a configurable size threshold are rendered as dense PNGs
via the pxpipe-proxy library API (transformAnthropicMessages) before
dispatch, cutting estimated input tokens by ~35-60% on token-dense
contexts. Integration follows the Headroom pattern: applied to the final
body in chatCore just before dispatch, fail-open on any error/timeout.

Managed npm install into DATA_DIR/pxpipe, dynamic loader with per-version
cache-bust, JSONL event log with rotation, /api/pxpipe/* endpoints, Token
Saver card (marked experimental) + /dashboard/pxpipe page, and per-request
Activated/Skipped annotation in Request Details. Disabled by default.
2026-07-10 16:10:42 +07:00
Nant361andCursor 8a664d619d feat(kimchi): add Kimchi OAuth provider support
Add Kimchi as a browser-token OAuth provider routed through its
OpenAI-compatible gateway. Discover live models for /v1/models and
provider models, normalize Claude-compatible requests, and wire up
provider connection tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 15:29:17 +07:00
Mink NguyenandCursor 0d21668917 Fix usage logging dedupe and reduce stats churn
- batch console log buffer events and support batched SSE log messages
- debounce usage stats update/pending events to reduce UI/runtime churn
- avoid awaiting request-success bookkeeping before returning provider responses
- deduplicate identical usage writes in usageHistory/daily aggregates
- reduce default logger verbosity from DEBUG to INFO (overridable via LOG_LEVEL)

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 10:22:20 +07:00
NautilaceaeandCursor 5306bd904e feat(antigravity): native image generation support
Add image generation for Antigravity provider via gemini-3.1-flash-image
and gemini-3-pro-image, exposed through Text to Image UI and
/v1/images/generations.

- registry: serviceKinds ['llm','image'] + image model entries
- executor: image model detection + image_gen request envelope
- chatCore: force stream=false for image models (generateContent)
- nonStreamingHandler: parse inlineData -> markdown image
- imageGenerationCore: useExecutor fast-path for executor delegation
- imageProviders/antigravity: image adapter with image input support
- usage/google: image models in quota whitelist

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-21 17:54:32 +07:00
fjiaandCursor 411a589781 fix(claude-to-openai): handle OpenAI-format responses in non-streaming path
Some providers (e.g. xiaomi-tokenplan -claude models) return OpenAI-format
responses even when request was translated to Claude. Early-return now detects
choices[]. Also strip reasoning_content only when content is non-empty so
thinking models keep their only output.

Closes #1836

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-17 09:43:58 +07:00
41f94ce8c8 fix(minimax): Bổ sung MiniMax-M3 + cập nhật Quota Tracker coding/CN
Squash-merge PR #1631 (decolua/9router) — chỉ lấy file code + test, bỏ docs.

- feat(minimax): add MiniMax-M3 to intl + cn provider models (targetFormat claude)
- feat(minimax): add MiniMax-M3 pricing entry
- fix(minimax): translate Claude body khi content=null (M3 thinking-only)
- fix(minimax): hiển thị quota M-series bucket "general"/"MiniMax-M*" + percent-only
- test: minimax usage / model registration / pricing

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-06 10:01:05 +07:00
decolua 5abc9e5c74 add GPT 5.5 model 2026-04-24 09:51:05 +07:00
Anurag Saxena a53ccf1343 fix: strip reasoning_content from non-streaming responses (closes #509) (#517) 2026-04-07 09:46:28 +07:00
Anurag Saxena e3a7733a08 fix: strip functionCall/functionResponse id and synthetic thoughtSignature for Vertex AI (closes #388) (#414) 2026-03-27 10:46:47 +07:00
Liam 01e4a28f0a fix: normalize finish_reason to 'tool_calls' when tool calls are present (#379)
Some upstream providers (e.g. Antigravity) return non-standard finish_reason
values like 'other' instead of the OpenAI-standard 'tool_calls' when the
model invokes tools. This causes downstream consumers (e.g. OpenClaw) to
fail to execute tool calls, breaking agentic sub-agent workflows.

Changes:
- nonStreamingHandler: post-translation guard that normalizes finish_reason
  to 'tool_calls' when message.tool_calls is present
- sseToJsonHandler: accumulate tool_calls from streaming deltas in
  parseSSEToOpenAIResponse; extract function_call items from Responses API
  output in handleForcedSSEToJson
- openai-responses translator: use toolCallIndex to choose between
  'tool_calls' and 'stop' in flush and response.completed events

Tested: 7 scenarios (non-stream text, single/multiple tool calls, stream
text/tool calls, multi-turn tool conversation, tools present but unused)
2026-03-23 09:35:25 +07:00
decolua adae2605bf Feat : Auto restart after crash 2026-03-14 09:37:29 +07:00
Nick Roth d12b14f411 feat: AI SDK compatibility - Accept header & JSON markdown stripping
- Respect Accept: application/json header to return non-streaming JSON
  instead of SSE, fixing AI SDK generateObject/generateText compatibility
- Strip markdown code block markers (```json...```) from Claude
  non-streaming responses to prevent JSON parse errors

Cherry-picked and adapted from PR #290 by @rothnic
https://github.com/decolua/9router/pull/290

Made-with: Cursor
2026-03-13 10:00:47 +07:00
decolua b0c6b61398 Refactor config 2026-03-12 16:20:46 +07:00
decolua 83d94daa82 feat(ollama): Enhance Ollama support by adding new models, updating API format handling, and integrating translation functionality. 2026-03-12 15:24:10 +07:00
decolua 5954b8f4eb - Refactor chatCore.js to streamline imports and remove unused functions.
- Fix streaming /v1/responses
2026-02-27 11:15:12 +07:00