Merge remote-tracking branch 'refs/remotes/upstream/main'
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s

# Conflicts:
#	ROADMAP.md
#	TODO.md
#	src/cli.tsx
#	src/config.ts
#	src/mcp.ts
#	src/permission.ts
#	src/prompt.ts
#	src/session.ts
#	src/snapshot.ts
#	src/subagent.ts
#	src/tools-extra.ts
#	src/tools.ts
#	src/ui/App.tsx
#	test/mcp.test.ts
#	test/prune.test.ts
#	test/session.test.ts
#	test/tools.test.ts
This commit is contained in:
asepharyana
2026-09-21 20:43:35 +07:00
59 changed files with 3481 additions and 463 deletions
+4
View File
@@ -147,6 +147,10 @@ running, the handler exits as usual.
`task` runs a nested `streamText` and returns one message. The subagent kinds hold different
tool sets: `explore` and `review` the read-only tools, `worker` those plus every write tool.
A `tasks` array on the call runs several of these nests concurrently — each gets its own
context window and stream, awaited together via `Promise.all` — so independent investigations
overlap instead of queueing. The single-prompt form is just the one-element case; the two paths
share the same nested-loop machinery and the same reporting bus.
The consequences follow from the tool set, not from policy:
+12 -6
View File
@@ -16,7 +16,7 @@ faster and the fallback path is exercised without it.
```bash
bun run shiro # run from source
bun run typecheck # tsc --noEmit
bun test # 713 tests
bun test # 895 tests
bun run build # single binary for this platform -> dist/shiro
bun run release # all five platforms -> dist/release + SHA256SUMS
bun run install:local # build, then copy onto PATH
@@ -100,7 +100,7 @@ mock-verification test:
fails if a builtin tool has no `_meta` or lives in no set, or a mutating tool is outside
`MUTATING_TOOLS` — both are derived from the definitions, not hand-lists.
Every tool costs roughly 550 characters of schema on every request. Nineteen built-in tools is
Every tool costs roughly 550 characters of schema on every request. Forty-one built-in tools is
past where selection accuracy starts to matter, which is why sets exist and why a new tool
needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why.
One set, `net`, is opt-in rather than on: `web_fetch` is the one tool that leaves the machine.
@@ -153,9 +153,9 @@ Windows host and rejected everywhere else — so a green local release is not pr
`buildArgs()` is unit-tested for both hosts because of exactly that.
`.github/workflows/release.yml` then runs typecheck and tests, cross-compiles all five
targets on one Ubuntu runner, asserts the built binary reports the expected version, and
publishes a GitHub release with the binaries and `SHA256SUMS`. A tag containing `-` is
published as a prerelease.
targets on one Ubuntu runner, asserts each built binary reports the expected version and is
non-empty, and publishes a GitHub release with the binaries and `SHA256SUMS`. A tag containing
`-` is published as a prerelease.
Bun cross-compiles from any host, which is why there is no build matrix. Verified: a working
`darwin-arm64` binary builds on Windows.
@@ -163,10 +163,16 @@ Bun cross-compiles from any host, which is why there is no build matrix. Verifie
Publishing is gated on a `v*` tag, so a manual `workflow_dispatch` run produces artifacts
without releasing.
The release body is composed by `scripts/make-release-notes.ts` from the matching `## [<version>]`
section of `CHANGELOG.md` (plus an artifact inventory), and it fails the release if that heading
is missing — so a tag with no changelog entry cannot ship an empty body. Write the changelog
entry first, then tag.
## CI
`.github/workflows/ci.yml` runs typecheck, tests, and a build on Ubuntu, macOS, and Windows
for every push and PR.
for every push and PR, caching the bun install store across runs so a no-change run skips the
dependency download.
All three are necessary. The tools shell out to `rg`, `git`, and a platform shell, and path
handling differs — a Windows-only break is invisible on Linux until someone hits it.
+6 -3
View File
@@ -51,7 +51,7 @@ $ shiro -p "count the tools" --json
{"type":"tool-start","id":"c1","name":"grep"}
{"type":"tool-call","id":"c1","name":"grep","input":{"pattern":"tool\\("}}
{"type":"tool-result","id":"c1","name":"grep","output":"src/tools.ts:26: ..."}
{"type":"text","text":"There are 16 built-in tools."}
{"type":"text","text":"There are 41 built-in tools."}
{"type":"done","inputTokens":4210,"outputTokens":88}
```
@@ -182,8 +182,11 @@ in the system prompt, and CI is exactly where nobody is watching what it says. S
## Cost control
Headless runs are unattended, so a runaway loop costs real money. `--agent quick` caps the
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. There is no
spend ceiling yet — see [TODO.md](../TODO.md).
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. `maxSpendUsd` in
the config is a session spend ceiling: it warns once at 80%, and past 100% the next turn is
refused naming the ceiling and the run exits non-zero. It is only enforced on priced models —
an unpriced model has no dollar figure to compare against. See
[configuration](configuration.md).
What a run actually costs is in the `done` event, so a wrapper can total it:
+58 -21
View File
@@ -6,7 +6,7 @@ command over stdio, and a remote http or sse endpoint.
## Adding one from the prompt
```
/mcp list what is configured, with the tool count each contributed
/mcp list what is configured, with the tool count (or "lazy") each contributed
/mcp add wizard: local or remote, then the fields that kind needs
/mcp remove <name>
```
@@ -42,7 +42,7 @@ list under a request that is already running.
mcp servers
/mcp add to add one
- `filesystem` (local) - 11 tools
- `filesystem` (local) - connected (lazy)
npx -y @modelcontextprotocol/server-filesystem .
- `api` (remote) - failed: fetch failed
https://example.com/mcp
@@ -50,6 +50,9 @@ mcp servers
configured in /home/you/.shiro-neko/config.json
```
Under `eager` the same list shows a tool count instead of `(lazy)`, because every server's tools
are registered up front.
## The config file
The wizard writes this; it is equally editable by hand.
@@ -93,6 +96,7 @@ schema check at load.
```json
{
"mcpMode": "eager",
"mcpServers": {
"fs": {
"command": "npx",
@@ -120,6 +124,10 @@ inherits your `PATH` unless you replace it.
**Remote** servers take `url`, and optionally `type` (`http` or `sse`, default `http`) and
`headers`.
A sibling key, `"mcpMode"`, picks how the configured servers' tools reach the model: `lazy`
(the default) or `eager`. It is hand-edited — the `/mcp add` wizard does not set it — and applies
to every server, so it lives beside `mcpServers`, not inside one.
A token in `headers` sits in `config.json` in plain text, same as `apiKey`. For anything beyond
a local dev token, prefer a stdio server that reads its own credential from the environment.
@@ -132,10 +140,35 @@ calling the binary directly is usually the difference between a noticeable wait
`--no-mcp` skips them all, which is also the quickest way to tell whether a slow start is MCP
or something else.
## How a server's tools reach the model
Two modes, switched with `"mcpMode"` in config. **`lazy` is the default**; `eager` is the opt-in.
- **Lazy** registers three meta-tools — `mcp_list`, `mcp_inspect`, `mcp_call` — instead of one
schema per server tool. The server's real tools are fetched only when `mcp_call` actually
invokes one, so a server exposing twenty tools costs almost nothing until one is used. This
is why the affordability paragraph in the README says a configured server no longer taxes
every request.
- **Eager** registers every server tool up front as in the old 1.0 behaviour. If a server
exposes only two tools, eager is cheaper because there is no list-then-inspect round trip.
The model is told the connected server names through the meta-tool descriptions and a prompt
line, then discovers each tool's schema on demand:
```
- mcp_list, mcp_inspect, mcp_call: MCP tools are fetched on demand. mcp_list names a
server's tools, mcp_inspect reads one tool's schema, mcp_call runs it. Never guess a
server or tool name: list first.
```
A `mcp_call` still routes through the same permission rules and guard as a built-in, so the
lazy path is not a way around approval.
## Naming
Tools arrive as `mcp__<server>__<tool>`. A server named `fs` exposing `read_file` becomes
`mcp__fs__read_file`.
In `lazy` mode (the default) tools are addressed as `mcp_call(server, toolName, args)`; the
server names are zero-ambiguity identifiers you list first. In `eager` mode tools arrive as
`mcp__<server>__<tool>`: a server named `fs` exposing `read_file` becomes `mcp__fs__read_file`.
The namespace is not cosmetic. Two servers both exposing `search` would otherwise silently
shadow each other, and the model would call one believing it was the other.
@@ -165,26 +198,29 @@ Python interpreter should not stop you from editing a file.
## Inspecting
`/tools` lists everything offered this turn, MCP tools included. The system prompt describes
them as a group:
`/tools` lists everything offered this turn. In `eager` mode that includes each MCP tool, named
`mcp__<server>__<tool>`, described with whatever the server sent. In `lazy` mode the three
meta-tools appear and the server's real tools are surfaced by `mcp_list` inside the session.
```
- mcp__api__query, mcp__fs__read_file: from MCP servers, named mcp__<server>__<tool>.
Each needs approval; read its own description before calling.
```
Their individual descriptions come from the server, so that is what the model reads before
calling one.
Into `mcp_inspect` or `mcp_list` goes the server name, not an `mcp__` path, so the prompt tells
the model which servers are connected and to list first before guessing a tool name.
## Cost
Each tool adds its name, description, and JSON schema to every request. The built-ins average
548 bytes; MCP tools vary with how verbose the server's schema is. A server exposing twenty
tools costs roughly 2,750 tokens per turn, sent whether or not the model uses any of them.
In **lazy** mode (the default) a configured server contributes three small meta-tool schemas to
every request, not one schema per tool. A server exposing twenty tools therefore costs a few
hundred tokens per turn rather than roughly 2,750, and it stays cheap whether the model uses
the tools or not. Browsing a server's tools and reading a schema still brings that server's
schema into view one tool at a time, but only when the model asks for it.
MCP tools are **not** covered by `toolSets` — that budget only governs the built-ins. There is
no per-server switch either, so the choice is a server or no server, and `--no-mcp` for all of
them. If one exposes many tools you never use, a narrower server is worth finding or writing.
In **eager** mode each tool adds its name, description, and JSON schema to every request, and
that cost is sent whether or not the model uses any of them. That is the right trade only for a
server with one or two tools, which is why eager exists.
MCP tools are **not** covered by `toolSets` in either mode — that budget only governs the
built-ins. There is no per-server switch beyond the global `mcpMode`, so choosing `eager` turns
every server eager; a server exposing many tools you never use is worth finding a narrower one
for.
`/tools` shows the count both ways:
@@ -210,8 +246,9 @@ that break in practice are the handshake and the framing, and a mock asserts nei
A server that starts but returns nothing useful is the harder case. In order of speed:
1. `/tools` — did the tools arrive at all? A server with no tools is a `tools/list` problem.
2. `shiro -p "call mcp__x__y with ..." --json --yolo` — the exact `tool-call` input and
`tool-result` output, one JSON object per line.
2. `shiro -p "run mcp_call(server, \"api\", \"query\", {...}) with ..." --json --yolo` — the
exact `tool-call` input and `tool-result` output, one JSON object per line. In eager mode the
address is `mcp__<server>__<tool>` instead.
3. Run the server by hand: `echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | your-server`.
If that is wrong, nothing above it can be right.