Merge remote-tracking branch 'refs/remotes/upstream/main'
# Conflicts: # ROADMAP.md # TODO.md # src/cli.tsx # src/config.ts # src/mcp.ts # src/permission.ts # src/prompt.ts # src/session.ts # src/snapshot.ts # src/subagent.ts # src/tools-extra.ts # src/tools.ts # src/ui/App.tsx # test/mcp.test.ts # test/prune.test.ts # test/session.test.ts # test/tools.test.ts
This commit is contained in:
@@ -147,6 +147,10 @@ running, the handler exits as usual.
|
||||
|
||||
`task` runs a nested `streamText` and returns one message. The subagent kinds hold different
|
||||
tool sets: `explore` and `review` the read-only tools, `worker` those plus every write tool.
|
||||
A `tasks` array on the call runs several of these nests concurrently — each gets its own
|
||||
context window and stream, awaited together via `Promise.all` — so independent investigations
|
||||
overlap instead of queueing. The single-prompt form is just the one-element case; the two paths
|
||||
share the same nested-loop machinery and the same reporting bus.
|
||||
|
||||
The consequences follow from the tool set, not from policy:
|
||||
|
||||
|
||||
+12
-6
@@ -16,7 +16,7 @@ faster and the fallback path is exercised without it.
|
||||
```bash
|
||||
bun run shiro # run from source
|
||||
bun run typecheck # tsc --noEmit
|
||||
bun test # 713 tests
|
||||
bun test # 895 tests
|
||||
bun run build # single binary for this platform -> dist/shiro
|
||||
bun run release # all five platforms -> dist/release + SHA256SUMS
|
||||
bun run install:local # build, then copy onto PATH
|
||||
@@ -100,7 +100,7 @@ mock-verification test:
|
||||
fails if a builtin tool has no `_meta` or lives in no set, or a mutating tool is outside
|
||||
`MUTATING_TOOLS` — both are derived from the definitions, not hand-lists.
|
||||
|
||||
Every tool costs roughly 550 characters of schema on every request. Nineteen built-in tools is
|
||||
Every tool costs roughly 550 characters of schema on every request. Forty-one built-in tools is
|
||||
past where selection accuracy starts to matter, which is why sets exist and why a new tool
|
||||
needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why.
|
||||
One set, `net`, is opt-in rather than on: `web_fetch` is the one tool that leaves the machine.
|
||||
@@ -153,9 +153,9 @@ Windows host and rejected everywhere else — so a green local release is not pr
|
||||
`buildArgs()` is unit-tested for both hosts because of exactly that.
|
||||
|
||||
`.github/workflows/release.yml` then runs typecheck and tests, cross-compiles all five
|
||||
targets on one Ubuntu runner, asserts the built binary reports the expected version, and
|
||||
publishes a GitHub release with the binaries and `SHA256SUMS`. A tag containing `-` is
|
||||
published as a prerelease.
|
||||
targets on one Ubuntu runner, asserts each built binary reports the expected version and is
|
||||
non-empty, and publishes a GitHub release with the binaries and `SHA256SUMS`. A tag containing
|
||||
`-` is published as a prerelease.
|
||||
|
||||
Bun cross-compiles from any host, which is why there is no build matrix. Verified: a working
|
||||
`darwin-arm64` binary builds on Windows.
|
||||
@@ -163,10 +163,16 @@ Bun cross-compiles from any host, which is why there is no build matrix. Verifie
|
||||
Publishing is gated on a `v*` tag, so a manual `workflow_dispatch` run produces artifacts
|
||||
without releasing.
|
||||
|
||||
The release body is composed by `scripts/make-release-notes.ts` from the matching `## [<version>]`
|
||||
section of `CHANGELOG.md` (plus an artifact inventory), and it fails the release if that heading
|
||||
is missing — so a tag with no changelog entry cannot ship an empty body. Write the changelog
|
||||
entry first, then tag.
|
||||
|
||||
## CI
|
||||
|
||||
`.github/workflows/ci.yml` runs typecheck, tests, and a build on Ubuntu, macOS, and Windows
|
||||
for every push and PR.
|
||||
for every push and PR, caching the bun install store across runs so a no-change run skips the
|
||||
dependency download.
|
||||
|
||||
All three are necessary. The tools shell out to `rg`, `git`, and a platform shell, and path
|
||||
handling differs — a Windows-only break is invisible on Linux until someone hits it.
|
||||
|
||||
+6
-3
@@ -51,7 +51,7 @@ $ shiro -p "count the tools" --json
|
||||
{"type":"tool-start","id":"c1","name":"grep"}
|
||||
{"type":"tool-call","id":"c1","name":"grep","input":{"pattern":"tool\\("}}
|
||||
{"type":"tool-result","id":"c1","name":"grep","output":"src/tools.ts:26: ..."}
|
||||
{"type":"text","text":"There are 16 built-in tools."}
|
||||
{"type":"text","text":"There are 41 built-in tools."}
|
||||
{"type":"done","inputTokens":4210,"outputTokens":88}
|
||||
```
|
||||
|
||||
@@ -182,8 +182,11 @@ in the system prompt, and CI is exactly where nobody is watching what it says. S
|
||||
## Cost control
|
||||
|
||||
Headless runs are unattended, so a runaway loop costs real money. `--agent quick` caps the
|
||||
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. There is no
|
||||
spend ceiling yet — see [TODO.md](../TODO.md).
|
||||
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. `maxSpendUsd` in
|
||||
the config is a session spend ceiling: it warns once at 80%, and past 100% the next turn is
|
||||
refused naming the ceiling and the run exits non-zero. It is only enforced on priced models —
|
||||
an unpriced model has no dollar figure to compare against. See
|
||||
[configuration](configuration.md).
|
||||
|
||||
What a run actually costs is in the `done` event, so a wrapper can total it:
|
||||
|
||||
|
||||
+58
-21
@@ -6,7 +6,7 @@ command over stdio, and a remote http or sse endpoint.
|
||||
## Adding one from the prompt
|
||||
|
||||
```
|
||||
/mcp list what is configured, with the tool count each contributed
|
||||
/mcp list what is configured, with the tool count (or "lazy") each contributed
|
||||
/mcp add wizard: local or remote, then the fields that kind needs
|
||||
/mcp remove <name>
|
||||
```
|
||||
@@ -42,7 +42,7 @@ list under a request that is already running.
|
||||
mcp servers
|
||||
/mcp add to add one
|
||||
|
||||
- `filesystem` (local) - 11 tools
|
||||
- `filesystem` (local) - connected (lazy)
|
||||
npx -y @modelcontextprotocol/server-filesystem .
|
||||
- `api` (remote) - failed: fetch failed
|
||||
https://example.com/mcp
|
||||
@@ -50,6 +50,9 @@ mcp servers
|
||||
configured in /home/you/.shiro-neko/config.json
|
||||
```
|
||||
|
||||
Under `eager` the same list shows a tool count instead of `(lazy)`, because every server's tools
|
||||
are registered up front.
|
||||
|
||||
## The config file
|
||||
|
||||
The wizard writes this; it is equally editable by hand.
|
||||
@@ -93,6 +96,7 @@ schema check at load.
|
||||
|
||||
```json
|
||||
{
|
||||
"mcpMode": "eager",
|
||||
"mcpServers": {
|
||||
"fs": {
|
||||
"command": "npx",
|
||||
@@ -120,6 +124,10 @@ inherits your `PATH` unless you replace it.
|
||||
**Remote** servers take `url`, and optionally `type` (`http` or `sse`, default `http`) and
|
||||
`headers`.
|
||||
|
||||
A sibling key, `"mcpMode"`, picks how the configured servers' tools reach the model: `lazy`
|
||||
(the default) or `eager`. It is hand-edited — the `/mcp add` wizard does not set it — and applies
|
||||
to every server, so it lives beside `mcpServers`, not inside one.
|
||||
|
||||
A token in `headers` sits in `config.json` in plain text, same as `apiKey`. For anything beyond
|
||||
a local dev token, prefer a stdio server that reads its own credential from the environment.
|
||||
|
||||
@@ -132,10 +140,35 @@ calling the binary directly is usually the difference between a noticeable wait
|
||||
`--no-mcp` skips them all, which is also the quickest way to tell whether a slow start is MCP
|
||||
or something else.
|
||||
|
||||
## How a server's tools reach the model
|
||||
|
||||
Two modes, switched with `"mcpMode"` in config. **`lazy` is the default**; `eager` is the opt-in.
|
||||
|
||||
- **Lazy** registers three meta-tools — `mcp_list`, `mcp_inspect`, `mcp_call` — instead of one
|
||||
schema per server tool. The server's real tools are fetched only when `mcp_call` actually
|
||||
invokes one, so a server exposing twenty tools costs almost nothing until one is used. This
|
||||
is why the affordability paragraph in the README says a configured server no longer taxes
|
||||
every request.
|
||||
- **Eager** registers every server tool up front as in the old 1.0 behaviour. If a server
|
||||
exposes only two tools, eager is cheaper because there is no list-then-inspect round trip.
|
||||
|
||||
The model is told the connected server names through the meta-tool descriptions and a prompt
|
||||
line, then discovers each tool's schema on demand:
|
||||
|
||||
```
|
||||
- mcp_list, mcp_inspect, mcp_call: MCP tools are fetched on demand. mcp_list names a
|
||||
server's tools, mcp_inspect reads one tool's schema, mcp_call runs it. Never guess a
|
||||
server or tool name: list first.
|
||||
```
|
||||
|
||||
A `mcp_call` still routes through the same permission rules and guard as a built-in, so the
|
||||
lazy path is not a way around approval.
|
||||
|
||||
## Naming
|
||||
|
||||
Tools arrive as `mcp__<server>__<tool>`. A server named `fs` exposing `read_file` becomes
|
||||
`mcp__fs__read_file`.
|
||||
In `lazy` mode (the default) tools are addressed as `mcp_call(server, toolName, args)`; the
|
||||
server names are zero-ambiguity identifiers you list first. In `eager` mode tools arrive as
|
||||
`mcp__<server>__<tool>`: a server named `fs` exposing `read_file` becomes `mcp__fs__read_file`.
|
||||
|
||||
The namespace is not cosmetic. Two servers both exposing `search` would otherwise silently
|
||||
shadow each other, and the model would call one believing it was the other.
|
||||
@@ -165,26 +198,29 @@ Python interpreter should not stop you from editing a file.
|
||||
|
||||
## Inspecting
|
||||
|
||||
`/tools` lists everything offered this turn, MCP tools included. The system prompt describes
|
||||
them as a group:
|
||||
`/tools` lists everything offered this turn. In `eager` mode that includes each MCP tool, named
|
||||
`mcp__<server>__<tool>`, described with whatever the server sent. In `lazy` mode the three
|
||||
meta-tools appear and the server's real tools are surfaced by `mcp_list` inside the session.
|
||||
|
||||
```
|
||||
- mcp__api__query, mcp__fs__read_file: from MCP servers, named mcp__<server>__<tool>.
|
||||
Each needs approval; read its own description before calling.
|
||||
```
|
||||
|
||||
Their individual descriptions come from the server, so that is what the model reads before
|
||||
calling one.
|
||||
Into `mcp_inspect` or `mcp_list` goes the server name, not an `mcp__` path, so the prompt tells
|
||||
the model which servers are connected and to list first before guessing a tool name.
|
||||
|
||||
## Cost
|
||||
|
||||
Each tool adds its name, description, and JSON schema to every request. The built-ins average
|
||||
548 bytes; MCP tools vary with how verbose the server's schema is. A server exposing twenty
|
||||
tools costs roughly 2,750 tokens per turn, sent whether or not the model uses any of them.
|
||||
In **lazy** mode (the default) a configured server contributes three small meta-tool schemas to
|
||||
every request, not one schema per tool. A server exposing twenty tools therefore costs a few
|
||||
hundred tokens per turn rather than roughly 2,750, and it stays cheap whether the model uses
|
||||
the tools or not. Browsing a server's tools and reading a schema still brings that server's
|
||||
schema into view one tool at a time, but only when the model asks for it.
|
||||
|
||||
MCP tools are **not** covered by `toolSets` — that budget only governs the built-ins. There is
|
||||
no per-server switch either, so the choice is a server or no server, and `--no-mcp` for all of
|
||||
them. If one exposes many tools you never use, a narrower server is worth finding or writing.
|
||||
In **eager** mode each tool adds its name, description, and JSON schema to every request, and
|
||||
that cost is sent whether or not the model uses any of them. That is the right trade only for a
|
||||
server with one or two tools, which is why eager exists.
|
||||
|
||||
MCP tools are **not** covered by `toolSets` in either mode — that budget only governs the
|
||||
built-ins. There is no per-server switch beyond the global `mcpMode`, so choosing `eager` turns
|
||||
every server eager; a server exposing many tools you never use is worth finding a narrower one
|
||||
for.
|
||||
|
||||
`/tools` shows the count both ways:
|
||||
|
||||
@@ -210,8 +246,9 @@ that break in practice are the handshake and the framing, and a mock asserts nei
|
||||
A server that starts but returns nothing useful is the harder case. In order of speed:
|
||||
|
||||
1. `/tools` — did the tools arrive at all? A server with no tools is a `tools/list` problem.
|
||||
2. `shiro -p "call mcp__x__y with ..." --json --yolo` — the exact `tool-call` input and
|
||||
`tool-result` output, one JSON object per line.
|
||||
2. `shiro -p "run mcp_call(server, \"api\", \"query\", {...}) with ..." --json --yolo` — the
|
||||
exact `tool-call` input and `tool-result` output, one JSON object per line. In eager mode the
|
||||
address is `mcp__<server>__<tool>` instead.
|
||||
3. Run the server by hand: `echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | your-server`.
|
||||
If that is wrong, nothing above it can be right.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user