Expand the documentation with measured figures and operational detail

Every guide catches up with the sixteen-tool registry: apply_patch and web_fetch sections, the delegation guide with the worker kind, bounded compaction replacing the three-message window description, net in the tool-set tables, six bundled skills, and the module map gains tools-net.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 16:43:27 +07:00
co-authored by Sisyphus
parent 55ebb40524
commit 0dc0819fad
11 changed files with 172 additions and 79 deletions
+37 -4
View File
@@ -33,10 +33,10 @@ pure latency.
and findings recorded with `remember` so they survive compaction.
**`plan`** and **`review`** are genuinely read-only. `write_file`, `edit_file`, `multi_edit`,
and `bash` are withheld from the model, not merely discouraged in prose — a model that cannot
see a tool cannot call it. They keep everything that only reads, including `read_many_files`,
`list_dir`, and the git tools. Their prompts also forbid describing edits as if they had been
made.
`apply_patch`, and `bash` are withheld from the model, not merely discouraged in prose — a
model that cannot see a tool cannot call it. They keep everything that only reads, including
`read_many_files`, `list_dir`, the git tools, and `web_fetch` when the `net` set is enabled.
Their prompts also forbid describing edits as if they had been made.
## Variants and tool sets
@@ -146,3 +146,36 @@ the full set — cheaper per turn as well as safer.
Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and
nothing about the history.
## Delegating with `task`
The `task` tool spans a separate axis from the variants: it runs a subagent with its own
context window, so the parent pays for one report rather than the whole search transcript. The
subagent sees none of the parent's conversation, so its prompt must stand alone.
| Kind | Tools | Approval | For |
|---|---|---|---|
| `explore` (default) | read and search only | never prompts — structurally read-only | a search spanning many files |
| `review` | read and search only | never prompts | a critique of code or a diff |
| `worker` | everything, including writes | every write and command asks, through the parent's gate | a self-contained change whose steps you do not need to watch |
Three properties of the `worker` kind are structural rather than policy:
**The gate is the parent's.** A worker routes each gated call back through the same permission
rules, the same guard plugins, and the same approval prompt as a direct call — flagged `a
worker subagent wants to run ...` so you can tell who is asking. Answering `always` grants the
pattern for the session exactly as it does for you. A subagent that could approve its own
writes would be a way to launder a tool call past you, so there is no separate, weaker gate.
**Denial stops the work.** The worker is told a denial is your decision: report it, do not work
around it. The tool descriptions say the same thing, so the rule survives compaction.
**No `worker` without a channel.** In headless runs there is no one to answer a prompt, so the
`worker` kind is not offered at all — an unattended write is not something to fall into by
accident. The read-only kinds work everywhere. No subagent holds `web_fetch`; network access
stays with the main agent, where the approval prompt says what it is for.
When not to delegate: a single grep, or anything you must supervise step by step — keep that in
your own turn, where every call is on screen. A worker wins when the intermediate steps are
noise: a mechanical rename across twenty files, a test scaffold written to match an existing
suite, a cleanup whose shape you already know.
+21 -12
View File
@@ -145,12 +145,16 @@ running, the handler exits as usual.
## Subagents
`task` runs a nested `streamText` with only `read_file`, `glob`, and `grep`. It returns one
message.
`task` runs a nested `streamText` and returns one message. The subagent kinds hold different
tool sets: `explore` and `review` the read-only tools, `worker` those plus every write tool.
Two consequences follow from the tool set, not from policy:
The consequences follow from the tool set, not from policy:
- It can never need approval, because it has no gated tools.
- `explore` and `review` can never need approval, because they hold no gated tool.
- `worker` needs approval for exactly the calls a direct one would, so the parent owns the
gate: the subagent's `toolApproval` callback routes back through the parent's permission
rules, guard plugins, and prompt. A subagent with its own approval would be a way to launder
a tool call past the user.
- The parent's context holds the findings, not the search transcript.
Progress is reported through a callback, wired to a bus the panel subscribes to. Without the
@@ -196,9 +200,9 @@ tool call carries an itemId, so after the first compaction the model could not s
already run, and re-ran the same tools until the step limit ended the turn. **Compaction may
shorten the history; it must not blank it.**
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts
*messages*, so the cut lands between an assistant `tool-call` and the `tool` message answering
it. What reaches the wire is a `function_call_output` with no `function_call`:
**A tool result without its tool call.** Tool pruning counts messages, so a cut can land between
an assistant `tool-call` and the `tool` message answering it. What reaches the wire is a
`function_call_output` with no `function_call`:
```
400 No tool call found for function call output with call_id call_…
@@ -208,6 +212,10 @@ it. What reaches the wire is a `function_call_output` with no `function_call`:
reverse pairing is deliberately left alone: a call still awaiting its result is exactly what a
suspended approval looks like, and dropping it would break resume.
The pruning ladder drops reasoning first and then keeps the widest recent tool tail that fits.
The SDK carries that returned message view into later steps, and the session reports compaction
once per turn rather than once per step.
## Registry
`/registry` fetches an index of external skills and plugins over https. Skills are prompt text
@@ -226,6 +234,7 @@ the reasoning.
| `session.ts` | the loop, approvals, compaction, event stream |
| `tools.ts` | file and shell tools, tool sets, ripgrep bridge, bash streaming and interrupt |
| `tools-git.ts` | read-only git tools, spawned with a fixed argv |
| `tools-net.ts` | `web_fetch`, private-address and redirect checks |
| `ignore.ts` | gitignore-aware walker, path jail |
| `complete.ts` | `@path` token extraction, ranking, insertion |
| `registry.ts` | external index, validation, install and removal |
@@ -255,11 +264,11 @@ Every module is pure of the UI except `ui/`, and `ui/` never touches the SDK. Th
## Testing
538 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop;
`ink-testing-library` drives the UI with real keystrokes; MCP is tested against a real stdio
server subprocess; provider wire formats and the registry are tested against a local HTTP
server; the interrupt path spawns a real subprocess and asserts it died early rather than ran
out.
538 tests became 647 as the suites grew; no mocking framework. `MockLanguageModelV4` from
`ai/test` drives the loop; `ink-testing-library` drives the UI with real keystrokes; MCP is
tested against a real stdio server subprocess; provider wire formats and the registry are
tested against a local HTTP server; the interrupt path spawns a real subprocess and asserts it
died early rather than ran out.
The pattern throughout is to assert on what actually crossed a boundary — what went on the
wire, what is on screen, what is on disk — rather than on internal calls.
+1 -1
View File
@@ -43,7 +43,7 @@ Written by `/provider`, editable by hand. Every field is optional.
| `thinking` | default level: `off`, `low`, `medium`, `high`, `max` |
| `maxRetries` | retries per model call for transient failures. Default 3 |
| `plugins` | which builtin plugins to enable. Omit for `["guard", "time"]` |
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`. Omit for all of them. See [tools](tools.md) |
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`, and `net`. Omit for the defaults; `net` is opt-in. See [tools](tools.md) |
| `permission` | which calls run, ask, or are refused, matched per command or path. See [permissions](permissions.md) |
| `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) |
| `mcpServers` | see [MCP](mcp.md) |
+4 -3
View File
@@ -16,7 +16,7 @@ faster and the fallback path is exercised without it.
```bash
bun run shiro # run from source
bun run typecheck # tsc --noEmit
bun test # 538 tests
bun test # 647 tests
bun run build # single binary for this platform -> dist/shiro
bun run release # all five platforms -> dist/release + SHA256SUMS
bun run install:local # build, then copy onto PATH
@@ -95,9 +95,10 @@ Steps 3 and 4 are two hand-maintained lists of tool names, which is a known weak
added to one and forgotten in the other is a silently ungated write. Deriving both from the
tool definitions is on [TODO.md](../TODO.md).
Every tool costs roughly 550 characters of schema on every request. Fourteen built-in tools is
well past where selection accuracy starts to matter, which is why sets exist and why a new tool
Every tool costs roughly 550 characters of schema on every request. Sixteen built-in tools is
past where selection accuracy starts to matter, which is why sets exist and why a new tool
needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why.
One set, `net`, is opt-in rather than on: `web_fetch` is the one tool that leaves the machine.
## Adding a slash command
+2 -2
View File
@@ -16,7 +16,7 @@ There is no terminal to approve on, so every gated tool is denied unless `--yolo
```
$ shiro -p "add a test for paginate()"
shiro: headless denies write_file, edit_file, multi_edit, bash and mcp tools unless --yolo is passed
shiro: headless denies write_file, edit_file, multi_edit, apply_patch, bash, web_fetch and mcp tools unless --yolo is passed
[tool] write_file {"path":"test/paginate.test.ts",...}
[denied] write_file (run with --yolo to allow tool use in headless mode)
```
@@ -51,7 +51,7 @@ $ shiro -p "count the tools" --json
{"type":"tool-start","id":"c1","name":"grep"}
{"type":"tool-call","id":"c1","name":"grep","input":{"pattern":"tool\\("}}
{"type":"tool-result","id":"c1","name":"grep","output":"src/tools.ts:26: ..."}
{"type":"text","text":"There are 14 built-in tools."}
{"type":"text","text":"There are 16 built-in tools."}
{"type":"done","inputTokens":4210,"outputTokens":88}
```
+10 -11
View File
@@ -4,7 +4,7 @@ Four kinds of state, each with a different lifetime.
| State | Lives in | Survives |
|---|---|---|
| transcript | the message array | until compaction or `/clear` |
| transcript | the message array | until `/compact` or `/clear` |
| task list | the system prompt, rebuilt each step | pruning and `/compact` |
| project memory | `~/.shiro-neko/memory/<hash>.json` | across sessions, forever |
| session record | `~/.shiro-neko/sessions/<uuid>.json` | until you delete it |
@@ -145,9 +145,10 @@ unless you pass it again.
Two mechanisms.
**Automatic**, at roughly 120k estimated tokens: `pruneMessages` strips reasoning and older
tool calls from what goes on the wire. Local history is untouched, so the transcript on your
screen stays complete. The turn reports it:
**Automatic**, at roughly 120k estimated tokens: reasoning is stripped first, then older tool
content is removed in a bounded ladder until the request fits. The SDK keeps that pruned view
for later steps in the turn; local session history remains complete. One `compacted` event is
reported per turn:
```
context compacted: 192 messages pruned to 15 on the wire
@@ -156,10 +157,8 @@ context compacted: 192 messages pruned to 15 on the wire
The status bar warns before that happens: context is shown as a percentage of the threshold,
amber from two thirds, red at 90.
What gets discarded, in order: reasoning items first, then tool calls and their results older
than the last three messages. Reasoning is the cheapest thing to lose — it was progress, not
conclusions — and tool results are the bulkiest. Recent exchanges are always kept, which is what
lets a turn continue rather than restart.
What gets discarded, in order: reasoning items first, then the oldest tool calls and results as
needed. Recent exchanges are kept by the ladder, which lets a turn continue rather than restart.
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
and outcomes, what remains — and it replaces the transcript entirely.
@@ -197,9 +196,9 @@ first compaction the model could no longer see what it had already run. It re-ra
tools until the step limit ended the turn. The history is the model's memory; compaction may
shorten it but must not blank it.
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts
*messages*, not pairs, so the cut can land between the assistant message holding a `tool-call`
and the `tool` message answering it:
**A tool result without its tool call.** Tool pruning counts messages, not call/result pairs, so
the cut can land between the assistant message holding a `tool-call` and the `tool` message
answering it:
```
400 No tool call found for function call output with call_id call_…
+4 -2
View File
@@ -34,6 +34,8 @@ remain are the ones worth reading.
|---|---|
| `bash` | the command, e.g. `git status --porcelain` |
| `read_file` `write_file` `edit_file` `multi_edit` `list_dir` | the path |
| `apply_patch` | every file marker path in the patch |
| `web_fetch` | the URL |
| `read_many_files` | every path in the batch; one match is enough |
| `glob` `grep` | the pattern |
| `git_diff` `git_log` `git_blame` | the path, when given |
@@ -96,7 +98,7 @@ With no `permission` config:
| `glob` `grep` `list_dir` | `allow` |
| the git tools | `allow` — they cannot mutate anything |
| `task`, and every session tool | `allow` — they touch the agent's own state |
| `write_file` `edit_file` `multi_edit` `bash` | `ask` |
| `write_file` `edit_file` `multi_edit` `apply_patch` `bash` `web_fetch` | `ask` |
| anything else, including every `mcp__*` tool | `ask` |
Credentials are denied on read rather than gated, because there is no recovery. A model that
@@ -228,7 +230,7 @@ unmatched and the tool on its default:
withholding the tools, which is stronger; use rules when you want the tools present but inert.
```json
{ "permission": { "write_file": "deny", "edit_file": "deny", "multi_edit": "deny", "bash": "deny" } }
{ "permission": { "write_file": "deny", "edit_file": "deny", "multi_edit": "deny", "apply_patch": "deny", "bash": "deny" } }
```
**An unattended job that may commit but never push.**
+2 -2
View File
@@ -125,7 +125,7 @@ export const noSecretsPlugin: Plugin = {
'The no-secrets plugin refuses writes to .env and credential files. Ask the user to ' +
'add secrets themselves rather than working around it.',
beforeToolCall: ({ toolName, input }) => {
if (toolName !== 'write_file' && toolName !== 'edit_file' && toolName !== 'multi_edit') return undefined;
if (!['write_file', 'edit_file', 'multi_edit', 'apply_patch'].includes(toolName)) return undefined;
const path = String((input as { path?: unknown } | null)?.path ?? '');
if (/(^|\/)\.env|credentials|\.pem$/.test(path)) {
return `refusing to write ${path}; add secrets yourself`;
@@ -137,7 +137,7 @@ export const noSecretsPlugin: Plugin = {
Then add it to `BUILTIN_PLUGINS` and, if it should be on by default, `DEFAULT_ENABLED`.
Note the three tool names. Every write tool has to be listed, and `multi_edit` is easy to miss
Note the four tool names. Every write tool has to be listed, and `multi_edit` is easy to miss
— a guard that only checks `write_file` and `edit_file` is bypassed by a batch edit.
Write the `appendix` whenever the plugin can block something. Without it the model hits a
+10 -2
View File
@@ -3,8 +3,8 @@
A skill is a markdown file with instructions for one kind of task. Only its name and
description sit in the system prompt; the body is loaded on demand.
That split matters. The four bundled skills are 5,284 characters of body against 681 characters
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt
That split matters. The six bundled skills are 8,900 characters of body against roughly 1,000
characters of catalogue — paid on every request. Putting every body in the prompt
would cost that on every turn, for instructions relevant to one turn in twenty.
## Format
@@ -69,6 +69,14 @@ each, do not fix bugs while refactoring, do not add abstraction for a single cal
implementation, never weaken an assertion to make a test pass, a flaky test is a shared-state
problem and not something to retry around.
**`verify`** — confirm a change works by running the artifact the way a user would, not by
reading the source. What counts as evidence, what to do with the failure path, and reporting
what was not verified.
**`commit`** — stage and commit work: look at the diff before staging, one commit one reason,
match the repository's message style, and the refusals — no amending pushed commits, no
`--no-verify`, no push unless asked.
They are string constants in `src/skills-builtin.ts` rather than files, because
`bun build --compile` only embeds modules reachable through imports. A directory of `.md`
files would be missing from the shipped binary.
+49 -20
View File
@@ -15,7 +15,8 @@ auto-approved.
reaches the context is on the wire and in the session file, and there is no taking it back.
`*.env.example` is allowed.
**Asked by default.** `write_file`, `edit_file`, `multi_edit`, `bash`, and every `mcp__*` tool.
**Asked by default.** `write_file`, `edit_file`, `multi_edit`, `apply_patch`, `bash`, `web_fetch`,
and every `mcp__*` tool.
```
bash wants to run
@@ -44,8 +45,9 @@ Three more things sit around the rules:
## Tool sets
Each tool costs its name, its description, and its JSON schema on **every request**. Measured
across the fourteen built-ins:
Each tool costs its name, its description, and its JSON schema on **every request**. The current
registry has sixteen built-ins. `/tools` shows the live set; disabling an optional set removes
its schemas from both the request and the system prompt.
| Tool | Bytes | Tool | Bytes |
|---|---|---|---|
@@ -57,24 +59,24 @@ across the fourteen built-ins:
| `read_file` | 526 | `git_status` | 292 |
| `glob` | 499 | `write_file` | 289 |
7,673 bytes for all fourteen, averaging 548. Roughly 1,900 tokens per request before your
prompt or the conversation. Selection accuracy also falls as the list grows: a model choosing
between six tools picks better than one choosing between twenty.
Selection accuracy also falls as the list grows: a model choosing between six tools picks better
than one choosing between twenty.
Sets let you switch off what a project does not need:
| Set | Tools | Cost |
|---|---|---|
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B |
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` | ~2,500 B |
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` `apply_patch` | patch included |
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` | ~2,180 B |
| `net` | `web_fetch` | opt in |
```json
{ "toolSets": ["edit-plus"] }
```
Omit `toolSets` for all of them. `core` is always on — without read, edit, and bash the
agent is not an agent. A disabled set reaches neither the wire nor the system prompt, since
Omit `toolSets` for the default sets. Add `net` when the agent should fetch public pages.
`core` is always on — without read, edit, and bash the agent is not an agent. A disabled set reaches neither the wire nor the system prompt, since
a prompt that names an absent tool teaches the model to attempt calls that cannot succeed.
Session, plugin, and MCP tools are not part of this budget and are never gated here.
@@ -190,6 +192,17 @@ edit 2: oldString not found in src/users.ts. No edits were applied.
The last sentence matters. Without it a model reading the error has to guess whether edit 1
landed, and its next move — retry the whole batch, or only what failed — depends on the answer.
### `apply_patch`
```
patch one envelope containing Add, Update, Move, and Delete file markers
```
All operations are validated before anything is written, so a failure leaves every file
unchanged. Use it when one change spans files that must land together; use `multi_edit` for
several edits to one file and `edit_file` for one edit. Paths stay inside the workspace and the
call asks for approval.
### `list_dir`
```
@@ -279,6 +292,18 @@ stdout:
The turn continues from there. `esc` still aborts everything, and `ctrl-c` with nothing
running quits as usual.
## `web_fetch`
```
url absolute HTTP(S) URL
maxChars returned characters, default 30,000, max 30,000
```
Fetches a public text page and converts HTML to markdown. HTTPS is required for public hosts;
private and loopback addresses are refused, redirects are checked one hop at a time, and the
body is capped. The result is untrusted page content, not an instruction, and the call asks for
approval. It belongs to the opt-in `net` set.
## Git tools
All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument
@@ -325,21 +350,25 @@ mutate anything and so never needs one.
```
description short label shown to you
prompt self-contained instructions
kind "explore" (default) or "review"
kind "explore" (default), "review", or "worker"
```
Spawns a read-only subagent with `read_file`, `glob`, and `grep` only. It returns one
report, so the parent pays for findings rather than the whole search transcript. It sees
none of the parent conversation, so its prompt has to stand alone.
Spawns a subagent with its own context window. It returns one report, so the parent pays for
the findings rather than the whole search transcript, and it sees none of the parent
conversation, so its prompt has to stand alone.
Two properties follow from that tool set rather than from policy: it can never trigger an
approval prompt, because it has no gated tools; and the parent's context holds the conclusion
instead of the search. A subagent reading forty files to answer one question costs the parent
the answer, not the forty files.
`explore` finds and reports; `review` critiques code in severity order. Both are structurally
read-only — they hold no gated tool at all, so they cannot trigger an approval prompt whatever
the config says. The parent's context holds the conclusion instead of the search: a subagent
reading forty files to answer one question costs the parent the answer, not the forty files.
`explore` finds and reports. `review` critiques code in severity order. Progress streams to
the subagent panel. Capped at 20 steps, and it shares the parent's model — an `explore` run
pays reasoning rates for what is really a search, which is [on the list](../TODO.md) to fix.
`worker` holds the write tools as well, and every write and command routes through the
parent's approval gate — the same rules and the same prompt as a direct call. The full
delegation trade-offs are in [agents](agents.md#delegating-with-task).
Progress streams to the subagent panel, with each call's outcome. Capped at 20 steps, and it
shares the parent's model — an `explore` run pays reasoning rates for what is really a search,
which is [on the list](../TODO.md) to fix.
Not worth delegating a single grep: the subagent is a whole extra model loop, so it wins on a
search spanning many files and loses on anything you could answer in one call.