Expand the documentation with measured figures and operational detail

Every guide catches up with the sixteen-tool registry: apply_patch and web_fetch sections, the delegation guide with the worker kind, bounded compaction replacing the three-message window description, net in the tool-set tables, six bundled skills, and the module map gains tools-net.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 16:43:27 +07:00
co-authored by Sisyphus
parent 55ebb40524
commit 0dc0819fad
11 changed files with 172 additions and 79 deletions
+32 -20
View File
@@ -49,9 +49,9 @@ from the models that endpoint actually reports. Settings land in
shiro-neko 0.1.0-beta.3 openai/gpt-5 session 0193ab2c shiro-neko 0.1.0-beta.3 openai/gpt-5 session 0193ab2c
agent: default thinking: medium agent: default thinking: medium
cwd: /home/you/project cwd: /home/you/project
skills: debug, refactor, review, test skills: commit, debug, refactor, review, test, verify
plugins: guard, time plugins: guard, time
approvals: ask for write_file, edit_file, multi_edit, bash, mcp__* approvals: ask for write_file, edit_file, multi_edit, apply_patch, bash, web_fetch, mcp__*
/help for commands /help for commands
> why does the pagination test fail? > why does the pagination test fail?
@@ -64,16 +64,19 @@ installed and honours `.gitignore`. `list_dir` gives an ignore-aware tree so it
blindly to orient, and `read_many_files` pulls a batch in one round trip. `read_file` refuses blindly to orient, and `read_many_files` pulls a batch in one round trip. `read_file` refuses
binaries rather than filling the context with mojibake. binaries rather than filling the context with mojibake.
**Edits with your approval, gated per command.** `write_file`, `edit_file`, `multi_edit`, and **Edits with your approval, gated per command.** `write_file`, `edit_file`, `multi_edit`,
`bash` stop for a `y`/`a`/`n` decision, with a coloured diff for edits. Rules match the command or `apply_patch`, and `bash` stop for a `y`/`a`/`n` decision, with a coloured diff for edits.
path rather than the tool, so `git *` can run unprompted while everything else still asks — `apply_patch` lands one atomic patch across files — add, update, move, delete — and nothing is
answering `a` whitelists that pattern, not the whole tool. `.env` and `.pem` files are refused on written if any part of it fails. Rules match the command or path rather than the tool, so
read outright. The `guard` plugin refuses irreversible commands ahead of any of it — `rm -rf`, `git *` can run unprompted while everything else still asks — answering `a` whitelists that
`git reset --hard`, force pushes, `DROP TABLE` — and `--yolo` cannot bypass it. pattern, not the whole tool. `.env` and `.pem` files are refused on read outright. The `guard`
plugin refuses irreversible commands ahead of any of it — `rm -rf`, `git reset --hard`, force
pushes, `DROP TABLE` — and `--yolo` cannot bypass it.
**Shows its work.** Reasoning streams to a collapsed panel you can expand with `ctrl-r`, the **Shows its work.** Reasoning streams to a collapsed panel you can expand with `ctrl-r`, the
tool in flight is named as it runs, and `bash` output streams live instead of arriving all at tool in flight is named as it runs with the arguments that identify the call, and `bash`
once when the command exits. `ctrl-c` kills a runaway command without ending the turn. output streams live instead of arriving all at once when the command exits. `ctrl-c` kills a
runaway command without ending the turn.
**Takes prompts while it works.** Type during a turn and it queues; the queue drains in order **Takes prompts while it works.** Type during a turn and it queues; the queue drains in order
when the turn ends. `esc` interrupts and clears it. `@` completes workspace paths. when the turn ends. `esc` interrupts and clears it. `@` completes workspace paths.
@@ -82,12 +85,18 @@ when the turn ends. `esc` interrupts and clears it. `@` completes workspace path
`git_blame` are approval-free, because they spawn git with a fixed argument list and cannot `git_blame` are approval-free, because they spawn git with a fixed argument list and cannot
mutate anything. mutate anything.
**Fetches docs when the codebase cannot answer.** `web_fetch` pulls a public page and returns
it as markdown — a changelog, an RFC, a migration guide — size-capped and stripped of anything
that is not text. It lives in the opt-in `net` tool set: the one tool that leaves the machine
is a decision rather than a default, and it asks before every call.
**Asks instead of guessing.** When a request has two readings that lead to different work, **Asks instead of guessing.** When a request has two readings that lead to different work,
the agent puts a question on screen with options. the agent puts a question on screen with options.
**Delegates searches.** `task` spawns a read-only subagent whose findings come back as one **Delegates work.** `task` spawns a subagent with its own context window whose findings come
message, so a search across forty files does not fill the main context. Its progress back as one message, so a search across forty files does not fill the main context. `explore`
streams to a panel. and `review` are read-only; `worker` also edits and runs commands, and every one of its writes
stops at the same approval prompt as yours. Progress streams to a panel.
**Extensible from the prompt.** `/registry` browses external skills and plugins and installs **Extensible from the prompt.** `/registry` browses external skills and plugins and installs
them with one confirmation. A skill is shown in full before its text joins your system prompt; them with one confirmation. A skill is shown in full before its text joins your system prompt;
@@ -97,11 +106,13 @@ a plugin is a manifest of refusal rules, never code.
memory that is injected at the start of every future session. memory that is injected at the start of every future session.
**Survives long tasks.** The task list and project memory live outside the message array, **Survives long tasks.** The task list and project memory live outside the message array,
so they survive both automatic pruning and `/compact`. so they survive both automatic pruning and `/compact`. Pruning itself is bounded: it drops
reasoning first and keeps the widest recent tool tail that fits, so the model keeps its
record of what it already ran instead of repeating it.
**Runs headless.** `shiro -p "review this diff" --json` for scripts and CI. **Runs headless.** `shiro -p "review this diff" --json` for scripts and CI.
**Keeps the tool list affordable.** Fourteen built-in tools, grouped into sets. Each costs **Keeps the tool list affordable.** Sixteen built-in tools, grouped into sets. Each costs
about 550 characters of schema on every request, so `{ "toolSets": [] }` trims back to the six about 550 characters of schema on every request, so `{ "toolSets": [] }` trims back to the six
core ones and a disabled set reaches neither the wire nor the prompt. core ones and a disabled set reaches neither the wire nor the prompt.
@@ -144,14 +155,15 @@ workspace path. Up and down recall earlier prompts.
## Status ## Status
Working: the agent loop, tool approvals, subagents, skills, plugins, per-project memory, Working: the agent loop, tool approvals, subagents including the gated `worker` kind, skills,
session persistence, MCP, markdown rendering, headless mode, five-platform builds, streaming plugins, per-project memory, session persistence, MCP, markdown rendering, headless mode,
reasoning display, the mid-turn prompt queue, gateable tool sets, read-only git tools, batch five-platform builds, streaming reasoning display, the mid-turn prompt queue, gateable tool
reads, `@file` completion, interruptible commands, and the external registry. sets, read-only git tools, batch reads, `apply_patch`, `web_fetch`, `@file` completion,
interruptible commands, and the external registry.
Next up is in [TODO.md](TODO.md); the longer view and what has been declined are in Next up is in [TODO.md](TODO.md); the longer view and what has been declined are in
[ROADMAP.md](ROADMAP.md). The short version of what is missing: a summary of what compaction [ROADMAP.md](ROADMAP.md). The short version of what is missing: a summary of what compaction
discarded, `web_fetch`, a spend ceiling, and a cheaper model for subagent searches. discarded, a spend ceiling, and a cheaper model for subagent searches.
## License ## License
+37 -4
View File
@@ -33,10 +33,10 @@ pure latency.
and findings recorded with `remember` so they survive compaction. and findings recorded with `remember` so they survive compaction.
**`plan`** and **`review`** are genuinely read-only. `write_file`, `edit_file`, `multi_edit`, **`plan`** and **`review`** are genuinely read-only. `write_file`, `edit_file`, `multi_edit`,
and `bash` are withheld from the model, not merely discouraged in prose — a model that cannot `apply_patch`, and `bash` are withheld from the model, not merely discouraged in prose — a
see a tool cannot call it. They keep everything that only reads, including `read_many_files`, model that cannot see a tool cannot call it. They keep everything that only reads, including
`list_dir`, and the git tools. Their prompts also forbid describing edits as if they had been `read_many_files`, `list_dir`, the git tools, and `web_fetch` when the `net` set is enabled.
made. Their prompts also forbid describing edits as if they had been made.
## Variants and tool sets ## Variants and tool sets
@@ -146,3 +146,36 @@ the full set — cheaper per turn as well as safer.
Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and
nothing about the history. nothing about the history.
## Delegating with `task`
The `task` tool spans a separate axis from the variants: it runs a subagent with its own
context window, so the parent pays for one report rather than the whole search transcript. The
subagent sees none of the parent's conversation, so its prompt must stand alone.
| Kind | Tools | Approval | For |
|---|---|---|---|
| `explore` (default) | read and search only | never prompts — structurally read-only | a search spanning many files |
| `review` | read and search only | never prompts | a critique of code or a diff |
| `worker` | everything, including writes | every write and command asks, through the parent's gate | a self-contained change whose steps you do not need to watch |
Three properties of the `worker` kind are structural rather than policy:
**The gate is the parent's.** A worker routes each gated call back through the same permission
rules, the same guard plugins, and the same approval prompt as a direct call — flagged `a
worker subagent wants to run ...` so you can tell who is asking. Answering `always` grants the
pattern for the session exactly as it does for you. A subagent that could approve its own
writes would be a way to launder a tool call past you, so there is no separate, weaker gate.
**Denial stops the work.** The worker is told a denial is your decision: report it, do not work
around it. The tool descriptions say the same thing, so the rule survives compaction.
**No `worker` without a channel.** In headless runs there is no one to answer a prompt, so the
`worker` kind is not offered at all — an unattended write is not something to fall into by
accident. The read-only kinds work everywhere. No subagent holds `web_fetch`; network access
stays with the main agent, where the approval prompt says what it is for.
When not to delegate: a single grep, or anything you must supervise step by step — keep that in
your own turn, where every call is on screen. A worker wins when the intermediate steps are
noise: a mechanical rename across twenty files, a test scaffold written to match an existing
suite, a cleanup whose shape you already know.
+21 -12
View File
@@ -145,12 +145,16 @@ running, the handler exits as usual.
## Subagents ## Subagents
`task` runs a nested `streamText` with only `read_file`, `glob`, and `grep`. It returns one `task` runs a nested `streamText` and returns one message. The subagent kinds hold different
message. tool sets: `explore` and `review` the read-only tools, `worker` those plus every write tool.
Two consequences follow from the tool set, not from policy: The consequences follow from the tool set, not from policy:
- It can never need approval, because it has no gated tools. - `explore` and `review` can never need approval, because they hold no gated tool.
- `worker` needs approval for exactly the calls a direct one would, so the parent owns the
gate: the subagent's `toolApproval` callback routes back through the parent's permission
rules, guard plugins, and prompt. A subagent with its own approval would be a way to launder
a tool call past the user.
- The parent's context holds the findings, not the search transcript. - The parent's context holds the findings, not the search transcript.
Progress is reported through a callback, wired to a bus the panel subscribes to. Without the Progress is reported through a callback, wired to a bus the panel subscribes to. Without the
@@ -196,9 +200,9 @@ tool call carries an itemId, so after the first compaction the model could not s
already run, and re-ran the same tools until the step limit ended the turn. **Compaction may already run, and re-ran the same tools until the step limit ended the turn. **Compaction may
shorten the history; it must not blank it.** shorten the history; it must not blank it.**
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts **A tool result without its tool call.** Tool pruning counts messages, so a cut can land between
*messages*, so the cut lands between an assistant `tool-call` and the `tool` message answering an assistant `tool-call` and the `tool` message answering it. What reaches the wire is a
it. What reaches the wire is a `function_call_output` with no `function_call`: `function_call_output` with no `function_call`:
``` ```
400 No tool call found for function call output with call_id call_… 400 No tool call found for function call output with call_id call_…
@@ -208,6 +212,10 @@ it. What reaches the wire is a `function_call_output` with no `function_call`:
reverse pairing is deliberately left alone: a call still awaiting its result is exactly what a reverse pairing is deliberately left alone: a call still awaiting its result is exactly what a
suspended approval looks like, and dropping it would break resume. suspended approval looks like, and dropping it would break resume.
The pruning ladder drops reasoning first and then keeps the widest recent tool tail that fits.
The SDK carries that returned message view into later steps, and the session reports compaction
once per turn rather than once per step.
## Registry ## Registry
`/registry` fetches an index of external skills and plugins over https. Skills are prompt text `/registry` fetches an index of external skills and plugins over https. Skills are prompt text
@@ -226,6 +234,7 @@ the reasoning.
| `session.ts` | the loop, approvals, compaction, event stream | | `session.ts` | the loop, approvals, compaction, event stream |
| `tools.ts` | file and shell tools, tool sets, ripgrep bridge, bash streaming and interrupt | | `tools.ts` | file and shell tools, tool sets, ripgrep bridge, bash streaming and interrupt |
| `tools-git.ts` | read-only git tools, spawned with a fixed argv | | `tools-git.ts` | read-only git tools, spawned with a fixed argv |
| `tools-net.ts` | `web_fetch`, private-address and redirect checks |
| `ignore.ts` | gitignore-aware walker, path jail | | `ignore.ts` | gitignore-aware walker, path jail |
| `complete.ts` | `@path` token extraction, ranking, insertion | | `complete.ts` | `@path` token extraction, ranking, insertion |
| `registry.ts` | external index, validation, install and removal | | `registry.ts` | external index, validation, install and removal |
@@ -255,11 +264,11 @@ Every module is pure of the UI except `ui/`, and `ui/` never touches the SDK. Th
## Testing ## Testing
538 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop; 538 tests became 647 as the suites grew; no mocking framework. `MockLanguageModelV4` from
`ink-testing-library` drives the UI with real keystrokes; MCP is tested against a real stdio `ai/test` drives the loop; `ink-testing-library` drives the UI with real keystrokes; MCP is
server subprocess; provider wire formats and the registry are tested against a local HTTP tested against a real stdio server subprocess; provider wire formats and the registry are
server; the interrupt path spawns a real subprocess and asserts it died early rather than ran tested against a local HTTP server; the interrupt path spawns a real subprocess and asserts it
out. died early rather than ran out.
The pattern throughout is to assert on what actually crossed a boundary — what went on the The pattern throughout is to assert on what actually crossed a boundary — what went on the
wire, what is on screen, what is on disk — rather than on internal calls. wire, what is on screen, what is on disk — rather than on internal calls.
+1 -1
View File
@@ -43,7 +43,7 @@ Written by `/provider`, editable by hand. Every field is optional.
| `thinking` | default level: `off`, `low`, `medium`, `high`, `max` | | `thinking` | default level: `off`, `low`, `medium`, `high`, `max` |
| `maxRetries` | retries per model call for transient failures. Default 3 | | `maxRetries` | retries per model call for transient failures. Default 3 |
| `plugins` | which builtin plugins to enable. Omit for `["guard", "time"]` | | `plugins` | which builtin plugins to enable. Omit for `["guard", "time"]` |
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`. Omit for all of them. See [tools](tools.md) | | `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`, and `net`. Omit for the defaults; `net` is opt-in. See [tools](tools.md) |
| `permission` | which calls run, ask, or are refused, matched per command or path. See [permissions](permissions.md) | | `permission` | which calls run, ask, or are refused, matched per command or path. See [permissions](permissions.md) |
| `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) | | `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) |
| `mcpServers` | see [MCP](mcp.md) | | `mcpServers` | see [MCP](mcp.md) |
+4 -3
View File
@@ -16,7 +16,7 @@ faster and the fallback path is exercised without it.
```bash ```bash
bun run shiro # run from source bun run shiro # run from source
bun run typecheck # tsc --noEmit bun run typecheck # tsc --noEmit
bun test # 538 tests bun test # 647 tests
bun run build # single binary for this platform -> dist/shiro bun run build # single binary for this platform -> dist/shiro
bun run release # all five platforms -> dist/release + SHA256SUMS bun run release # all five platforms -> dist/release + SHA256SUMS
bun run install:local # build, then copy onto PATH bun run install:local # build, then copy onto PATH
@@ -95,9 +95,10 @@ Steps 3 and 4 are two hand-maintained lists of tool names, which is a known weak
added to one and forgotten in the other is a silently ungated write. Deriving both from the added to one and forgotten in the other is a silently ungated write. Deriving both from the
tool definitions is on [TODO.md](../TODO.md). tool definitions is on [TODO.md](../TODO.md).
Every tool costs roughly 550 characters of schema on every request. Fourteen built-in tools is Every tool costs roughly 550 characters of schema on every request. Sixteen built-in tools is
well past where selection accuracy starts to matter, which is why sets exist and why a new tool past where selection accuracy starts to matter, which is why sets exist and why a new tool
needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why. needs to earn its place — see [ROADMAP.md](../ROADMAP.md) for what has been declined and why.
One set, `net`, is opt-in rather than on: `web_fetch` is the one tool that leaves the machine.
## Adding a slash command ## Adding a slash command
+2 -2
View File
@@ -16,7 +16,7 @@ There is no terminal to approve on, so every gated tool is denied unless `--yolo
``` ```
$ shiro -p "add a test for paginate()" $ shiro -p "add a test for paginate()"
shiro: headless denies write_file, edit_file, multi_edit, bash and mcp tools unless --yolo is passed shiro: headless denies write_file, edit_file, multi_edit, apply_patch, bash, web_fetch and mcp tools unless --yolo is passed
[tool] write_file {"path":"test/paginate.test.ts",...} [tool] write_file {"path":"test/paginate.test.ts",...}
[denied] write_file (run with --yolo to allow tool use in headless mode) [denied] write_file (run with --yolo to allow tool use in headless mode)
``` ```
@@ -51,7 +51,7 @@ $ shiro -p "count the tools" --json
{"type":"tool-start","id":"c1","name":"grep"} {"type":"tool-start","id":"c1","name":"grep"}
{"type":"tool-call","id":"c1","name":"grep","input":{"pattern":"tool\\("}} {"type":"tool-call","id":"c1","name":"grep","input":{"pattern":"tool\\("}}
{"type":"tool-result","id":"c1","name":"grep","output":"src/tools.ts:26: ..."} {"type":"tool-result","id":"c1","name":"grep","output":"src/tools.ts:26: ..."}
{"type":"text","text":"There are 14 built-in tools."} {"type":"text","text":"There are 16 built-in tools."}
{"type":"done","inputTokens":4210,"outputTokens":88} {"type":"done","inputTokens":4210,"outputTokens":88}
``` ```
+10 -11
View File
@@ -4,7 +4,7 @@ Four kinds of state, each with a different lifetime.
| State | Lives in | Survives | | State | Lives in | Survives |
|---|---|---| |---|---|---|
| transcript | the message array | until compaction or `/clear` | | transcript | the message array | until `/compact` or `/clear` |
| task list | the system prompt, rebuilt each step | pruning and `/compact` | | task list | the system prompt, rebuilt each step | pruning and `/compact` |
| project memory | `~/.shiro-neko/memory/<hash>.json` | across sessions, forever | | project memory | `~/.shiro-neko/memory/<hash>.json` | across sessions, forever |
| session record | `~/.shiro-neko/sessions/<uuid>.json` | until you delete it | | session record | `~/.shiro-neko/sessions/<uuid>.json` | until you delete it |
@@ -145,9 +145,10 @@ unless you pass it again.
Two mechanisms. Two mechanisms.
**Automatic**, at roughly 120k estimated tokens: `pruneMessages` strips reasoning and older **Automatic**, at roughly 120k estimated tokens: reasoning is stripped first, then older tool
tool calls from what goes on the wire. Local history is untouched, so the transcript on your content is removed in a bounded ladder until the request fits. The SDK keeps that pruned view
screen stays complete. The turn reports it: for later steps in the turn; local session history remains complete. One `compacted` event is
reported per turn:
``` ```
context compacted: 192 messages pruned to 15 on the wire context compacted: 192 messages pruned to 15 on the wire
@@ -156,10 +157,8 @@ context compacted: 192 messages pruned to 15 on the wire
The status bar warns before that happens: context is shown as a percentage of the threshold, The status bar warns before that happens: context is shown as a percentage of the threshold,
amber from two thirds, red at 90. amber from two thirds, red at 90.
What gets discarded, in order: reasoning items first, then tool calls and their results older What gets discarded, in order: reasoning items first, then the oldest tool calls and results as
than the last three messages. Reasoning is the cheapest thing to lose — it was progress, not needed. Recent exchanges are kept by the ladder, which lets a turn continue rather than restart.
conclusions — and tool results are the bulkiest. Recent exchanges are always kept, which is what
lets a turn continue rather than restart.
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands **Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
and outcomes, what remains — and it replaces the transcript entirely. and outcomes, what remains — and it replaces the transcript entirely.
@@ -197,9 +196,9 @@ first compaction the model could no longer see what it had already run. It re-ra
tools until the step limit ended the turn. The history is the model's memory; compaction may tools until the step limit ended the turn. The history is the model's memory; compaction may
shorten it but must not blank it. shorten it but must not blank it.
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts **A tool result without its tool call.** Tool pruning counts messages, not call/result pairs, so
*messages*, not pairs, so the cut can land between the assistant message holding a `tool-call` the cut can land between the assistant message holding a `tool-call` and the `tool` message
and the `tool` message answering it: answering it:
``` ```
400 No tool call found for function call output with call_id call_… 400 No tool call found for function call output with call_id call_…
+4 -2
View File
@@ -34,6 +34,8 @@ remain are the ones worth reading.
|---|---| |---|---|
| `bash` | the command, e.g. `git status --porcelain` | | `bash` | the command, e.g. `git status --porcelain` |
| `read_file` `write_file` `edit_file` `multi_edit` `list_dir` | the path | | `read_file` `write_file` `edit_file` `multi_edit` `list_dir` | the path |
| `apply_patch` | every file marker path in the patch |
| `web_fetch` | the URL |
| `read_many_files` | every path in the batch; one match is enough | | `read_many_files` | every path in the batch; one match is enough |
| `glob` `grep` | the pattern | | `glob` `grep` | the pattern |
| `git_diff` `git_log` `git_blame` | the path, when given | | `git_diff` `git_log` `git_blame` | the path, when given |
@@ -96,7 +98,7 @@ With no `permission` config:
| `glob` `grep` `list_dir` | `allow` | | `glob` `grep` `list_dir` | `allow` |
| the git tools | `allow` — they cannot mutate anything | | the git tools | `allow` — they cannot mutate anything |
| `task`, and every session tool | `allow` — they touch the agent's own state | | `task`, and every session tool | `allow` — they touch the agent's own state |
| `write_file` `edit_file` `multi_edit` `bash` | `ask` | | `write_file` `edit_file` `multi_edit` `apply_patch` `bash` `web_fetch` | `ask` |
| anything else, including every `mcp__*` tool | `ask` | | anything else, including every `mcp__*` tool | `ask` |
Credentials are denied on read rather than gated, because there is no recovery. A model that Credentials are denied on read rather than gated, because there is no recovery. A model that
@@ -228,7 +230,7 @@ unmatched and the tool on its default:
withholding the tools, which is stronger; use rules when you want the tools present but inert. withholding the tools, which is stronger; use rules when you want the tools present but inert.
```json ```json
{ "permission": { "write_file": "deny", "edit_file": "deny", "multi_edit": "deny", "bash": "deny" } } { "permission": { "write_file": "deny", "edit_file": "deny", "multi_edit": "deny", "apply_patch": "deny", "bash": "deny" } }
``` ```
**An unattended job that may commit but never push.** **An unattended job that may commit but never push.**
+2 -2
View File
@@ -125,7 +125,7 @@ export const noSecretsPlugin: Plugin = {
'The no-secrets plugin refuses writes to .env and credential files. Ask the user to ' + 'The no-secrets plugin refuses writes to .env and credential files. Ask the user to ' +
'add secrets themselves rather than working around it.', 'add secrets themselves rather than working around it.',
beforeToolCall: ({ toolName, input }) => { beforeToolCall: ({ toolName, input }) => {
if (toolName !== 'write_file' && toolName !== 'edit_file' && toolName !== 'multi_edit') return undefined; if (!['write_file', 'edit_file', 'multi_edit', 'apply_patch'].includes(toolName)) return undefined;
const path = String((input as { path?: unknown } | null)?.path ?? ''); const path = String((input as { path?: unknown } | null)?.path ?? '');
if (/(^|\/)\.env|credentials|\.pem$/.test(path)) { if (/(^|\/)\.env|credentials|\.pem$/.test(path)) {
return `refusing to write ${path}; add secrets yourself`; return `refusing to write ${path}; add secrets yourself`;
@@ -137,7 +137,7 @@ export const noSecretsPlugin: Plugin = {
Then add it to `BUILTIN_PLUGINS` and, if it should be on by default, `DEFAULT_ENABLED`. Then add it to `BUILTIN_PLUGINS` and, if it should be on by default, `DEFAULT_ENABLED`.
Note the three tool names. Every write tool has to be listed, and `multi_edit` is easy to miss Note the four tool names. Every write tool has to be listed, and `multi_edit` is easy to miss
— a guard that only checks `write_file` and `edit_file` is bypassed by a batch edit. — a guard that only checks `write_file` and `edit_file` is bypassed by a batch edit.
Write the `appendix` whenever the plugin can block something. Without it the model hits a Write the `appendix` whenever the plugin can block something. Without it the model hits a
+10 -2
View File
@@ -3,8 +3,8 @@
A skill is a markdown file with instructions for one kind of task. Only its name and A skill is a markdown file with instructions for one kind of task. Only its name and
description sit in the system prompt; the body is loaded on demand. description sit in the system prompt; the body is loaded on demand.
That split matters. The four bundled skills are 5,284 characters of body against 681 characters That split matters. The six bundled skills are 8,900 characters of body against roughly 1,000
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt characters of catalogue — paid on every request. Putting every body in the prompt
would cost that on every turn, for instructions relevant to one turn in twenty. would cost that on every turn, for instructions relevant to one turn in twenty.
## Format ## Format
@@ -69,6 +69,14 @@ each, do not fix bugs while refactoring, do not add abstraction for a single cal
implementation, never weaken an assertion to make a test pass, a flaky test is a shared-state implementation, never weaken an assertion to make a test pass, a flaky test is a shared-state
problem and not something to retry around. problem and not something to retry around.
**`verify`** — confirm a change works by running the artifact the way a user would, not by
reading the source. What counts as evidence, what to do with the failure path, and reporting
what was not verified.
**`commit`** — stage and commit work: look at the diff before staging, one commit one reason,
match the repository's message style, and the refusals — no amending pushed commits, no
`--no-verify`, no push unless asked.
They are string constants in `src/skills-builtin.ts` rather than files, because They are string constants in `src/skills-builtin.ts` rather than files, because
`bun build --compile` only embeds modules reachable through imports. A directory of `.md` `bun build --compile` only embeds modules reachable through imports. A directory of `.md`
files would be missing from the shipped binary. files would be missing from the shipped binary.
+49 -20
View File
@@ -15,7 +15,8 @@ auto-approved.
reaches the context is on the wire and in the session file, and there is no taking it back. reaches the context is on the wire and in the session file, and there is no taking it back.
`*.env.example` is allowed. `*.env.example` is allowed.
**Asked by default.** `write_file`, `edit_file`, `multi_edit`, `bash`, and every `mcp__*` tool. **Asked by default.** `write_file`, `edit_file`, `multi_edit`, `apply_patch`, `bash`, `web_fetch`,
and every `mcp__*` tool.
``` ```
bash wants to run bash wants to run
@@ -44,8 +45,9 @@ Three more things sit around the rules:
## Tool sets ## Tool sets
Each tool costs its name, its description, and its JSON schema on **every request**. Measured Each tool costs its name, its description, and its JSON schema on **every request**. The current
across the fourteen built-ins: registry has sixteen built-ins. `/tools` shows the live set; disabling an optional set removes
its schemas from both the request and the system prompt.
| Tool | Bytes | Tool | Bytes | | Tool | Bytes | Tool | Bytes |
|---|---|---|---| |---|---|---|---|
@@ -57,24 +59,24 @@ across the fourteen built-ins:
| `read_file` | 526 | `git_status` | 292 | | `read_file` | 526 | `git_status` | 292 |
| `glob` | 499 | `write_file` | 289 | | `glob` | 499 | `write_file` | 289 |
7,673 bytes for all fourteen, averaging 548. Roughly 1,900 tokens per request before your Selection accuracy also falls as the list grows: a model choosing between six tools picks better
prompt or the conversation. Selection accuracy also falls as the list grows: a model choosing than one choosing between twenty.
between six tools picks better than one choosing between twenty.
Sets let you switch off what a project does not need: Sets let you switch off what a project does not need:
| Set | Tools | Cost | | Set | Tools | Cost |
|---|---|---| |---|---|---|
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B | | `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B |
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` | ~2,500 B | | `edit-plus` | `multi_edit` `list_dir` `read_many_files` `apply_patch` | patch included |
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` | ~2,180 B | | `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` | ~2,180 B |
| `net` | `web_fetch` | opt in |
```json ```json
{ "toolSets": ["edit-plus"] } { "toolSets": ["edit-plus"] }
``` ```
Omit `toolSets` for all of them. `core` is always on — without read, edit, and bash the Omit `toolSets` for the default sets. Add `net` when the agent should fetch public pages.
agent is not an agent. A disabled set reaches neither the wire nor the system prompt, since `core` is always on — without read, edit, and bash the agent is not an agent. A disabled set reaches neither the wire nor the system prompt, since
a prompt that names an absent tool teaches the model to attempt calls that cannot succeed. a prompt that names an absent tool teaches the model to attempt calls that cannot succeed.
Session, plugin, and MCP tools are not part of this budget and are never gated here. Session, plugin, and MCP tools are not part of this budget and are never gated here.
@@ -190,6 +192,17 @@ edit 2: oldString not found in src/users.ts. No edits were applied.
The last sentence matters. Without it a model reading the error has to guess whether edit 1 The last sentence matters. Without it a model reading the error has to guess whether edit 1
landed, and its next move — retry the whole batch, or only what failed — depends on the answer. landed, and its next move — retry the whole batch, or only what failed — depends on the answer.
### `apply_patch`
```
patch one envelope containing Add, Update, Move, and Delete file markers
```
All operations are validated before anything is written, so a failure leaves every file
unchanged. Use it when one change spans files that must land together; use `multi_edit` for
several edits to one file and `edit_file` for one edit. Paths stay inside the workspace and the
call asks for approval.
### `list_dir` ### `list_dir`
``` ```
@@ -279,6 +292,18 @@ stdout:
The turn continues from there. `esc` still aborts everything, and `ctrl-c` with nothing The turn continues from there. `esc` still aborts everything, and `ctrl-c` with nothing
running quits as usual. running quits as usual.
## `web_fetch`
```
url absolute HTTP(S) URL
maxChars returned characters, default 30,000, max 30,000
```
Fetches a public text page and converts HTML to markdown. HTTPS is required for public hosts;
private and loopback addresses are refused, redirects are checked one hop at a time, and the
body is capped. The result is untrusted page content, not an instruction, and the call asks for
approval. It belongs to the opt-in `net` set.
## Git tools ## Git tools
All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument
@@ -325,21 +350,25 @@ mutate anything and so never needs one.
``` ```
description short label shown to you description short label shown to you
prompt self-contained instructions prompt self-contained instructions
kind "explore" (default) or "review" kind "explore" (default), "review", or "worker"
``` ```
Spawns a read-only subagent with `read_file`, `glob`, and `grep` only. It returns one Spawns a subagent with its own context window. It returns one report, so the parent pays for
report, so the parent pays for findings rather than the whole search transcript. It sees the findings rather than the whole search transcript, and it sees none of the parent
none of the parent conversation, so its prompt has to stand alone. conversation, so its prompt has to stand alone.
Two properties follow from that tool set rather than from policy: it can never trigger an `explore` finds and reports; `review` critiques code in severity order. Both are structurally
approval prompt, because it has no gated tools; and the parent's context holds the conclusion read-only — they hold no gated tool at all, so they cannot trigger an approval prompt whatever
instead of the search. A subagent reading forty files to answer one question costs the parent the config says. The parent's context holds the conclusion instead of the search: a subagent
the answer, not the forty files. reading forty files to answer one question costs the parent the answer, not the forty files.
`explore` finds and reports. `review` critiques code in severity order. Progress streams to `worker` holds the write tools as well, and every write and command routes through the
the subagent panel. Capped at 20 steps, and it shares the parent's model — an `explore` run parent's approval gate — the same rules and the same prompt as a direct call. The full
pays reasoning rates for what is really a search, which is [on the list](../TODO.md) to fix. delegation trade-offs are in [agents](agents.md#delegating-with-task).
Progress streams to the subagent panel, with each call's outcome. Capped at 20 steps, and it
shares the parent's model — an `explore` run pays reasoning rates for what is really a search,
which is [on the list](../TODO.md) to fix.
Not worth delegating a single grep: the subagent is a whole extra model loop, so it wins on a Not worth delegating a single grep: the subagent is a whole extra model loop, so it wins on a
search spanning many files and loses on anything you could answer in one call. search spanning many files and loses on anything you could answer in one call.