Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual per-tool measurements, and fills in the parts a reader hits after the happy path: what a specific error means, what a setting costs, what is not covered. Measured rather than estimated: - Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the full set, averaging 548. - Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the argument for loading bodies on demand. - Full system prompt 3,571 chars, core-only 2,045. New sections: - tools: which sets to keep and why, the jail function itself, an output-cap table, and the real error strings for edit_file and multi_edit. - configuration: env var per provider preset, cost-estimate limits, what each --no-* flag isolates, and three settings that do more than they look like. - agents: step caps per variant, which variant to reach for, and the fact that reasoning is charged as output and discarded first by compaction. - headless: exit code 0 means "the turn completed", not "the answer was yes" — with the jq pattern for gating on content. Timeouts, concurrent -c runs fighting over one session, CI recipes for --no-skills. - mcp: parallel connect, startup cost, a debugging ladder, and that toolSets does not gate MCP tools. - registry: publishing, local testing over http://localhost, and a troubleshooting section keyed on the actual validator messages. - memory: what compaction discards in what order, /compact versus automatic pruning, and that -c matches on cwd. - skills: the frontmatter reader's limits, and how to verify a skill loaded. Corrections found while cross-checking against the source: - The guard table was missing --force-with-lease and > /dev/sd… - The done event's token fields are optional, so the jq example filters on one rather than assuming it. Two honest limits now written down: the guard matches command strings, so a base64-decoded or script-wrapped command is not caught; and a registry index is trusted for its contents, not its authorship. Verified: all internal links and heading anchors resolve, every docs/ page is reachable from the README, 538 tests pass, typecheck clean.
This commit is contained in:
@@ -106,21 +106,24 @@ core ones and a disabled set reaches neither the wire nor the prompt.
|
|||||||
|
|
||||||
## Documentation
|
## Documentation
|
||||||
|
|
||||||
|
Start with whichever question you have. Each guide says what it decided and why, not just what
|
||||||
|
the flags are.
|
||||||
|
|
||||||
| Guide | Contents |
|
| Guide | Contents |
|
||||||
|---|---|
|
|---|---|
|
||||||
| [Configuration](docs/configuration.md) | config file, environment variables, every flag |
|
| [Configuration](docs/configuration.md) | config file, provider presets, environment, every flag |
|
||||||
| [Tools](docs/tools.md) | every tool, tool sets, the approval model, the guard |
|
| [Tools](docs/tools.md) | every tool, tool sets and what they cost, the approval model |
|
||||||
| [Agents and thinking](docs/agents.md) | variants, thinking levels, read-only modes |
|
| [Agents and thinking](docs/agents.md) | variants, thinking levels, step caps, which to reach for |
|
||||||
| [Skills](docs/skills.md) | the bundled skills and writing your own |
|
| [Skills](docs/skills.md) | the bundled skills, writing your own, why the catalogue is split |
|
||||||
| [Plugins](docs/plugins.md) | the plugin interface and the builtins |
|
| [Plugins](docs/plugins.md) | the interface, the guard and its limits, builtin versus installed |
|
||||||
| [Registry](docs/registry.md) | installing external skills and plugins |
|
| [Registry](docs/registry.md) | installing external skills and plugins, publishing your own |
|
||||||
| [Memory and state](docs/memory.md) | memory, task lists, sessions, compaction |
|
| [Memory and state](docs/memory.md) | memory, task lists, sessions, compaction and its repair |
|
||||||
| [MCP](docs/mcp.md) | connecting Model Context Protocol servers |
|
| [MCP](docs/mcp.md) | connecting servers, namespacing, cost, debugging one |
|
||||||
| [Headless mode](docs/headless.md) | `-p`, JSON events, exit codes, CI recipes |
|
| [Headless mode](docs/headless.md) | `-p`, JSON events, exit codes, CI recipes |
|
||||||
| [Architecture](docs/architecture.md) | how the loop works and why it is built this way |
|
| [Architecture](docs/architecture.md) | how the loop works and why it is built this way |
|
||||||
| [Development](docs/development.md) | building, testing, releasing |
|
| [Development](docs/development.md) | building, testing, adding a tool, releasing |
|
||||||
| [Roadmap](ROADMAP.md) | what is next and what has been declined |
|
| [Roadmap](ROADMAP.md) | what is next and what has been declined |
|
||||||
| [TODO](TODO.md) | the current work list |
|
| [TODO](TODO.md) | the current work list, with known rough edges |
|
||||||
|
|
||||||
## Commands
|
## Commands
|
||||||
|
|
||||||
|
|||||||
+50
-1
@@ -62,12 +62,35 @@ from.
|
|||||||
| `high` | `high` | `budget_tokens: 38400` |
|
| `high` | `high` | `budget_tokens: 38400` |
|
||||||
| `max` | `xhigh` | maximum budget |
|
| `max` | `xhigh` | maximum budget |
|
||||||
|
|
||||||
Verified against both wire formats rather than assumed.
|
Verified against both wire formats rather than assumed. The vocabulary is deliberately ours:
|
||||||
|
`off` through `max` means the same thing whichever provider is configured, and switching
|
||||||
|
providers mid-session does not change what `/think high` asks for.
|
||||||
|
|
||||||
Higher costs more and takes longer. `off` on a hard problem produces confident wrong
|
Higher costs more and takes longer. `off` on a hard problem produces confident wrong
|
||||||
answers; `max` on a rename wastes a few cents and several seconds. The variants pick
|
answers; `max` on a rename wastes a few cents and several seconds. The variants pick
|
||||||
sensible defaults, so reach for `/think` only when a specific turn needs something else.
|
sensible defaults, so reach for `/think` only when a specific turn needs something else.
|
||||||
|
|
||||||
|
Reasoning is also charged as output tokens, so `max` shows up in `/cost` even on a turn where
|
||||||
|
the model wrote two lines. And reasoning is the **first thing compaction discards** — see
|
||||||
|
[memory](memory.md#compaction) — so a long turn at `max` pays for thinking that will not be on
|
||||||
|
the wire by the end of it.
|
||||||
|
|
||||||
|
## Steps
|
||||||
|
|
||||||
|
`maxSteps` caps how many model calls one turn may make. A step is one request: a tool call and
|
||||||
|
its result, or the final text.
|
||||||
|
|
||||||
|
| Variant | Steps |
|
||||||
|
|---|---|
|
||||||
|
| `quick` | 12 |
|
||||||
|
| `default`, `plan`, `review` | 50 |
|
||||||
|
| `deep` | 80 |
|
||||||
|
|
||||||
|
The cap is a backstop against a loop, not a budget to spend. A turn that hits it stops
|
||||||
|
mid-work with whatever it has, which is why `quick`'s 12 suits a rename and would strand a
|
||||||
|
refactor. If turns regularly hit the cap on the same kind of task, the task wants `deep`
|
||||||
|
rather than a higher number.
|
||||||
|
|
||||||
## Overriding
|
## Overriding
|
||||||
|
|
||||||
`--agent deep --think low` gives you `deep`'s tools, steps, and appendix with a low thinking
|
`--agent deep --think low` gives you `deep`'s tools, steps, and appendix with a low thinking
|
||||||
@@ -97,3 +120,29 @@ told which tools need approval and to verify with the project's tests.
|
|||||||
|
|
||||||
A prompt that describes a withheld tool teaches the model to attempt calls that cannot
|
A prompt that describes a withheld tool teaches the model to attempt calls that cannot
|
||||||
succeed, which is why the description is generated from the live tool set.
|
succeed, which is why the description is generated from the live tool set.
|
||||||
|
|
||||||
|
Three rules flip on what is available:
|
||||||
|
|
||||||
|
| Condition | `default` says | `plan` says |
|
||||||
|
|---|---|---|
|
||||||
|
| can edit | "these need approval; if denied, stop and ask" | "you have no tools that change anything" |
|
||||||
|
| can run commands | "verify with the project's build or tests" | "say what should be run rather than claiming it passed" |
|
||||||
|
| can ask | "ask rather than guess when two readings differ" | (same, unless headless) |
|
||||||
|
|
||||||
|
The read-only variants are around 2,000 characters of system prompt against roughly 3,600 for
|
||||||
|
the full set — cheaper per turn as well as safer.
|
||||||
|
|
||||||
|
## Which to reach for
|
||||||
|
|
||||||
|
- **`default`** for anything you have not thought about. It is the right answer most of the time.
|
||||||
|
- **`quick`** for a rename, a typo, a one-line fix. Its value is not the model being cheaper but
|
||||||
|
the absence of deliberation latency on work that needs none.
|
||||||
|
- **`deep`** when the first attempt already failed, or the cause is unclear. Asking for more than
|
||||||
|
one hypothesis is the actual difference; the thinking budget is secondary.
|
||||||
|
- **`plan`** before a change you are not sure about. Read-only means the plan cannot quietly
|
||||||
|
become a half-applied edit.
|
||||||
|
- **`review`** on a diff or a module. In headless CI this is the one that needs no `--yolo`,
|
||||||
|
because it holds no tool that can modify anything — see [headless](headless.md).
|
||||||
|
|
||||||
|
Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and
|
||||||
|
nothing about the history.
|
||||||
|
|||||||
+100
-19
@@ -48,21 +48,48 @@ Written by `/provider`, editable by hand. Every field is optional.
|
|||||||
|
|
||||||
`/provider` offers these. Each sets `baseURL` and the wire protocol for you.
|
`/provider` offers these. Each sets `baseURL` and the wire protocol for you.
|
||||||
|
|
||||||
| Preset | Protocol | Endpoint |
|
| Preset | Protocol | Endpoint | Env var checked |
|
||||||
|---|---|---|
|
|---|---|---|---|
|
||||||
| Anthropic | `anthropic` | `api.anthropic.com/v1` |
|
| Anthropic | `anthropic` | `api.anthropic.com/v1` | `ANTHROPIC_API_KEY` |
|
||||||
| OpenAI | `openai` | `api.openai.com/v1` |
|
| OpenAI | `openai` | `api.openai.com/v1` | `OPENAI_API_KEY` |
|
||||||
| OpenRouter | `openai` | `openrouter.ai/api/v1` |
|
| OpenRouter | `openai` | `openrouter.ai/api/v1` | `OPENROUTER_API_KEY` |
|
||||||
| Groq | `openai` | `api.groq.com/openai/v1` |
|
| Groq | `openai` | `api.groq.com/openai/v1` | `GROQ_API_KEY` |
|
||||||
| DeepSeek | `openai` | `api.deepseek.com/v1` |
|
| DeepSeek | `openai` | `api.deepseek.com/v1` | `DEEPSEEK_API_KEY` |
|
||||||
| xAI | `openai` | `api.x.ai/v1` |
|
| xAI | `openai` | `api.x.ai/v1` | `XAI_API_KEY` |
|
||||||
| Ollama | `openai` | `localhost:11434/v1` |
|
| Ollama | `openai` | `localhost:11434/v1` | none, keyless |
|
||||||
| LM Studio | `openai` | `localhost:1234/v1` |
|
| LM Studio | `openai` | `localhost:1234/v1` | none, keyless |
|
||||||
| Custom OpenAI-compatible | `openai` | you supply it |
|
| Custom OpenAI-compatible | `openai` | you supply it | none |
|
||||||
| Custom Anthropic-compatible | `anthropic` | you supply it |
|
| Custom Anthropic-compatible | `anthropic` | you supply it | none |
|
||||||
|
|
||||||
After the key is entered, `GET /v1/models` is called and the list becomes a picker. If the
|
`provider` is the **wire protocol**, not the vendor. Groq, DeepSeek, xAI, OpenRouter, Ollama,
|
||||||
endpoint does not implement it, you type the model id instead — the setup still completes.
|
and LM Studio all speak `openai`; only Anthropic speaks `anthropic`. Two things differ between
|
||||||
|
them: the auth header (`Authorization: Bearer` versus `x-api-key`), and how thinking levels map.
|
||||||
|
|
||||||
|
After the key is entered, `GET /v1/models` is called and the list becomes a picker. Both
|
||||||
|
protocols expose that endpoint with the same `data[].id` shape, so one code path handles both.
|
||||||
|
If the endpoint does not implement it, a preset with a known model list falls back to that;
|
||||||
|
otherwise you type the model id and setup still completes.
|
||||||
|
|
||||||
|
Anything the picker offers is a model the endpoint actually reports, which is more reliable than
|
||||||
|
a hard-coded list — that is why the fallback lists are short and only exist for Anthropic and
|
||||||
|
OpenAI.
|
||||||
|
|
||||||
|
## Cost estimates
|
||||||
|
|
||||||
|
`/cost` and the status bar price a turn from a table in `src/pricing.ts`, matched by longest
|
||||||
|
prefix on the model id, so `claude-sonnet-4-5-20250929` resolves via `claude-sonnet-4-5`. An
|
||||||
|
OpenRouter-style `anthropic/claude-sonnet-4-5` has its vendor prefix stripped first.
|
||||||
|
|
||||||
|
An unknown model is reported as unpriced rather than guessed:
|
||||||
|
|
||||||
|
```
|
||||||
|
4210 in / 88 out tokens (llama-3.3-70b is unpriced)
|
||||||
|
```
|
||||||
|
|
||||||
|
Two limits worth knowing. The rates are hand-entered and drift as vendors change them, so treat
|
||||||
|
the figure as an estimate, not a bill. And the token counts come from the provider's usage
|
||||||
|
report, while `~ctx` in the status bar is `JSON.stringify(messages).length / 4` — good enough to
|
||||||
|
decide when to compact, wrong enough that it should not be read as a token count.
|
||||||
|
|
||||||
## Environment variables
|
## Environment variables
|
||||||
|
|
||||||
@@ -74,12 +101,21 @@ endpoint does not implement it, you type the model id instead — the setup stil
|
|||||||
| `SHIRO_API_KEY` | overrides `apiKey` |
|
| `SHIRO_API_KEY` | overrides `apiKey` |
|
||||||
| `ANTHROPIC_API_KEY` | used when `provider` is `anthropic` and no key is set |
|
| `ANTHROPIC_API_KEY` | used when `provider` is `anthropic` and no key is set |
|
||||||
| `OPENAI_API_KEY` | used when `provider` is `openai` and no key is set |
|
| `OPENAI_API_KEY` | used when `provider` is `openai` and no key is set |
|
||||||
| `SHIRO_HOME` | relocates config, sessions, memory, history, and user skills |
|
| `SHIRO_HOME` | relocates config, sessions, memory, history, user skills, and installs |
|
||||||
| `SHIRO_INSTALL_DIR` | where `install:local` and the installers put the binary |
|
| `SHIRO_INSTALL_DIR` | where `install:local` and the installers put the binary |
|
||||||
| `SHIRO_REPO` | which GitHub repo the installers download from |
|
| `SHIRO_REPO` | which GitHub repo the installers download from |
|
||||||
| `SHIRO_VERSION` | pins the version the installers fetch |
|
| `SHIRO_VERSION` | pins the version the installers fetch |
|
||||||
|
|
||||||
`SHIRO_HOME` is what the test suite uses to keep a run out of your real config.
|
`SHIRO_HOME` is what the test suite uses to keep a run out of your real config. It is also the
|
||||||
|
way to run two isolated setups side by side — a work profile and a personal one — since it moves
|
||||||
|
every piece of state at once:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
SHIRO_HOME=~/work-shiro shiro
|
||||||
|
```
|
||||||
|
|
||||||
|
A key on the command line ends up in your shell history and in `ps`. `SHIRO_API_KEY` in front of
|
||||||
|
one command is better; `/provider` writing to `config.json` is better still.
|
||||||
|
|
||||||
## Flags
|
## Flags
|
||||||
|
|
||||||
@@ -103,13 +139,31 @@ cat file | shiro -p prompt read from stdin
|
|||||||
| `--no-mcp` | skip MCP servers |
|
| `--no-mcp` | skip MCP servers |
|
||||||
| `--no-subagent` | omit the `task` tool |
|
| `--no-subagent` | omit the `task` tool |
|
||||||
| `--no-instructions` | ignore `AGENTS.md` and friends |
|
| `--no-instructions` | ignore `AGENTS.md` and friends |
|
||||||
| `--no-skills` | ignore builtin and project skills |
|
| `--no-skills` | ignore builtin, installed, and project skills |
|
||||||
| `--no-plugins` | disable all plugins, including the guard |
|
| `--no-plugins` | disable all plugins, builtin and installed, including the guard |
|
||||||
| `--no-memory` | do not load or write project memory |
|
| `--no-memory` | do not load or write project memory |
|
||||||
| `--yolo` | skip every approval prompt |
|
| `--yolo` | skip every approval prompt |
|
||||||
| `-v`, `--version` | version, bun version, platform, source or compiled |
|
| `-v`, `--version` | version, bun version, platform, source or compiled |
|
||||||
| `-h`, `--help` | usage |
|
| `-h`, `--help` | usage |
|
||||||
|
|
||||||
|
The `--no-*` flags exist for isolating a problem. All six together strip the agent to its
|
||||||
|
built-in tools and nothing else, which answers "is this the loop or something layered on it?"
|
||||||
|
in one run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
shiro --no-plugins --no-skills --no-memory --no-instructions --no-subagent --no-mcp
|
||||||
|
```
|
||||||
|
|
||||||
|
`--no-plugins` also disables the guard, so `rm -rf` becomes an ordinary approval prompt.
|
||||||
|
Reasonable while debugging, not something to leave on.
|
||||||
|
|
||||||
|
An unknown value fails at startup with the valid list rather than falling back silently:
|
||||||
|
|
||||||
|
```
|
||||||
|
$ shiro --agent turbo
|
||||||
|
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review
|
||||||
|
```
|
||||||
|
|
||||||
## Where things live
|
## Where things live
|
||||||
|
|
||||||
```
|
```
|
||||||
@@ -133,7 +187,8 @@ Project files:
|
|||||||
```
|
```
|
||||||
|
|
||||||
Memory and history file names are SHA-256 prefixes of the absolute project path, because a
|
Memory and history file names are SHA-256 prefixes of the absolute project path, because a
|
||||||
path is not a safe filename.
|
path is not a safe filename. Two consequences: moving a project loses its memory and history,
|
||||||
|
and two checkouts of the same repo at different paths keep separate ones.
|
||||||
|
|
||||||
## OpenAI reasoning models
|
## OpenAI reasoning models
|
||||||
|
|
||||||
@@ -142,5 +197,31 @@ Newer OpenAI models reject function tools on `/v1/chat/completions` and require
|
|||||||
on the first switches to the second, sticks for the rest of the session, and prints one
|
on the first switches to the second, sticks for the rest of the session, and prints one
|
||||||
notice. Retryable failures — 429 and 5xx — are left to the SDK's backoff instead.
|
notice. Retryable failures — 429 and 5xx — are left to the SDK's backoff instead.
|
||||||
|
|
||||||
|
Only those six codes qualify, because they mean "this endpoint cannot serve this request shape".
|
||||||
|
A 401 is a wrong key and switching endpoints would only produce a second 401 with a more
|
||||||
|
confusing message.
|
||||||
|
|
||||||
|
The switch is sticky on purpose: once an endpoint rejects the shape it rejects every later step
|
||||||
|
too, so re-probing would waste a round trip per step of every turn.
|
||||||
|
|
||||||
Third-party endpoints get a plain chat-completions model with no fallback probe, since they
|
Third-party endpoints get a plain chat-completions model with no fallback probe, since they
|
||||||
do not implement `/v1/responses`.
|
do not implement `/v1/responses`.
|
||||||
|
|
||||||
|
The two endpoints also differ in how they carry assistant history, which is where compaction gets
|
||||||
|
interesting — see [memory](memory.md#the-pruning-repair).
|
||||||
|
|
||||||
|
## Config that changes behaviour subtly
|
||||||
|
|
||||||
|
Three fields do more than they look like they do.
|
||||||
|
|
||||||
|
**`thinking`** costs money and latency on every turn, not just hard ones. `off` on a hard problem
|
||||||
|
produces confident wrong answers; `max` on a rename wastes cents and seconds. The agent variants
|
||||||
|
already pick sensible levels — see [agents](agents.md).
|
||||||
|
|
||||||
|
**`toolSets`** removes tools from the model's view entirely. If the agent stops using a tool you
|
||||||
|
expected, check the startup header for which sets loaded: an unrecognised name is dropped
|
||||||
|
silently, so `"gti"` reads as "git is off". See [tools](tools.md#tool-sets).
|
||||||
|
|
||||||
|
**`registryUrl`** is the whole trust decision for installed skills and plugins. There are no
|
||||||
|
signatures, so pointing it at an index means trusting whoever controls that URL — including for
|
||||||
|
whatever they publish later. See [registry](registry.md).
|
||||||
|
|||||||
@@ -74,6 +74,29 @@ told and can respond to it.
|
|||||||
if shiro -p "does this build?" --yolo; then echo ok; else echo failed; fi
|
if shiro -p "does this build?" --yolo; then echo ok; else echo failed; fi
|
||||||
```
|
```
|
||||||
|
|
||||||
|
That distinction is deliberate and it has a consequence: **a successful run says nothing about
|
||||||
|
whether the answer was yes.** `0` means the turn completed, not that the build passed. To gate CI
|
||||||
|
on the content, read the output:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
shiro -p "Does this build? Answer only YES or NO." --json --yolo \
|
||||||
|
| jq -r 'select(.type=="text") | .text' | grep -q YES
|
||||||
|
```
|
||||||
|
|
||||||
|
Anything that must fail the build has to be asserted on text or, better, on the exit code of a
|
||||||
|
real command the agent ran.
|
||||||
|
|
||||||
|
## Timeouts
|
||||||
|
|
||||||
|
There is no wall-clock limit on a headless run. Three things bound it:
|
||||||
|
|
||||||
|
- `maxSteps` per variant — 12 for `quick`, 50 by default, 80 for `deep`.
|
||||||
|
- The `timeout` the model passes to `bash`, 120 s by default and 600 s at most.
|
||||||
|
- Whatever your CI runner enforces, which is the only hard stop.
|
||||||
|
|
||||||
|
Interactively `ctrl-c` kills one command and keeps the turn. Headless has no terminal for that, so
|
||||||
|
a signal ends the run. In CI, prefer `--agent quick` and a runner timeout over hoping.
|
||||||
|
|
||||||
## Sessions
|
## Sessions
|
||||||
|
|
||||||
Headless runs save like interactive ones, so `-c` picks up where one left off:
|
Headless runs save like interactive ones, so `-c` picks up where one left off:
|
||||||
@@ -83,6 +106,14 @@ shiro -p "start the refactor" --yolo
|
|||||||
shiro -p "now update the tests" --yolo -c
|
shiro -p "now update the tests" --yolo -c
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Useful, and worth knowing the shape of: each `-p` run is **one turn**, and `-c` resumes the newest
|
||||||
|
session for that directory. Two concurrent runs in the same directory therefore fight over the
|
||||||
|
same session, and the second overwrites the first. Pass `-r <id>` to keep parallel runs separate,
|
||||||
|
or point them at different `SHIRO_HOME` directories.
|
||||||
|
|
||||||
|
Memory also accumulates. An unattended loop calling `remember` writes to the project store like
|
||||||
|
any other run, so `--no-memory` is worth considering for a job that runs on every push.
|
||||||
|
|
||||||
## What is withheld
|
## What is withheld
|
||||||
|
|
||||||
The `ask` tool is not offered at all, rather than being offered and left to hang. The model
|
The `ask` tool is not offered at all, rather than being offered and left to hang. The model
|
||||||
@@ -131,8 +162,34 @@ env:
|
|||||||
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Two more worth having in a workflow. Trim the tool schema to what the job needs, since a CI run
|
||||||
|
pays for it on every step:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
- run: echo '{ "toolSets": [] }' > ~/.shiro-neko/config.json
|
||||||
|
```
|
||||||
|
|
||||||
|
And keep an unattended job from inheriting an installed skill nobody reviewed:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
- run: shiro -p "..." --agent review --no-skills --no-plugins
|
||||||
|
```
|
||||||
|
|
||||||
|
`--no-skills` matters more in CI than locally: a skill installed from a registry is instructions
|
||||||
|
in the system prompt, and CI is exactly where nobody is watching what it says. See
|
||||||
|
[registry](registry.md).
|
||||||
|
|
||||||
## Cost control
|
## Cost control
|
||||||
|
|
||||||
Headless runs are unattended, so a runaway loop costs real money. `--agent quick` caps the
|
Headless runs are unattended, so a runaway loop costs real money. `--agent quick` caps the
|
||||||
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. There is no
|
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. There is no
|
||||||
spend ceiling yet — see [TODO.md](../TODO.md).
|
spend ceiling yet — see [TODO.md](../TODO.md).
|
||||||
|
|
||||||
|
What a run actually costs is in the `done` event, so a wrapper can total it:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
shiro -p "..." --json --yolo | jq -r 'select(.type=="done" and .inputTokens) | "\(.inputTokens) in, \(.outputTokens) out"'
|
||||||
|
```
|
||||||
|
|
||||||
|
The token fields are optional: an aborted turn emits `done` with neither, which is why the filter
|
||||||
|
checks for one rather than assuming it.
|
||||||
|
|||||||
+46
-5
@@ -28,11 +28,24 @@ them in `~/.shiro-neko/config.json` and they appear alongside the builtins.
|
|||||||
```
|
```
|
||||||
|
|
||||||
**stdio** servers take `command`, and optionally `args`, `env`, `cwd`. The process is spawned
|
**stdio** servers take `command`, and optionally `args`, `env`, `cwd`. The process is spawned
|
||||||
at startup and closed on exit.
|
at startup and closed on exit. `env` is merged over the inherited environment, so a server
|
||||||
|
inherits your `PATH` unless you replace it.
|
||||||
|
|
||||||
**Remote** servers take `url`, and optionally `type` (`http` or `sse`, default `http`) and
|
**Remote** servers take `url`, and optionally `type` (`http` or `sse`, default `http`) and
|
||||||
`headers`.
|
`headers`.
|
||||||
|
|
||||||
|
A token in `headers` sits in `config.json` in plain text, same as `apiKey`. For anything beyond
|
||||||
|
a local dev token, prefer a stdio server that reads its own credential from the environment.
|
||||||
|
|
||||||
|
## Startup cost
|
||||||
|
|
||||||
|
Servers connect **in parallel**, so the slowest one sets how long startup takes rather than the
|
||||||
|
sum of them. `npx -y some-server` re-resolves the package on each launch; installing it and
|
||||||
|
calling the binary directly is usually the difference between a noticeable wait and none.
|
||||||
|
|
||||||
|
`--no-mcp` skips them all, which is also the quickest way to tell whether a slow start is MCP
|
||||||
|
or something else.
|
||||||
|
|
||||||
## Naming
|
## Naming
|
||||||
|
|
||||||
Tools arrive as `mcp__<server>__<tool>`. A server named `fs` exposing `read_file` becomes
|
Tools arrive as `mcp__<server>__<tool>`. A server named `fs` exposing `read_file` becomes
|
||||||
@@ -79,14 +92,42 @@ calling one.
|
|||||||
|
|
||||||
## Cost
|
## Cost
|
||||||
|
|
||||||
Each tool adds roughly 550 characters of schema to every request. A server exposing twenty
|
Each tool adds its name, description, and JSON schema to every request. The built-ins average
|
||||||
tools costs about 2,750 tokens per turn, sent whether or not the model uses any of them.
|
548 bytes; MCP tools vary with how verbose the server's schema is. A server exposing twenty
|
||||||
|
tools costs roughly 2,750 tokens per turn, sent whether or not the model uses any of them.
|
||||||
|
|
||||||
Prefer servers with a focused tool set. If one exposes many tools you never use, it is worth
|
MCP tools are **not** covered by `toolSets` — that budget only governs the built-ins. There is
|
||||||
finding a narrower server or writing one.
|
no per-server switch either, so the choice is a server or no server, and `--no-mcp` for all of
|
||||||
|
them. If one exposes many tools you never use, a narrower server is worth finding or writing.
|
||||||
|
|
||||||
|
`/tools` shows the count both ways:
|
||||||
|
|
||||||
|
```
|
||||||
|
tools
|
||||||
|
26 offered this turn of 26 registered
|
||||||
|
```
|
||||||
|
|
||||||
|
A gap between the two numbers means a tool set or a read-only agent variant is withholding
|
||||||
|
something. MCP tools never appear in that gap.
|
||||||
|
|
||||||
## Writing a server
|
## Writing a server
|
||||||
|
|
||||||
Any MCP-compliant server works. A minimal stdio one needs three methods: `initialize`,
|
Any MCP-compliant server works. A minimal stdio one needs three methods: `initialize`,
|
||||||
`tools/list`, and `tools/call`. The test suite includes one at
|
`tools/list`, and `tools/call`. The test suite includes one at
|
||||||
`test/fixtures/mcp-stub.ts` — about 50 lines, and useful as a starting point.
|
`test/fixtures/mcp-stub.ts` — about 50 lines, and useful as a starting point.
|
||||||
|
|
||||||
|
The suite runs it as a **real subprocess** rather than mocking the transport, because the parts
|
||||||
|
that break in practice are the handshake and the framing, and a mock asserts neither.
|
||||||
|
|
||||||
|
## Debugging a server
|
||||||
|
|
||||||
|
A server that starts but returns nothing useful is the harder case. In order of speed:
|
||||||
|
|
||||||
|
1. `/tools` — did the tools arrive at all? A server with no tools is a `tools/list` problem.
|
||||||
|
2. `shiro -p "call mcp__x__y with ..." --json --yolo` — the exact `tool-call` input and
|
||||||
|
`tool-result` output, one JSON object per line.
|
||||||
|
3. Run the server by hand: `echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | your-server`.
|
||||||
|
If that is wrong, nothing above it can be right.
|
||||||
|
|
||||||
|
For an HTTP server, `curl -X POST $URL -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'`
|
||||||
|
answers the same question without shiro in the way.
|
||||||
|
|||||||
+49
-1
@@ -13,6 +13,9 @@ The split exists because compaction is destructive. `pruneMessages` deletes tool
|
|||||||
`/compact` deletes the whole transcript, so anything recorded only in messages is lost
|
`/compact` deletes the whole transcript, so anything recorded only in messages is lost
|
||||||
exactly when a long task needs it most.
|
exactly when a long task needs it most.
|
||||||
|
|
||||||
|
The practical rule: if it should survive this turn, `todo_write` it. If it should survive this
|
||||||
|
session, `remember` it. The transcript is for the conversation, not for storage.
|
||||||
|
|
||||||
## Project memory
|
## Project memory
|
||||||
|
|
||||||
Durable notes about the codebase, injected at the start of every session.
|
Durable notes about the codebase, injected at the start of every session.
|
||||||
@@ -32,11 +35,21 @@ text one self-contained line
|
|||||||
|
|
||||||
Duplicates are refused. Text is capped at 400 characters, the store at 300 entries.
|
Duplicates are refused. Text is capped at 400 characters, the store at 300 entries.
|
||||||
|
|
||||||
|
The kinds are not decoration: they are what the model reads back at boot, and they set how much
|
||||||
|
to trust a note. A `command` is verifiable in one run. A `decision` explains why the obvious
|
||||||
|
alternative was not taken, which is the thing a newcomer most often gets wrong.
|
||||||
|
|
||||||
|
"One self-contained line" is the part that matters most. A note reading "use the new approach"
|
||||||
|
is worthless next session — there is no conversation left to say which approach.
|
||||||
|
|
||||||
### `recall`
|
### `recall`
|
||||||
|
|
||||||
Every term must appear. A match increments that entry's hit count, which protects it from
|
Every term must appear. A match increments that entry's hit count, which protects it from
|
||||||
compaction later — an entry the agent actually uses is worth keeping verbatim.
|
compaction later — an entry the agent actually uses is worth keeping verbatim.
|
||||||
|
|
||||||
|
AND rather than OR, on purpose: "migration seed order" should find the one note about that,
|
||||||
|
not every note mentioning any of the three words. Returns the 15 most recent matches.
|
||||||
|
|
||||||
### `forget`
|
### `forget`
|
||||||
|
|
||||||
Removes by substring, for a note that turned out wrong.
|
Removes by substring, for a note that turned out wrong.
|
||||||
@@ -61,7 +74,9 @@ memory goes stale and a confidently wrong note is worse than none.
|
|||||||
- Entries with at least one recall are kept verbatim and never merged.
|
- Entries with at least one recall are kept verbatim and never merged.
|
||||||
- A model returning nothing parseable leaves the store untouched.
|
- A model returning nothing parseable leaves the store untouched.
|
||||||
|
|
||||||
Without the second rule a bad response wipes everything the agent has learned.
|
Without the second rule a bad response wipes everything the agent has learned. The store also
|
||||||
|
has to be past 60 entries with at least two unused ones before anything happens, so `/memory`
|
||||||
|
on a small store is a deliberate no-op rather than a rewrite.
|
||||||
|
|
||||||
`/notes` lists the store with hit counts. `--no-memory` disables loading and writing.
|
`/notes` lists the store with hit counts. `--no-memory` disables loading and writing.
|
||||||
|
|
||||||
@@ -118,6 +133,14 @@ shiro -r 0193ab2c # by id or unique prefix
|
|||||||
|
|
||||||
A corrupt session file is skipped rather than crashing the list.
|
A corrupt session file is skipped rather than crashing the list.
|
||||||
|
|
||||||
|
Ids are UUIDv7, so they sort by creation time and a prefix is usually enough to identify one.
|
||||||
|
`-c` matches on `cwd`, so it picks up the newest session **for this directory** rather than the
|
||||||
|
newest overall — two projects side by side do not steal each other's `-c`.
|
||||||
|
|
||||||
|
Resuming restores the messages and the task list, but not the model or agent variant: those come
|
||||||
|
from the current config and flags. A session started with `--agent deep` resumes as `default`
|
||||||
|
unless you pass it again.
|
||||||
|
|
||||||
## Compaction
|
## Compaction
|
||||||
|
|
||||||
Two mechanisms.
|
Two mechanisms.
|
||||||
@@ -130,9 +153,24 @@ screen stays complete. The turn reports it:
|
|||||||
context compacted: 192 messages pruned to 15 on the wire
|
context compacted: 192 messages pruned to 15 on the wire
|
||||||
```
|
```
|
||||||
|
|
||||||
|
The status bar warns before that happens: context is shown as a percentage of the threshold,
|
||||||
|
amber from two thirds, red at 90.
|
||||||
|
|
||||||
|
What gets discarded, in order: reasoning items first, then tool calls and their results older
|
||||||
|
than the last three messages. Reasoning is the cheapest thing to lose — it was progress, not
|
||||||
|
conclusions — and tool results are the bulkiest. Recent exchanges are always kept, which is what
|
||||||
|
lets a turn continue rather than restart.
|
||||||
|
|
||||||
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
|
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
|
||||||
and outcomes, what remains — and it replaces the transcript entirely.
|
and outcomes, what remains — and it replaces the transcript entirely.
|
||||||
|
|
||||||
|
The difference is which history is destroyed. Automatic pruning touches the wire only, so
|
||||||
|
scrolling back still shows everything and `/save` records everything. `/compact` replaces the
|
||||||
|
real message array, so it is irreversible for that session.
|
||||||
|
|
||||||
|
Use `/compact` when a session has drifted across several unrelated tasks and the early part is
|
||||||
|
noise. Let automatic pruning handle a single long task, since it keeps the recent work intact.
|
||||||
|
|
||||||
### The pruning repair
|
### The pruning repair
|
||||||
|
|
||||||
Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both were real 400s
|
Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both were real 400s
|
||||||
@@ -171,6 +209,16 @@ and the `tool` message answering it:
|
|||||||
alone deliberately: a tool call still waiting for its result is what a suspended approval looks
|
alone deliberately: a tool call still waiting for its result is what a suspended approval looks
|
||||||
like, and dropping it would break `/resume`.
|
like, and dropping it would break `/resume`.
|
||||||
|
|
||||||
|
### What compaction still does not do
|
||||||
|
|
||||||
|
It tells the model the history was pruned but not what was in it. A decision from forty messages
|
||||||
|
ago can be contradicted with confidence, because from the model's side that span never existed.
|
||||||
|
Summarising the discarded part is [next on the list](../TODO.md).
|
||||||
|
|
||||||
|
The threshold is also measured with `JSON.stringify(messages).length / 4`, which is an estimate.
|
||||||
|
It is fine for deciding when to prune and wrong enough that it should not be read as a token
|
||||||
|
count — the real numbers in `/cost` come from the provider.
|
||||||
|
|
||||||
## Prompt history
|
## Prompt history
|
||||||
|
|
||||||
Per-directory, capped at 200, deduplicated against the previous entry. Up and down in the
|
Per-directory, capped at 200, deduplicated against the previous entry. Up and down in the
|
||||||
|
|||||||
+9
-2
@@ -70,10 +70,10 @@ holding `a` through a batch of edits will approve one of these without reading i
|
|||||||
| `rm -rf`, `rm -f` | recursive or forced delete |
|
| `rm -rf`, `rm -f` | recursive or forced delete |
|
||||||
| `git reset --hard` | discards uncommitted work |
|
| `git reset --hard` | discards uncommitted work |
|
||||||
| `git clean -f` | deletes untracked files |
|
| `git clean -f` | deletes untracked files |
|
||||||
| `git push --force`, `-f` | rewrites remote history |
|
| `git push --force`, `--force-with-lease`, `-f` | rewrites remote history |
|
||||||
| `git branch -D` | deletes a branch without a merge check |
|
| `git branch -D` | deletes a branch without a merge check |
|
||||||
| `DROP TABLE`, `TRUNCATE` | destroys database data |
|
| `DROP TABLE`, `TRUNCATE` | destroys database data |
|
||||||
| `mkfs`, `dd of=/dev/…` | writes to a raw device |
|
| `mkfs`, `dd of=/dev/…`, `> /dev/sd…` | writes to a raw device |
|
||||||
| `chmod 777` | makes files world-writable |
|
| `chmod 777` | makes files world-writable |
|
||||||
| `shutdown`, `reboot`, `halt` | affects the whole machine |
|
| `shutdown`, `reboot`, `halt` | affects the whole machine |
|
||||||
| `:(){ :\|:& };:` | fork bomb |
|
| `:(){ :\|:& };:` | fork bomb |
|
||||||
@@ -88,6 +88,13 @@ The model is told to relay the command rather than work around it. `rm build/one
|
|||||||
`git push origin feature`, and `git commit` all pass — the patterns target irreversibility,
|
`git push origin feature`, and `git commit` all pass — the patterns target irreversibility,
|
||||||
not the commands themselves.
|
not the commands themselves.
|
||||||
|
|
||||||
|
Two honest limits. The patterns match the command **string**, so `bash -c "$(echo cm0gLXJm | base64 -d)"`
|
||||||
|
is not caught, and neither is a script the agent wrote and then ran. And it only inspects `bash`:
|
||||||
|
a `write_file` overwriting something important is an approval question, not a guard question.
|
||||||
|
|
||||||
|
The guard is the last line before a command runs; `ctrl-c` is the one after. A pattern the guard
|
||||||
|
does not know about is still interruptible by hand — see [tools](tools.md#bash).
|
||||||
|
|
||||||
### `time` (default on)
|
### `time` (default on)
|
||||||
|
|
||||||
Adds `current_time`, returning ISO 8601 plus the local string. Auto-approved; it reads
|
Adds `current_time`, returning ISO 8601 plus the local string. Auto-approved; it reads
|
||||||
|
|||||||
@@ -126,3 +126,53 @@ add skill:review` disambiguates, and an ambiguous name is refused rather than gu
|
|||||||
|
|
||||||
A private index is just a URL you control. There is no account, no token, and no telemetry —
|
A private index is just a URL you control. There is no account, no token, and no telemetry —
|
||||||
`/registry` makes exactly one GET for the index and one for the entry you install.
|
`/registry` makes exactly one GET for the index and one for the entry you install.
|
||||||
|
|
||||||
|
## Publishing
|
||||||
|
|
||||||
|
Two files and a static host. GitHub raw works, and so does anything that serves JSON over https.
|
||||||
|
|
||||||
|
```
|
||||||
|
your-registry/
|
||||||
|
index.json
|
||||||
|
skills/migration.md
|
||||||
|
plugins/no-secrets.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Three rules the validator enforces, so worth getting right first:
|
||||||
|
|
||||||
|
- The name in `index.json` must match the name inside the file. A skill's frontmatter `name` and a
|
||||||
|
plugin manifest's `name` are both checked against the index entry.
|
||||||
|
- Names are `^[a-z0-9][a-z0-9-]*$`. No uppercase, no dots, no slashes.
|
||||||
|
- A plugin needs at least one deny rule. A manifest with an `appendix` and no rules is prompt
|
||||||
|
text, which is what a skill is for.
|
||||||
|
|
||||||
|
Test it locally before publishing. `registryUrl` accepts `http://localhost`, so:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd your-registry && python -m http.server 8000
|
||||||
|
```
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "registryUrl": "http://localhost:8000/index.json" }
|
||||||
|
```
|
||||||
|
|
||||||
|
`/registry` then exercises the real fetch, the real validation, and the real install path against
|
||||||
|
your files. That is the whole loop, without pushing anything.
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
**"the registry index is malformed: …"** — the message names the first failing field. The usual
|
||||||
|
causes are an uppercase name, a `url` that is not https, or a plugin entry with no `deny`.
|
||||||
|
|
||||||
|
**"X calls itself Y but the index calls it X"** — the file's own name disagrees with the index.
|
||||||
|
Fix one of the two; the check exists so an index cannot serve something else under a name you
|
||||||
|
trusted.
|
||||||
|
|
||||||
|
**"invalid pattern …"** — a `pathPattern` or `commandPattern` is not a valid regex. Remember it is
|
||||||
|
JSON, so a backslash needs doubling: `\\.env$`, not `\.env$`.
|
||||||
|
|
||||||
|
**Installed but nothing happens** — installs load at startup. Restart, then check `/skills` or
|
||||||
|
`/plugins` for the entry and its origin.
|
||||||
|
|
||||||
|
**In `/plugins` with an error beside it** — the manifest on disk no longer validates. It is skipped
|
||||||
|
rather than fatal, so the agent still starts; `/registry remove` and reinstall.
|
||||||
|
|||||||
+29
-3
@@ -3,9 +3,9 @@
|
|||||||
A skill is a markdown file with instructions for one kind of task. Only its name and
|
A skill is a markdown file with instructions for one kind of task. Only its name and
|
||||||
description sit in the system prompt; the body is loaded on demand.
|
description sit in the system prompt; the body is loaded on demand.
|
||||||
|
|
||||||
That split matters. Four bundled skills are 4,659 characters of body but 681 characters of
|
That split matters. The four bundled skills are 5,284 characters of body against 681 characters
|
||||||
catalogue. Putting every body in the prompt would cost that on every request, for
|
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt
|
||||||
instructions that are relevant to one turn in twenty.
|
would cost that on every turn, for instructions relevant to one turn in twenty.
|
||||||
|
|
||||||
## Format
|
## Format
|
||||||
|
|
||||||
@@ -28,6 +28,14 @@ Never deploy from a dirty working tree.
|
|||||||
description is what the model matches against, so write it as a trigger — "use when asked
|
description is what the model matches against, so write it as a trigger — "use when asked
|
||||||
to X" — not as a summary.
|
to X" — not as a summary.
|
||||||
|
|
||||||
|
The frontmatter reader handles those two fields and nothing else. A real YAML parser would be a
|
||||||
|
dependency for two strings, so lists, nesting, and multi-line values are not supported: keep both
|
||||||
|
on one line. Quotes around a value are stripped. A body over 20,000 characters is truncated.
|
||||||
|
|
||||||
|
A file that fails to parse is skipped silently rather than reported, which is worth knowing when
|
||||||
|
a skill you wrote does not appear in `/skills` — the usual cause is a missing `---` fence or a
|
||||||
|
description spilling onto a second line.
|
||||||
|
|
||||||
## Where they load from
|
## Where they load from
|
||||||
|
|
||||||
Four sources, later overriding earlier by name:
|
Four sources, later overriding earlier by name:
|
||||||
@@ -79,6 +87,24 @@ before you start working, and follow it as if the user had written it:
|
|||||||
When the model calls `skill({ name: "debug" })` it gets the full body back and is told to
|
When the model calls `skill({ name: "debug" })` it gets the full body back and is told to
|
||||||
follow it for this task. The call needs no approval — it reads nothing outside the binary.
|
follow it for this task. The call needs no approval — it reads nothing outside the binary.
|
||||||
|
|
||||||
|
"Before you start working" is the load-bearing phrase. A skill loaded after the work is done is
|
||||||
|
wasted tokens, and the failure mode in practice is a model that reads the catalogue, decides it
|
||||||
|
already knows, and never calls the tool. A description written as a trigger is what prevents that.
|
||||||
|
|
||||||
|
Loading one costs its body, once, in that turn's context. A 3,000-character skill is cheaper than
|
||||||
|
one wrong approach it prevents, and more expensive than the catalogue line that would have been
|
||||||
|
enough.
|
||||||
|
|
||||||
|
## Verifying a skill loaded
|
||||||
|
|
||||||
|
```bash
|
||||||
|
shiro -p "fix the failing pagination test" --json --yolo | grep skill
|
||||||
|
```
|
||||||
|
|
||||||
|
`--json` shows the `tool-call` for `skill` with the name it chose, or its absence. If the model
|
||||||
|
never calls it on a task the skill was written for, the description is the thing to change — not
|
||||||
|
the body.
|
||||||
|
|
||||||
## Writing a good one
|
## Writing a good one
|
||||||
|
|
||||||
Skills work when they encode what a newcomer to *your* project would get wrong. The bundled
|
Skills work when they encode what a newcomer to *your* project would get wrong. The bundled
|
||||||
|
|||||||
+180
-31
@@ -27,19 +27,42 @@ y allow once | a always allow edit_file | n deny
|
|||||||
`a` whitelists that tool for the rest of the session. `n` tells the model it was denied and
|
`a` whitelists that tool for the rest of the session. `n` tells the model it was denied and
|
||||||
to ask what to do instead. `--yolo` skips all prompts.
|
to ask what to do instead. `--yolo` skips all prompts.
|
||||||
|
|
||||||
|
The approval is enforced by the SDK, not by the tools. A denied call **provably never
|
||||||
|
executes**: the SDK never reaches the tool's `execute`, so a tool cannot forget to honour a
|
||||||
|
denial or opt out of the check. See [architecture](architecture.md#why-approval-goes-through-the-sdk).
|
||||||
|
|
||||||
|
MCP tools are gated as a group because they are third-party code with unknown side effects —
|
||||||
|
`mcp__fs__read_file` sounds harmless and might not be. See [MCP](mcp.md).
|
||||||
|
|
||||||
**The guard runs before all of this.** It is not an approval — it is a refusal, and `--yolo`
|
**The guard runs before all of this.** It is not an approval — it is a refusal, and `--yolo`
|
||||||
does not reach it. See [plugins](plugins.md).
|
does not reach it. See [plugins](plugins.md).
|
||||||
|
|
||||||
## Tool sets
|
## Tool sets
|
||||||
|
|
||||||
Each tool costs roughly 550 characters of JSON schema on every request, and selection
|
Each tool costs its name, its description, and its JSON schema on **every request**. Measured
|
||||||
accuracy drops as the list grows. Sets let you switch off what a project does not need:
|
across the fourteen built-ins:
|
||||||
|
|
||||||
| Set | Tools |
|
| Tool | Bytes | Tool | Bytes |
|
||||||
|---|---|
|
|---|---|---|---|
|
||||||
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` |
|
| `read_many_files` | 972 | `git_blame` | 499 |
|
||||||
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` |
|
| `multi_edit` | 934 | `git_log` | 484 |
|
||||||
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` |
|
| `edit_file` | 618 | `git_diff` | 473 |
|
||||||
|
| `grep` | 595 | `bash` | 466 |
|
||||||
|
| `list_dir` | 594 | `git_show` | 432 |
|
||||||
|
| `read_file` | 526 | `git_status` | 292 |
|
||||||
|
| `glob` | 499 | `write_file` | 289 |
|
||||||
|
|
||||||
|
7,673 bytes for all fourteen, averaging 548. Roughly 1,900 tokens per request before your
|
||||||
|
prompt or the conversation. Selection accuracy also falls as the list grows: a model choosing
|
||||||
|
between six tools picks better than one choosing between twenty.
|
||||||
|
|
||||||
|
Sets let you switch off what a project does not need:
|
||||||
|
|
||||||
|
| Set | Tools | Cost |
|
||||||
|
|---|---|---|
|
||||||
|
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B |
|
||||||
|
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` | ~2,500 B |
|
||||||
|
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` | ~2,180 B |
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{ "toolSets": ["edit-plus"] }
|
{ "toolSets": ["edit-plus"] }
|
||||||
@@ -50,7 +73,32 @@ agent is not an agent. A disabled set reaches neither the wire nor the system pr
|
|||||||
a prompt that names an absent tool teaches the model to attempt calls that cannot succeed.
|
a prompt that names an absent tool teaches the model to attempt calls that cannot succeed.
|
||||||
Session, plugin, and MCP tools are not part of this budget and are never gated here.
|
Session, plugin, and MCP tools are not part of this budget and are never gated here.
|
||||||
|
|
||||||
`/tools` shows which set each live tool came from.
|
An unrecognised set name is dropped silently. The header line at startup shows which sets
|
||||||
|
actually loaded, so a typo reads as "that set is off" rather than as an error — worth checking
|
||||||
|
if a tool you expected is missing.
|
||||||
|
|
||||||
|
`/tools` shows which set each live tool came from:
|
||||||
|
|
||||||
|
```
|
||||||
|
tools
|
||||||
|
20 offered this turn of 22 registered
|
||||||
|
- `bash` core
|
||||||
|
- `git_diff` git
|
||||||
|
- `list_dir` edit-plus
|
||||||
|
- `remember`
|
||||||
|
```
|
||||||
|
|
||||||
|
A tool with no set is a session, plugin, or MCP tool.
|
||||||
|
|
||||||
|
### Which sets to keep
|
||||||
|
|
||||||
|
Both extra sets earn their place in most projects, but not all:
|
||||||
|
|
||||||
|
- **No git in the repo?** `git` is 2,180 bytes the model can never use. Switch it off.
|
||||||
|
- **A model that handles many tools badly?** `{ "toolSets": [] }` trims to six, which is the
|
||||||
|
smallest set that still lets the agent work.
|
||||||
|
- **Reading a lot, editing rarely?** Keep `edit-plus` for `list_dir` and `read_many_files`
|
||||||
|
alone; they pay for themselves in round trips saved.
|
||||||
|
|
||||||
## File tools
|
## File tools
|
||||||
|
|
||||||
@@ -107,7 +155,14 @@ replaceAll replace every occurrence instead of requiring exactly one
|
|||||||
|
|
||||||
`oldString` must match byte-for-byte and appear exactly once unless `replaceAll` is set.
|
`oldString` must match byte-for-byte and appear exactly once unless `replaceAll` is set.
|
||||||
An ambiguous match is an error naming the count, which pushes the model to add surrounding
|
An ambiguous match is an error naming the count, which pushes the model to add surrounding
|
||||||
context rather than guessing which occurrence it meant.
|
context rather than guessing which occurrence it meant:
|
||||||
|
|
||||||
|
```
|
||||||
|
oldString appears 3 times in src/users.ts. Add surrounding context or set replaceAll.
|
||||||
|
```
|
||||||
|
|
||||||
|
That error is deliberately specific. `edit failed` would leave the model to retry blind; the
|
||||||
|
count tells it what to do next.
|
||||||
|
|
||||||
### `multi_edit`
|
### `multi_edit`
|
||||||
|
|
||||||
@@ -121,7 +176,14 @@ the previous one, so edits may build on each other.
|
|||||||
|
|
||||||
Atomic: every edit is validated and applied in memory first, so a failure on the third edit
|
Atomic: every edit is validated and applied in memory first, so a failure on the third edit
|
||||||
leaves the file exactly as it was rather than half-changed. The same uniqueness rule as
|
leaves the file exactly as it was rather than half-changed. The same uniqueness rule as
|
||||||
`edit_file` applies per edit, and the error names which edit failed.
|
`edit_file` applies per edit, and the error names which edit failed:
|
||||||
|
|
||||||
|
```
|
||||||
|
edit 2: oldString not found in src/users.ts. No edits were applied.
|
||||||
|
```
|
||||||
|
|
||||||
|
The last sentence matters. Without it a model reading the error has to guess whether edit 1
|
||||||
|
landed, and its next move — retry the whole batch, or only what failed — depends on the answer.
|
||||||
|
|
||||||
### `list_dir`
|
### `list_dir`
|
||||||
|
|
||||||
@@ -131,9 +193,19 @@ depth levels to descend, 1-6, default 2
|
|||||||
includeIgnored also show files git ignores
|
includeIgnored also show files git ignores
|
||||||
```
|
```
|
||||||
|
|
||||||
Tree view honouring `.gitignore`. Directories end with `/`, files show their size. Past the
|
Tree view honouring `.gitignore`. Directories end with `/`, files show their size:
|
||||||
depth limit the containing directory is still listed, so the shape of the tree stays visible
|
|
||||||
without its contents. Capped at 300 entries.
|
```
|
||||||
|
.
|
||||||
|
README.md 2B
|
||||||
|
src/
|
||||||
|
app.ts 2K
|
||||||
|
ui/
|
||||||
|
```
|
||||||
|
|
||||||
|
Past the depth limit the containing directory is still listed, so the shape of the tree stays
|
||||||
|
visible without its contents — `src/ui/` above appears at `depth: 2` even though its files do
|
||||||
|
not. Capped at 300 entries.
|
||||||
|
|
||||||
### `glob`
|
### `glob`
|
||||||
|
|
||||||
@@ -145,9 +217,12 @@ includeIgnored also return files git ignores
|
|||||||
|
|
||||||
Walks the tree honouring `.gitignore` and `.shiroignore`, skipping `.git` and
|
Walks the tree honouring `.gitignore` and `.shiroignore`, skipping `.git` and
|
||||||
`node_modules` unconditionally. Nested ignore files apply only within their own directory,
|
`node_modules` unconditionally. Nested ignore files apply only within their own directory,
|
||||||
as git does. Returns posix paths relative to the workspace root. A symlinked directory is
|
as git does. Returns posix paths relative to the workspace root.
|
||||||
classified as a directory and not descended into, since it can point anywhere including
|
|
||||||
back into the tree.
|
A symlinked directory is classified as a directory and not descended into. Both halves matter:
|
||||||
|
`readdir` reports a junction as a non-directory, so without the extra `stat` a symlinked
|
||||||
|
directory leaked past `dir/` ignore rules and was yielded as a file with a nonsense size. Not
|
||||||
|
descending is separate — a link can point anywhere, including back into the tree.
|
||||||
|
|
||||||
### `grep`
|
### `grep`
|
||||||
|
|
||||||
@@ -162,6 +237,15 @@ Shells out to ripgrep when it is on PATH — roughly 15x faster on a real repo
|
|||||||
back to a JavaScript walker otherwise. Output is `path:line: text` either way, so the model
|
back to a JavaScript walker otherwise. Output is `path:line: text` either way, so the model
|
||||||
sees one format regardless. Skips binaries. Caps at 200 hits.
|
sees one format regardless. Skips binaries. Caps at 200 hits.
|
||||||
|
|
||||||
|
Two details keep the two paths in agreement. ripgrep is passed `--no-require-git`, because it
|
||||||
|
otherwise ignores `.gitignore` outside a repository while the JavaScript fallback always honours
|
||||||
|
it. And an rg exit code above 1 means rg could not run the search at all, so the fallback takes
|
||||||
|
over; exit 1 is simply "no matches" and is reported as such.
|
||||||
|
|
||||||
|
Regex syntax differs between the two: ripgrep is Rust regex, the fallback is JavaScript. A
|
||||||
|
pattern using look-around works in the fallback and fails under rg. An invalid pattern is
|
||||||
|
reported as `Invalid regex: <reason>` rather than returning an empty result set.
|
||||||
|
|
||||||
### `bash`
|
### `bash`
|
||||||
|
|
||||||
```
|
```
|
||||||
@@ -194,12 +278,25 @@ running quits as usual.
|
|||||||
|
|
||||||
All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument
|
All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument
|
||||||
array rather than a shell string, so an argument like `--author="; rm -rf /"` can only ever
|
array rather than a shell string, so an argument like `--author="; rm -rf /"` can only ever
|
||||||
be a literal argument — which is what makes auto-approval safe.
|
be a literal argument — which is what makes auto-approval safe. A test asserts exactly that:
|
||||||
|
`git_log` with the path `; touch pwned.txt` creates no file.
|
||||||
|
|
||||||
Output is described rather than raw porcelain: `git_status` names the branch and says
|
Output is described rather than raw porcelain. `git_status` names the branch and says
|
||||||
`staged modified` or `untracked` per file instead of leaving the model to decode two columns
|
`staged modified` or `untracked` per file instead of leaving the model to decode porcelain's two
|
||||||
of flags. Outside a repository they fail with `<cwd> is not a git repository.` rather than
|
leading columns:
|
||||||
passing git's own error text through.
|
|
||||||
|
```
|
||||||
|
On main, 2 changed:
|
||||||
|
src/app.ts (staged modified, modified)
|
||||||
|
new.ts (untracked)
|
||||||
|
```
|
||||||
|
|
||||||
|
That file has a staged change *and* a later unstaged one, which the raw `MM` prefix conveys only
|
||||||
|
to a reader who knows the format.
|
||||||
|
|
||||||
|
Outside a repository they fail with `<cwd> is not a git repository.` rather than passing git's
|
||||||
|
own error text through. Other git failures do pass through, on purpose: `git_show no-such-ref`
|
||||||
|
reports what git said, because git's own message is the most useful thing available.
|
||||||
|
|
||||||
```
|
```
|
||||||
git_status branch, staged, modified, untracked
|
git_status branch, staged, modified, untracked
|
||||||
@@ -209,6 +306,13 @@ git_show ref path? one commit: message, author, diff
|
|||||||
git_blame path startLine? endLine? who last changed each line
|
git_blame path startLine? endLine? who last changed each line
|
||||||
```
|
```
|
||||||
|
|
||||||
|
`git_log` defaults to 15 commits and caps at 40. `git_blame` without a range blames the whole
|
||||||
|
file; with `startLine` and no `endLine` it covers 40 lines from there.
|
||||||
|
|
||||||
|
Everything here is also reachable through `bash`. The reason the set exists anyway is the
|
||||||
|
approval boundary: `bash git diff` stops for a decision on every call, while `git_diff` cannot
|
||||||
|
mutate anything and so never needs one.
|
||||||
|
|
||||||
## Agent tools
|
## Agent tools
|
||||||
|
|
||||||
### `task`
|
### `task`
|
||||||
@@ -223,8 +327,17 @@ Spawns a read-only subagent with `read_file`, `glob`, and `grep` only. It return
|
|||||||
report, so the parent pays for findings rather than the whole search transcript. It sees
|
report, so the parent pays for findings rather than the whole search transcript. It sees
|
||||||
none of the parent conversation, so its prompt has to stand alone.
|
none of the parent conversation, so its prompt has to stand alone.
|
||||||
|
|
||||||
|
Two properties follow from that tool set rather than from policy: it can never trigger an
|
||||||
|
approval prompt, because it has no gated tools; and the parent's context holds the conclusion
|
||||||
|
instead of the search. A subagent reading forty files to answer one question costs the parent
|
||||||
|
the answer, not the forty files.
|
||||||
|
|
||||||
`explore` finds and reports. `review` critiques code in severity order. Progress streams to
|
`explore` finds and reports. `review` critiques code in severity order. Progress streams to
|
||||||
the subagent panel.
|
the subagent panel. Capped at 20 steps, and it shares the parent's model — an `explore` run
|
||||||
|
pays reasoning rates for what is really a search, which is [on the list](../TODO.md) to fix.
|
||||||
|
|
||||||
|
Not worth delegating a single grep: the subagent is a whole extra model loop, so it wins on a
|
||||||
|
search spanning many files and loses on anything you could answer in one call.
|
||||||
|
|
||||||
### `ask`
|
### `ask`
|
||||||
|
|
||||||
@@ -235,9 +348,11 @@ multiple allow more than one
|
|||||||
```
|
```
|
||||||
|
|
||||||
Stops the turn and puts the question on screen. With options it is a picker; without, free
|
Stops the turn and puts the question on screen. With options it is a picker; without, free
|
||||||
text. `esc` skips, which tells the model to decide and state its assumption.
|
text. `esc` skips, which returns "the user dismissed this; use your best judgement" — so a
|
||||||
|
dismissal is an instruction to decide, not a dead end.
|
||||||
|
|
||||||
Withheld entirely in headless mode — a question with no one to answer it would hang.
|
Withheld entirely in headless mode — a question with no one to answer it would hang. The system
|
||||||
|
prompt says so, and tells the model to decide and state its assumption instead.
|
||||||
|
|
||||||
### `todo_write`
|
### `todo_write`
|
||||||
|
|
||||||
@@ -249,6 +364,9 @@ Statuses: `pending`, `in_progress`, `done`, `blocked`. Send the whole list each
|
|||||||
replaces the previous one. Warns when more than one task is `in_progress`, when nothing is
|
replaces the previous one. Warns when more than one task is `in_progress`, when nothing is
|
||||||
`in_progress` while work remains, or when a `blocked` task has no note.
|
`in_progress` while work remains, or when a `blocked` task has no note.
|
||||||
|
|
||||||
|
The list lives in the system prompt, rebuilt every step, so it survives both pruning and
|
||||||
|
`/compact`. See [memory](memory.md).
|
||||||
|
|
||||||
### `remember`, `recall`, `forget`
|
### `remember`, `recall`, `forget`
|
||||||
|
|
||||||
Durable per-project notes. See [memory](memory.md).
|
Durable per-project notes. See [memory](memory.md).
|
||||||
@@ -264,15 +382,46 @@ Loads the body of a skill. See [skills](skills.md).
|
|||||||
## Path safety
|
## Path safety
|
||||||
|
|
||||||
Every path a tool receives goes through a jail: resolved against the workspace root, then
|
Every path a tool receives goes through a jail: resolved against the workspace root, then
|
||||||
checked that it did not escape. `../../etc/passwd` and absolute paths outside the root are
|
checked that it did not escape.
|
||||||
both refused before any filesystem call.
|
|
||||||
|
```ts
|
||||||
|
export function jail(p: string, root = process.cwd()): string {
|
||||||
|
const abs = isAbsolute(p) ? resolve(p) : resolve(root, p);
|
||||||
|
const rel = relative(resolve(root), abs);
|
||||||
|
if (rel.startsWith('..') || isAbsolute(rel)) throw new Error(`Path escapes workspace: ${p}`);
|
||||||
|
return abs;
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
`../../etc/passwd`, `a/../../secret`, and absolute paths outside the root are all refused
|
||||||
|
before any filesystem call. Resolving first and comparing after is what catches the middle
|
||||||
|
case: string-prefix checks on the raw input miss `a/../../secret` entirely.
|
||||||
|
|
||||||
The model's output is a trust boundary. It can emit any string, so the check happens on
|
The model's output is a trust boundary. It can emit any string, so the check happens on
|
||||||
every call rather than being assumed.
|
every call rather than being assumed. That includes `read_many_files`, where a bad path is
|
||||||
|
reported in its block like any other unreadable file.
|
||||||
|
|
||||||
|
`jail` guards the workspace, not the shell. `bash` runs whatever it is given, which is why
|
||||||
|
every call needs approval and why the guard plugin exists — see [plugins](plugins.md).
|
||||||
|
|
||||||
## Output caps
|
## Output caps
|
||||||
|
|
||||||
Any single tool result is truncated at 30,000 characters with a note saying how much was
|
Any single tool result is truncated at 30,000 characters with a note saying how much was cut:
|
||||||
cut. `grep` stops at 200 hits, `glob` at 200 paths, `list_dir` at 300 entries,
|
|
||||||
`read_many_files` at 20 files, `read_file` at 2000 lines by default. Without caps one `grep`
|
```
|
||||||
for `function` can end a session.
|
... [truncated 41,233 chars]
|
||||||
|
```
|
||||||
|
|
||||||
|
| Tool | Cap |
|
||||||
|
|---|---|
|
||||||
|
| any result | 30,000 characters |
|
||||||
|
| `grep` | 200 hits, each line cut at 300 chars |
|
||||||
|
| `glob` | 200 paths |
|
||||||
|
| `list_dir` | 300 entries |
|
||||||
|
| `read_many_files` | 20 files |
|
||||||
|
| `read_file` | 2,000 lines by default |
|
||||||
|
| `bash` | 120 s default timeout, 600 s max |
|
||||||
|
|
||||||
|
Without caps one `grep` for `function` can end a session. The caps are per call, so a model
|
||||||
|
that needs more can narrow and ask again — which is cheaper than one call that fills the
|
||||||
|
context and forces compaction.
|
||||||
|
|||||||
Reference in New Issue
Block a user