Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual per-tool measurements, and fills in the parts a reader hits after the happy path: what a specific error means, what a setting costs, what is not covered. Measured rather than estimated: - Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the full set, averaging 548. - Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the argument for loading bodies on demand. - Full system prompt 3,571 chars, core-only 2,045. New sections: - tools: which sets to keep and why, the jail function itself, an output-cap table, and the real error strings for edit_file and multi_edit. - configuration: env var per provider preset, cost-estimate limits, what each --no-* flag isolates, and three settings that do more than they look like. - agents: step caps per variant, which variant to reach for, and the fact that reasoning is charged as output and discarded first by compaction. - headless: exit code 0 means "the turn completed", not "the answer was yes" — with the jq pattern for gating on content. Timeouts, concurrent -c runs fighting over one session, CI recipes for --no-skills. - mcp: parallel connect, startup cost, a debugging ladder, and that toolSets does not gate MCP tools. - registry: publishing, local testing over http://localhost, and a troubleshooting section keyed on the actual validator messages. - memory: what compaction discards in what order, /compact versus automatic pruning, and that -c matches on cwd. - skills: the frontmatter reader's limits, and how to verify a skill loaded. Corrections found while cross-checking against the source: - The guard table was missing --force-with-lease and > /dev/sd… - The done event's token fields are optional, so the jq example filters on one rather than assuming it. Two honest limits now written down: the guard matches command strings, so a base64-decoded or script-wrapped command is not caught; and a registry index is trusted for its contents, not its authorship. Verified: all internal links and heading anchors resolve, every docs/ page is reachable from the README, 538 tests pass, typecheck clean.
This commit is contained in:
+50
-1
@@ -62,12 +62,35 @@ from.
|
||||
| `high` | `high` | `budget_tokens: 38400` |
|
||||
| `max` | `xhigh` | maximum budget |
|
||||
|
||||
Verified against both wire formats rather than assumed.
|
||||
Verified against both wire formats rather than assumed. The vocabulary is deliberately ours:
|
||||
`off` through `max` means the same thing whichever provider is configured, and switching
|
||||
providers mid-session does not change what `/think high` asks for.
|
||||
|
||||
Higher costs more and takes longer. `off` on a hard problem produces confident wrong
|
||||
answers; `max` on a rename wastes a few cents and several seconds. The variants pick
|
||||
sensible defaults, so reach for `/think` only when a specific turn needs something else.
|
||||
|
||||
Reasoning is also charged as output tokens, so `max` shows up in `/cost` even on a turn where
|
||||
the model wrote two lines. And reasoning is the **first thing compaction discards** — see
|
||||
[memory](memory.md#compaction) — so a long turn at `max` pays for thinking that will not be on
|
||||
the wire by the end of it.
|
||||
|
||||
## Steps
|
||||
|
||||
`maxSteps` caps how many model calls one turn may make. A step is one request: a tool call and
|
||||
its result, or the final text.
|
||||
|
||||
| Variant | Steps |
|
||||
|---|---|
|
||||
| `quick` | 12 |
|
||||
| `default`, `plan`, `review` | 50 |
|
||||
| `deep` | 80 |
|
||||
|
||||
The cap is a backstop against a loop, not a budget to spend. A turn that hits it stops
|
||||
mid-work with whatever it has, which is why `quick`'s 12 suits a rename and would strand a
|
||||
refactor. If turns regularly hit the cap on the same kind of task, the task wants `deep`
|
||||
rather than a higher number.
|
||||
|
||||
## Overriding
|
||||
|
||||
`--agent deep --think low` gives you `deep`'s tools, steps, and appendix with a low thinking
|
||||
@@ -97,3 +120,29 @@ told which tools need approval and to verify with the project's tests.
|
||||
|
||||
A prompt that describes a withheld tool teaches the model to attempt calls that cannot
|
||||
succeed, which is why the description is generated from the live tool set.
|
||||
|
||||
Three rules flip on what is available:
|
||||
|
||||
| Condition | `default` says | `plan` says |
|
||||
|---|---|---|
|
||||
| can edit | "these need approval; if denied, stop and ask" | "you have no tools that change anything" |
|
||||
| can run commands | "verify with the project's build or tests" | "say what should be run rather than claiming it passed" |
|
||||
| can ask | "ask rather than guess when two readings differ" | (same, unless headless) |
|
||||
|
||||
The read-only variants are around 2,000 characters of system prompt against roughly 3,600 for
|
||||
the full set — cheaper per turn as well as safer.
|
||||
|
||||
## Which to reach for
|
||||
|
||||
- **`default`** for anything you have not thought about. It is the right answer most of the time.
|
||||
- **`quick`** for a rename, a typo, a one-line fix. Its value is not the model being cheaper but
|
||||
the absence of deliberation latency on work that needs none.
|
||||
- **`deep`** when the first attempt already failed, or the cause is unclear. Asking for more than
|
||||
one hypothesis is the actual difference; the thinking budget is secondary.
|
||||
- **`plan`** before a change you are not sure about. Read-only means the plan cannot quietly
|
||||
become a half-applied edit.
|
||||
- **`review`** on a diff or a module. In headless CI this is the one that needs no `--yolo`,
|
||||
because it holds no tool that can modify anything — see [headless](headless.md).
|
||||
|
||||
Switching mid-session is fine and cheap: `/agent` changes the next turn's tools and prompt, and
|
||||
nothing about the history.
|
||||
|
||||
+100
-19
@@ -48,21 +48,48 @@ Written by `/provider`, editable by hand. Every field is optional.
|
||||
|
||||
`/provider` offers these. Each sets `baseURL` and the wire protocol for you.
|
||||
|
||||
| Preset | Protocol | Endpoint |
|
||||
|---|---|---|
|
||||
| Anthropic | `anthropic` | `api.anthropic.com/v1` |
|
||||
| OpenAI | `openai` | `api.openai.com/v1` |
|
||||
| OpenRouter | `openai` | `openrouter.ai/api/v1` |
|
||||
| Groq | `openai` | `api.groq.com/openai/v1` |
|
||||
| DeepSeek | `openai` | `api.deepseek.com/v1` |
|
||||
| xAI | `openai` | `api.x.ai/v1` |
|
||||
| Ollama | `openai` | `localhost:11434/v1` |
|
||||
| LM Studio | `openai` | `localhost:1234/v1` |
|
||||
| Custom OpenAI-compatible | `openai` | you supply it |
|
||||
| Custom Anthropic-compatible | `anthropic` | you supply it |
|
||||
| Preset | Protocol | Endpoint | Env var checked |
|
||||
|---|---|---|---|
|
||||
| Anthropic | `anthropic` | `api.anthropic.com/v1` | `ANTHROPIC_API_KEY` |
|
||||
| OpenAI | `openai` | `api.openai.com/v1` | `OPENAI_API_KEY` |
|
||||
| OpenRouter | `openai` | `openrouter.ai/api/v1` | `OPENROUTER_API_KEY` |
|
||||
| Groq | `openai` | `api.groq.com/openai/v1` | `GROQ_API_KEY` |
|
||||
| DeepSeek | `openai` | `api.deepseek.com/v1` | `DEEPSEEK_API_KEY` |
|
||||
| xAI | `openai` | `api.x.ai/v1` | `XAI_API_KEY` |
|
||||
| Ollama | `openai` | `localhost:11434/v1` | none, keyless |
|
||||
| LM Studio | `openai` | `localhost:1234/v1` | none, keyless |
|
||||
| Custom OpenAI-compatible | `openai` | you supply it | none |
|
||||
| Custom Anthropic-compatible | `anthropic` | you supply it | none |
|
||||
|
||||
After the key is entered, `GET /v1/models` is called and the list becomes a picker. If the
|
||||
endpoint does not implement it, you type the model id instead — the setup still completes.
|
||||
`provider` is the **wire protocol**, not the vendor. Groq, DeepSeek, xAI, OpenRouter, Ollama,
|
||||
and LM Studio all speak `openai`; only Anthropic speaks `anthropic`. Two things differ between
|
||||
them: the auth header (`Authorization: Bearer` versus `x-api-key`), and how thinking levels map.
|
||||
|
||||
After the key is entered, `GET /v1/models` is called and the list becomes a picker. Both
|
||||
protocols expose that endpoint with the same `data[].id` shape, so one code path handles both.
|
||||
If the endpoint does not implement it, a preset with a known model list falls back to that;
|
||||
otherwise you type the model id and setup still completes.
|
||||
|
||||
Anything the picker offers is a model the endpoint actually reports, which is more reliable than
|
||||
a hard-coded list — that is why the fallback lists are short and only exist for Anthropic and
|
||||
OpenAI.
|
||||
|
||||
## Cost estimates
|
||||
|
||||
`/cost` and the status bar price a turn from a table in `src/pricing.ts`, matched by longest
|
||||
prefix on the model id, so `claude-sonnet-4-5-20250929` resolves via `claude-sonnet-4-5`. An
|
||||
OpenRouter-style `anthropic/claude-sonnet-4-5` has its vendor prefix stripped first.
|
||||
|
||||
An unknown model is reported as unpriced rather than guessed:
|
||||
|
||||
```
|
||||
4210 in / 88 out tokens (llama-3.3-70b is unpriced)
|
||||
```
|
||||
|
||||
Two limits worth knowing. The rates are hand-entered and drift as vendors change them, so treat
|
||||
the figure as an estimate, not a bill. And the token counts come from the provider's usage
|
||||
report, while `~ctx` in the status bar is `JSON.stringify(messages).length / 4` — good enough to
|
||||
decide when to compact, wrong enough that it should not be read as a token count.
|
||||
|
||||
## Environment variables
|
||||
|
||||
@@ -74,12 +101,21 @@ endpoint does not implement it, you type the model id instead — the setup stil
|
||||
| `SHIRO_API_KEY` | overrides `apiKey` |
|
||||
| `ANTHROPIC_API_KEY` | used when `provider` is `anthropic` and no key is set |
|
||||
| `OPENAI_API_KEY` | used when `provider` is `openai` and no key is set |
|
||||
| `SHIRO_HOME` | relocates config, sessions, memory, history, and user skills |
|
||||
| `SHIRO_HOME` | relocates config, sessions, memory, history, user skills, and installs |
|
||||
| `SHIRO_INSTALL_DIR` | where `install:local` and the installers put the binary |
|
||||
| `SHIRO_REPO` | which GitHub repo the installers download from |
|
||||
| `SHIRO_VERSION` | pins the version the installers fetch |
|
||||
|
||||
`SHIRO_HOME` is what the test suite uses to keep a run out of your real config.
|
||||
`SHIRO_HOME` is what the test suite uses to keep a run out of your real config. It is also the
|
||||
way to run two isolated setups side by side — a work profile and a personal one — since it moves
|
||||
every piece of state at once:
|
||||
|
||||
```bash
|
||||
SHIRO_HOME=~/work-shiro shiro
|
||||
```
|
||||
|
||||
A key on the command line ends up in your shell history and in `ps`. `SHIRO_API_KEY` in front of
|
||||
one command is better; `/provider` writing to `config.json` is better still.
|
||||
|
||||
## Flags
|
||||
|
||||
@@ -103,13 +139,31 @@ cat file | shiro -p prompt read from stdin
|
||||
| `--no-mcp` | skip MCP servers |
|
||||
| `--no-subagent` | omit the `task` tool |
|
||||
| `--no-instructions` | ignore `AGENTS.md` and friends |
|
||||
| `--no-skills` | ignore builtin and project skills |
|
||||
| `--no-plugins` | disable all plugins, including the guard |
|
||||
| `--no-skills` | ignore builtin, installed, and project skills |
|
||||
| `--no-plugins` | disable all plugins, builtin and installed, including the guard |
|
||||
| `--no-memory` | do not load or write project memory |
|
||||
| `--yolo` | skip every approval prompt |
|
||||
| `-v`, `--version` | version, bun version, platform, source or compiled |
|
||||
| `-h`, `--help` | usage |
|
||||
|
||||
The `--no-*` flags exist for isolating a problem. All six together strip the agent to its
|
||||
built-in tools and nothing else, which answers "is this the loop or something layered on it?"
|
||||
in one run:
|
||||
|
||||
```bash
|
||||
shiro --no-plugins --no-skills --no-memory --no-instructions --no-subagent --no-mcp
|
||||
```
|
||||
|
||||
`--no-plugins` also disables the guard, so `rm -rf` becomes an ordinary approval prompt.
|
||||
Reasonable while debugging, not something to leave on.
|
||||
|
||||
An unknown value fails at startup with the valid list rather than falling back silently:
|
||||
|
||||
```
|
||||
$ shiro --agent turbo
|
||||
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review
|
||||
```
|
||||
|
||||
## Where things live
|
||||
|
||||
```
|
||||
@@ -133,7 +187,8 @@ Project files:
|
||||
```
|
||||
|
||||
Memory and history file names are SHA-256 prefixes of the absolute project path, because a
|
||||
path is not a safe filename.
|
||||
path is not a safe filename. Two consequences: moving a project loses its memory and history,
|
||||
and two checkouts of the same repo at different paths keep separate ones.
|
||||
|
||||
## OpenAI reasoning models
|
||||
|
||||
@@ -142,5 +197,31 @@ Newer OpenAI models reject function tools on `/v1/chat/completions` and require
|
||||
on the first switches to the second, sticks for the rest of the session, and prints one
|
||||
notice. Retryable failures — 429 and 5xx — are left to the SDK's backoff instead.
|
||||
|
||||
Only those six codes qualify, because they mean "this endpoint cannot serve this request shape".
|
||||
A 401 is a wrong key and switching endpoints would only produce a second 401 with a more
|
||||
confusing message.
|
||||
|
||||
The switch is sticky on purpose: once an endpoint rejects the shape it rejects every later step
|
||||
too, so re-probing would waste a round trip per step of every turn.
|
||||
|
||||
Third-party endpoints get a plain chat-completions model with no fallback probe, since they
|
||||
do not implement `/v1/responses`.
|
||||
|
||||
The two endpoints also differ in how they carry assistant history, which is where compaction gets
|
||||
interesting — see [memory](memory.md#the-pruning-repair).
|
||||
|
||||
## Config that changes behaviour subtly
|
||||
|
||||
Three fields do more than they look like they do.
|
||||
|
||||
**`thinking`** costs money and latency on every turn, not just hard ones. `off` on a hard problem
|
||||
produces confident wrong answers; `max` on a rename wastes cents and seconds. The agent variants
|
||||
already pick sensible levels — see [agents](agents.md).
|
||||
|
||||
**`toolSets`** removes tools from the model's view entirely. If the agent stops using a tool you
|
||||
expected, check the startup header for which sets loaded: an unrecognised name is dropped
|
||||
silently, so `"gti"` reads as "git is off". See [tools](tools.md#tool-sets).
|
||||
|
||||
**`registryUrl`** is the whole trust decision for installed skills and plugins. There are no
|
||||
signatures, so pointing it at an index means trusting whoever controls that URL — including for
|
||||
whatever they publish later. See [registry](registry.md).
|
||||
|
||||
@@ -74,6 +74,29 @@ told and can respond to it.
|
||||
if shiro -p "does this build?" --yolo; then echo ok; else echo failed; fi
|
||||
```
|
||||
|
||||
That distinction is deliberate and it has a consequence: **a successful run says nothing about
|
||||
whether the answer was yes.** `0` means the turn completed, not that the build passed. To gate CI
|
||||
on the content, read the output:
|
||||
|
||||
```bash
|
||||
shiro -p "Does this build? Answer only YES or NO." --json --yolo \
|
||||
| jq -r 'select(.type=="text") | .text' | grep -q YES
|
||||
```
|
||||
|
||||
Anything that must fail the build has to be asserted on text or, better, on the exit code of a
|
||||
real command the agent ran.
|
||||
|
||||
## Timeouts
|
||||
|
||||
There is no wall-clock limit on a headless run. Three things bound it:
|
||||
|
||||
- `maxSteps` per variant — 12 for `quick`, 50 by default, 80 for `deep`.
|
||||
- The `timeout` the model passes to `bash`, 120 s by default and 600 s at most.
|
||||
- Whatever your CI runner enforces, which is the only hard stop.
|
||||
|
||||
Interactively `ctrl-c` kills one command and keeps the turn. Headless has no terminal for that, so
|
||||
a signal ends the run. In CI, prefer `--agent quick` and a runner timeout over hoping.
|
||||
|
||||
## Sessions
|
||||
|
||||
Headless runs save like interactive ones, so `-c` picks up where one left off:
|
||||
@@ -83,6 +106,14 @@ shiro -p "start the refactor" --yolo
|
||||
shiro -p "now update the tests" --yolo -c
|
||||
```
|
||||
|
||||
Useful, and worth knowing the shape of: each `-p` run is **one turn**, and `-c` resumes the newest
|
||||
session for that directory. Two concurrent runs in the same directory therefore fight over the
|
||||
same session, and the second overwrites the first. Pass `-r <id>` to keep parallel runs separate,
|
||||
or point them at different `SHIRO_HOME` directories.
|
||||
|
||||
Memory also accumulates. An unattended loop calling `remember` writes to the project store like
|
||||
any other run, so `--no-memory` is worth considering for a job that runs on every push.
|
||||
|
||||
## What is withheld
|
||||
|
||||
The `ask` tool is not offered at all, rather than being offered and left to hang. The model
|
||||
@@ -131,8 +162,34 @@ env:
|
||||
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
||||
```
|
||||
|
||||
Two more worth having in a workflow. Trim the tool schema to what the job needs, since a CI run
|
||||
pays for it on every step:
|
||||
|
||||
```yaml
|
||||
- run: echo '{ "toolSets": [] }' > ~/.shiro-neko/config.json
|
||||
```
|
||||
|
||||
And keep an unattended job from inheriting an installed skill nobody reviewed:
|
||||
|
||||
```yaml
|
||||
- run: shiro -p "..." --agent review --no-skills --no-plugins
|
||||
```
|
||||
|
||||
`--no-skills` matters more in CI than locally: a skill installed from a registry is instructions
|
||||
in the system prompt, and CI is exactly where nobody is watching what it says. See
|
||||
[registry](registry.md).
|
||||
|
||||
## Cost control
|
||||
|
||||
Headless runs are unattended, so a runaway loop costs real money. `--agent quick` caps the
|
||||
step count at 12, and `{ "toolSets": [] }` trims the schema sent every request. There is no
|
||||
spend ceiling yet — see [TODO.md](../TODO.md).
|
||||
|
||||
What a run actually costs is in the `done` event, so a wrapper can total it:
|
||||
|
||||
```bash
|
||||
shiro -p "..." --json --yolo | jq -r 'select(.type=="done" and .inputTokens) | "\(.inputTokens) in, \(.outputTokens) out"'
|
||||
```
|
||||
|
||||
The token fields are optional: an aborted turn emits `done` with neither, which is why the filter
|
||||
checks for one rather than assuming it.
|
||||
|
||||
+46
-5
@@ -28,11 +28,24 @@ them in `~/.shiro-neko/config.json` and they appear alongside the builtins.
|
||||
```
|
||||
|
||||
**stdio** servers take `command`, and optionally `args`, `env`, `cwd`. The process is spawned
|
||||
at startup and closed on exit.
|
||||
at startup and closed on exit. `env` is merged over the inherited environment, so a server
|
||||
inherits your `PATH` unless you replace it.
|
||||
|
||||
**Remote** servers take `url`, and optionally `type` (`http` or `sse`, default `http`) and
|
||||
`headers`.
|
||||
|
||||
A token in `headers` sits in `config.json` in plain text, same as `apiKey`. For anything beyond
|
||||
a local dev token, prefer a stdio server that reads its own credential from the environment.
|
||||
|
||||
## Startup cost
|
||||
|
||||
Servers connect **in parallel**, so the slowest one sets how long startup takes rather than the
|
||||
sum of them. `npx -y some-server` re-resolves the package on each launch; installing it and
|
||||
calling the binary directly is usually the difference between a noticeable wait and none.
|
||||
|
||||
`--no-mcp` skips them all, which is also the quickest way to tell whether a slow start is MCP
|
||||
or something else.
|
||||
|
||||
## Naming
|
||||
|
||||
Tools arrive as `mcp__<server>__<tool>`. A server named `fs` exposing `read_file` becomes
|
||||
@@ -79,14 +92,42 @@ calling one.
|
||||
|
||||
## Cost
|
||||
|
||||
Each tool adds roughly 550 characters of schema to every request. A server exposing twenty
|
||||
tools costs about 2,750 tokens per turn, sent whether or not the model uses any of them.
|
||||
Each tool adds its name, description, and JSON schema to every request. The built-ins average
|
||||
548 bytes; MCP tools vary with how verbose the server's schema is. A server exposing twenty
|
||||
tools costs roughly 2,750 tokens per turn, sent whether or not the model uses any of them.
|
||||
|
||||
Prefer servers with a focused tool set. If one exposes many tools you never use, it is worth
|
||||
finding a narrower server or writing one.
|
||||
MCP tools are **not** covered by `toolSets` — that budget only governs the built-ins. There is
|
||||
no per-server switch either, so the choice is a server or no server, and `--no-mcp` for all of
|
||||
them. If one exposes many tools you never use, a narrower server is worth finding or writing.
|
||||
|
||||
`/tools` shows the count both ways:
|
||||
|
||||
```
|
||||
tools
|
||||
26 offered this turn of 26 registered
|
||||
```
|
||||
|
||||
A gap between the two numbers means a tool set or a read-only agent variant is withholding
|
||||
something. MCP tools never appear in that gap.
|
||||
|
||||
## Writing a server
|
||||
|
||||
Any MCP-compliant server works. A minimal stdio one needs three methods: `initialize`,
|
||||
`tools/list`, and `tools/call`. The test suite includes one at
|
||||
`test/fixtures/mcp-stub.ts` — about 50 lines, and useful as a starting point.
|
||||
|
||||
The suite runs it as a **real subprocess** rather than mocking the transport, because the parts
|
||||
that break in practice are the handshake and the framing, and a mock asserts neither.
|
||||
|
||||
## Debugging a server
|
||||
|
||||
A server that starts but returns nothing useful is the harder case. In order of speed:
|
||||
|
||||
1. `/tools` — did the tools arrive at all? A server with no tools is a `tools/list` problem.
|
||||
2. `shiro -p "call mcp__x__y with ..." --json --yolo` — the exact `tool-call` input and
|
||||
`tool-result` output, one JSON object per line.
|
||||
3. Run the server by hand: `echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | your-server`.
|
||||
If that is wrong, nothing above it can be right.
|
||||
|
||||
For an HTTP server, `curl -X POST $URL -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'`
|
||||
answers the same question without shiro in the way.
|
||||
|
||||
+49
-1
@@ -13,6 +13,9 @@ The split exists because compaction is destructive. `pruneMessages` deletes tool
|
||||
`/compact` deletes the whole transcript, so anything recorded only in messages is lost
|
||||
exactly when a long task needs it most.
|
||||
|
||||
The practical rule: if it should survive this turn, `todo_write` it. If it should survive this
|
||||
session, `remember` it. The transcript is for the conversation, not for storage.
|
||||
|
||||
## Project memory
|
||||
|
||||
Durable notes about the codebase, injected at the start of every session.
|
||||
@@ -32,11 +35,21 @@ text one self-contained line
|
||||
|
||||
Duplicates are refused. Text is capped at 400 characters, the store at 300 entries.
|
||||
|
||||
The kinds are not decoration: they are what the model reads back at boot, and they set how much
|
||||
to trust a note. A `command` is verifiable in one run. A `decision` explains why the obvious
|
||||
alternative was not taken, which is the thing a newcomer most often gets wrong.
|
||||
|
||||
"One self-contained line" is the part that matters most. A note reading "use the new approach"
|
||||
is worthless next session — there is no conversation left to say which approach.
|
||||
|
||||
### `recall`
|
||||
|
||||
Every term must appear. A match increments that entry's hit count, which protects it from
|
||||
compaction later — an entry the agent actually uses is worth keeping verbatim.
|
||||
|
||||
AND rather than OR, on purpose: "migration seed order" should find the one note about that,
|
||||
not every note mentioning any of the three words. Returns the 15 most recent matches.
|
||||
|
||||
### `forget`
|
||||
|
||||
Removes by substring, for a note that turned out wrong.
|
||||
@@ -61,7 +74,9 @@ memory goes stale and a confidently wrong note is worse than none.
|
||||
- Entries with at least one recall are kept verbatim and never merged.
|
||||
- A model returning nothing parseable leaves the store untouched.
|
||||
|
||||
Without the second rule a bad response wipes everything the agent has learned.
|
||||
Without the second rule a bad response wipes everything the agent has learned. The store also
|
||||
has to be past 60 entries with at least two unused ones before anything happens, so `/memory`
|
||||
on a small store is a deliberate no-op rather than a rewrite.
|
||||
|
||||
`/notes` lists the store with hit counts. `--no-memory` disables loading and writing.
|
||||
|
||||
@@ -118,6 +133,14 @@ shiro -r 0193ab2c # by id or unique prefix
|
||||
|
||||
A corrupt session file is skipped rather than crashing the list.
|
||||
|
||||
Ids are UUIDv7, so they sort by creation time and a prefix is usually enough to identify one.
|
||||
`-c` matches on `cwd`, so it picks up the newest session **for this directory** rather than the
|
||||
newest overall — two projects side by side do not steal each other's `-c`.
|
||||
|
||||
Resuming restores the messages and the task list, but not the model or agent variant: those come
|
||||
from the current config and flags. A session started with `--agent deep` resumes as `default`
|
||||
unless you pass it again.
|
||||
|
||||
## Compaction
|
||||
|
||||
Two mechanisms.
|
||||
@@ -130,9 +153,24 @@ screen stays complete. The turn reports it:
|
||||
context compacted: 192 messages pruned to 15 on the wire
|
||||
```
|
||||
|
||||
The status bar warns before that happens: context is shown as a percentage of the threshold,
|
||||
amber from two thirds, red at 90.
|
||||
|
||||
What gets discarded, in order: reasoning items first, then tool calls and their results older
|
||||
than the last three messages. Reasoning is the cheapest thing to lose — it was progress, not
|
||||
conclusions — and tool results are the bulkiest. Recent exchanges are always kept, which is what
|
||||
lets a turn continue rather than restart.
|
||||
|
||||
**Manual**, `/compact`: the model writes a summary — goal, files touched, decisions, commands
|
||||
and outcomes, what remains — and it replaces the transcript entirely.
|
||||
|
||||
The difference is which history is destroyed. Automatic pruning touches the wire only, so
|
||||
scrolling back still shows everything and `/save` records everything. `/compact` replaces the
|
||||
real message array, so it is irreversible for that session.
|
||||
|
||||
Use `/compact` when a session has drifted across several unrelated tasks and the early part is
|
||||
noise. Let automatic pruning handle a single long task, since it keeps the recent work intact.
|
||||
|
||||
### The pruning repair
|
||||
|
||||
Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both were real 400s
|
||||
@@ -171,6 +209,16 @@ and the `tool` message answering it:
|
||||
alone deliberately: a tool call still waiting for its result is what a suspended approval looks
|
||||
like, and dropping it would break `/resume`.
|
||||
|
||||
### What compaction still does not do
|
||||
|
||||
It tells the model the history was pruned but not what was in it. A decision from forty messages
|
||||
ago can be contradicted with confidence, because from the model's side that span never existed.
|
||||
Summarising the discarded part is [next on the list](../TODO.md).
|
||||
|
||||
The threshold is also measured with `JSON.stringify(messages).length / 4`, which is an estimate.
|
||||
It is fine for deciding when to prune and wrong enough that it should not be read as a token
|
||||
count — the real numbers in `/cost` come from the provider.
|
||||
|
||||
## Prompt history
|
||||
|
||||
Per-directory, capped at 200, deduplicated against the previous entry. Up and down in the
|
||||
|
||||
+9
-2
@@ -70,10 +70,10 @@ holding `a` through a batch of edits will approve one of these without reading i
|
||||
| `rm -rf`, `rm -f` | recursive or forced delete |
|
||||
| `git reset --hard` | discards uncommitted work |
|
||||
| `git clean -f` | deletes untracked files |
|
||||
| `git push --force`, `-f` | rewrites remote history |
|
||||
| `git push --force`, `--force-with-lease`, `-f` | rewrites remote history |
|
||||
| `git branch -D` | deletes a branch without a merge check |
|
||||
| `DROP TABLE`, `TRUNCATE` | destroys database data |
|
||||
| `mkfs`, `dd of=/dev/…` | writes to a raw device |
|
||||
| `mkfs`, `dd of=/dev/…`, `> /dev/sd…` | writes to a raw device |
|
||||
| `chmod 777` | makes files world-writable |
|
||||
| `shutdown`, `reboot`, `halt` | affects the whole machine |
|
||||
| `:(){ :\|:& };:` | fork bomb |
|
||||
@@ -88,6 +88,13 @@ The model is told to relay the command rather than work around it. `rm build/one
|
||||
`git push origin feature`, and `git commit` all pass — the patterns target irreversibility,
|
||||
not the commands themselves.
|
||||
|
||||
Two honest limits. The patterns match the command **string**, so `bash -c "$(echo cm0gLXJm | base64 -d)"`
|
||||
is not caught, and neither is a script the agent wrote and then ran. And it only inspects `bash`:
|
||||
a `write_file` overwriting something important is an approval question, not a guard question.
|
||||
|
||||
The guard is the last line before a command runs; `ctrl-c` is the one after. A pattern the guard
|
||||
does not know about is still interruptible by hand — see [tools](tools.md#bash).
|
||||
|
||||
### `time` (default on)
|
||||
|
||||
Adds `current_time`, returning ISO 8601 plus the local string. Auto-approved; it reads
|
||||
|
||||
@@ -126,3 +126,53 @@ add skill:review` disambiguates, and an ambiguous name is refused rather than gu
|
||||
|
||||
A private index is just a URL you control. There is no account, no token, and no telemetry —
|
||||
`/registry` makes exactly one GET for the index and one for the entry you install.
|
||||
|
||||
## Publishing
|
||||
|
||||
Two files and a static host. GitHub raw works, and so does anything that serves JSON over https.
|
||||
|
||||
```
|
||||
your-registry/
|
||||
index.json
|
||||
skills/migration.md
|
||||
plugins/no-secrets.json
|
||||
```
|
||||
|
||||
Three rules the validator enforces, so worth getting right first:
|
||||
|
||||
- The name in `index.json` must match the name inside the file. A skill's frontmatter `name` and a
|
||||
plugin manifest's `name` are both checked against the index entry.
|
||||
- Names are `^[a-z0-9][a-z0-9-]*$`. No uppercase, no dots, no slashes.
|
||||
- A plugin needs at least one deny rule. A manifest with an `appendix` and no rules is prompt
|
||||
text, which is what a skill is for.
|
||||
|
||||
Test it locally before publishing. `registryUrl` accepts `http://localhost`, so:
|
||||
|
||||
```bash
|
||||
cd your-registry && python -m http.server 8000
|
||||
```
|
||||
|
||||
```json
|
||||
{ "registryUrl": "http://localhost:8000/index.json" }
|
||||
```
|
||||
|
||||
`/registry` then exercises the real fetch, the real validation, and the real install path against
|
||||
your files. That is the whole loop, without pushing anything.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**"the registry index is malformed: …"** — the message names the first failing field. The usual
|
||||
causes are an uppercase name, a `url` that is not https, or a plugin entry with no `deny`.
|
||||
|
||||
**"X calls itself Y but the index calls it X"** — the file's own name disagrees with the index.
|
||||
Fix one of the two; the check exists so an index cannot serve something else under a name you
|
||||
trusted.
|
||||
|
||||
**"invalid pattern …"** — a `pathPattern` or `commandPattern` is not a valid regex. Remember it is
|
||||
JSON, so a backslash needs doubling: `\\.env$`, not `\.env$`.
|
||||
|
||||
**Installed but nothing happens** — installs load at startup. Restart, then check `/skills` or
|
||||
`/plugins` for the entry and its origin.
|
||||
|
||||
**In `/plugins` with an error beside it** — the manifest on disk no longer validates. It is skipped
|
||||
rather than fatal, so the agent still starts; `/registry remove` and reinstall.
|
||||
|
||||
+29
-3
@@ -3,9 +3,9 @@
|
||||
A skill is a markdown file with instructions for one kind of task. Only its name and
|
||||
description sit in the system prompt; the body is loaded on demand.
|
||||
|
||||
That split matters. Four bundled skills are 4,659 characters of body but 681 characters of
|
||||
catalogue. Putting every body in the prompt would cost that on every request, for
|
||||
instructions that are relevant to one turn in twenty.
|
||||
That split matters. The four bundled skills are 5,284 characters of body against 681 characters
|
||||
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt
|
||||
would cost that on every turn, for instructions relevant to one turn in twenty.
|
||||
|
||||
## Format
|
||||
|
||||
@@ -28,6 +28,14 @@ Never deploy from a dirty working tree.
|
||||
description is what the model matches against, so write it as a trigger — "use when asked
|
||||
to X" — not as a summary.
|
||||
|
||||
The frontmatter reader handles those two fields and nothing else. A real YAML parser would be a
|
||||
dependency for two strings, so lists, nesting, and multi-line values are not supported: keep both
|
||||
on one line. Quotes around a value are stripped. A body over 20,000 characters is truncated.
|
||||
|
||||
A file that fails to parse is skipped silently rather than reported, which is worth knowing when
|
||||
a skill you wrote does not appear in `/skills` — the usual cause is a missing `---` fence or a
|
||||
description spilling onto a second line.
|
||||
|
||||
## Where they load from
|
||||
|
||||
Four sources, later overriding earlier by name:
|
||||
@@ -79,6 +87,24 @@ before you start working, and follow it as if the user had written it:
|
||||
When the model calls `skill({ name: "debug" })` it gets the full body back and is told to
|
||||
follow it for this task. The call needs no approval — it reads nothing outside the binary.
|
||||
|
||||
"Before you start working" is the load-bearing phrase. A skill loaded after the work is done is
|
||||
wasted tokens, and the failure mode in practice is a model that reads the catalogue, decides it
|
||||
already knows, and never calls the tool. A description written as a trigger is what prevents that.
|
||||
|
||||
Loading one costs its body, once, in that turn's context. A 3,000-character skill is cheaper than
|
||||
one wrong approach it prevents, and more expensive than the catalogue line that would have been
|
||||
enough.
|
||||
|
||||
## Verifying a skill loaded
|
||||
|
||||
```bash
|
||||
shiro -p "fix the failing pagination test" --json --yolo | grep skill
|
||||
```
|
||||
|
||||
`--json` shows the `tool-call` for `skill` with the name it chose, or its absence. If the model
|
||||
never calls it on a task the skill was written for, the description is the thing to change — not
|
||||
the body.
|
||||
|
||||
## Writing a good one
|
||||
|
||||
Skills work when they encode what a newcomer to *your* project would get wrong. The bundled
|
||||
|
||||
+180
-31
@@ -27,19 +27,42 @@ y allow once | a always allow edit_file | n deny
|
||||
`a` whitelists that tool for the rest of the session. `n` tells the model it was denied and
|
||||
to ask what to do instead. `--yolo` skips all prompts.
|
||||
|
||||
The approval is enforced by the SDK, not by the tools. A denied call **provably never
|
||||
executes**: the SDK never reaches the tool's `execute`, so a tool cannot forget to honour a
|
||||
denial or opt out of the check. See [architecture](architecture.md#why-approval-goes-through-the-sdk).
|
||||
|
||||
MCP tools are gated as a group because they are third-party code with unknown side effects —
|
||||
`mcp__fs__read_file` sounds harmless and might not be. See [MCP](mcp.md).
|
||||
|
||||
**The guard runs before all of this.** It is not an approval — it is a refusal, and `--yolo`
|
||||
does not reach it. See [plugins](plugins.md).
|
||||
|
||||
## Tool sets
|
||||
|
||||
Each tool costs roughly 550 characters of JSON schema on every request, and selection
|
||||
accuracy drops as the list grows. Sets let you switch off what a project does not need:
|
||||
Each tool costs its name, its description, and its JSON schema on **every request**. Measured
|
||||
across the fourteen built-ins:
|
||||
|
||||
| Set | Tools |
|
||||
|---|---|
|
||||
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` |
|
||||
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` |
|
||||
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` |
|
||||
| Tool | Bytes | Tool | Bytes |
|
||||
|---|---|---|---|
|
||||
| `read_many_files` | 972 | `git_blame` | 499 |
|
||||
| `multi_edit` | 934 | `git_log` | 484 |
|
||||
| `edit_file` | 618 | `git_diff` | 473 |
|
||||
| `grep` | 595 | `bash` | 466 |
|
||||
| `list_dir` | 594 | `git_show` | 432 |
|
||||
| `read_file` | 526 | `git_status` | 292 |
|
||||
| `glob` | 499 | `write_file` | 289 |
|
||||
|
||||
7,673 bytes for all fourteen, averaging 548. Roughly 1,900 tokens per request before your
|
||||
prompt or the conversation. Selection accuracy also falls as the list grows: a model choosing
|
||||
between six tools picks better than one choosing between twenty.
|
||||
|
||||
Sets let you switch off what a project does not need:
|
||||
|
||||
| Set | Tools | Cost |
|
||||
|---|---|---|
|
||||
| `core` | `read_file` `write_file` `edit_file` `glob` `grep` `bash` | ~2,993 B |
|
||||
| `edit-plus` | `multi_edit` `list_dir` `read_many_files` | ~2,500 B |
|
||||
| `git` | `git_status` `git_diff` `git_log` `git_show` `git_blame` | ~2,180 B |
|
||||
|
||||
```json
|
||||
{ "toolSets": ["edit-plus"] }
|
||||
@@ -50,7 +73,32 @@ agent is not an agent. A disabled set reaches neither the wire nor the system pr
|
||||
a prompt that names an absent tool teaches the model to attempt calls that cannot succeed.
|
||||
Session, plugin, and MCP tools are not part of this budget and are never gated here.
|
||||
|
||||
`/tools` shows which set each live tool came from.
|
||||
An unrecognised set name is dropped silently. The header line at startup shows which sets
|
||||
actually loaded, so a typo reads as "that set is off" rather than as an error — worth checking
|
||||
if a tool you expected is missing.
|
||||
|
||||
`/tools` shows which set each live tool came from:
|
||||
|
||||
```
|
||||
tools
|
||||
20 offered this turn of 22 registered
|
||||
- `bash` core
|
||||
- `git_diff` git
|
||||
- `list_dir` edit-plus
|
||||
- `remember`
|
||||
```
|
||||
|
||||
A tool with no set is a session, plugin, or MCP tool.
|
||||
|
||||
### Which sets to keep
|
||||
|
||||
Both extra sets earn their place in most projects, but not all:
|
||||
|
||||
- **No git in the repo?** `git` is 2,180 bytes the model can never use. Switch it off.
|
||||
- **A model that handles many tools badly?** `{ "toolSets": [] }` trims to six, which is the
|
||||
smallest set that still lets the agent work.
|
||||
- **Reading a lot, editing rarely?** Keep `edit-plus` for `list_dir` and `read_many_files`
|
||||
alone; they pay for themselves in round trips saved.
|
||||
|
||||
## File tools
|
||||
|
||||
@@ -107,7 +155,14 @@ replaceAll replace every occurrence instead of requiring exactly one
|
||||
|
||||
`oldString` must match byte-for-byte and appear exactly once unless `replaceAll` is set.
|
||||
An ambiguous match is an error naming the count, which pushes the model to add surrounding
|
||||
context rather than guessing which occurrence it meant.
|
||||
context rather than guessing which occurrence it meant:
|
||||
|
||||
```
|
||||
oldString appears 3 times in src/users.ts. Add surrounding context or set replaceAll.
|
||||
```
|
||||
|
||||
That error is deliberately specific. `edit failed` would leave the model to retry blind; the
|
||||
count tells it what to do next.
|
||||
|
||||
### `multi_edit`
|
||||
|
||||
@@ -121,7 +176,14 @@ the previous one, so edits may build on each other.
|
||||
|
||||
Atomic: every edit is validated and applied in memory first, so a failure on the third edit
|
||||
leaves the file exactly as it was rather than half-changed. The same uniqueness rule as
|
||||
`edit_file` applies per edit, and the error names which edit failed.
|
||||
`edit_file` applies per edit, and the error names which edit failed:
|
||||
|
||||
```
|
||||
edit 2: oldString not found in src/users.ts. No edits were applied.
|
||||
```
|
||||
|
||||
The last sentence matters. Without it a model reading the error has to guess whether edit 1
|
||||
landed, and its next move — retry the whole batch, or only what failed — depends on the answer.
|
||||
|
||||
### `list_dir`
|
||||
|
||||
@@ -131,9 +193,19 @@ depth levels to descend, 1-6, default 2
|
||||
includeIgnored also show files git ignores
|
||||
```
|
||||
|
||||
Tree view honouring `.gitignore`. Directories end with `/`, files show their size. Past the
|
||||
depth limit the containing directory is still listed, so the shape of the tree stays visible
|
||||
without its contents. Capped at 300 entries.
|
||||
Tree view honouring `.gitignore`. Directories end with `/`, files show their size:
|
||||
|
||||
```
|
||||
.
|
||||
README.md 2B
|
||||
src/
|
||||
app.ts 2K
|
||||
ui/
|
||||
```
|
||||
|
||||
Past the depth limit the containing directory is still listed, so the shape of the tree stays
|
||||
visible without its contents — `src/ui/` above appears at `depth: 2` even though its files do
|
||||
not. Capped at 300 entries.
|
||||
|
||||
### `glob`
|
||||
|
||||
@@ -145,9 +217,12 @@ includeIgnored also return files git ignores
|
||||
|
||||
Walks the tree honouring `.gitignore` and `.shiroignore`, skipping `.git` and
|
||||
`node_modules` unconditionally. Nested ignore files apply only within their own directory,
|
||||
as git does. Returns posix paths relative to the workspace root. A symlinked directory is
|
||||
classified as a directory and not descended into, since it can point anywhere including
|
||||
back into the tree.
|
||||
as git does. Returns posix paths relative to the workspace root.
|
||||
|
||||
A symlinked directory is classified as a directory and not descended into. Both halves matter:
|
||||
`readdir` reports a junction as a non-directory, so without the extra `stat` a symlinked
|
||||
directory leaked past `dir/` ignore rules and was yielded as a file with a nonsense size. Not
|
||||
descending is separate — a link can point anywhere, including back into the tree.
|
||||
|
||||
### `grep`
|
||||
|
||||
@@ -162,6 +237,15 @@ Shells out to ripgrep when it is on PATH — roughly 15x faster on a real repo
|
||||
back to a JavaScript walker otherwise. Output is `path:line: text` either way, so the model
|
||||
sees one format regardless. Skips binaries. Caps at 200 hits.
|
||||
|
||||
Two details keep the two paths in agreement. ripgrep is passed `--no-require-git`, because it
|
||||
otherwise ignores `.gitignore` outside a repository while the JavaScript fallback always honours
|
||||
it. And an rg exit code above 1 means rg could not run the search at all, so the fallback takes
|
||||
over; exit 1 is simply "no matches" and is reported as such.
|
||||
|
||||
Regex syntax differs between the two: ripgrep is Rust regex, the fallback is JavaScript. A
|
||||
pattern using look-around works in the fallback and fails under rg. An invalid pattern is
|
||||
reported as `Invalid regex: <reason>` rather than returning an empty result set.
|
||||
|
||||
### `bash`
|
||||
|
||||
```
|
||||
@@ -194,12 +278,25 @@ running quits as usual.
|
||||
|
||||
All five are read-only and therefore approval-free. Each spawns `git` with a fixed argument
|
||||
array rather than a shell string, so an argument like `--author="; rm -rf /"` can only ever
|
||||
be a literal argument — which is what makes auto-approval safe.
|
||||
be a literal argument — which is what makes auto-approval safe. A test asserts exactly that:
|
||||
`git_log` with the path `; touch pwned.txt` creates no file.
|
||||
|
||||
Output is described rather than raw porcelain: `git_status` names the branch and says
|
||||
`staged modified` or `untracked` per file instead of leaving the model to decode two columns
|
||||
of flags. Outside a repository they fail with `<cwd> is not a git repository.` rather than
|
||||
passing git's own error text through.
|
||||
Output is described rather than raw porcelain. `git_status` names the branch and says
|
||||
`staged modified` or `untracked` per file instead of leaving the model to decode porcelain's two
|
||||
leading columns:
|
||||
|
||||
```
|
||||
On main, 2 changed:
|
||||
src/app.ts (staged modified, modified)
|
||||
new.ts (untracked)
|
||||
```
|
||||
|
||||
That file has a staged change *and* a later unstaged one, which the raw `MM` prefix conveys only
|
||||
to a reader who knows the format.
|
||||
|
||||
Outside a repository they fail with `<cwd> is not a git repository.` rather than passing git's
|
||||
own error text through. Other git failures do pass through, on purpose: `git_show no-such-ref`
|
||||
reports what git said, because git's own message is the most useful thing available.
|
||||
|
||||
```
|
||||
git_status branch, staged, modified, untracked
|
||||
@@ -209,6 +306,13 @@ git_show ref path? one commit: message, author, diff
|
||||
git_blame path startLine? endLine? who last changed each line
|
||||
```
|
||||
|
||||
`git_log` defaults to 15 commits and caps at 40. `git_blame` without a range blames the whole
|
||||
file; with `startLine` and no `endLine` it covers 40 lines from there.
|
||||
|
||||
Everything here is also reachable through `bash`. The reason the set exists anyway is the
|
||||
approval boundary: `bash git diff` stops for a decision on every call, while `git_diff` cannot
|
||||
mutate anything and so never needs one.
|
||||
|
||||
## Agent tools
|
||||
|
||||
### `task`
|
||||
@@ -223,8 +327,17 @@ Spawns a read-only subagent with `read_file`, `glob`, and `grep` only. It return
|
||||
report, so the parent pays for findings rather than the whole search transcript. It sees
|
||||
none of the parent conversation, so its prompt has to stand alone.
|
||||
|
||||
Two properties follow from that tool set rather than from policy: it can never trigger an
|
||||
approval prompt, because it has no gated tools; and the parent's context holds the conclusion
|
||||
instead of the search. A subagent reading forty files to answer one question costs the parent
|
||||
the answer, not the forty files.
|
||||
|
||||
`explore` finds and reports. `review` critiques code in severity order. Progress streams to
|
||||
the subagent panel.
|
||||
the subagent panel. Capped at 20 steps, and it shares the parent's model — an `explore` run
|
||||
pays reasoning rates for what is really a search, which is [on the list](../TODO.md) to fix.
|
||||
|
||||
Not worth delegating a single grep: the subagent is a whole extra model loop, so it wins on a
|
||||
search spanning many files and loses on anything you could answer in one call.
|
||||
|
||||
### `ask`
|
||||
|
||||
@@ -235,9 +348,11 @@ multiple allow more than one
|
||||
```
|
||||
|
||||
Stops the turn and puts the question on screen. With options it is a picker; without, free
|
||||
text. `esc` skips, which tells the model to decide and state its assumption.
|
||||
text. `esc` skips, which returns "the user dismissed this; use your best judgement" — so a
|
||||
dismissal is an instruction to decide, not a dead end.
|
||||
|
||||
Withheld entirely in headless mode — a question with no one to answer it would hang.
|
||||
Withheld entirely in headless mode — a question with no one to answer it would hang. The system
|
||||
prompt says so, and tells the model to decide and state its assumption instead.
|
||||
|
||||
### `todo_write`
|
||||
|
||||
@@ -249,6 +364,9 @@ Statuses: `pending`, `in_progress`, `done`, `blocked`. Send the whole list each
|
||||
replaces the previous one. Warns when more than one task is `in_progress`, when nothing is
|
||||
`in_progress` while work remains, or when a `blocked` task has no note.
|
||||
|
||||
The list lives in the system prompt, rebuilt every step, so it survives both pruning and
|
||||
`/compact`. See [memory](memory.md).
|
||||
|
||||
### `remember`, `recall`, `forget`
|
||||
|
||||
Durable per-project notes. See [memory](memory.md).
|
||||
@@ -264,15 +382,46 @@ Loads the body of a skill. See [skills](skills.md).
|
||||
## Path safety
|
||||
|
||||
Every path a tool receives goes through a jail: resolved against the workspace root, then
|
||||
checked that it did not escape. `../../etc/passwd` and absolute paths outside the root are
|
||||
both refused before any filesystem call.
|
||||
checked that it did not escape.
|
||||
|
||||
```ts
|
||||
export function jail(p: string, root = process.cwd()): string {
|
||||
const abs = isAbsolute(p) ? resolve(p) : resolve(root, p);
|
||||
const rel = relative(resolve(root), abs);
|
||||
if (rel.startsWith('..') || isAbsolute(rel)) throw new Error(`Path escapes workspace: ${p}`);
|
||||
return abs;
|
||||
}
|
||||
```
|
||||
|
||||
`../../etc/passwd`, `a/../../secret`, and absolute paths outside the root are all refused
|
||||
before any filesystem call. Resolving first and comparing after is what catches the middle
|
||||
case: string-prefix checks on the raw input miss `a/../../secret` entirely.
|
||||
|
||||
The model's output is a trust boundary. It can emit any string, so the check happens on
|
||||
every call rather than being assumed.
|
||||
every call rather than being assumed. That includes `read_many_files`, where a bad path is
|
||||
reported in its block like any other unreadable file.
|
||||
|
||||
`jail` guards the workspace, not the shell. `bash` runs whatever it is given, which is why
|
||||
every call needs approval and why the guard plugin exists — see [plugins](plugins.md).
|
||||
|
||||
## Output caps
|
||||
|
||||
Any single tool result is truncated at 30,000 characters with a note saying how much was
|
||||
cut. `grep` stops at 200 hits, `glob` at 200 paths, `list_dir` at 300 entries,
|
||||
`read_many_files` at 20 files, `read_file` at 2000 lines by default. Without caps one `grep`
|
||||
for `function` can end a session.
|
||||
Any single tool result is truncated at 30,000 characters with a note saying how much was cut:
|
||||
|
||||
```
|
||||
... [truncated 41,233 chars]
|
||||
```
|
||||
|
||||
| Tool | Cap |
|
||||
|---|---|
|
||||
| any result | 30,000 characters |
|
||||
| `grep` | 200 hits, each line cut at 300 chars |
|
||||
| `glob` | 200 paths |
|
||||
| `list_dir` | 300 entries |
|
||||
| `read_many_files` | 20 files |
|
||||
| `read_file` | 2,000 lines by default |
|
||||
| `bash` | 120 s default timeout, 600 s max |
|
||||
|
||||
Without caps one `grep` for `function` can end a session. The caps are per call, so a model
|
||||
that needs more can narrow and ask again — which is cheaper than one call that fills the
|
||||
context and forces compaction.
|
||||
|
||||
Reference in New Issue
Block a user