Expand the documentation with measured figures and operational detail

Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.

Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
  full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
  argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.

New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
  table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
  --no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
  reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
  with the jq pattern for gating on content. Timeouts, concurrent -c runs
  fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
  does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
  troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
  pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.

Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
  rather than assuming it.

Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.

Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
This commit is contained in:
Muhammad Zakir Ramadhan
2026-09-03 09:26:37 +07:00
parent 84c60f2022
commit 9b978fdbe1
10 changed files with 583 additions and 72 deletions
+29 -3
View File
@@ -3,9 +3,9 @@
A skill is a markdown file with instructions for one kind of task. Only its name and
description sit in the system prompt; the body is loaded on demand.
That split matters. Four bundled skills are 4,659 characters of body but 681 characters of
catalogue. Putting every body in the prompt would cost that on every request, for
instructions that are relevant to one turn in twenty.
That split matters. The four bundled skills are 5,284 characters of body against 681 characters
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt
would cost that on every turn, for instructions relevant to one turn in twenty.
## Format
@@ -28,6 +28,14 @@ Never deploy from a dirty working tree.
description is what the model matches against, so write it as a trigger — "use when asked
to X" — not as a summary.
The frontmatter reader handles those two fields and nothing else. A real YAML parser would be a
dependency for two strings, so lists, nesting, and multi-line values are not supported: keep both
on one line. Quotes around a value are stripped. A body over 20,000 characters is truncated.
A file that fails to parse is skipped silently rather than reported, which is worth knowing when
a skill you wrote does not appear in `/skills` — the usual cause is a missing `---` fence or a
description spilling onto a second line.
## Where they load from
Four sources, later overriding earlier by name:
@@ -79,6 +87,24 @@ before you start working, and follow it as if the user had written it:
When the model calls `skill({ name: "debug" })` it gets the full body back and is told to
follow it for this task. The call needs no approval — it reads nothing outside the binary.
"Before you start working" is the load-bearing phrase. A skill loaded after the work is done is
wasted tokens, and the failure mode in practice is a model that reads the catalogue, decides it
already knows, and never calls the tool. A description written as a trigger is what prevents that.
Loading one costs its body, once, in that turn's context. A 3,000-character skill is cheaper than
one wrong approach it prevents, and more expensive than the catalogue line that would have been
enough.
## Verifying a skill loaded
```bash
shiro -p "fix the failing pagination test" --json --yolo | grep skill
```
`--json` shows the `tool-call` for `skill` with the name it chose, or its absence. If the model
never calls it on a task the skill was written for, the description is the thing to change — not
the body.
## Writing a good one
Skills work when they encode what a newcomer to *your* project would get wrong. The bundled