Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual per-tool measurements, and fills in the parts a reader hits after the happy path: what a specific error means, what a setting costs, what is not covered. Measured rather than estimated: - Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the full set, averaging 548. - Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the argument for loading bodies on demand. - Full system prompt 3,571 chars, core-only 2,045. New sections: - tools: which sets to keep and why, the jail function itself, an output-cap table, and the real error strings for edit_file and multi_edit. - configuration: env var per provider preset, cost-estimate limits, what each --no-* flag isolates, and three settings that do more than they look like. - agents: step caps per variant, which variant to reach for, and the fact that reasoning is charged as output and discarded first by compaction. - headless: exit code 0 means "the turn completed", not "the answer was yes" — with the jq pattern for gating on content. Timeouts, concurrent -c runs fighting over one session, CI recipes for --no-skills. - mcp: parallel connect, startup cost, a debugging ladder, and that toolSets does not gate MCP tools. - registry: publishing, local testing over http://localhost, and a troubleshooting section keyed on the actual validator messages. - memory: what compaction discards in what order, /compact versus automatic pruning, and that -c matches on cwd. - skills: the frontmatter reader's limits, and how to verify a skill loaded. Corrections found while cross-checking against the source: - The guard table was missing --force-with-lease and > /dev/sd… - The done event's token fields are optional, so the jq example filters on one rather than assuming it. Two honest limits now written down: the guard matches command strings, so a base64-decoded or script-wrapped command is not caught; and a registry index is trusted for its contents, not its authorship. Verified: all internal links and heading anchors resolve, every docs/ page is reachable from the README, 538 tests pass, typecheck clean.
This commit is contained in:
+29
-3
@@ -3,9 +3,9 @@
|
||||
A skill is a markdown file with instructions for one kind of task. Only its name and
|
||||
description sit in the system prompt; the body is loaded on demand.
|
||||
|
||||
That split matters. Four bundled skills are 4,659 characters of body but 681 characters of
|
||||
catalogue. Putting every body in the prompt would cost that on every request, for
|
||||
instructions that are relevant to one turn in twenty.
|
||||
That split matters. The four bundled skills are 5,284 characters of body against 681 characters
|
||||
of catalogue — an eightfold difference, paid on every request. Putting every body in the prompt
|
||||
would cost that on every turn, for instructions relevant to one turn in twenty.
|
||||
|
||||
## Format
|
||||
|
||||
@@ -28,6 +28,14 @@ Never deploy from a dirty working tree.
|
||||
description is what the model matches against, so write it as a trigger — "use when asked
|
||||
to X" — not as a summary.
|
||||
|
||||
The frontmatter reader handles those two fields and nothing else. A real YAML parser would be a
|
||||
dependency for two strings, so lists, nesting, and multi-line values are not supported: keep both
|
||||
on one line. Quotes around a value are stripped. A body over 20,000 characters is truncated.
|
||||
|
||||
A file that fails to parse is skipped silently rather than reported, which is worth knowing when
|
||||
a skill you wrote does not appear in `/skills` — the usual cause is a missing `---` fence or a
|
||||
description spilling onto a second line.
|
||||
|
||||
## Where they load from
|
||||
|
||||
Four sources, later overriding earlier by name:
|
||||
@@ -79,6 +87,24 @@ before you start working, and follow it as if the user had written it:
|
||||
When the model calls `skill({ name: "debug" })` it gets the full body back and is told to
|
||||
follow it for this task. The call needs no approval — it reads nothing outside the binary.
|
||||
|
||||
"Before you start working" is the load-bearing phrase. A skill loaded after the work is done is
|
||||
wasted tokens, and the failure mode in practice is a model that reads the catalogue, decides it
|
||||
already knows, and never calls the tool. A description written as a trigger is what prevents that.
|
||||
|
||||
Loading one costs its body, once, in that turn's context. A 3,000-character skill is cheaper than
|
||||
one wrong approach it prevents, and more expensive than the catalogue line that would have been
|
||||
enough.
|
||||
|
||||
## Verifying a skill loaded
|
||||
|
||||
```bash
|
||||
shiro -p "fix the failing pagination test" --json --yolo | grep skill
|
||||
```
|
||||
|
||||
`--json` shows the `tool-call` for `skill` with the name it chose, or its absence. If the model
|
||||
never calls it on a task the skill was written for, the description is the thing to change — not
|
||||
the body.
|
||||
|
||||
## Writing a good one
|
||||
|
||||
Skills work when they encode what a newcomer to *your* project would get wrong. The bundled
|
||||
|
||||
Reference in New Issue
Block a user