Files
shiro-neko/docs/agents.md
T
Muhammad Zakir RamadhanandSisyphus 0dc0819fad Expand the documentation with measured figures and operational detail
Every guide catches up with the sixteen-tool registry: apply_patch and web_fetch sections, the delegation guide with the worker kind, bounded compaction replacing the three-message window description, net in the tool-set tables, six bundled skills, and the module map gains tools-net.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-09-03 16:43:27 +07:00

8.3 KiB

Agents and thinking

An agent variant sets three things: how much the model deliberates, which tools it is offered, and a behaviour appendix in the system prompt.

shiro --agent deep              # at launch
shiro --agent plan --think low  # variant with an overridden level
/agent          picker
/agent review   direct
/think          picker
/think max      direct

The variants

Variant Thinking Tools Steps For
default medium all 50 ordinary work
quick off all 12 small, well-scoped edits
deep max all 80 hard problems, unclear causes
plan high read-only 50 investigate and propose
review high read-only 50 critique a change

quick tells the model not to deliberate, not to write a task list, and not to explore beyond what the change needs. Good for a rename or a one-line fix where thinking budget is pure latency.

deep asks for more than one hypothesis before acting, more reading before concluding, and findings recorded with remember so they survive compaction.

plan and review are genuinely read-only. write_file, edit_file, multi_edit, apply_patch, and bash are withheld from the model, not merely discouraged in prose — a model that cannot see a tool cannot call it. They keep everything that only reads, including read_many_files, list_dir, the git tools, and web_fetch when the net set is enabled. Their prompts also forbid describing edits as if they had been made.

Variants and tool sets

Two separate things narrow the tool list, and they compose.

A variant withholds tools by capability: plan cannot write, whatever the config says. toolSets withholds them by cost: a project that never wants the git tools switches that set off for every variant. See tools.

Both go through one function, so a withheld tool is missing from the wire and from the system prompt together. /tools lists what is actually offered this turn, with the set each tool came from.

Thinking levels

off, low, medium, high, max. They map to whatever the provider actually supports:

Level OpenAI reasoning_effort Anthropic thinking
off none { type: "disabled" }
low low budget_tokens: 6400
medium medium proportional budget
high high budget_tokens: 38400
max xhigh maximum budget

Verified against both wire formats rather than assumed. The vocabulary is deliberately ours: off through max means the same thing whichever provider is configured, and switching providers mid-session does not change what /think high asks for.

Higher costs more and takes longer. off on a hard problem produces confident wrong answers; max on a rename wastes a few cents and several seconds. The variants pick sensible defaults, so reach for /think only when a specific turn needs something else.

Reasoning is also charged as output tokens, so max shows up in /cost even on a turn where the model wrote two lines. And reasoning is the first thing compaction discards — see memory — so a long turn at max pays for thinking that will not be on the wire by the end of it.

Steps

maxSteps caps how many model calls one turn may make. A step is one request: a tool call and its result, or the final text.

Variant Steps
quick 12
default, plan, review 50
deep 80

The cap is a backstop against a loop, not a budget to spend. A turn that hits it stops mid-work with whatever it has, which is why quick's 12 suits a rename and would strand a refactor. If turns regularly hit the cap on the same kind of task, the task wants deep rather than a higher number.

Overriding

--agent deep --think low gives you deep's tools, steps, and appendix with a low thinking budget. The override clones the preset rather than mutating it, so a later /agent deep in the same session still gets max.

Defaults in config

{ "agent": "deep", "thinking": "high" }

A flag beats the config file. An unknown name fails at startup with the valid list rather than silently falling back:

$ shiro --agent turbo
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review

What the variant changes in the prompt

The system prompt describes only the tools actually offered, and the workflow rules adapt. Under plan the model is told it has no tools that change anything and that it cannot run commands, so it should say what to run rather than claim it passed. Under default it is told which tools need approval and to verify with the project's tests.

A prompt that describes a withheld tool teaches the model to attempt calls that cannot succeed, which is why the description is generated from the live tool set.

Three rules flip on what is available:

Condition default says plan says
can edit "these need approval; if denied, stop and ask" "you have no tools that change anything"
can run commands "verify with the project's build or tests" "say what should be run rather than claiming it passed"
can ask "ask rather than guess when two readings differ" (same, unless headless)

The read-only variants are around 2,000 characters of system prompt against roughly 3,600 for the full set — cheaper per turn as well as safer.

Which to reach for

  • default for anything you have not thought about. It is the right answer most of the time.
  • quick for a rename, a typo, a one-line fix. Its value is not the model being cheaper but the absence of deliberation latency on work that needs none.
  • deep when the first attempt already failed, or the cause is unclear. Asking for more than one hypothesis is the actual difference; the thinking budget is secondary.
  • plan before a change you are not sure about. Read-only means the plan cannot quietly become a half-applied edit.
  • review on a diff or a module. In headless CI this is the one that needs no --yolo, because it holds no tool that can modify anything — see headless.

Switching mid-session is fine and cheap: /agent changes the next turn's tools and prompt, and nothing about the history.

Delegating with task

The task tool spans a separate axis from the variants: it runs a subagent with its own context window, so the parent pays for one report rather than the whole search transcript. The subagent sees none of the parent's conversation, so its prompt must stand alone.

Kind Tools Approval For
explore (default) read and search only never prompts — structurally read-only a search spanning many files
review read and search only never prompts a critique of code or a diff
worker everything, including writes every write and command asks, through the parent's gate a self-contained change whose steps you do not need to watch

Three properties of the worker kind are structural rather than policy:

The gate is the parent's. A worker routes each gated call back through the same permission rules, the same guard plugins, and the same approval prompt as a direct call — flagged a worker subagent wants to run ... so you can tell who is asking. Answering always grants the pattern for the session exactly as it does for you. A subagent that could approve its own writes would be a way to launder a tool call past you, so there is no separate, weaker gate.

Denial stops the work. The worker is told a denial is your decision: report it, do not work around it. The tool descriptions say the same thing, so the rule survives compaction.

No worker without a channel. In headless runs there is no one to answer a prompt, so the worker kind is not offered at all — an unattended write is not something to fall into by accident. The read-only kinds work everywhere. No subagent holds web_fetch; network access stays with the main agent, where the approval prompt says what it is for.

When not to delegate: a single grep, or anything you must supervise step by step — keep that in your own turn, where every call is on screen. A worker wins when the intermediate steps are noise: a mechanical rename across twenty files, a test scaffold written to match an existing suite, a cleanup whose shape you already know.