Files
shiro-neko/docs/configuration.md
T
Muhammad Zakir Ramadhan 9b978fdbe1 Expand the documentation with measured figures and operational detail
Most of this replaces "roughly 550 characters per tool" with the actual
per-tool measurements, and fills in the parts a reader hits after the happy
path: what a specific error means, what a setting costs, what is not covered.

Measured rather than estimated:
- Per-tool byte cost, all fourteen, and the per-set totals. 7,673 B for the
  full set, averaging 548.
- Builtin skill bodies at 5,284 B against a 681 B catalogue, which is the
  argument for loading bodies on demand.
- Full system prompt 3,571 chars, core-only 2,045.

New sections:
- tools: which sets to keep and why, the jail function itself, an output-cap
  table, and the real error strings for edit_file and multi_edit.
- configuration: env var per provider preset, cost-estimate limits, what each
  --no-* flag isolates, and three settings that do more than they look like.
- agents: step caps per variant, which variant to reach for, and the fact that
  reasoning is charged as output and discarded first by compaction.
- headless: exit code 0 means "the turn completed", not "the answer was yes" —
  with the jq pattern for gating on content. Timeouts, concurrent -c runs
  fighting over one session, CI recipes for --no-skills.
- mcp: parallel connect, startup cost, a debugging ladder, and that toolSets
  does not gate MCP tools.
- registry: publishing, local testing over http://localhost, and a
  troubleshooting section keyed on the actual validator messages.
- memory: what compaction discards in what order, /compact versus automatic
  pruning, and that -c matches on cwd.
- skills: the frontmatter reader's limits, and how to verify a skill loaded.

Corrections found while cross-checking against the source:
- The guard table was missing --force-with-lease and > /dev/sd…
- The done event's token fields are optional, so the jq example filters on one
  rather than assuming it.

Two honest limits now written down: the guard matches command strings, so a
base64-decoded or script-wrapped command is not caught; and a registry index is
trusted for its contents, not its authorship.

Verified: all internal links and heading anchors resolve, every docs/ page is
reachable from the README, 538 tests pass, typecheck clean.
2026-09-03 09:26:37 +07:00

9.7 KiB

Configuration

Settings come from three places. Later wins:

  1. ~/.shiro-neko/config.json
  2. environment variables
  3. command-line flags

The config file

Written by /provider, editable by hand. Every field is optional.

{
  "provider": "openai",
  "model": "gpt-5",
  "baseURL": "https://api.openai.com/v1",
  "apiKey": "sk-...",
  "presetId": "openai",
  "agent": "default",
  "thinking": "medium",
  "maxRetries": 3,
  "plugins": ["guard", "time"],
  "toolSets": ["edit-plus", "git"],
  "registryUrl": "https://example.com/my-registry/index.json",
  "mcpServers": {
    "fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] }
  }
}
Field Meaning
provider wire protocol: anthropic or openai. Not the vendor — Groq, OpenRouter, and Ollama all speak openai
model model id as the endpoint names it
baseURL API root. Defaults to the official endpoint for the provider
apiKey sent as Authorization: Bearer for openai, x-api-key for anthropic
presetId which preset /provider chose, so it can show what is configured
agent default variant: default, quick, deep, plan, review
thinking default level: off, low, medium, high, max
maxRetries retries per model call for transient failures. Default 3
plugins which builtin plugins to enable. Omit for ["guard", "time"]
toolSets optional tool sets beyond core: edit-plus, git. Omit for all of them. See tools
registryUrl index for /registry. Omit for the default. See registry
mcpServers see MCP

Provider presets

/provider offers these. Each sets baseURL and the wire protocol for you.

Preset Protocol Endpoint Env var checked
Anthropic anthropic api.anthropic.com/v1 ANTHROPIC_API_KEY
OpenAI openai api.openai.com/v1 OPENAI_API_KEY
OpenRouter openai openrouter.ai/api/v1 OPENROUTER_API_KEY
Groq openai api.groq.com/openai/v1 GROQ_API_KEY
DeepSeek openai api.deepseek.com/v1 DEEPSEEK_API_KEY
xAI openai api.x.ai/v1 XAI_API_KEY
Ollama openai localhost:11434/v1 none, keyless
LM Studio openai localhost:1234/v1 none, keyless
Custom OpenAI-compatible openai you supply it none
Custom Anthropic-compatible anthropic you supply it none

provider is the wire protocol, not the vendor. Groq, DeepSeek, xAI, OpenRouter, Ollama, and LM Studio all speak openai; only Anthropic speaks anthropic. Two things differ between them: the auth header (Authorization: Bearer versus x-api-key), and how thinking levels map.

After the key is entered, GET /v1/models is called and the list becomes a picker. Both protocols expose that endpoint with the same data[].id shape, so one code path handles both. If the endpoint does not implement it, a preset with a known model list falls back to that; otherwise you type the model id and setup still completes.

Anything the picker offers is a model the endpoint actually reports, which is more reliable than a hard-coded list — that is why the fallback lists are short and only exist for Anthropic and OpenAI.

Cost estimates

/cost and the status bar price a turn from a table in src/pricing.ts, matched by longest prefix on the model id, so claude-sonnet-4-5-20250929 resolves via claude-sonnet-4-5. An OpenRouter-style anthropic/claude-sonnet-4-5 has its vendor prefix stripped first.

An unknown model is reported as unpriced rather than guessed:

4210 in / 88 out tokens (llama-3.3-70b is unpriced)

Two limits worth knowing. The rates are hand-entered and drift as vendors change them, so treat the figure as an estimate, not a bill. And the token counts come from the provider's usage report, while ~ctx in the status bar is JSON.stringify(messages).length / 4 — good enough to decide when to compact, wrong enough that it should not be read as a token count.

Environment variables

Variable Effect
SHIRO_PROVIDER overrides provider
SHIRO_MODEL overrides model
SHIRO_BASE_URL overrides baseURL
SHIRO_API_KEY overrides apiKey
ANTHROPIC_API_KEY used when provider is anthropic and no key is set
OPENAI_API_KEY used when provider is openai and no key is set
SHIRO_HOME relocates config, sessions, memory, history, user skills, and installs
SHIRO_INSTALL_DIR where install:local and the installers put the binary
SHIRO_REPO which GitHub repo the installers download from
SHIRO_VERSION pins the version the installers fetch

SHIRO_HOME is what the test suite uses to keep a run out of your real config. It is also the way to run two isolated setups side by side — a work profile and a personal one — since it moves every piece of state at once:

SHIRO_HOME=~/work-shiro shiro

A key on the command line ends up in your shell history and in ps. SHIRO_API_KEY in front of one command is better; /provider writing to config.json is better still.

Flags

shiro [options]
shiro -p "prompt"          headless, prints to stdout
cat file | shiro -p        prompt read from stdin
Flag Effect
-p, --print [prompt] headless mode. Needs --yolo for tool use
--json with -p, one JSON event per line
-c, --continue resume the newest session for this directory
-r, --resume <id> resume by session id or unique prefix
--agent <name> default, quick, deep, plan, review
--think <level> off, low, medium, high, max
--provider <name> anthropic or openai
--model <id> model id
--base-url <url> API root
--no-mcp skip MCP servers
--no-subagent omit the task tool
--no-instructions ignore AGENTS.md and friends
--no-skills ignore builtin, installed, and project skills
--no-plugins disable all plugins, builtin and installed, including the guard
--no-memory do not load or write project memory
--yolo skip every approval prompt
-v, --version version, bun version, platform, source or compiled
-h, --help usage

The --no-* flags exist for isolating a problem. All six together strip the agent to its built-in tools and nothing else, which answers "is this the loop or something layered on it?" in one run:

shiro --no-plugins --no-skills --no-memory --no-instructions --no-subagent --no-mcp

--no-plugins also disables the guard, so rm -rf becomes an ordinary approval prompt. Reasonable while debugging, not something to leave on.

An unknown value fails at startup with the valid list rather than falling back silently:

$ shiro --agent turbo
shiro: Unknown agent "turbo". Available: default, quick, deep, plan, review

Where things live

~/.shiro-neko/
  config.json                 provider, model, key, defaults
  sessions/<uuid>.json        transcripts, token counts, cost, task list
  memory/<hash>.json          durable per-project notes
  history/<hash>.json         prompt history for up-arrow recall
  skills/*.md                 your own skills
  registry/skills/*.md        skills installed with /registry
  registry/plugins/*.json     plugin manifests installed with /registry

Project files:

<project>/
  AGENTS.md                   instructions injected into the system prompt
  .shiro/skills/*.md          project skills, override user and builtin
  .shiroignore                extra ignore rules on top of .gitignore

Memory and history file names are SHA-256 prefixes of the absolute project path, because a path is not a safe filename. Two consequences: moving a project loses its memory and history, and two checkouts of the same repo at different paths keep separate ones.

OpenAI reasoning models

Newer OpenAI models reject function tools on /v1/chat/completions and require /v1/responses. For api.openai.com both are chained: a 400, 404, 405, 415, 422, or 501 on the first switches to the second, sticks for the rest of the session, and prints one notice. Retryable failures — 429 and 5xx — are left to the SDK's backoff instead.

Only those six codes qualify, because they mean "this endpoint cannot serve this request shape". A 401 is a wrong key and switching endpoints would only produce a second 401 with a more confusing message.

The switch is sticky on purpose: once an endpoint rejects the shape it rejects every later step too, so re-probing would waste a round trip per step of every turn.

Third-party endpoints get a plain chat-completions model with no fallback probe, since they do not implement /v1/responses.

The two endpoints also differ in how they carry assistant history, which is where compaction gets interesting — see memory.

Config that changes behaviour subtly

Three fields do more than they look like they do.

thinking costs money and latency on every turn, not just hard ones. off on a hard problem produces confident wrong answers; max on a rename wastes cents and seconds. The agent variants already pick sensible levels — see agents.

toolSets removes tools from the model's view entirely. If the agent stops using a tool you expected, check the startup header for which sets loaded: an unrecognised name is dropped silently, so "gti" reads as "git is off". See tools.

registryUrl is the whole trust decision for installed skills and plugins. There are no signatures, so pointing it at an index means trusting whoever controls that URL — including for whatever they publish later. See registry.