Fix the loop stalling after compaction, add an external registry
The compaction bug, which is the important one:
beta.2 taught the pruner to drop any assistant part whose reasoning item it had
removed. That was right about the 400 and wrong about everything else. On a
reasoning model every tool call carries a provider itemId, so past the threshold
the model could no longer see what it had already run, and re-ran the same tools
until maxSteps ended the turn. Reproduced at 12 model calls for a job needing 4,
with nothing but the user message reaching the wire.
The dependency is not the part, it is the itemId. A part carrying one is
serialised as `{ type: 'item_reference', id }`, a pointer to an item stored
provider-side that depends on its reasoning item. Without the itemId the same
content goes out inline and carries no dependency at all. Verified against the
provider's own serialiser: `text` with an itemId becomes item_reference, the
identical part without one becomes output_text.
So `dropOrphanedItems` becomes `detachOrphanedItems`: strip the itemId, keep the
content. Compaction may shorten the history; it must not blank it. The new test
asserts behaviour rather than shape — the loop must end because the model chose
to, and every call after the first must still carry the earlier exchange. A shape
assertion passed the whole time the model was losing its memory.
Registry, via `/registry [list|search|add|remove|installed]`:
Skills and plugins are treated differently on purpose. A skill is prompt text, so
installing one puts a stranger's words into the system prompt of every future
session in this project; the install shows the body first and the origin is
recorded, so /skills always says where an instruction came from. A plugin is a
JSON manifest of deny rules, evaluated by compiled code identical for every
install. Loading TypeScript from a URL is declined outright: a plugin that can
block tool calls could otherwise lie about blocking them.
Validated before anything is written: https only (file: and data: rejected), name
matched against ^[a-z0-9][a-z0-9-]*$ so it cannot escape its directory, size
caps on index and body, every regex compiled, pattern length capped since it runs
on every tool call, and the body's own name checked against the index. Installed
skills rank below your own, so an install can never shadow a skill you wrote.
Interface:
- Context is a percentage of the compaction threshold, amber from two thirds and
red at 90. A turn about to lose history now says so beforehand.
- Aligned command menu and registry tables; /skills and /plugins name origins.
538 tests, up from 488. The registry is tested against a real local HTTP server,
and the guard is proven to refuse a .env write end to end rather than assumed to.
This commit is contained in:
@@ -88,6 +88,10 @@ the agent puts a question on screen with options.
|
||||
message, so a search across forty files does not fill the main context. Its progress
|
||||
streams to a panel.
|
||||
|
||||
**Extensible from the prompt.** `/registry` browses external skills and plugins and installs
|
||||
them with one confirmation. A skill is shown in full before its text joins your system prompt;
|
||||
a plugin is a manifest of refusal rules, never code.
|
||||
|
||||
**Remembers between sessions.** Decisions, working commands, and traps go into per-project
|
||||
memory that is injected at the start of every future session.
|
||||
|
||||
@@ -109,6 +113,7 @@ core ones and a disabled set reaches neither the wire nor the prompt.
|
||||
| [Agents and thinking](docs/agents.md) | variants, thinking levels, read-only modes |
|
||||
| [Skills](docs/skills.md) | the bundled skills and writing your own |
|
||||
| [Plugins](docs/plugins.md) | the plugin interface and the builtins |
|
||||
| [Registry](docs/registry.md) | installing external skills and plugins |
|
||||
| [Memory and state](docs/memory.md) | memory, task lists, sessions, compaction |
|
||||
| [MCP](docs/mcp.md) | connecting Model Context Protocol servers |
|
||||
| [Headless mode](docs/headless.md) | `-p`, JSON events, exit codes, CI recipes |
|
||||
@@ -123,8 +128,9 @@ Type `/` and a menu appears, narrowing as you type.
|
||||
|
||||
```
|
||||
/help /agent [name] /think [level] /provider /models /model <id>
|
||||
/skills /plugins /init /context /todos /notes /memory
|
||||
/tools /compact /cost /sessions /resume <id> /save /clear /exit
|
||||
/skills /plugins /registry [search|add|remove] /init /context
|
||||
/todos /notes /memory /tools /compact /cost
|
||||
/sessions /resume <id> /save /clear /exit
|
||||
```
|
||||
|
||||
`esc` dismisses a panel, interrupts a running turn, and clears the queue. `ctrl-c` kills the
|
||||
@@ -136,11 +142,11 @@ workspace path. Up and down recall earlier prompts.
|
||||
Working: the agent loop, tool approvals, subagents, skills, plugins, per-project memory,
|
||||
session persistence, MCP, markdown rendering, headless mode, five-platform builds, streaming
|
||||
reasoning display, the mid-turn prompt queue, gateable tool sets, read-only git tools, batch
|
||||
reads, `@file` completion, and interruptible commands.
|
||||
reads, `@file` completion, interruptible commands, and the external registry.
|
||||
|
||||
Next up is in [TODO.md](TODO.md); the longer view and what has been declined are in
|
||||
[ROADMAP.md](ROADMAP.md). The short version of what is missing: a summary of what compaction
|
||||
discarded, `web_fetch`, and a cheaper model for subagent searches.
|
||||
discarded, `web_fetch`, a spend ceiling, and a cheaper model for subagent searches.
|
||||
|
||||
## License
|
||||
|
||||
|
||||
+33
-5
@@ -95,15 +95,35 @@ fails with a message saying the command did not finish and its effects are unkno
|
||||
model takes its next step from there. The kill takes the whole process tree, because killing
|
||||
`cmd /c` alone leaves the real command holding both pipes open and the read never returns.
|
||||
|
||||
### 0.1.0-beta.4 (unreleased)
|
||||
|
||||
**Compaction no longer stops the loop.** The beta.2 repair dropped any assistant part whose
|
||||
reasoning item pruning had removed. On a reasoning model that is every tool call, so past the
|
||||
threshold the model could no longer see what it had already run — and re-ran it until the step
|
||||
limit ended the turn. The fix strips the provider `itemId` rather than the part: without one the
|
||||
same content is serialised inline instead of as an `item_reference`, so the dependency on the
|
||||
pruned reasoning item disappears while the history survives. Compaction may shorten the
|
||||
history; it must not blank it.
|
||||
|
||||
**External registry.** `/registry` browses, searches, installs, and removes skills and plugins
|
||||
from an index over https. The two kinds are treated differently on purpose: a skill is prompt
|
||||
text and is shown in full before it joins your system prompt, while a plugin is a validated
|
||||
manifest of deny rules that the compiled guard evaluates. Loading code from a URL is declined
|
||||
outright — a plugin that can block tool calls could otherwise lie about blocking them.
|
||||
|
||||
**Interface.** Context shown as a percentage of the compaction threshold, amber from two
|
||||
thirds and red at 90, so a turn about to lose history says so first. Aligned command menu and
|
||||
registry tables, and `/skills` and `/plugins` name the origin of every entry.
|
||||
|
||||
---
|
||||
|
||||
## Next
|
||||
|
||||
### Lossless-enough compaction
|
||||
|
||||
Compaction says the history was pruned but not what was in it, so the model can contradict
|
||||
its own earlier decision with confidence. A summary of the discarded span costs one cheap
|
||||
call and removes the whole class of problem.
|
||||
Compaction keeps the model's memory of a turn now, but it still says nothing about the messages
|
||||
it discarded, so the model can contradict its own earlier decision with confidence. A summary of
|
||||
the discarded span costs one cheap call and removes the whole class of problem.
|
||||
|
||||
### Cost control
|
||||
|
||||
@@ -117,6 +137,12 @@ and a per-session ceiling are both small changes on top of the pricing that alre
|
||||
and forgotten in the other is a silently ungated write. Marking each tool where it is defined,
|
||||
and checking the coverage in the suite, removes the failure mode rather than documenting it.
|
||||
|
||||
### Registry trust
|
||||
|
||||
An index is trusted for its contents, not its authorship: `registryUrl` is the whole trust
|
||||
decision, and there are no signatures. Publisher keys and a pinned digest per entry would make
|
||||
"install this skill" a decision about a specific artifact rather than about a URL.
|
||||
|
||||
### `web_fetch`
|
||||
|
||||
URL to markdown, in a `net` set that is off by default — it is the one tool that leaves the
|
||||
@@ -139,8 +165,10 @@ losing the original.
|
||||
**Structured diff review.** Approve or reject individual hunks of an `edit_file` call rather
|
||||
than the whole thing.
|
||||
|
||||
**Plugin loading from disk.** Plugins are compiled in. Loading `.shiro/plugins/*.ts` needs a
|
||||
sandbox story first — a plugin that can block tool calls can also lie about blocking them.
|
||||
**Plugin code from disk.** Declarative manifests ship in beta.4, and that is the whole of it
|
||||
for now. Loading `.shiro/plugins/*.ts` needs a sandbox story first — a plugin that can block
|
||||
tool calls can also lie about blocking them, and one that can execute can read whatever the
|
||||
agent can read.
|
||||
|
||||
**Prompt caching.** Anthropic and OpenAI both support it. The system prompt is rebuilt every
|
||||
step for task-list freshness, which defeats a naive cache; splitting the stable prefix from
|
||||
|
||||
@@ -10,23 +10,15 @@ Longer-term direction lives in [ROADMAP.md](ROADMAP.md).
|
||||
|
||||
### Summarize the pruned span
|
||||
|
||||
Compaction tells the model the history was pruned but not what was in it, so it can
|
||||
confidently contradict a decision it made forty messages ago.
|
||||
Compaction now keeps the model's memory of a turn, but it still tells the model nothing about
|
||||
the messages it dropped, so a decision from forty messages ago can be contradicted with
|
||||
confidence.
|
||||
|
||||
- [ ] Summarize the discarded messages before dropping them
|
||||
- [ ] Inject the summary in place of the count
|
||||
- [ ] Budget it: a summary that grows with the session defeats the point
|
||||
- [ ] Test: a pruned decision is still recoverable from the summary
|
||||
|
||||
### A cheaper model for subagents
|
||||
|
||||
The subagent shares the parent's model. An `explore` run is search, not reasoning, and it
|
||||
currently pays the parent's per-token rate.
|
||||
|
||||
- [ ] `subagentModel` in config, defaulting to the parent
|
||||
- [ ] `/cost` separates parent from subagent spend
|
||||
- [ ] Test: the subagent's calls go to the configured model, the parent's do not
|
||||
|
||||
### A spend ceiling
|
||||
|
||||
A headless run that loops costs real money with nothing to stop it.
|
||||
@@ -36,20 +28,28 @@ A headless run that loops costs real money with nothing to stop it.
|
||||
- [ ] Headless exits non-zero with the ceiling named, rather than stopping silently
|
||||
- [ ] Test: a session past its ceiling refuses the next turn and says why
|
||||
|
||||
### A cheaper model for subagents
|
||||
|
||||
The subagent shares the parent's model. An `explore` run is search, not reasoning, and it
|
||||
currently pays the parent's per-token rate.
|
||||
|
||||
- [ ] `subagentModel` in config, defaulting to the parent
|
||||
- [ ] `/cost` separates parent from subagent spend
|
||||
- [ ] Test: the subagent's calls go to the configured model, the parent's do not
|
||||
|
||||
### Hot-reload an installed entry
|
||||
|
||||
`/registry add` writes the file and says to restart. The skill catalogue and the guard chain
|
||||
are both assembled at boot, so a mid-session install does nothing until then.
|
||||
|
||||
- [ ] Rebuild the skill list and plugin host after an install or removal
|
||||
- [ ] Leave a turn in flight alone: its rules must not change underneath it
|
||||
- [ ] Test: a skill installed mid-session is callable in the next turn without a restart
|
||||
|
||||
---
|
||||
|
||||
## Next
|
||||
|
||||
### Summarize the pruned span
|
||||
|
||||
Compaction tells the model the history was pruned but not what was in it, so it can
|
||||
confidently contradict a decision it made forty messages ago.
|
||||
|
||||
- [ ] Summarize the discarded messages before dropping them
|
||||
- [ ] Inject the summary in place of the count
|
||||
- [ ] Budget it: a summary that grows with the session defeats the point
|
||||
- [ ] Test: a pruned decision is still recoverable from the summary
|
||||
|
||||
### `web_fetch`
|
||||
|
||||
- [ ] URL to markdown, size-capped
|
||||
@@ -106,6 +106,11 @@ Not bugs exactly, but things that will bite someone.
|
||||
how far a half-run migration got.
|
||||
- **`@` completion lists files, not directories.** `@src/` narrows correctly, but you cannot
|
||||
complete to `src/` itself, because the walker only yields files.
|
||||
- **An installed skill is a stranger's words in your system prompt.** The install shows the
|
||||
body first and `/skills` records the origin, but nothing re-checks it later: a registry that
|
||||
changes a URL's contents affects the next install, not one already on disk.
|
||||
- **A registry index is trusted for its contents, not its authorship.** There are no
|
||||
signatures. `registryUrl` is the whole trust decision.
|
||||
|
||||
---
|
||||
|
||||
@@ -127,3 +132,10 @@ Kept for one release, then deleted.
|
||||
- [x] `ctrl-c` kills the running command and keeps the turn. The kill takes the whole process
|
||||
tree: killing `cmd /c` alone left the real command holding both pipes open, so the
|
||||
interrupt appeared to do nothing for 19 seconds
|
||||
- [x] **Compaction no longer stops the loop.** Pruning used to drop any assistant part whose
|
||||
reasoning item it removed, which on a reasoning model is every tool call. The model lost
|
||||
its record of what it had run and re-ran it until the step limit. The repair strips the
|
||||
provider `itemId` instead of the part, so the same content is sent inline
|
||||
- [x] `/registry`: browse, search, install, and remove external skills and plugins. Skills are
|
||||
shown in full before install; plugins are a validated manifest of deny rules, never code
|
||||
- [x] Context shown as a percentage of the compaction threshold, amber at two thirds, red at 90
|
||||
|
||||
+35
-9
@@ -176,16 +176,25 @@ Only `api.openai.com` gets the chain. Third-party endpoints do not implement `/v
|
||||
|
||||
## Compaction and its repair
|
||||
|
||||
Pruning breaks two different provider invariants, and `src/prune.ts` repairs both.
|
||||
Pruning breaks two provider invariants, and `src/prune.ts` repairs both.
|
||||
|
||||
**A message without its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a
|
||||
reasoning item and keeps the message item from the same response. The responses API treats the
|
||||
message as that reasoning item's dependent and returns 400.
|
||||
**A message detached from its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a
|
||||
reasoning item and keeps the message item from the same response. A part carrying a provider
|
||||
`itemId` is not sent inline: the responses provider serialises it as
|
||||
`{ type: 'item_reference', id }`, pointing at an item stored on their side, and that stored
|
||||
item depends on the reasoning item that pruning just removed. The result is a 400.
|
||||
|
||||
The two carry different ids, so they cannot be matched by id. What links them is the assistant
|
||||
message they arrived in: one message is one response, and its reasoning item covers every
|
||||
other item in it. `dropOrphanedItems` drops the dependent parts of any turn whose reasoning was
|
||||
removed — which costs nothing, since pruning was already discarding those turns.
|
||||
other item in it. `detachOrphanedItems` strips the `itemId` from those parts, which is what
|
||||
sends the same content inline instead — verified against the provider's own serialiser, where
|
||||
a `text` part with an itemId goes out as `item_reference` and the identical part without one
|
||||
goes out as `output_text`.
|
||||
|
||||
Dropping the parts was the first attempt and it broke the loop. On a reasoning model every
|
||||
tool call carries an itemId, so after the first compaction the model could not see what it had
|
||||
already run, and re-ran the same tools until the step limit ended the turn. **Compaction may
|
||||
shorten the history; it must not blank it.**
|
||||
|
||||
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts
|
||||
*messages*, so the cut lands between an assistant `tool-call` and the `tool` message answering
|
||||
@@ -199,6 +208,17 @@ it. What reaches the wire is a `function_call_output` with no `function_call`:
|
||||
reverse pairing is deliberately left alone: a call still awaiting its result is exactly what a
|
||||
suspended approval looks like, and dropping it would break resume.
|
||||
|
||||
## Registry
|
||||
|
||||
`/registry` fetches an index of external skills and plugins over https. Skills are prompt text
|
||||
and are shown in full before install; plugins are a JSON manifest of refusal rules, never code.
|
||||
The guard evaluating those rules is compiled, identical for every installed plugin, so an entry
|
||||
from a registry cannot execute anything. See [registry](registry.md) for the validation and
|
||||
the reasoning.
|
||||
|
||||
`src/registry.ts` has no UI and no side effects until `install()` is called, which is what lets
|
||||
`stage()` show a body before it becomes part of every future prompt.
|
||||
|
||||
## Module map
|
||||
|
||||
| Module | Responsibility |
|
||||
@@ -208,6 +228,7 @@ suspended approval looks like, and dropping it would break resume.
|
||||
| `tools-git.ts` | read-only git tools, spawned with a fixed argv |
|
||||
| `ignore.ts` | gitignore-aware walker, path jail |
|
||||
| `complete.ts` | `@path` token extraction, ranking, insertion |
|
||||
| `registry.ts` | external index, validation, install and removal |
|
||||
| `prompt.ts` | system prompt assembly from live state |
|
||||
| `agents.ts` | variants, thinking levels |
|
||||
| `skills.ts` | discovery, catalogue, `skill` tool |
|
||||
@@ -234,10 +255,15 @@ Every module is pure of the UI except `ui/`, and `ui/` never touches the SDK. Th
|
||||
|
||||
## Testing
|
||||
|
||||
482 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop;
|
||||
538 tests, no mocking framework. `MockLanguageModelV4` from `ai/test` drives the loop;
|
||||
`ink-testing-library` drives the UI with real keystrokes; MCP is tested against a real stdio
|
||||
server subprocess; provider wire formats are tested against a local HTTP server; the interrupt
|
||||
path spawns a real subprocess and asserts it died early rather than ran out.
|
||||
server subprocess; provider wire formats and the registry are tested against a local HTTP
|
||||
server; the interrupt path spawns a real subprocess and asserts it died early rather than ran
|
||||
out.
|
||||
|
||||
The pattern throughout is to assert on what actually crossed a boundary — what went on the
|
||||
wire, what is on screen, what is on disk — rather than on internal calls.
|
||||
|
||||
Compaction is asserted on **behaviour**, not shape: the loop must terminate because the model
|
||||
chose to, and every call after the first must still carry the earlier exchange. A shape
|
||||
assertion would have passed while the model was losing its memory.
|
||||
|
||||
@@ -22,6 +22,7 @@ Written by `/provider`, editable by hand. Every field is optional.
|
||||
"maxRetries": 3,
|
||||
"plugins": ["guard", "time"],
|
||||
"toolSets": ["edit-plus", "git"],
|
||||
"registryUrl": "https://example.com/my-registry/index.json",
|
||||
"mcpServers": {
|
||||
"fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] }
|
||||
}
|
||||
@@ -38,8 +39,9 @@ Written by `/provider`, editable by hand. Every field is optional.
|
||||
| `agent` | default variant: `default`, `quick`, `deep`, `plan`, `review` |
|
||||
| `thinking` | default level: `off`, `low`, `medium`, `high`, `max` |
|
||||
| `maxRetries` | retries per model call for transient failures. Default 3 |
|
||||
| `plugins` | which plugins to enable. Omit for `["guard", "time"]` |
|
||||
| `plugins` | which builtin plugins to enable. Omit for `["guard", "time"]` |
|
||||
| `toolSets` | optional tool sets beyond `core`: `edit-plus`, `git`. Omit for all of them. See [tools](tools.md) |
|
||||
| `registryUrl` | index for `/registry`. Omit for the default. See [registry](registry.md) |
|
||||
| `mcpServers` | see [MCP](mcp.md) |
|
||||
|
||||
## Provider presets
|
||||
@@ -117,6 +119,8 @@ cat file | shiro -p prompt read from stdin
|
||||
memory/<hash>.json durable per-project notes
|
||||
history/<hash>.json prompt history for up-arrow recall
|
||||
skills/*.md your own skills
|
||||
registry/skills/*.md skills installed with /registry
|
||||
registry/plugins/*.json plugin manifests installed with /registry
|
||||
```
|
||||
|
||||
Project files:
|
||||
|
||||
+4
-1
@@ -16,7 +16,7 @@ faster and the fallback path is exercised without it.
|
||||
```bash
|
||||
bun run shiro # run from source
|
||||
bun run typecheck # tsc --noEmit
|
||||
bun test # 482 tests
|
||||
bun test # 538 tests
|
||||
bun run build # single binary for this platform -> dist/shiro
|
||||
bun run release # all five platforms -> dist/release + SHA256SUMS
|
||||
bun run install:local # build, then copy onto PATH
|
||||
@@ -71,6 +71,9 @@ mock-verification test:
|
||||
request body
|
||||
- `pruneMessages` leaving a tool result without its tool call — same, and it took a stub
|
||||
endpoint that rejected the pairing to prove the fix
|
||||
- Compaction blanking the model's memory of its own tool calls — invisible in any single
|
||||
request, and visible only as "the loop ran to its step limit". Caught by asserting the loop
|
||||
terminated because the model chose to, not that the messages had a particular shape
|
||||
- `--json` serialising `Error` as `{}` — visible only in the printed output
|
||||
- Automatic approval requests prompting the user — visible only in the event sequence
|
||||
- `ctrl-c` killing `cmd /c` but not the command under it — visible only as elapsed time, since
|
||||
|
||||
+13
-5
@@ -139,17 +139,25 @@ Pruning breaks two provider invariants. `src/prune.ts` repairs both, and both we
|
||||
before it did.
|
||||
|
||||
**A message without its reasoning item.** `pruneMessages({ reasoning: 'all' })` strips a
|
||||
reasoning item and keeps the message item from the same response. The OpenAI responses API
|
||||
treats the message as a dependent of that reasoning item and rejects the request:
|
||||
reasoning item and keeps the message item from the same response. That message carries a
|
||||
provider `itemId`, and the OpenAI responses provider serialises anything with one as
|
||||
`{ type: 'item_reference', id }` — a pointer to an item stored on their side, which depends on
|
||||
the reasoning item that is now gone:
|
||||
|
||||
```
|
||||
400 Item 'msg_…' of type 'message' was provided without its required 'reasoning' item: 'rs_…'
|
||||
```
|
||||
|
||||
The two carry different ids, so they cannot be matched by id. What links them is the
|
||||
assistant message they arrived in — one message is one response. `dropOrphanedItems` drops the
|
||||
dependent parts of any turn whose reasoning was removed. That costs nothing, because pruning
|
||||
was already discarding those turns.
|
||||
assistant message they arrived in — one message is one response. `detachOrphanedItems` strips
|
||||
the `itemId` from those parts. Without one the same content is serialised **inline**, which
|
||||
carries no dependency on anything stored, so the turn survives intact.
|
||||
|
||||
Dropping the parts instead was the first attempt, and it was wrong in a way that only showed
|
||||
up over a long turn: on a reasoning model every tool call carries an itemId, so after the
|
||||
first compaction the model could no longer see what it had already run. It re-ran the same
|
||||
tools until the step limit ended the turn. The history is the model's memory; compaction may
|
||||
shorten it but must not blank it.
|
||||
|
||||
**A tool result without its tool call.** `toolCalls: 'before-last-3-messages'` counts
|
||||
*messages*, not pairs, so the cut can land between the assistant message holding a `tool-call`
|
||||
|
||||
+25
-9
@@ -4,8 +4,17 @@ A plugin extends the agent in four ways: it can add tools, mark tools auto-appro
|
||||
a tool call before it runs, and append to the system prompt. It can also run something after
|
||||
each turn.
|
||||
|
||||
Plugins are compiled into the binary. Loading them from disk is deliberately not supported
|
||||
yet — see [ROADMAP.md](../ROADMAP.md).
|
||||
Two kinds exist, and only one can contain code:
|
||||
|
||||
- **Builtin** plugins are compiled into the binary and may do anything in the interface below.
|
||||
- **Installed** plugins come from a registry as a JSON manifest of refusal rules. They are
|
||||
data: the guard evaluating them is compiled code, identical for every install. See
|
||||
[registry](registry.md).
|
||||
|
||||
Loading TypeScript from disk or a URL is deliberately not supported. A plugin that can block
|
||||
tool calls can also lie about blocking them, and one that could execute could read every file
|
||||
the agent can read. That is a sandbox problem, not a loader problem — see
|
||||
[ROADMAP.md](../ROADMAP.md).
|
||||
|
||||
## Enabling
|
||||
|
||||
@@ -13,9 +22,12 @@ yet — see [ROADMAP.md](../ROADMAP.md).
|
||||
{ "plugins": ["guard", "time"] }
|
||||
```
|
||||
|
||||
That is also the default when the field is absent. `--no-plugins` disables all of them,
|
||||
including the guard. `/plugins` lists what is active and reports any name that did not
|
||||
resolve.
|
||||
That is also the default when the field is absent, and it lists **builtin** plugins only.
|
||||
Installed plugins are always active once present, because installing one was the decision to
|
||||
enable it; remove it with `/registry remove <name>`.
|
||||
|
||||
`--no-plugins` disables everything, builtin and installed, including the guard. `/plugins`
|
||||
lists what is active, marks installed entries, and reports any name that did not resolve.
|
||||
|
||||
## The interface
|
||||
|
||||
@@ -92,7 +104,11 @@ intrusive, but it is genuinely useful when a turn takes minutes.
|
||||
|
||||
## Writing one
|
||||
|
||||
Plugins live in `src/plugins-builtin.ts` and are registered in `BUILTIN_PLUGINS`.
|
||||
A refusal rule is usually better as an installed manifest: no rebuild, and nothing to review.
|
||||
See [registry](registry.md) for the manifest shape. Reach for a builtin only when the plugin
|
||||
needs to contribute a tool or run something after a turn.
|
||||
|
||||
Builtin plugins live in `src/plugins-builtin.ts` and are registered in `BUILTIN_PLUGINS`.
|
||||
|
||||
```ts
|
||||
export const noSecretsPlugin: Plugin = {
|
||||
@@ -122,6 +138,6 @@ refusal it was never told about and tries to route around it.
|
||||
|
||||
## Ordering
|
||||
|
||||
Plugins run in the order they are enabled. The first `beforeToolCall` to block wins;
|
||||
later hooks are not consulted. `afterTurn` runs every hook, and one throwing does not stop
|
||||
the rest.
|
||||
Builtin plugins run first, in the order they are enabled, then installed ones. The first
|
||||
`beforeToolCall` to block wins; later hooks are not consulted. `afterTurn` runs every hook, and
|
||||
one throwing does not stop the rest.
|
||||
|
||||
@@ -0,0 +1,128 @@
|
||||
# Registry
|
||||
|
||||
External skills and plugins, browsed and installed from the CLI.
|
||||
|
||||
```
|
||||
/registry everything in the index
|
||||
/registry search <query> narrow by name or description
|
||||
/registry installed what is already here
|
||||
/registry add <name> fetch, show, then install on confirmation
|
||||
/registry remove <name> delete an installed entry
|
||||
```
|
||||
|
||||
```
|
||||
registry
|
||||
3 of 3 available
|
||||
S migration Write or run a database migration
|
||||
S commit-style Write commits the way this team does
|
||||
P no-secrets Refuses writes to credential files installed
|
||||
S skill P plugin | /registry add <name>
|
||||
```
|
||||
|
||||
`esc` dismisses the panel.
|
||||
|
||||
## The two kinds are not equally safe
|
||||
|
||||
**A skill is instructions.** Installing one puts a stranger's words into the system prompt of
|
||||
every future session in this project. That is prompt injection by invitation, so the install
|
||||
shows the body first and `/skills` always says the origin:
|
||||
|
||||
```
|
||||
install skill "migration"?
|
||||
https://raw.githubusercontent.com/example/registry/main/skills/migration.md
|
||||
|
||||
Migrations live in `db/migrations/` and are timestamped, never renumbered.
|
||||
|
||||
Run `bun run db:migrate` locally first. Staging runs them on deploy.
|
||||
|
||||
A skill is instructions the agent follows. This text joins your system prompt.
|
||||
y install | n cancel
|
||||
```
|
||||
|
||||
**A plugin is data.** Never code. A manifest declares refusal rules; the guard that evaluates
|
||||
them is the same compiled code for every installed plugin:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "no-secrets",
|
||||
"description": "Refuses writes to credential files",
|
||||
"appendix": "The no-secrets plugin refuses writes to .env and credential files.",
|
||||
"deny": [
|
||||
{
|
||||
"tools": ["write_file", "edit_file", "multi_edit"],
|
||||
"pathPattern": "(^|/)\\.env|credentials|\\.pem$",
|
||||
"reason": "refusing to write a credential file; add secrets yourself"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Loading TypeScript from a URL is not offered at any price. A plugin can block tool calls, so
|
||||
one that could also execute could read every file the agent can read and lie about blocking
|
||||
anything. See [plugins](plugins.md) for why this boundary exists.
|
||||
|
||||
## What is validated before anything is written
|
||||
|
||||
| Check | Why |
|
||||
|---|---|
|
||||
| `https` only, `localhost` for tests | `file:` would read a local path, `data:` would inline a payload |
|
||||
| Name matches `^[a-z0-9][a-z0-9-]*$` | the name becomes a filename, so `../evil` must not parse |
|
||||
| Index at most 256 KB, body at most 64 KB | a hostile index should not exhaust memory |
|
||||
| Manifest against a strict schema | extra keys like `beforeToolCall` are dropped, not honoured |
|
||||
| Every pattern compiles as a regex | a broken pattern would fail on the first tool call instead |
|
||||
| Pattern at most 200 characters | it runs on every tool call; a pathological one is a denial of service |
|
||||
| Body name matches the index name | an index entry cannot serve something else under a trusted name |
|
||||
| At least one deny rule | a plugin with no rules is only prompt text, which is what a skill is for |
|
||||
|
||||
A malformed installed plugin is reported by `/plugins` and skipped. One bad install does not
|
||||
stop the agent from starting.
|
||||
|
||||
## Where installs land
|
||||
|
||||
```
|
||||
~/.shiro-neko/registry/
|
||||
skills/<name>.md loaded as origin "registry"
|
||||
plugins/<name>.json loaded as a declarative plugin
|
||||
```
|
||||
|
||||
Precedence for skills, low to high: **builtin → registry → user → project**. A skill you wrote
|
||||
in `~/.shiro-neko/skills/` or `.shiro/skills/` always beats one fetched from a registry, so an
|
||||
install can never silently shadow your own work.
|
||||
|
||||
Installs take effect on the next start. A skill joins the system prompt and a plugin joins the
|
||||
guard chain, and both are assembled once at boot; hot-swapping either mid-session would mean a
|
||||
turn whose rules changed underneath it.
|
||||
|
||||
## Pointing at your own index
|
||||
|
||||
```json
|
||||
{ "registryUrl": "https://example.com/my-registry/index.json" }
|
||||
```
|
||||
|
||||
The index is one JSON document:
|
||||
|
||||
```json
|
||||
{
|
||||
"skills": [
|
||||
{
|
||||
"name": "migration",
|
||||
"description": "Write or run a database migration",
|
||||
"url": "https://example.com/skills/migration.md",
|
||||
"author": "you"
|
||||
}
|
||||
],
|
||||
"plugins": [
|
||||
{
|
||||
"name": "no-secrets",
|
||||
"description": "Refuses writes to credential files",
|
||||
"url": "https://example.com/plugins/no-secrets.json"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Both arrays are optional. A name may appear once as a skill and once as a plugin; `/registry
|
||||
add skill:review` disambiguates, and an ambiguous name is refused rather than guessed.
|
||||
|
||||
A private index is just a URL you control. There is no account, no token, and no telemetry —
|
||||
`/registry` makes exactly one GET for the index and one for the entry you install.
|
||||
+6
-4
@@ -30,14 +30,16 @@ to X" — not as a summary.
|
||||
|
||||
## Where they load from
|
||||
|
||||
Three sources, later overriding earlier by name:
|
||||
Four sources, later overriding earlier by name:
|
||||
|
||||
1. **builtin** — compiled into the binary
|
||||
2. **user** — `~/.shiro-neko/skills/*.md`
|
||||
3. **project** — `.shiro/skills/*.md`
|
||||
2. **registry** — `~/.shiro-neko/registry/skills/*.md`, installed with `/registry add`
|
||||
3. **user** — `~/.shiro-neko/skills/*.md`
|
||||
4. **project** — `.shiro/skills/*.md`
|
||||
|
||||
A project skill named `debug` replaces the bundled one entirely. `/skills` shows what
|
||||
loaded and where each came from.
|
||||
loaded and where each came from — which matters most for `registry`, since that body came
|
||||
from someone else and is now in your system prompt. See [registry](registry.md).
|
||||
|
||||
`--no-skills` skips all of them, builtin included.
|
||||
|
||||
|
||||
+97
-7
@@ -14,6 +14,7 @@ import { costOf } from './pricing';
|
||||
import { BUILTIN_PLUGINS, DEFAULT_ENABLED } from './plugins-builtin';
|
||||
import { createHost } from './plugins';
|
||||
import { fetchModels, presetById } from './providers';
|
||||
import * as registry from './registry';
|
||||
import { Session } from './session';
|
||||
import { loadSkills } from './skills';
|
||||
import * as store from './store';
|
||||
@@ -21,6 +22,7 @@ import { createTaskTool } from './subagent';
|
||||
import { VERSION, versionLine } from './version';
|
||||
import { createAskBridge } from './ui/Ask';
|
||||
import { App, createApprovalBridge, createNoticeBus, createSubagentBus, type AppHooks } from './ui/App';
|
||||
import type { RegistryRow as AppRegistryRow } from './ui/Panels';
|
||||
|
||||
// SDK warnings go straight to stderr, which tears up the Ink render.
|
||||
(globalThis as { AI_SDK_LOG_WARNINGS?: boolean }).AI_SDK_LOG_WARNINGS = false;
|
||||
@@ -56,13 +58,14 @@ first run: start shiro with no key and it opens provider setup, or use /provider
|
||||
config: ${configPath()}
|
||||
{ "provider": "openai", "model": "gpt-5", "apiKey": "...",
|
||||
"agent": "default", "thinking": "medium", "plugins": ["guard", "time"],
|
||||
"toolSets": ["edit-plus", "git"],
|
||||
"toolSets": ["edit-plus", "git"], "registryUrl": "https://...",
|
||||
"mcpServers": { "fs": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."] } } }
|
||||
|
||||
env: SHIRO_PROVIDER SHIRO_MODEL SHIRO_BASE_URL SHIRO_API_KEY
|
||||
ANTHROPIC_API_KEY OPENAI_API_KEY
|
||||
|
||||
skills: builtin, plus ~/.shiro-neko/skills/*.md and .shiro/skills/*.md
|
||||
registry: /registry to browse and install external skills and plugins
|
||||
sessions: ${store.sessionsDir()}
|
||||
in-session: /help for the command list`;
|
||||
|
||||
@@ -150,6 +153,40 @@ const instructions = has('--no-instructions') ? [] : await loadInstructions();
|
||||
const skills = has('--no-skills') ? [] : await loadSkills();
|
||||
const promptHistory = await store.loadHistory();
|
||||
|
||||
const installedPlugins = has('--no-plugins') ? { plugins: [], errors: [] } : await registry.loadInstalledPlugins();
|
||||
|
||||
/** Installed entries, as `kind:name`, so the registry list can mark what is already here. */
|
||||
async function installedNames(): Promise<Set<string>> {
|
||||
const names = new Set<string>();
|
||||
for (const s of skills) if (s.origin === 'registry') names.add(`skill:${s.name}`);
|
||||
for (const p of installedPlugins.plugins) names.add(`plugin:${p.name}`);
|
||||
return names;
|
||||
}
|
||||
|
||||
/**
|
||||
* One entry by name, from the index.
|
||||
*
|
||||
* A name alone is ambiguous when a skill and a plugin share it, so `skill:name`
|
||||
* disambiguates. Asking rather than guessing would be worse here: the two differ in
|
||||
* what they can do, and picking one silently is the wrong default.
|
||||
*/
|
||||
async function findEntry(name: string): Promise<registry.RegistryEntry> {
|
||||
const parsed = /^(skill|plugin):(.+)$/.exec(name.trim());
|
||||
const kind = parsed?.[1] as registry.RegistryKind | undefined;
|
||||
const bare = (parsed?.[2] ?? name).trim().toLowerCase();
|
||||
|
||||
const entries = await registry.fetchIndex(cfg.registryUrl);
|
||||
const hits = entries.filter((e) => e.name === bare && (kind === undefined || e.kind === kind));
|
||||
|
||||
if (hits.length === 0) throw new Error(`no registry entry named "${bare}". Try /registry search ${bare}`);
|
||||
if (hits.length > 1) {
|
||||
throw new Error(
|
||||
`"${bare}" is both a skill and a plugin. Use /registry add skill:${bare} or /registry add plugin:${bare}`,
|
||||
);
|
||||
}
|
||||
return hits[0]!;
|
||||
}
|
||||
|
||||
let agentVariant: AgentVariant;
|
||||
try {
|
||||
agentVariant = resolveAgent(flag('--agent') || cfg.agent, flag('--think') || cfg.thinking);
|
||||
@@ -163,8 +200,8 @@ const pluginErrors = enabledPlugins
|
||||
.filter((name) => !BUILTIN_PLUGINS.some((p) => p.name === name))
|
||||
.map((name) => ({ plugin: name, message: 'no such plugin' }));
|
||||
const plugins = createHost(
|
||||
BUILTIN_PLUGINS.filter((p) => enabledPlugins.includes(p.name)),
|
||||
pluginErrors,
|
||||
[...BUILTIN_PLUGINS.filter((p) => enabledPlugins.includes(p.name)), ...installedPlugins.plugins],
|
||||
[...pluginErrors, ...installedPlugins.errors],
|
||||
);
|
||||
|
||||
const memory = has('--no-memory') ? undefined : new Memory(process.cwd(), languageModel);
|
||||
@@ -280,6 +317,55 @@ const hooks: AppHooks = {
|
||||
for await (const rel of walk({ limit: 5000 })) found.push(rel);
|
||||
return found;
|
||||
},
|
||||
registry: {
|
||||
list: async () => {
|
||||
const entries = await registry.fetchIndex(cfg.registryUrl);
|
||||
const installed = await installedNames();
|
||||
return entries.map((e) => ({
|
||||
name: e.name,
|
||||
kind: e.kind,
|
||||
description: e.description,
|
||||
...(e.author ? { author: e.author } : {}),
|
||||
installed: installed.has(`${e.kind}:${e.name}`),
|
||||
}));
|
||||
},
|
||||
installed: async () => {
|
||||
const rows: AppRegistryRow[] = [];
|
||||
for (const s of skills) {
|
||||
if (s.origin === 'registry') rows.push({ name: s.name, kind: 'skill', description: s.description, installed: true });
|
||||
}
|
||||
for (const p of installedPlugins.plugins) {
|
||||
rows.push({ name: p.name, kind: 'plugin', description: p.description, installed: true });
|
||||
}
|
||||
return rows.sort((a, b) => a.name.localeCompare(b.name));
|
||||
},
|
||||
stage: async (name) => {
|
||||
const entry = await findEntry(name);
|
||||
const { preview } = await registry.stage(entry);
|
||||
return {
|
||||
row: { name: entry.name, kind: entry.kind, description: entry.description },
|
||||
url: entry.url,
|
||||
preview,
|
||||
};
|
||||
},
|
||||
install: async (name) => {
|
||||
const entry = await findEntry(name);
|
||||
const { path } = await registry.install(entry);
|
||||
// Loaded on the next start rather than hot-swapped: a skill joins the system
|
||||
// prompt and a plugin joins the guard chain, and both are built once at boot.
|
||||
return `installed ${entry.kind} ${entry.name} to ${path}\nrestart shiro to load it`;
|
||||
},
|
||||
remove: async (name) => {
|
||||
const parsed = /^(skill|plugin):(.+)$/.exec(name);
|
||||
const kinds: registry.RegistryKind[] = parsed ? [parsed[1] as registry.RegistryKind] : ['skill', 'plugin'];
|
||||
const bare = parsed ? parsed[2]! : name;
|
||||
|
||||
for (const kind of kinds) {
|
||||
if (await registry.uninstall(kind, bare)) return `removed ${kind} ${bare}\nrestart shiro to unload it`;
|
||||
}
|
||||
throw new Error(`nothing installed under the name "${bare}"`);
|
||||
},
|
||||
},
|
||||
initPrompt: INIT_PROMPT,
|
||||
history: promptHistory,
|
||||
recordPrompt: (text) => void store.appendHistory(text),
|
||||
@@ -302,13 +388,17 @@ const hooks: AppHooks = {
|
||||
},
|
||||
listSkills: () => {
|
||||
if (skills.length === 0) return 'no skills loaded';
|
||||
return skills.map((s) => `${s.name.padEnd(10)} ${s.origin.padEnd(8)} ${s.description}`).join('\n');
|
||||
return [
|
||||
...skills.map((s) => `${s.name.padEnd(12)} ${s.origin.padEnd(9)} ${s.description}`),
|
||||
'',
|
||||
'`/registry` to browse and install more',
|
||||
].join('\n');
|
||||
},
|
||||
listPlugins: () => {
|
||||
const active = plugins.plugins.map((p) => `${p.name.padEnd(8)} ${p.description}`);
|
||||
const failed = plugins.errors.map((e) => `${e.plugin.padEnd(8)} ${e.message}`);
|
||||
const active = plugins.plugins.map((p) => `${p.name.padEnd(12)} ${p.description}`);
|
||||
const failed = plugins.errors.map((e) => `${e.plugin.padEnd(12)} ${e.message}`);
|
||||
if (active.length === 0 && failed.length === 0) return 'no plugins active';
|
||||
return [...active, ...failed].join('\n');
|
||||
return [...active, ...failed, '', '`/registry` to browse and install more'].join('\n');
|
||||
},
|
||||
listMemory: async () => {
|
||||
if (!memory) return 'memory is disabled (--no-memory)';
|
||||
|
||||
+47
-7
@@ -16,6 +16,7 @@ export type CommandAction =
|
||||
| { type: 'notes' }
|
||||
| { type: 'skills' }
|
||||
| { type: 'plugins' }
|
||||
| { type: 'registry'; action: 'list' | 'search' | 'add' | 'remove' | 'installed'; arg?: string }
|
||||
| { type: 'memory' }
|
||||
| { type: 'agent'; agent?: string }
|
||||
| { type: 'think'; level?: string }
|
||||
@@ -42,6 +43,7 @@ export const COMMANDS: CommandSpec[] = [
|
||||
{ name: 'model', arg: '<id>', summary: 'switch model by name' },
|
||||
{ name: 'skills', summary: 'list loaded skills' },
|
||||
{ name: 'plugins', summary: 'list active plugins' },
|
||||
{ name: 'registry', arg: '[search|add|remove] [name]', summary: 'browse and install external skills and plugins' },
|
||||
{ name: 'init', summary: 'have the agent write AGENTS.md for this project' },
|
||||
{ name: 'context', summary: 'show which instruction files are loaded' },
|
||||
{ name: 'todos', summary: "show the agent's task list" },
|
||||
@@ -60,14 +62,14 @@ export const COMMANDS: CommandSpec[] = [
|
||||
const usage = (c: CommandSpec) => `/${c.name}${c.arg ? ` ${c.arg}` : ''}`;
|
||||
|
||||
export const HELP = [
|
||||
...COMMANDS.map((c) => `${usage(c).padEnd(18)} ${c.summary}`),
|
||||
...COMMANDS.map((c) => `${usage(c).padEnd(22)} ${c.summary}`),
|
||||
'',
|
||||
'esc interrupt the running turn and clear the queue',
|
||||
'ctrl-c kill the running command, keeping the turn',
|
||||
'ctrl-r expand or collapse the reasoning panel',
|
||||
'tab complete the highlighted command or file path',
|
||||
'up / down recall earlier prompts, or move in an open menu',
|
||||
'@ complete a workspace path',
|
||||
'esc interrupt the running turn and clear the queue',
|
||||
'ctrl-c kill the running command, keeping the turn',
|
||||
'ctrl-r expand or collapse the reasoning panel',
|
||||
'tab complete the highlighted command or file path',
|
||||
'up / down recall earlier prompts, or move in an open menu',
|
||||
'@ complete a workspace path',
|
||||
'',
|
||||
'typing during a turn queues the prompt; queued prompts run in order afterwards',
|
||||
].join('\n');
|
||||
@@ -89,6 +91,42 @@ export function matchCommands(input: string): CommandSpec[] {
|
||||
/** True while the input is a bare command name being typed, so the menu should show. */
|
||||
export const isMenuOpen = (input: string) => input.startsWith('/') && !input.includes(' ');
|
||||
|
||||
/**
|
||||
* `/registry [list|installed|search <q>|add <name>|remove <name>]`.
|
||||
*
|
||||
* A bare `/registry` lists everything. `add` and `remove` need a name, and saying
|
||||
* so beats fetching the whole index to then complain.
|
||||
*/
|
||||
function parseRegistry(arg: string): CommandAction {
|
||||
const [verb = '', ...rest] = arg.split(/\s+/).filter(Boolean);
|
||||
const name = rest.join(' ').trim();
|
||||
|
||||
switch (verb) {
|
||||
case '':
|
||||
case 'list':
|
||||
return { type: 'registry', action: 'list' };
|
||||
case 'installed':
|
||||
return { type: 'registry', action: 'installed' };
|
||||
case 'search':
|
||||
return name
|
||||
? { type: 'registry', action: 'search', arg: name }
|
||||
: { type: 'info', text: 'usage: /registry search <query>' };
|
||||
case 'add':
|
||||
case 'install':
|
||||
return name
|
||||
? { type: 'registry', action: 'add', arg: name }
|
||||
: { type: 'info', text: 'usage: /registry add <name>' };
|
||||
case 'remove':
|
||||
case 'uninstall':
|
||||
return name
|
||||
? { type: 'registry', action: 'remove', arg: name }
|
||||
: { type: 'info', text: 'usage: /registry remove <name>' };
|
||||
default:
|
||||
// A bare word is almost always a search, and guessing beats a usage line.
|
||||
return { type: 'registry', action: 'search', arg: arg.trim() };
|
||||
}
|
||||
}
|
||||
|
||||
/** Pure parser: no IO, so the TUI and headless mode share one definition. */
|
||||
export function parseCommand(raw: string): CommandAction {
|
||||
const input = raw.trim();
|
||||
@@ -134,6 +172,8 @@ export function parseCommand(raw: string): CommandAction {
|
||||
return { type: 'skills' };
|
||||
case 'plugins':
|
||||
return { type: 'plugins' };
|
||||
case 'registry':
|
||||
return parseRegistry(arg);
|
||||
case 'memory':
|
||||
return { type: 'memory' };
|
||||
case 'agent':
|
||||
|
||||
@@ -27,6 +27,8 @@ export type Config = {
|
||||
plugins?: string[];
|
||||
/** Optional tool sets to offer beyond `core`; omit for all of them. */
|
||||
toolSets?: ToolSetName[];
|
||||
/** Index for `/registry`. Omit for the default one. */
|
||||
registryUrl?: string;
|
||||
mcpServers?: Record<string, McpServerConfig>;
|
||||
};
|
||||
|
||||
@@ -89,6 +91,7 @@ export async function loadConfig(): Promise<Config> {
|
||||
...(file.thinking ? { thinking: file.thinking } : {}),
|
||||
...(Array.isArray(file.plugins) ? { plugins: file.plugins } : {}),
|
||||
...(Array.isArray(file.toolSets) ? { toolSets: file.toolSets.filter(isToolSetName) } : {}),
|
||||
...(typeof file.registryUrl === 'string' ? { registryUrl: file.registryUrl } : {}),
|
||||
...(file.mcpServers ? { mcpServers: file.mcpServers } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
+41
-34
@@ -6,9 +6,6 @@ type Part = {
|
||||
providerOptions?: Record<string, Record<string, unknown>>;
|
||||
};
|
||||
|
||||
/** Parts the OpenAI responses API refuses to accept without their reasoning item. */
|
||||
const DEPENDENT = new Set(['text', 'tool-call']);
|
||||
|
||||
function itemId(part: Part): string | undefined {
|
||||
for (const options of Object.values(part.providerOptions ?? {})) {
|
||||
const id = options['itemId'];
|
||||
@@ -20,23 +17,39 @@ function itemId(part: Part): string | undefined {
|
||||
const partsOf = (message: ModelMessage): Part[] =>
|
||||
message.role === 'assistant' && Array.isArray(message.content) ? (message.content as Part[]) : [];
|
||||
|
||||
/** The part again, with every provider `itemId` removed. */
|
||||
function withoutItemId(part: Part): Part {
|
||||
const providerOptions: Record<string, Record<string, unknown>> = {};
|
||||
for (const [provider, options] of Object.entries(part.providerOptions ?? {})) {
|
||||
const { itemId: _dropped, ...rest } = options;
|
||||
if (Object.keys(rest).length > 0) providerOptions[provider] = rest;
|
||||
}
|
||||
const next: Part = { ...part };
|
||||
if (Object.keys(providerOptions).length > 0) next.providerOptions = providerOptions;
|
||||
else delete next.providerOptions;
|
||||
return next;
|
||||
}
|
||||
|
||||
/**
|
||||
* Drops assistant parts left orphaned by reasoning removal.
|
||||
* Detaches assistant parts from reasoning items that pruning removed.
|
||||
*
|
||||
* The OpenAI responses API treats a `message` item as a dependent of the `reasoning`
|
||||
* item from the same response: send the message without its reasoning and the request
|
||||
* is rejected with 400 "was provided without its required 'reasoning' item".
|
||||
* `pruneMessages({ reasoning: 'all' })` strips the reasoning and keeps the message,
|
||||
* producing exactly that request.
|
||||
* A part carrying a provider `itemId` is not sent inline. The OpenAI responses
|
||||
* provider serialises it as `{ type: 'item_reference', id }`, pointing at an item
|
||||
* stored on their side, and that stored item depends on the `reasoning` item from
|
||||
* the same response. Send the reference without the reasoning and the request is
|
||||
* rejected with 400 "was provided without its required 'reasoning' item".
|
||||
*
|
||||
* The two carry different item ids, so they cannot be matched by id. What links them
|
||||
* is the assistant message they arrived in: one message is one response, and its
|
||||
* reasoning item covers every other item in it.
|
||||
* The repair is to drop the `itemId`, not the part. Without it the same content is
|
||||
* serialised inline — a plain assistant message, a plain `function_call` — which
|
||||
* carries no dependency on anything stored. Verified against the provider's own
|
||||
* serialiser: `text` with an itemId goes out as `item_reference`, and the identical
|
||||
* part without one goes out as `output_text`.
|
||||
*
|
||||
* Reasoning only disappears from turns pruning is already discarding, so dropping the
|
||||
* orphaned text costs nothing pruning was not already spending.
|
||||
* Dropping the part instead, which is what this used to do, cost the model its
|
||||
* memory of the turn: after compaction it could no longer see the tool results it
|
||||
* had just collected, so it called the same tools again until it hit the step limit.
|
||||
*/
|
||||
export function dropOrphanedItems(before: ModelMessage[], after: ModelMessage[]): ModelMessage[] {
|
||||
export function detachOrphanedItems(before: ModelMessage[], after: ModelMessage[]): ModelMessage[] {
|
||||
const survivingReasoning = new Set<string>();
|
||||
for (const message of after) {
|
||||
for (const part of partsOf(message)) {
|
||||
@@ -54,7 +67,6 @@ export function dropOrphanedItems(before: ModelMessage[], after: ModelMessage[])
|
||||
if (reasoning.some((id) => id !== undefined && survivingReasoning.has(id))) continue;
|
||||
|
||||
for (const part of parts) {
|
||||
if (!DEPENDENT.has(part.type)) continue;
|
||||
const id = itemId(part);
|
||||
if (id) orphaned.add(id);
|
||||
}
|
||||
@@ -62,23 +74,20 @@ export function dropOrphanedItems(before: ModelMessage[], after: ModelMessage[])
|
||||
|
||||
if (orphaned.size === 0) return after;
|
||||
|
||||
const cleaned: ModelMessage[] = [];
|
||||
for (const message of after) {
|
||||
return after.map((message) => {
|
||||
const parts = partsOf(message);
|
||||
if (parts.length === 0) {
|
||||
cleaned.push(message);
|
||||
continue;
|
||||
}
|
||||
if (parts.length === 0) return message;
|
||||
|
||||
const kept = parts.filter((part) => {
|
||||
let changed = false;
|
||||
const next = parts.map((part) => {
|
||||
const id = itemId(part);
|
||||
return id === undefined || !orphaned.has(id);
|
||||
if (id === undefined || !orphaned.has(id)) return part;
|
||||
changed = true;
|
||||
return withoutItemId(part);
|
||||
});
|
||||
|
||||
if (kept.length > 0) cleaned.push({ ...message, content: kept } as ModelMessage);
|
||||
}
|
||||
|
||||
return cleaned;
|
||||
return changed ? ({ ...message, content: next } as ModelMessage) : message;
|
||||
});
|
||||
}
|
||||
|
||||
export type PruneOptions = Parameters<typeof pruneMessages>[0];
|
||||
@@ -93,11 +102,9 @@ const anyParts = (message: ModelMessage): Part[] =>
|
||||
*
|
||||
* The OpenAI responses API rejects a `function_call_output` with no `function_call`
|
||||
* carrying the same call id: 400 "No tool call found for function call output with
|
||||
* call_id ...". Two things strand a result that way, and both happen on a long turn:
|
||||
* `pruneMessages({ toolCalls: 'before-last-3-messages' })` counts messages, so the
|
||||
* cut can land between an assistant tool-call and the tool message answering it, and
|
||||
* `dropOrphanedItems` removes a tool-call whose reasoning item did not survive while
|
||||
* the result sits in a separate message it never looks at.
|
||||
* call_id ...". `pruneMessages({ toolCalls: 'before-last-3-messages' })` counts
|
||||
* messages, so the cut can land between an assistant tool-call and the tool message
|
||||
* answering it.
|
||||
*
|
||||
* The reverse pairing is left alone on purpose: a call still awaiting its result is
|
||||
* exactly what a suspended approval looks like, and dropping it would break resume.
|
||||
@@ -132,5 +139,5 @@ export function dropOrphanedResults(messages: ModelMessage[]): ModelMessage[] {
|
||||
/** pruneMessages, then repair the provider-item dependencies it breaks. */
|
||||
export function prunePreservingItems(options: PruneOptions): ModelMessage[] {
|
||||
const pruned = pruneMessages(options);
|
||||
return dropOrphanedResults(dropOrphanedItems(options.messages, pruned));
|
||||
return dropOrphanedResults(detachOrphanedItems(options.messages, pruned));
|
||||
}
|
||||
|
||||
+295
@@ -0,0 +1,295 @@
|
||||
import { homedir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { z } from 'zod';
|
||||
import type { Plugin } from './plugins';
|
||||
import { parseSkill, type Skill } from './skills';
|
||||
|
||||
/**
|
||||
* External skills and plugins, fetched from an index over HTTPS.
|
||||
*
|
||||
* Two kinds, and they are not equally safe.
|
||||
*
|
||||
* A **skill** is prompt text. Installing one puts a stranger's words into the
|
||||
* system prompt of every future session in this project, which is prompt injection
|
||||
* by invitation. The command shows the body before writing it and the origin is
|
||||
* recorded, so `/skills` always says where an instruction came from.
|
||||
*
|
||||
* A **plugin** is declarative: a name, a prompt appendix, and deny rules matched
|
||||
* against tool input. Never code. Loading arbitrary TypeScript from a URL would let
|
||||
* an entry read every file the agent can read and lie about blocking anything, so
|
||||
* that is not offered at any price.
|
||||
*/
|
||||
|
||||
const MAX_INDEX_BYTES = 256 * 1024;
|
||||
const MAX_BODY_BYTES = 64 * 1024;
|
||||
/** A pattern from the internet runs on every tool call; a huge one is a denial of service. */
|
||||
const MAX_PATTERN = 200;
|
||||
|
||||
/**
|
||||
* https only, with localhost allowed so the suite can serve a real index.
|
||||
*
|
||||
* `file:` would read a local path and `data:` would inline a payload, neither of
|
||||
* which is what "fetch this from a registry" means to the person typing it.
|
||||
*/
|
||||
export function isFetchable(url: string): boolean {
|
||||
let parsed: URL;
|
||||
try {
|
||||
parsed = new URL(url);
|
||||
} catch {
|
||||
return false;
|
||||
}
|
||||
if (parsed.protocol === 'https:') return true;
|
||||
return parsed.protocol === 'http:' && (parsed.hostname === 'localhost' || parsed.hostname === '127.0.0.1');
|
||||
}
|
||||
|
||||
const entrySchema = z.object({
|
||||
name: z
|
||||
.string()
|
||||
.min(1)
|
||||
.max(40)
|
||||
// The name becomes a filename. Anything else is a path traversal.
|
||||
.regex(/^[a-z0-9][a-z0-9-]*$/, 'lowercase letters, digits, and dashes only'),
|
||||
description: z.string().min(1).max(300),
|
||||
// z.url() accepts file: and data:, which would read a local path or inline a
|
||||
// payload. Only https reaches the network the way the user expects.
|
||||
url: z
|
||||
.string()
|
||||
.url()
|
||||
.refine((u) => isFetchable(u), 'must be an https URL'),
|
||||
author: z.string().max(80).optional(),
|
||||
});
|
||||
|
||||
const indexSchema = z.object({
|
||||
skills: z.array(entrySchema).max(500).optional(),
|
||||
plugins: z.array(entrySchema).max(500).optional(),
|
||||
});
|
||||
|
||||
export type RegistryKind = 'skill' | 'plugin';
|
||||
export type RegistryEntry = z.infer<typeof entrySchema> & { kind: RegistryKind };
|
||||
|
||||
const denyRuleSchema = z
|
||||
.object({
|
||||
tools: z.array(z.string().max(60)).min(1).max(30),
|
||||
pathPattern: z.string().max(MAX_PATTERN).optional(),
|
||||
commandPattern: z.string().max(MAX_PATTERN).optional(),
|
||||
reason: z.string().min(1).max(300),
|
||||
})
|
||||
.refine((r) => r.pathPattern !== undefined || r.commandPattern !== undefined, {
|
||||
message: 'a deny rule needs pathPattern or commandPattern',
|
||||
});
|
||||
|
||||
const manifestSchema = z.object({
|
||||
name: z.string().min(1).max(40).regex(/^[a-z0-9][a-z0-9-]*$/),
|
||||
description: z.string().min(1).max(300),
|
||||
appendix: z.string().max(2000).optional(),
|
||||
deny: z.array(denyRuleSchema).min(1).max(50),
|
||||
});
|
||||
|
||||
export type PluginManifest = z.infer<typeof manifestSchema>;
|
||||
|
||||
export const DEFAULT_INDEX_URL =
|
||||
'https://raw.githubusercontent.com/zakirkun/shiro-neko-registry/main/index.json';
|
||||
|
||||
const home = () => process.env['SHIRO_HOME'] ?? homedir();
|
||||
|
||||
export const skillsDir = () => join(home(), '.shiro-neko', 'registry', 'skills');
|
||||
export const pluginsDir = () => join(home(), '.shiro-neko', 'registry', 'plugins');
|
||||
|
||||
async function fetchText(url: string, limit: number): Promise<string> {
|
||||
if (!isFetchable(url)) throw new Error(`refusing a non-https URL: ${url}`);
|
||||
const res = await fetch(url, { redirect: 'follow', signal: AbortSignal.timeout(20_000) });
|
||||
if (!res.ok) throw new Error(`${url} returned ${res.status} ${res.statusText}`);
|
||||
|
||||
const declared = Number(res.headers.get('content-length') ?? 0);
|
||||
if (declared > limit) throw new Error(`${url} is ${declared} bytes, over the ${limit} byte limit`);
|
||||
|
||||
const text = await res.text();
|
||||
if (text.length > limit) throw new Error(`${url} is over the ${limit} byte limit`);
|
||||
return text;
|
||||
}
|
||||
|
||||
/** Parses an index document. Exported so the shape can be tested without a network. */
|
||||
export function parseIndex(source: string): RegistryEntry[] {
|
||||
let raw: unknown;
|
||||
try {
|
||||
raw = JSON.parse(source);
|
||||
} catch {
|
||||
throw new Error('the registry index is not valid JSON');
|
||||
}
|
||||
|
||||
const parsed = indexSchema.safeParse(raw);
|
||||
if (!parsed.success) {
|
||||
throw new Error(`the registry index is malformed: ${parsed.error.issues[0]?.message ?? 'unknown reason'}`);
|
||||
}
|
||||
|
||||
const seen = new Set<string>();
|
||||
const entries: RegistryEntry[] = [];
|
||||
for (const kind of ['skill', 'plugin'] as const) {
|
||||
for (const entry of parsed.data[`${kind}s`] ?? []) {
|
||||
const key = `${kind}:${entry.name}`;
|
||||
if (seen.has(key)) continue;
|
||||
seen.add(key);
|
||||
entries.push({ ...entry, kind });
|
||||
}
|
||||
}
|
||||
return entries;
|
||||
}
|
||||
|
||||
export async function fetchIndex(url = DEFAULT_INDEX_URL): Promise<RegistryEntry[]> {
|
||||
return parseIndex(await fetchText(url, MAX_INDEX_BYTES));
|
||||
}
|
||||
|
||||
export const searchEntries = (entries: RegistryEntry[], query: string): RegistryEntry[] => {
|
||||
const needle = query.trim().toLowerCase();
|
||||
if (!needle) return entries;
|
||||
return entries.filter(
|
||||
(e) => e.name.includes(needle) || e.description.toLowerCase().includes(needle),
|
||||
);
|
||||
};
|
||||
|
||||
/** Validates a plugin manifest, including that every pattern is a usable regex. */
|
||||
export function parseManifest(source: string): PluginManifest {
|
||||
let raw: unknown;
|
||||
try {
|
||||
raw = JSON.parse(source);
|
||||
} catch {
|
||||
throw new Error('the plugin manifest is not valid JSON');
|
||||
}
|
||||
|
||||
const parsed = manifestSchema.safeParse(raw);
|
||||
if (!parsed.success) {
|
||||
throw new Error(`the plugin manifest is malformed: ${parsed.error.issues[0]?.message ?? 'unknown reason'}`);
|
||||
}
|
||||
|
||||
for (const rule of parsed.data.deny) {
|
||||
for (const pattern of [rule.pathPattern, rule.commandPattern]) {
|
||||
if (pattern === undefined) continue;
|
||||
try {
|
||||
new RegExp(pattern, 'i');
|
||||
} catch (e) {
|
||||
throw new Error(`invalid pattern "${pattern}": ${(e as Error).message}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return parsed.data;
|
||||
}
|
||||
|
||||
/**
|
||||
* A manifest as a Plugin.
|
||||
*
|
||||
* Rules are data, so the guard is the same code for every installed plugin: match
|
||||
* the tool name, then the path or command against a compiled regex. Nothing from
|
||||
* the manifest is ever evaluated.
|
||||
*/
|
||||
export function manifestToPlugin(manifest: PluginManifest): Plugin {
|
||||
const rules = manifest.deny.map((rule) => ({
|
||||
tools: new Set(rule.tools),
|
||||
path: rule.pathPattern ? new RegExp(rule.pathPattern, 'i') : undefined,
|
||||
command: rule.commandPattern ? new RegExp(rule.commandPattern, 'i') : undefined,
|
||||
reason: rule.reason,
|
||||
}));
|
||||
|
||||
return {
|
||||
name: manifest.name,
|
||||
description: `${manifest.description} (installed)`,
|
||||
...(manifest.appendix ? { appendix: manifest.appendix } : {}),
|
||||
beforeToolCall: ({ toolName, input }) => {
|
||||
const o = (input ?? {}) as Record<string, unknown>;
|
||||
const path = typeof o['path'] === 'string' ? o['path'] : '';
|
||||
const command = typeof o['command'] === 'string' ? o['command'] : '';
|
||||
|
||||
for (const rule of rules) {
|
||||
if (!rule.tools.has(toolName)) continue;
|
||||
if (rule.path && path && rule.path.test(path)) return rule.reason;
|
||||
if (rule.command && command && rule.command.test(command)) return rule.reason;
|
||||
}
|
||||
return undefined;
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
export type Installed = { name: string; kind: RegistryKind; path: string };
|
||||
|
||||
/**
|
||||
* Downloads an entry and returns what would be written, without writing it.
|
||||
*
|
||||
* Separated from the write so the caller can show the user a skill body before it
|
||||
* becomes part of every future prompt.
|
||||
*/
|
||||
export async function stage(entry: RegistryEntry): Promise<{ path: string; content: string; preview: string }> {
|
||||
const body = await fetchText(entry.url, MAX_BODY_BYTES);
|
||||
|
||||
if (entry.kind === 'plugin') {
|
||||
const manifest = parseManifest(body);
|
||||
if (manifest.name !== entry.name) {
|
||||
throw new Error(`the manifest calls itself "${manifest.name}" but the index calls it "${entry.name}"`);
|
||||
}
|
||||
const rules = manifest.deny
|
||||
.map((r) => `- denies ${r.tools.join(', ')}: ${r.pathPattern ?? r.commandPattern}`)
|
||||
.join('\n');
|
||||
return {
|
||||
path: join(pluginsDir(), `${entry.name}.json`),
|
||||
content: JSON.stringify(manifest, null, 2),
|
||||
preview: `${manifest.description}\n\n${rules}`,
|
||||
};
|
||||
}
|
||||
|
||||
const skill = parseSkill(body, 'registry');
|
||||
if (!skill) throw new Error('that skill has no name/description frontmatter, so it cannot be loaded');
|
||||
if (skill.name !== entry.name) {
|
||||
throw new Error(`the skill calls itself "${skill.name}" but the index calls it "${entry.name}"`);
|
||||
}
|
||||
|
||||
return { path: join(skillsDir(), `${entry.name}.md`), content: body, preview: skill.body };
|
||||
}
|
||||
|
||||
export async function install(entry: RegistryEntry): Promise<Installed> {
|
||||
const { path, content } = await stage(entry);
|
||||
await Bun.write(path, content);
|
||||
return { name: entry.name, kind: entry.kind, path };
|
||||
}
|
||||
|
||||
/** Removes an installed entry. Returns false when there was nothing to remove. */
|
||||
export async function uninstall(kind: RegistryKind, name: string): Promise<boolean> {
|
||||
if (!/^[a-z0-9][a-z0-9-]*$/.test(name)) throw new Error(`"${name}" is not a valid entry name`);
|
||||
const path = kind === 'plugin' ? join(pluginsDir(), `${name}.json`) : join(skillsDir(), `${name}.md`);
|
||||
const file = Bun.file(path);
|
||||
if (!(await file.exists())) return false;
|
||||
await file.delete();
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Declarative plugins from disk.
|
||||
*
|
||||
* A malformed file is reported and skipped rather than fatal: one bad install
|
||||
* should not stop the agent from starting.
|
||||
*/
|
||||
export async function loadInstalledPlugins(): Promise<{ plugins: Plugin[]; errors: { plugin: string; message: string }[] }> {
|
||||
const plugins: Plugin[] = [];
|
||||
const errors: { plugin: string; message: string }[] = [];
|
||||
|
||||
let files: string[] = [];
|
||||
try {
|
||||
for await (const f of new Bun.Glob('*.json').scan({ cwd: pluginsDir(), onlyFiles: true })) files.push(f);
|
||||
} catch {
|
||||
return { plugins, errors };
|
||||
}
|
||||
|
||||
for (const file of files.sort()) {
|
||||
const name = file.replace(/\.json$/, '');
|
||||
try {
|
||||
plugins.push(manifestToPlugin(parseManifest(await Bun.file(join(pluginsDir(), file)).text())));
|
||||
} catch (e) {
|
||||
errors.push({ plugin: name, message: e instanceof Error ? e.message : String(e) });
|
||||
}
|
||||
}
|
||||
|
||||
return { plugins, errors };
|
||||
}
|
||||
|
||||
/** Installed skills, for `/registry list` — the loader already merges them by name. */
|
||||
export function installedSkillNames(skills: Skill[]): string[] {
|
||||
return skills.filter((s) => s.origin === 'registry').map((s) => s.name);
|
||||
}
|
||||
+9
-1
@@ -83,6 +83,9 @@ export type SessionOptions = {
|
||||
|
||||
const estimateTokens = (messages: ModelMessage[]) => Math.round(JSON.stringify(messages).length / 4);
|
||||
|
||||
/** Estimated tokens at which the wire history is pruned. */
|
||||
const DEFAULT_COMPACT_THRESHOLD = 120_000;
|
||||
|
||||
export class Session {
|
||||
readonly messages: ModelMessage[];
|
||||
readonly tools: ToolSet;
|
||||
@@ -164,6 +167,11 @@ export class Session {
|
||||
return estimateTokens(this.messages);
|
||||
}
|
||||
|
||||
/** Where compaction kicks in, so the status bar can show how close it is. */
|
||||
compactThreshold(): number {
|
||||
return this.opts.compactThreshold ?? DEFAULT_COMPACT_THRESHOLD;
|
||||
}
|
||||
|
||||
private systemFor(): string {
|
||||
return systemPrompt({
|
||||
cwd: this.opts.cwd ?? process.cwd(),
|
||||
@@ -234,7 +242,7 @@ export class Session {
|
||||
this.opts.onChange?.(this.messages);
|
||||
this.controller = new AbortController();
|
||||
const signal = this.controller.signal;
|
||||
const threshold = this.opts.compactThreshold ?? 120_000;
|
||||
const threshold = this.compactThreshold();
|
||||
|
||||
const outputs: Extract<AgentEvent, { type: 'tool-output' }>[] = [];
|
||||
onBashOutput(({ toolCallId, chunk }) => {
|
||||
|
||||
+16
-7
@@ -4,7 +4,7 @@ import { join } from 'node:path';
|
||||
import { z } from 'zod';
|
||||
import { BUILTIN_SKILLS } from './skills-builtin';
|
||||
|
||||
export type SkillOrigin = 'builtin' | 'user' | 'project';
|
||||
export type SkillOrigin = 'builtin' | 'registry' | 'user' | 'project';
|
||||
|
||||
export type Skill = {
|
||||
name: string;
|
||||
@@ -37,14 +37,23 @@ export function parseSkill(source: string, origin: SkillOrigin, path?: string):
|
||||
return { name, description, origin, ...(path ? { path } : {}), body: match[2]!.trim().slice(0, MAX_BODY) };
|
||||
}
|
||||
|
||||
const skillDirs = (cwd: string) => [
|
||||
{ dir: join(process.env['SHIRO_HOME'] ?? homedir(), '.shiro-neko', 'skills'), origin: 'user' as const },
|
||||
{ dir: join(cwd, '.shiro', 'skills'), origin: 'project' as const },
|
||||
];
|
||||
/**
|
||||
* Precedence, low to high. Installed skills sit below your own on purpose: a skill
|
||||
* you wrote must never be shadowed by one fetched from a registry.
|
||||
*/
|
||||
const skillDirs = (cwd: string) => {
|
||||
const home = join(process.env['SHIRO_HOME'] ?? homedir(), '.shiro-neko');
|
||||
return [
|
||||
{ dir: join(home, 'registry', 'skills'), origin: 'registry' as const },
|
||||
{ dir: join(home, 'skills'), origin: 'user' as const },
|
||||
{ dir: join(cwd, '.shiro', 'skills'), origin: 'project' as const },
|
||||
];
|
||||
};
|
||||
|
||||
/**
|
||||
* Builtin, then user, then project. Later wins, so a project can override a
|
||||
* bundled skill by using the same name.
|
||||
* Builtin, then installed, then user, then project. Later wins, so a project can
|
||||
* override a bundled or installed skill by using the same name — and a skill you
|
||||
* wrote yourself always beats one fetched from a registry.
|
||||
*/
|
||||
export async function loadSkills(cwd = process.cwd()): Promise<Skill[]> {
|
||||
const byName = new Map<string, Skill>();
|
||||
|
||||
+116
-7
@@ -15,7 +15,7 @@ import { AskPanel, type AskBridge, type AskPending } from './Ask';
|
||||
import { Diff } from './Diff';
|
||||
import { Markdown } from './Markdown';
|
||||
import { Onboard, type OnboardResult } from './Onboard';
|
||||
import { InfoPanel, OutputPanel, QueuePanel, StatusBar, SubagentPanel, ThinkingPanel, TodoPanel, ActiveTool, FileMenu, type SubagentView } from './Panels';
|
||||
import { InfoPanel, OutputPanel, QueuePanel, RegistryPanel, InstallPrompt, StatusBar, SubagentPanel, ThinkingPanel, TodoPanel, ActiveTool, FileMenu, type RegistryRow, type SubagentView } from './Panels';
|
||||
import { PromptInput } from './PromptInput';
|
||||
|
||||
type Line =
|
||||
@@ -139,6 +139,15 @@ export type AppHooks = {
|
||||
instructionFiles: () => string[];
|
||||
/** Ignore-aware workspace paths for `@` completion, loaded on first use. */
|
||||
listPaths: () => Promise<string[]>;
|
||||
/** Registry index, installed set, and the install/remove actions. */
|
||||
registry: {
|
||||
list: () => Promise<RegistryRow[]>;
|
||||
installed: () => Promise<RegistryRow[]>;
|
||||
/** Fetches and validates without writing, so the body can be shown first. */
|
||||
stage: (name: string) => Promise<{ row: RegistryRow; url: string; preview: string }>;
|
||||
install: (name: string) => Promise<string>;
|
||||
remove: (name: string) => Promise<string>;
|
||||
};
|
||||
/** Prompt to hand the model for /init. */
|
||||
initPrompt: string;
|
||||
history: string[];
|
||||
@@ -198,19 +207,42 @@ function Approval({ pending }: { pending: Pending }) {
|
||||
}
|
||||
|
||||
function CommandMenu({ matches, index }: { matches: CommandSpec[]; index: number }) {
|
||||
const width = Math.max(...matches.map((c) => `/${c.name}${c.arg ? ` ${c.arg}` : ''}`.length)) + 1;
|
||||
return (
|
||||
<Box flexDirection="column" marginTop={1}>
|
||||
{matches.map((c, i) => (
|
||||
<Text key={c.name} color={i === index ? 'cyan' : undefined} dimColor={i !== index}>
|
||||
{i === index ? '> ' : ' '}
|
||||
{`/${c.name}${c.arg ? ` ${c.arg}` : ''}`.padEnd(18)} {c.summary}
|
||||
</Text>
|
||||
<Box key={c.name}>
|
||||
<Text color={i === index ? 'cyan' : undefined}>{i === index ? '> ' : ' '}</Text>
|
||||
<Text color={i === index ? 'cyan' : undefined} bold={i === index}>
|
||||
{`/${c.name}${c.arg ? ` ${c.arg}` : ''}`.padEnd(width)}
|
||||
</Text>
|
||||
<Text dimColor>{c.summary}</Text>
|
||||
</Box>
|
||||
))}
|
||||
<Text dimColor>up/down move | tab complete | enter run | esc dismiss</Text>
|
||||
</Box>
|
||||
);
|
||||
}
|
||||
|
||||
/** Keyboard wrapper around InstallPrompt, so the prompt itself stays presentational. */
|
||||
function InstallConfirm({
|
||||
staged,
|
||||
onDone,
|
||||
}: {
|
||||
staged: { row: RegistryRow; url: string; preview: string };
|
||||
onDone: (yes: boolean) => void;
|
||||
}) {
|
||||
useInput((input, key) => {
|
||||
const c = input.toLowerCase();
|
||||
if (c === 'y' || key.return) onDone(true);
|
||||
else if (c === 'n' || key.escape) onDone(false);
|
||||
});
|
||||
|
||||
return (
|
||||
<InstallPrompt name={staged.row.name} kind={staged.row.kind} url={staged.url} preview={staged.preview} />
|
||||
);
|
||||
}
|
||||
|
||||
export function App({
|
||||
session,
|
||||
bridge,
|
||||
@@ -260,8 +292,12 @@ export function App({
|
||||
const [notebook, setNotebook] = useState<NotebookState>(session.notebook.state());
|
||||
const [agents, setAgents] = useState<SubagentView[]>([]);
|
||||
const [panel, setPanel] = useState<{ title: string; hint?: string; body: string } | undefined>();
|
||||
const [registry, setRegistry] = useState<{ title: string; hint?: string; rows: RegistryRow[] } | undefined>();
|
||||
const [installing, setInstalling] = useState<
|
||||
{ row: RegistryRow; url: string; preview: string } | undefined
|
||||
>();
|
||||
|
||||
const modal = pending !== undefined || asking !== undefined || onboarding;
|
||||
const modal = pending !== undefined || asking !== undefined || onboarding || installing !== undefined;
|
||||
const anyPicker = modelPicker !== undefined || agentPicker || thinkPicker;
|
||||
const matches = matchCommands(draft);
|
||||
const menuOpen = matches.length > 0 && !menuDismissed && !busy && !modal && !anyPicker && !panel;
|
||||
@@ -371,6 +407,10 @@ export function App({
|
||||
setPanel(undefined);
|
||||
return true;
|
||||
}
|
||||
if (key.escape && registry) {
|
||||
setRegistry(undefined);
|
||||
return true;
|
||||
}
|
||||
|
||||
// The file picker gets first refusal: while an `@` token is open its keys
|
||||
// mean navigation, not history recall or command completion.
|
||||
@@ -425,7 +465,7 @@ export function App({
|
||||
}
|
||||
return false;
|
||||
},
|
||||
[draft, fileMatches.length, fileOpen, highlighted, highlightedPath, matches.length, menuOpen, panel, token],
|
||||
[draft, fileMatches.length, fileOpen, highlighted, highlightedPath, matches.length, menuOpen, panel, registry, token],
|
||||
);
|
||||
|
||||
const onDraftChange = useCallback((value: string, at: number) => {
|
||||
@@ -536,6 +576,7 @@ export function App({
|
||||
setFileIndex(0);
|
||||
setFileDismissed(false);
|
||||
setPanel(undefined);
|
||||
setRegistry(undefined);
|
||||
|
||||
// Enter on an open menu runs the highlighted entry, so `/mo` + enter works.
|
||||
const chosen = menuOpen && highlighted ? `/${highlighted.name}` : raw;
|
||||
@@ -676,6 +717,50 @@ export function App({
|
||||
push({ kind: 'user', text: chosen.trim() });
|
||||
setPanel({ title: 'plugins', body: hooks.listPlugins() });
|
||||
return;
|
||||
case 'registry': {
|
||||
push({ kind: 'user', text: chosen.trim() });
|
||||
setWorking(true);
|
||||
try {
|
||||
if (action.action === 'add') {
|
||||
// Staged, not installed: nothing is written until the prompt is answered.
|
||||
setInstalling(await hooks.registry.stage(action.arg!));
|
||||
return;
|
||||
}
|
||||
if (action.action === 'remove') {
|
||||
push({ kind: 'info', text: await hooks.registry.remove(action.arg!) });
|
||||
return;
|
||||
}
|
||||
if (action.action === 'installed') {
|
||||
const rows = await hooks.registry.installed();
|
||||
setRegistry({
|
||||
title: 'installed',
|
||||
hint: rows.length === 0 ? 'nothing installed yet' : `${rows.length} from the registry`,
|
||||
rows,
|
||||
});
|
||||
return;
|
||||
}
|
||||
|
||||
const all = await hooks.registry.list();
|
||||
const rows =
|
||||
action.action === 'search' && action.arg
|
||||
? all.filter(
|
||||
(r) =>
|
||||
r.name.includes(action.arg!.toLowerCase()) ||
|
||||
r.description.toLowerCase().includes(action.arg!.toLowerCase()),
|
||||
)
|
||||
: all;
|
||||
setRegistry({
|
||||
title: action.arg ? `registry: ${action.arg}` : 'registry',
|
||||
hint: `${rows.length} of ${all.length} available`,
|
||||
rows,
|
||||
});
|
||||
} catch (e) {
|
||||
push({ kind: 'error', text: e instanceof Error ? e.message : String(e) });
|
||||
} finally {
|
||||
setWorking(false);
|
||||
}
|
||||
return;
|
||||
}
|
||||
case 'memory': {
|
||||
push({ kind: 'user', text: chosen.trim() });
|
||||
setWorking(true);
|
||||
@@ -799,6 +884,29 @@ export function App({
|
||||
<InfoPanel title={panel.title} {...(panel.hint ? { hint: panel.hint } : {})} lines={panel.body} />
|
||||
)}
|
||||
|
||||
{registry && (
|
||||
<RegistryPanel title={registry.title} {...(registry.hint ? { hint: registry.hint } : {})} rows={registry.rows} />
|
||||
)}
|
||||
|
||||
{installing && (
|
||||
<InstallConfirm
|
||||
staged={installing}
|
||||
onDone={async (yes) => {
|
||||
const staged = installing;
|
||||
setInstalling(undefined);
|
||||
if (!yes) {
|
||||
push({ kind: 'info', text: `install cancelled: ${staged.row.name}` });
|
||||
return;
|
||||
}
|
||||
try {
|
||||
push({ kind: 'info', text: await hooks.registry.install(staged.row.name) });
|
||||
} catch (e) {
|
||||
push({ kind: 'error', text: e instanceof Error ? e.message : String(e) });
|
||||
}
|
||||
}}
|
||||
/>
|
||||
)}
|
||||
|
||||
{asking && <AskPanel pending={asking} />}
|
||||
|
||||
{pending && <Approval pending={pending} />}
|
||||
@@ -937,6 +1045,7 @@ export function App({
|
||||
agent={hooks.agentName()}
|
||||
thinking={hooks.thinkingLevel()}
|
||||
contextTokens={session.estimatedTokens()}
|
||||
contextLimit={session.compactThreshold()}
|
||||
cost={(() => {
|
||||
const spend = costOf(hooks.config().model, session.inputTokens, session.outputTokens);
|
||||
return spend === undefined ? 'unpriced' : formatUsd(spend);
|
||||
|
||||
+115
-1
@@ -216,6 +216,7 @@ export function StatusBar({
|
||||
agent,
|
||||
thinking,
|
||||
contextTokens,
|
||||
contextLimit,
|
||||
cost,
|
||||
toolCount,
|
||||
}: {
|
||||
@@ -223,14 +224,25 @@ export function StatusBar({
|
||||
agent: string;
|
||||
thinking: string;
|
||||
contextTokens: number;
|
||||
/** Threshold compaction fires at, so the bar means something. */
|
||||
contextLimit?: number;
|
||||
cost: string;
|
||||
toolCount: number;
|
||||
}) {
|
||||
const pct = contextLimit ? Math.min(100, Math.round((contextTokens / contextLimit) * 100)) : undefined;
|
||||
// Amber from two thirds, red once compaction is imminent: the point is to warn
|
||||
// before a turn silently loses its history, not after.
|
||||
const contextColor = pct === undefined ? undefined : pct >= 90 ? 'red' : pct >= 66 ? 'yellow' : undefined;
|
||||
|
||||
return (
|
||||
<Box>
|
||||
<Text dimColor>{`${model} `}</Text>
|
||||
<Text color="cyan">{agent}</Text>
|
||||
<Text dimColor>{`/${thinking} ${toolCount} tools ~${contextTokens} ctx ${cost}`}</Text>
|
||||
<Text dimColor>{`/${thinking} ${toolCount} tools `}</Text>
|
||||
<Text color={contextColor} dimColor={contextColor === undefined}>
|
||||
{pct === undefined ? `~${contextTokens} ctx` : `${pct}% ctx`}
|
||||
</Text>
|
||||
<Text dimColor>{` ${cost}`}</Text>
|
||||
</Box>
|
||||
);
|
||||
}
|
||||
@@ -258,3 +270,105 @@ export function InfoPanel({ title, hint, lines }: { title: string; hint?: string
|
||||
</Box>
|
||||
);
|
||||
}
|
||||
|
||||
export type RegistryRow = {
|
||||
name: string;
|
||||
kind: 'skill' | 'plugin';
|
||||
description: string;
|
||||
author?: string;
|
||||
installed?: boolean;
|
||||
};
|
||||
|
||||
const KIND_COLOR: Record<RegistryRow['kind'], string> = { skill: 'green', plugin: 'magenta' };
|
||||
|
||||
/**
|
||||
* The registry index as a table.
|
||||
*
|
||||
* `kind` is coloured rather than spelled out on every row: skill and plugin carry
|
||||
* very different risk, and a reader scanning the list should see that at a glance.
|
||||
*/
|
||||
export function RegistryPanel({
|
||||
rows,
|
||||
hint,
|
||||
title = 'registry',
|
||||
}: {
|
||||
rows: readonly RegistryRow[];
|
||||
hint?: string;
|
||||
title?: string;
|
||||
}) {
|
||||
if (rows.length === 0) {
|
||||
return <InfoPanel title={title} {...(hint ? { hint } : {})} lines="nothing found" />;
|
||||
}
|
||||
|
||||
const width = Math.min(22, Math.max(...rows.map((r) => r.name.length)) + 1);
|
||||
|
||||
return (
|
||||
<Box flexDirection="column" borderStyle="round" borderColor="cyan" paddingX={1} marginBottom={1}>
|
||||
<Text color="cyan" bold>
|
||||
{title}
|
||||
</Text>
|
||||
{hint && <Text dimColor>{hint}</Text>}
|
||||
{rows.map((r) => (
|
||||
<Box key={`${r.kind}:${r.name}`}>
|
||||
<Text color={KIND_COLOR[r.kind]}>{r.kind === 'skill' ? 'S' : 'P'} </Text>
|
||||
<Text bold>{r.name.padEnd(width)}</Text>
|
||||
<Text dimColor>{r.description.length > 58 ? `${r.description.slice(0, 58)}...` : r.description}</Text>
|
||||
{r.installed && <Text color="green">{' installed'}</Text>}
|
||||
</Box>
|
||||
))}
|
||||
<Text dimColor>{'S skill P plugin | /registry add <name>'}</Text>
|
||||
</Box>
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* Confirmation before an install writes anything.
|
||||
*
|
||||
* A skill body becomes part of the system prompt of every future session in this
|
||||
* project, so it is shown in full first. The wording says that plainly rather than
|
||||
* asking a generic "are you sure".
|
||||
*/
|
||||
export function InstallPrompt({
|
||||
name,
|
||||
kind,
|
||||
url,
|
||||
preview,
|
||||
lines = 14,
|
||||
}: {
|
||||
name: string;
|
||||
kind: 'skill' | 'plugin';
|
||||
url: string;
|
||||
preview: string;
|
||||
lines?: number;
|
||||
}) {
|
||||
const body = preview.split('\n');
|
||||
const shown = body.slice(0, lines);
|
||||
const hidden = body.length - shown.length;
|
||||
|
||||
return (
|
||||
<Box flexDirection="column" borderStyle="round" borderColor="yellow" paddingX={1}>
|
||||
<Text color="yellow" bold>
|
||||
{`install ${kind} "${name}"?`}
|
||||
</Text>
|
||||
<Text dimColor>{url}</Text>
|
||||
<Box marginTop={1} flexDirection="column">
|
||||
{shown.map((l, i) => (
|
||||
<Text key={i} dimColor>
|
||||
{` ${l}`}
|
||||
</Text>
|
||||
))}
|
||||
{hidden > 0 && <Text dimColor>{` ... ${hidden} more lines`}</Text>}
|
||||
</Box>
|
||||
<Box marginTop={1} flexDirection="column">
|
||||
<Text color="yellow">
|
||||
{kind === 'skill'
|
||||
? 'A skill is instructions the agent follows. This text joins your system prompt.'
|
||||
: 'A plugin adds refusal rules. It is data, not code: nothing here is executed.'}
|
||||
</Text>
|
||||
<Text>
|
||||
<Text color="green">y</Text> install | <Text color="red">n</Text> cancel
|
||||
</Text>
|
||||
</Box>
|
||||
</Box>
|
||||
);
|
||||
}
|
||||
|
||||
@@ -112,3 +112,49 @@ test('aliases are hidden from the menu but still parse', () => {
|
||||
expect(matchCommands('/lo')).toEqual([]);
|
||||
expect(parseCommand('/login').type).toBe('provider');
|
||||
});
|
||||
|
||||
test('/registry with no verb lists everything', () => {
|
||||
expect(parseCommand('/registry')).toEqual({ type: 'registry', action: 'list' });
|
||||
expect(parseCommand('/registry list')).toEqual({ type: 'registry', action: 'list' });
|
||||
});
|
||||
|
||||
test('/registry installed asks for what is already here', () => {
|
||||
expect(parseCommand('/registry installed')).toEqual({ type: 'registry', action: 'installed' });
|
||||
});
|
||||
|
||||
test('/registry search carries the query', () => {
|
||||
expect(parseCommand('/registry search migration')).toEqual({
|
||||
type: 'registry',
|
||||
action: 'search',
|
||||
arg: 'migration',
|
||||
});
|
||||
});
|
||||
|
||||
test('/registry add and remove carry the name, and their aliases work', () => {
|
||||
expect(parseCommand('/registry add migration')).toEqual({ type: 'registry', action: 'add', arg: 'migration' });
|
||||
expect(parseCommand('/registry install migration')).toEqual({ type: 'registry', action: 'add', arg: 'migration' });
|
||||
expect(parseCommand('/registry remove migration')).toEqual({ type: 'registry', action: 'remove', arg: 'migration' });
|
||||
expect(parseCommand('/registry uninstall migration')).toEqual({
|
||||
type: 'registry',
|
||||
action: 'remove',
|
||||
arg: 'migration',
|
||||
});
|
||||
});
|
||||
|
||||
test('a kind-qualified name survives parsing, since a name can be both', () => {
|
||||
expect(parseCommand('/registry add plugin:review')).toEqual({
|
||||
type: 'registry',
|
||||
action: 'add',
|
||||
arg: 'plugin:review',
|
||||
});
|
||||
});
|
||||
|
||||
test('/registry add with no name returns usage rather than fetching anything', () => {
|
||||
expect(parseCommand('/registry add')).toEqual({ type: 'info', text: 'usage: /registry add <name>' });
|
||||
expect(parseCommand('/registry remove')).toEqual({ type: 'info', text: 'usage: /registry remove <name>' });
|
||||
expect(parseCommand('/registry search')).toEqual({ type: 'info', text: 'usage: /registry search <query>' });
|
||||
});
|
||||
|
||||
test('a bare word after /registry is treated as a search', () => {
|
||||
expect(parseCommand('/registry migration')).toEqual({ type: 'registry', action: 'search', arg: 'migration' });
|
||||
});
|
||||
|
||||
@@ -2,6 +2,9 @@ import { expect, test } from 'bun:test';
|
||||
import { MockLanguageModelV4, simulateReadableStream } from 'ai/test';
|
||||
import type { LanguageModelV4CallOptions, LanguageModelV4StreamPart } from '@ai-sdk/provider';
|
||||
import type { ModelMessage } from 'ai';
|
||||
import { mkdtempSync, rmSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import { Session, type AgentEvent } from '../src/session';
|
||||
|
||||
const usage = {
|
||||
@@ -173,3 +176,129 @@ test('onChange fires for every history mutation so autosave stays current', asyn
|
||||
|
||||
expect(snapshots).toEqual([1, 2]);
|
||||
});
|
||||
|
||||
/** A reasoning model's turn: a reasoning item, then the tool call, both with item ids. */
|
||||
const reasoningToolStep = (n: number): LanguageModelV4StreamPart[] => [
|
||||
{ type: 'reasoning-start', id: `r${n}`, providerMetadata: { openai: { itemId: `rs_${n}` } } } as never,
|
||||
{ type: 'reasoning-delta', id: `r${n}`, delta: 'deciding what to read next' } as never,
|
||||
{ type: 'reasoning-end', id: `r${n}`, providerMetadata: { openai: { itemId: `rs_${n}` } } } as never,
|
||||
{ type: 'tool-input-start', id: `c${n}`, toolName: 'read_file' },
|
||||
{ type: 'tool-input-end', id: `c${n}` },
|
||||
{
|
||||
type: 'tool-call',
|
||||
toolCallId: `c${n}`,
|
||||
toolName: 'read_file',
|
||||
input: JSON.stringify({ path: 'big.txt' }),
|
||||
providerMetadata: { openai: { itemId: `fc_${n}` } },
|
||||
} as never,
|
||||
{ type: 'finish', finishReason: { unified: 'tool-calls', raw: 'tool_use' }, usage },
|
||||
];
|
||||
|
||||
function inTempDir<T>(fn: () => Promise<T>): Promise<T> {
|
||||
const orig = process.cwd();
|
||||
const dir = mkdtempSync(join(tmpdir(), 'shiro-compact-'));
|
||||
process.chdir(dir);
|
||||
return fn().finally(() => {
|
||||
process.chdir(orig);
|
||||
rmSync(dir, { recursive: true, force: true });
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* The bug this guards against: compaction used to drop any assistant part whose
|
||||
* reasoning item it had pruned. On a reasoning model that is every tool call, so
|
||||
* after the first compaction the model could no longer see what it had already
|
||||
* run — and kept re-running it until the step limit stopped the turn.
|
||||
*/
|
||||
test('a compacted turn still shows the model the tool calls it already made', async () =>
|
||||
inTempDir(async () => {
|
||||
await Bun.write(join(process.cwd(), 'big.txt'), 'lorem ipsum dolor sit amet\n'.repeat(1500));
|
||||
|
||||
const seen: LanguageModelV4CallOptions[] = [];
|
||||
let call = 0;
|
||||
const session = new Session({
|
||||
compactThreshold: 4000,
|
||||
maxSteps: 8,
|
||||
model: new MockLanguageModelV4({
|
||||
doStream: async (o) => {
|
||||
seen.push(o);
|
||||
const n = call++;
|
||||
return n < 3 ? stream(reasoningToolStep(n)) : stream(text('read it three times'));
|
||||
},
|
||||
}),
|
||||
askApproval: async () => 'deny',
|
||||
});
|
||||
|
||||
const events: string[] = [];
|
||||
for await (const ev of session.send('read big.txt a few times')) events.push(ev.type);
|
||||
|
||||
expect(events).toContain('compacted');
|
||||
expect(events).toContain('text');
|
||||
expect(events.at(-1)).toBe('done');
|
||||
// Four calls, not eight: the loop ended because the model chose to, not
|
||||
// because maxSteps cut it off.
|
||||
expect(call).toBe(4);
|
||||
|
||||
const shapeOf = (o: LanguageModelV4CallOptions) =>
|
||||
o.prompt
|
||||
.filter((m) => m.role !== 'system')
|
||||
.flatMap((m) => (Array.isArray(m.content) ? (m.content as { type: string }[]).map((p) => p.type) : ['str']));
|
||||
|
||||
// Every call after the first has to carry the earlier exchange.
|
||||
for (const o of seen.slice(1)) {
|
||||
const shape = shapeOf(o);
|
||||
expect(shape).toContain('tool-call');
|
||||
expect(shape).toContain('tool-result');
|
||||
}
|
||||
}), 20_000);
|
||||
|
||||
test('a compacted turn sends no assistant item reference whose reasoning was pruned', async () =>
|
||||
inTempDir(async () => {
|
||||
await Bun.write(join(process.cwd(), 'big.txt'), 'lorem ipsum dolor sit amet\n'.repeat(1500));
|
||||
|
||||
const seen: LanguageModelV4CallOptions[] = [];
|
||||
let call = 0;
|
||||
const session = new Session({
|
||||
compactThreshold: 4000,
|
||||
model: new MockLanguageModelV4({
|
||||
doStream: async (o) => {
|
||||
seen.push(o);
|
||||
const n = call++;
|
||||
return n < 3 ? stream(reasoningToolStep(n)) : stream(text('done'));
|
||||
},
|
||||
}),
|
||||
askApproval: async () => 'deny',
|
||||
});
|
||||
|
||||
for await (const _ of session.send('read it')) void _;
|
||||
|
||||
// An itemId on an assistant part is serialised as `item_reference`, which the
|
||||
// responses API resolves against a stored item that depends on its reasoning
|
||||
// item. Send one without that reasoning and the request is a 400. An itemId on
|
||||
// a tool result is harmless: it goes out as a plain function_call_output.
|
||||
for (const [i, o] of seen.entries()) {
|
||||
const assistant = o.prompt.filter((m) => m.role === 'assistant');
|
||||
const reasoningIds = new Set<string>();
|
||||
for (const m of assistant) {
|
||||
if (!Array.isArray(m.content)) continue;
|
||||
for (const p of m.content as { type: string; providerOptions?: Record<string, Record<string, unknown>> }[]) {
|
||||
if (p.type !== 'reasoning') continue;
|
||||
for (const options of Object.values(p.providerOptions ?? {})) {
|
||||
if (typeof options['itemId'] === 'string') reasoningIds.add(options['itemId']);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
for (const m of assistant) {
|
||||
if (!Array.isArray(m.content)) continue;
|
||||
for (const p of m.content as { type: string; providerOptions?: Record<string, Record<string, unknown>> }[]) {
|
||||
if (p.type === 'reasoning') continue;
|
||||
for (const options of Object.values(p.providerOptions ?? {})) {
|
||||
const id = options['itemId'];
|
||||
if (typeof id !== 'string') continue;
|
||||
expect(reasoningIds.size, `call ${i}: ${p.type} references ${id} with no reasoning item`).toBeGreaterThan(0);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}), 20_000);
|
||||
|
||||
@@ -21,6 +21,15 @@ export function testHooks(over: Partial<AppHooks> = {}): AppHooks {
|
||||
saveSession: async () => 'saved',
|
||||
instructionFiles: () => [],
|
||||
listPaths: async () => [],
|
||||
registry: {
|
||||
list: async () => [],
|
||||
installed: async () => [],
|
||||
stage: async () => {
|
||||
throw new Error('no registry in tests unless a test provides one');
|
||||
},
|
||||
install: async () => 'installed',
|
||||
remove: async () => 'removed',
|
||||
},
|
||||
initPrompt: 'write AGENTS.md',
|
||||
history: [],
|
||||
recordPrompt: () => {},
|
||||
|
||||
+100
-26
@@ -1,10 +1,24 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import type { ModelMessage } from 'ai';
|
||||
import { dropOrphanedItems, dropOrphanedResults, prunePreservingItems } from '../src/prune';
|
||||
import { detachOrphanedItems, dropOrphanedResults, prunePreservingItems } from '../src/prune';
|
||||
|
||||
const kinds = (messages: ModelMessage[]) =>
|
||||
messages.map((m) => (Array.isArray(m.content) ? `${m.role}:${m.content.map((p) => p.type).join('+')}` : m.role));
|
||||
|
||||
/** Every provider itemId in a message tree, which is what the repair strips. */
|
||||
const itemIds = (messages: ModelMessage[]): string[] => {
|
||||
const found: string[] = [];
|
||||
for (const m of messages) {
|
||||
if (!Array.isArray(m.content)) continue;
|
||||
for (const p of m.content as { providerOptions?: Record<string, Record<string, unknown>> }[]) {
|
||||
for (const options of Object.values(p.providerOptions ?? {})) {
|
||||
if (typeof options['itemId'] === 'string') found.push(options['itemId']);
|
||||
}
|
||||
}
|
||||
}
|
||||
return found;
|
||||
};
|
||||
|
||||
/** An assistant turn as the OpenAI responses API returns it. */
|
||||
const reasoningTurn = (rs: string, msg: string, text = 'answer'): ModelMessage => ({
|
||||
role: 'assistant',
|
||||
@@ -28,24 +42,33 @@ const toolTurn = (rs: string, call: string): ModelMessage => ({
|
||||
],
|
||||
});
|
||||
|
||||
test('a message left without its reasoning item is dropped', () => {
|
||||
/**
|
||||
* A part carrying an itemId is serialised as `{ type: 'item_reference', id }`,
|
||||
* which depends on the stored reasoning item. Stripping the id sends the same
|
||||
* content inline instead, so the turn survives without the dependency.
|
||||
*/
|
||||
test('a message left without its reasoning item keeps its text and loses its item id', () => {
|
||||
const before = [{ role: 'user' as const, content: 'q' }, reasoningTurn('rs_1', 'msg_1')];
|
||||
const after = [{ role: 'user' as const, content: 'q' }, { role: 'assistant' as const, content: [{ type: 'text' as const, text: 'answer', providerOptions: { openai: { itemId: 'msg_1' } } }] }];
|
||||
const after: ModelMessage[] = [
|
||||
{ role: 'user', content: 'q' },
|
||||
{ role: 'assistant', content: [{ type: 'text', text: 'answer', providerOptions: { openai: { itemId: 'msg_1' } } }] },
|
||||
];
|
||||
|
||||
const cleaned = dropOrphanedItems(before, after);
|
||||
expect(JSON.stringify(cleaned)).not.toContain('msg_1');
|
||||
expect(kinds(cleaned)).toEqual(['user']);
|
||||
const cleaned = detachOrphanedItems(before, after);
|
||||
expect(itemIds(cleaned)).toEqual([]);
|
||||
expect(kinds(cleaned)).toEqual(['user', 'assistant:text']);
|
||||
expect(JSON.stringify(cleaned)).toContain('answer');
|
||||
});
|
||||
|
||||
test('a tool call left without its reasoning item is dropped too', () => {
|
||||
test('a tool call left without its reasoning item survives, detached', () => {
|
||||
const before = [{ role: 'user' as const, content: 'q' }, toolTurn('rs_1', 'fc_1')];
|
||||
const after = [
|
||||
{ role: 'user' as const, content: 'q' },
|
||||
const after: ModelMessage[] = [
|
||||
{ role: 'user', content: 'q' },
|
||||
{
|
||||
role: 'assistant' as const,
|
||||
role: 'assistant',
|
||||
content: [
|
||||
{
|
||||
type: 'tool-call' as const,
|
||||
type: 'tool-call',
|
||||
toolCallId: 'tc1',
|
||||
toolName: 'grep',
|
||||
input: { pattern: 'x' },
|
||||
@@ -55,12 +78,60 @@ test('a tool call left without its reasoning item is dropped too', () => {
|
||||
},
|
||||
];
|
||||
|
||||
expect(JSON.stringify(dropOrphanedItems(before, after))).not.toContain('fc_1');
|
||||
const cleaned = detachOrphanedItems(before, after);
|
||||
expect(itemIds(cleaned)).toEqual([]);
|
||||
// The call itself has to stay, or its result is orphaned and the model loses
|
||||
// any record of what it already ran.
|
||||
expect(JSON.stringify(cleaned)).toContain('tc1');
|
||||
expect(kinds(cleaned)).toEqual(['user', 'assistant:tool-call']);
|
||||
});
|
||||
|
||||
test('an empty providerOptions object is removed rather than left behind', () => {
|
||||
const before = [toolTurn('rs_1', 'fc_1')];
|
||||
const after: ModelMessage[] = [
|
||||
{
|
||||
role: 'assistant',
|
||||
content: [
|
||||
{
|
||||
type: 'tool-call',
|
||||
toolCallId: 'tc1',
|
||||
toolName: 'grep',
|
||||
input: { pattern: 'x' },
|
||||
providerOptions: { openai: { itemId: 'fc_1' } },
|
||||
},
|
||||
],
|
||||
},
|
||||
];
|
||||
|
||||
const part = (detachOrphanedItems(before, after)[0]!.content as Record<string, unknown>[])[0]!;
|
||||
expect('providerOptions' in part).toBe(false);
|
||||
});
|
||||
|
||||
test('other provider options are kept when the item id is stripped', () => {
|
||||
const before: ModelMessage[] = [
|
||||
{
|
||||
role: 'assistant',
|
||||
content: [
|
||||
{ type: 'reasoning', text: 't', providerOptions: { openai: { itemId: 'rs_1' } } },
|
||||
{ type: 'text', text: 'a', providerOptions: { openai: { itemId: 'msg_1', phase: 'final' } } },
|
||||
],
|
||||
},
|
||||
];
|
||||
const after: ModelMessage[] = [
|
||||
{
|
||||
role: 'assistant',
|
||||
content: [{ type: 'text', text: 'a', providerOptions: { openai: { itemId: 'msg_1', phase: 'final' } } }],
|
||||
},
|
||||
];
|
||||
|
||||
const json = JSON.stringify(detachOrphanedItems(before, after));
|
||||
expect(json).not.toContain('msg_1');
|
||||
expect(json).toContain('final');
|
||||
});
|
||||
|
||||
test('a turn whose reasoning survived is left alone', () => {
|
||||
const messages = [{ role: 'user' as const, content: 'q' }, reasoningTurn('rs_1', 'msg_1')];
|
||||
expect(dropOrphanedItems(messages, messages)).toEqual(messages);
|
||||
expect(detachOrphanedItems(messages, messages)).toEqual(messages);
|
||||
});
|
||||
|
||||
test('nothing is touched when no reasoning was removed', () => {
|
||||
@@ -68,13 +139,13 @@ test('nothing is touched when no reasoning was removed', () => {
|
||||
{ role: 'user', content: 'q' },
|
||||
{ role: 'assistant', content: 'plain answer' },
|
||||
];
|
||||
expect(dropOrphanedItems(before, before)).toEqual(before);
|
||||
expect(detachOrphanedItems(before, before)).toEqual(before);
|
||||
});
|
||||
|
||||
test('parts with no provider itemId are always kept', () => {
|
||||
test('parts with no provider itemId are returned unchanged', () => {
|
||||
const before = [reasoningTurn('rs_1', 'msg_1')];
|
||||
const after: ModelMessage[] = [{ role: 'assistant', content: [{ type: 'text', text: 'no item id here' }] }];
|
||||
expect(dropOrphanedItems(before, after)).toEqual(after);
|
||||
expect(detachOrphanedItems(before, after)).toEqual(after);
|
||||
});
|
||||
|
||||
test('user and tool messages are never affected', () => {
|
||||
@@ -83,10 +154,10 @@ test('user and tool messages are never affected', () => {
|
||||
{ role: 'user', content: 'q' },
|
||||
{ role: 'tool', content: [{ type: 'tool-result', toolCallId: 't1', toolName: 'grep', output: { type: 'text', value: 'hit' } }] },
|
||||
];
|
||||
expect(dropOrphanedItems(before, after)).toEqual(after);
|
||||
expect(detachOrphanedItems(before, after)).toEqual(after);
|
||||
});
|
||||
|
||||
test('one orphaned turn does not take a healthy one with it', () => {
|
||||
test('one orphaned turn does not detach a healthy one with it', () => {
|
||||
const before = [
|
||||
{ role: 'user' as const, content: 'q1' },
|
||||
reasoningTurn('rs_1', 'msg_1', 'old answer'),
|
||||
@@ -100,14 +171,15 @@ test('one orphaned turn does not take a healthy one with it', () => {
|
||||
reasoningTurn('rs_2', 'msg_2', 'new answer'),
|
||||
];
|
||||
|
||||
const cleaned = dropOrphanedItems(before, after);
|
||||
const cleaned = detachOrphanedItems(before, after);
|
||||
const json = JSON.stringify(cleaned);
|
||||
expect(json).not.toContain('msg_1');
|
||||
expect(json).toContain('old answer');
|
||||
expect(json).toContain('msg_2');
|
||||
expect(json).toContain('rs_2');
|
||||
});
|
||||
|
||||
test('prunePreservingItems leaves no orphan behind on a real prune', () => {
|
||||
test('prunePreservingItems leaves no item reference behind on a real prune', () => {
|
||||
const messages: ModelMessage[] = [];
|
||||
for (let i = 0; i < 6; i++) {
|
||||
messages.push({ role: 'user', content: `question ${i} ${'x'.repeat(3000)}` });
|
||||
@@ -121,12 +193,11 @@ test('prunePreservingItems leaves no orphan behind on a real prune', () => {
|
||||
emptyMessages: 'remove',
|
||||
});
|
||||
|
||||
// Every surviving text part must either have no item id or belong to a turn
|
||||
// whose reasoning also survived. Since reasoning: 'all' removes them all, no
|
||||
// itemId-bearing assistant part may remain.
|
||||
const survivingIds = JSON.stringify(pruned);
|
||||
for (let i = 0; i < 6; i++) expect(survivingIds).not.toContain(`msg_${i}`);
|
||||
// reasoning: 'all' removes every reasoning item, so no surviving part may still
|
||||
// reference one. The text itself stays: that is the model's memory of the turn.
|
||||
expect(itemIds(pruned)).toEqual([]);
|
||||
expect(pruned.filter((m) => m.role === 'user')).toHaveLength(6);
|
||||
expect(JSON.stringify(pruned)).toContain('answer 5');
|
||||
});
|
||||
|
||||
test('prunePreservingItems is a no-op when nothing needs pruning', () => {
|
||||
@@ -150,7 +221,10 @@ test('a provider other than openai is handled the same way', () => {
|
||||
const after: ModelMessage[] = [
|
||||
{ role: 'assistant', content: [{ type: 'text', text: 'a', providerOptions: { someProvider: { itemId: 'm1' } } }] },
|
||||
];
|
||||
expect(dropOrphanedItems(before, after)).toEqual([]);
|
||||
|
||||
const cleaned = detachOrphanedItems(before, after);
|
||||
expect(itemIds(cleaned)).toEqual([]);
|
||||
expect(JSON.stringify(cleaned)).toContain('"text":"a"');
|
||||
});
|
||||
|
||||
/** The assistant tool-call plus the tool message answering it, as one exchange. */
|
||||
|
||||
@@ -0,0 +1,266 @@
|
||||
import { expect, test } from 'bun:test';
|
||||
import { render } from 'ink-testing-library';
|
||||
import React from 'react';
|
||||
import { MockLanguageModelV4, simulateReadableStream } from 'ai/test';
|
||||
import { Session } from '../src/session';
|
||||
import { App, createApprovalBridge, type AppHooks } from '../src/ui/App';
|
||||
import { InstallPrompt, RegistryPanel, type RegistryRow } from '../src/ui/Panels';
|
||||
import { testHooks } from './helpers';
|
||||
|
||||
const usage = { inputTokens: { total: 3, noCache: 3, cacheRead: 0, cacheWrite: 0 }, outputTokens: { total: 1 } } as any;
|
||||
|
||||
const model = new MockLanguageModelV4({
|
||||
doStream: async () =>
|
||||
({
|
||||
stream: simulateReadableStream({
|
||||
chunks: [
|
||||
{ type: 'text-start', id: '0' },
|
||||
{ type: 'text-delta', id: '0', delta: 'reply' },
|
||||
{ type: 'text-end', id: '0' },
|
||||
{ type: 'finish', finishReason: { unified: 'stop', raw: 'stop' }, usage },
|
||||
],
|
||||
chunkDelayInMs: null,
|
||||
initialDelayInMs: null,
|
||||
}),
|
||||
}) as any,
|
||||
});
|
||||
|
||||
const rows: RegistryRow[] = [
|
||||
{ name: 'migration', kind: 'skill', description: 'Write a database migration' },
|
||||
{ name: 'no-secrets', kind: 'plugin', description: 'Refuses credential writes', installed: true },
|
||||
];
|
||||
|
||||
const wait = (ms: number) => new Promise((r) => setTimeout(r, ms));
|
||||
|
||||
function mount(over: Partial<AppHooks> = {}) {
|
||||
const bridge = createApprovalBridge();
|
||||
const session = new Session({ model, askApproval: bridge.ask });
|
||||
const app = render(<App session={session} bridge={bridge} header="hdr" hooks={testHooks(over)} />);
|
||||
return { app, session };
|
||||
}
|
||||
|
||||
async function run(app: ReturnType<typeof render>, command: string, settle = 400) {
|
||||
for (const ch of command) {
|
||||
app.stdin.write(ch);
|
||||
await wait(30);
|
||||
}
|
||||
app.stdin.write('\r');
|
||||
await wait(settle);
|
||||
}
|
||||
|
||||
test('the registry panel marks skill and plugin differently and flags what is installed', () => {
|
||||
const app = render(<RegistryPanel rows={rows} hint="2 available" />);
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('migration');
|
||||
expect(frame).toContain('no-secrets');
|
||||
expect(frame).toContain('installed');
|
||||
expect(frame).toContain('S skill P plugin');
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('an empty registry panel says so rather than rendering an empty box', () => {
|
||||
const app = render(<RegistryPanel rows={[]} />);
|
||||
expect(app.lastFrame()).toContain('nothing found');
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('the install prompt shows the body and says what a skill actually is', () => {
|
||||
const app = render(
|
||||
<InstallPrompt
|
||||
name="migration"
|
||||
kind="skill"
|
||||
url="https://example.com/m.md"
|
||||
preview={'Migrations live in db/migrations.'}
|
||||
/>,
|
||||
);
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('install skill "migration"?');
|
||||
expect(frame).toContain('https://example.com/m.md');
|
||||
expect(frame).toContain('Migrations live in db/migrations.');
|
||||
expect(frame).toContain('joins your system prompt');
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('the install prompt says a plugin is data, not code', () => {
|
||||
const app = render(
|
||||
<InstallPrompt name="no-secrets" kind="plugin" url="https://e.com/p.json" preview="- denies write_file" />,
|
||||
);
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('nothing here is executed');
|
||||
expect(frame).toContain('denies write_file');
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('a long body is truncated with a count rather than flooding the screen', () => {
|
||||
const preview = Array.from({ length: 40 }, (_, i) => `line ${i}`).join('\n');
|
||||
const app = render(<InstallPrompt name="x" kind="skill" url="https://e.com/x.md" preview={preview} lines={5} />);
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('line 4');
|
||||
expect(frame).not.toContain('line 20');
|
||||
expect(frame).toContain('35 more lines');
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('/registry lists the index', async () => {
|
||||
const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } });
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry');
|
||||
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('migration');
|
||||
expect(frame).toContain('2 of 2 available');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('/registry search narrows the list', async () => {
|
||||
const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } });
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry search migra');
|
||||
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('migration');
|
||||
expect(frame).not.toContain('no-secrets');
|
||||
expect(frame).toContain('1 of 2');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('a failing index surfaces the error instead of an empty panel', async () => {
|
||||
const { app } = mount({
|
||||
registry: {
|
||||
...testHooks().registry,
|
||||
list: async () => {
|
||||
throw new Error('the registry index is not valid JSON');
|
||||
},
|
||||
},
|
||||
});
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry');
|
||||
expect(app.lastFrame()).toContain('not valid JSON');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('/registry add stages, shows the body, and installs only on y', async () => {
|
||||
const installed: string[] = [];
|
||||
const { app } = mount({
|
||||
registry: {
|
||||
...testHooks().registry,
|
||||
stage: async (name) => ({
|
||||
row: { name, kind: 'skill', description: 'd' },
|
||||
url: 'https://example.com/m.md',
|
||||
preview: 'Migrations live in db/migrations.',
|
||||
}),
|
||||
install: async (name) => {
|
||||
installed.push(name);
|
||||
return `installed skill ${name}`;
|
||||
},
|
||||
},
|
||||
});
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry add migration');
|
||||
expect(app.lastFrame()).toContain('install skill "migration"?');
|
||||
expect(installed).toEqual([]);
|
||||
|
||||
app.stdin.write('y');
|
||||
await wait(400);
|
||||
|
||||
expect(installed).toEqual(['migration']);
|
||||
expect(app.lastFrame()).toContain('installed skill migration');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('n cancels the install and nothing is written', async () => {
|
||||
const installed: string[] = [];
|
||||
const { app } = mount({
|
||||
registry: {
|
||||
...testHooks().registry,
|
||||
stage: async (name) => ({
|
||||
row: { name, kind: 'skill', description: 'd' },
|
||||
url: 'https://example.com/m.md',
|
||||
preview: 'body',
|
||||
}),
|
||||
install: async (name) => {
|
||||
installed.push(name);
|
||||
return 'installed';
|
||||
},
|
||||
},
|
||||
});
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry add migration');
|
||||
app.stdin.write('n');
|
||||
await wait(400);
|
||||
|
||||
expect(installed).toEqual([]);
|
||||
expect(app.lastFrame()).toContain('install cancelled: migration');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('a staging failure never reaches the confirmation prompt', async () => {
|
||||
const { app } = mount({
|
||||
registry: {
|
||||
...testHooks().registry,
|
||||
stage: async () => {
|
||||
throw new Error('no registry entry named "nope"');
|
||||
},
|
||||
},
|
||||
});
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry add nope');
|
||||
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('no registry entry named');
|
||||
expect(frame).not.toContain('install skill');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('/registry remove reports what it removed', async () => {
|
||||
const { app } = mount({
|
||||
registry: { ...testHooks().registry, remove: async (name) => `removed skill ${name}` },
|
||||
});
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry remove migration');
|
||||
expect(app.lastFrame()).toContain('removed skill migration');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('/registry installed shows what is already here', async () => {
|
||||
const { app } = mount({
|
||||
registry: { ...testHooks().registry, installed: async () => [rows[1]!] },
|
||||
});
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry installed');
|
||||
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('no-secrets');
|
||||
expect(frame).toContain('1 from the registry');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
|
||||
test('esc dismisses the registry panel', async () => {
|
||||
const { app } = mount({ registry: { ...testHooks().registry, list: async () => rows } });
|
||||
await wait(150);
|
||||
|
||||
await run(app, '/registry');
|
||||
expect(app.lastFrame()).toContain('migration');
|
||||
|
||||
app.stdin.write('\u001B');
|
||||
await wait(300);
|
||||
expect(app.lastFrame()).not.toContain('S skill P plugin');
|
||||
|
||||
app.unmount();
|
||||
}, 20_000);
|
||||
@@ -0,0 +1,314 @@
|
||||
import { afterEach, beforeEach, expect, test } from 'bun:test';
|
||||
import { mkdtempSync, rmSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import {
|
||||
DEFAULT_INDEX_URL,
|
||||
fetchIndex,
|
||||
install,
|
||||
loadInstalledPlugins,
|
||||
manifestToPlugin,
|
||||
parseIndex,
|
||||
parseManifest,
|
||||
pluginsDir,
|
||||
searchEntries,
|
||||
skillsDir,
|
||||
stage,
|
||||
uninstall,
|
||||
type RegistryEntry,
|
||||
} from '../src/registry';
|
||||
|
||||
let home: string;
|
||||
let origHome: string | undefined;
|
||||
|
||||
beforeEach(() => {
|
||||
origHome = process.env['SHIRO_HOME'];
|
||||
home = mkdtempSync(join(tmpdir(), 'shiro-reg-'));
|
||||
process.env['SHIRO_HOME'] = home;
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
if (origHome === undefined) delete process.env['SHIRO_HOME'];
|
||||
else process.env['SHIRO_HOME'] = origHome;
|
||||
rmSync(home, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
const index = (over: Record<string, unknown> = {}) =>
|
||||
JSON.stringify({
|
||||
skills: [{ name: 'migration', description: 'Write a database migration', url: 'https://example.com/m.md' }],
|
||||
plugins: [{ name: 'no-secrets', description: 'Refuses credential writes', url: 'https://example.com/p.json' }],
|
||||
...over,
|
||||
});
|
||||
|
||||
const manifest = (over: Record<string, unknown> = {}) =>
|
||||
JSON.stringify({
|
||||
name: 'no-secrets',
|
||||
description: 'Refuses credential writes',
|
||||
deny: [{ tools: ['write_file', 'edit_file'], pathPattern: '\\.env$', reason: 'refusing to write a .env file' }],
|
||||
...over,
|
||||
});
|
||||
|
||||
test('the default index is an https URL', () => {
|
||||
expect(DEFAULT_INDEX_URL.startsWith('https://')).toBe(true);
|
||||
});
|
||||
|
||||
test('an index parses into entries tagged with their kind', () => {
|
||||
const entries = parseIndex(index());
|
||||
expect(entries).toHaveLength(2);
|
||||
expect(entries.find((e) => e.name === 'migration')?.kind).toBe('skill');
|
||||
expect(entries.find((e) => e.name === 'no-secrets')?.kind).toBe('plugin');
|
||||
});
|
||||
|
||||
test('a malformed index is reported rather than half-loaded', () => {
|
||||
expect(() => parseIndex('not json')).toThrow(/not valid JSON/);
|
||||
expect(() => parseIndex(JSON.stringify({ skills: [{ name: 'x' }] }))).toThrow(/malformed/);
|
||||
});
|
||||
|
||||
test('a name that could escape its directory is rejected', () => {
|
||||
for (const name of ['../evil', 'a/b', 'UPPER', '.hidden', 'with space']) {
|
||||
const bad = JSON.stringify({ skills: [{ name, description: 'd', url: 'https://e.com/x.md' }] });
|
||||
expect(() => parseIndex(bad), name).toThrow(/malformed/);
|
||||
}
|
||||
});
|
||||
|
||||
test('a non-http url is rejected by the schema', () => {
|
||||
const bad = JSON.stringify({ skills: [{ name: 'x', description: 'd', url: 'file:///etc/passwd' }] });
|
||||
expect(() => parseIndex(bad)).toThrow(/malformed/);
|
||||
});
|
||||
|
||||
test('a duplicate name within one kind keeps the first', () => {
|
||||
const dup = JSON.stringify({
|
||||
skills: [
|
||||
{ name: 'x', description: 'first', url: 'https://e.com/1.md' },
|
||||
{ name: 'x', description: 'second', url: 'https://e.com/2.md' },
|
||||
],
|
||||
});
|
||||
const entries = parseIndex(dup);
|
||||
expect(entries).toHaveLength(1);
|
||||
expect(entries[0]?.description).toBe('first');
|
||||
});
|
||||
|
||||
test('the same name may exist as both a skill and a plugin', () => {
|
||||
const both = JSON.stringify({
|
||||
skills: [{ name: 'review', description: 's', url: 'https://e.com/s.md' }],
|
||||
plugins: [{ name: 'review', description: 'p', url: 'https://e.com/p.json' }],
|
||||
});
|
||||
expect(parseIndex(both)).toHaveLength(2);
|
||||
});
|
||||
|
||||
test('search matches name and description, case-insensitively', () => {
|
||||
const entries = parseIndex(index());
|
||||
expect(searchEntries(entries, 'migr').map((e) => e.name)).toEqual(['migration']);
|
||||
expect(searchEntries(entries, 'CREDENTIAL').map((e) => e.name)).toEqual(['no-secrets']);
|
||||
expect(searchEntries(entries, '')).toHaveLength(2);
|
||||
expect(searchEntries(entries, 'zzz')).toEqual([]);
|
||||
});
|
||||
|
||||
test('a manifest parses and every pattern must be a real regex', () => {
|
||||
expect(parseManifest(manifest()).name).toBe('no-secrets');
|
||||
expect(() => parseManifest(manifest({ deny: [{ tools: ['bash'], commandPattern: '([', reason: 'r' }] }))).toThrow(
|
||||
/invalid pattern/,
|
||||
);
|
||||
});
|
||||
|
||||
test('a deny rule needs at least one pattern', () => {
|
||||
expect(() => parseManifest(manifest({ deny: [{ tools: ['bash'], reason: 'r' }] }))).toThrow(/malformed/);
|
||||
});
|
||||
|
||||
test('a manifest with no deny rules is refused: it could only be prompt text', () => {
|
||||
expect(() => parseManifest(manifest({ deny: [] }))).toThrow(/malformed/);
|
||||
});
|
||||
|
||||
test('a manifest cannot smuggle code past the schema', () => {
|
||||
const sneaky = manifest({ beforeToolCall: 'process.exit(1)', tools: { evil: {} } });
|
||||
const parsed = parseManifest(sneaky);
|
||||
expect('beforeToolCall' in parsed).toBe(false);
|
||||
expect('tools' in parsed).toBe(false);
|
||||
});
|
||||
|
||||
test('a manifest becomes a plugin whose guard blocks by tool and path', async () => {
|
||||
const plugin = manifestToPlugin(parseManifest(manifest()));
|
||||
const cwd = process.cwd();
|
||||
|
||||
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 'app/.env' }, cwd })).toContain(
|
||||
'refusing to write',
|
||||
);
|
||||
// A tool the rule does not name is untouched.
|
||||
expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { path: 'app/.env' }, cwd })).toBeUndefined();
|
||||
// A path the pattern does not match is untouched.
|
||||
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 'src/app.ts' }, cwd })).toBeUndefined();
|
||||
});
|
||||
|
||||
test('a command rule matches the command, not the path', async () => {
|
||||
const plugin = manifestToPlugin(
|
||||
parseManifest(manifest({ deny: [{ tools: ['bash'], commandPattern: 'curl.*\\| *sh', reason: 'no pipe to shell' }] })),
|
||||
);
|
||||
const cwd = process.cwd();
|
||||
expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { command: 'curl x.sh | sh' }, cwd })).toBe(
|
||||
'no pipe to shell',
|
||||
);
|
||||
expect(await plugin.beforeToolCall!({ toolName: 'bash', input: { command: 'echo hi' }, cwd })).toBeUndefined();
|
||||
});
|
||||
|
||||
test('a guard given nothing to match on allows the call', async () => {
|
||||
const plugin = manifestToPlugin(parseManifest(manifest()));
|
||||
const cwd = process.cwd();
|
||||
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: undefined, cwd })).toBeUndefined();
|
||||
expect(await plugin.beforeToolCall!({ toolName: 'write_file', input: { path: 42 }, cwd })).toBeUndefined();
|
||||
});
|
||||
|
||||
test('installed plugins load from disk, and a broken one is reported not fatal', async () => {
|
||||
await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest());
|
||||
await Bun.write(join(pluginsDir(), 'broken.json'), '{ not json');
|
||||
|
||||
const { plugins, errors } = await loadInstalledPlugins();
|
||||
expect(plugins.map((p) => p.name)).toEqual(['no-secrets']);
|
||||
expect(errors.map((e) => e.plugin)).toEqual(['broken']);
|
||||
expect(errors[0]?.message).toContain('not valid JSON');
|
||||
});
|
||||
|
||||
test('no plugins directory yields nothing rather than throwing', async () => {
|
||||
expect(await loadInstalledPlugins()).toEqual({ plugins: [], errors: [] });
|
||||
});
|
||||
|
||||
test('an installed plugin says it was installed, so /plugins can be trusted', async () => {
|
||||
await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest());
|
||||
const { plugins } = await loadInstalledPlugins();
|
||||
expect(plugins[0]?.description).toContain('installed');
|
||||
});
|
||||
|
||||
test('uninstall removes an installed entry and reports when there was none', async () => {
|
||||
await Bun.write(join(pluginsDir(), 'no-secrets.json'), manifest());
|
||||
expect(await uninstall('plugin', 'no-secrets')).toBe(true);
|
||||
expect(await Bun.file(join(pluginsDir(), 'no-secrets.json')).exists()).toBe(false);
|
||||
expect(await uninstall('plugin', 'no-secrets')).toBe(false);
|
||||
});
|
||||
|
||||
test('uninstall refuses a name that is not a plain entry name', async () => {
|
||||
expect(uninstall('skill', '../../../etc/passwd')).rejects.toThrow(/not a valid entry name/);
|
||||
});
|
||||
|
||||
/** A local server stands in for the registry: no network, real HTTP. */
|
||||
function serve(routes: Record<string, { body: string; status?: number }>) {
|
||||
return Bun.serve({
|
||||
port: 0,
|
||||
fetch(req) {
|
||||
const path = new URL(req.url).pathname;
|
||||
const hit = routes[path];
|
||||
if (!hit) return new Response('not found', { status: 404 });
|
||||
return new Response(hit.body, { status: hit.status ?? 200 });
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
test('fetchIndex reads an index over http', async () => {
|
||||
const server = serve({ '/index.json': { body: index() } });
|
||||
try {
|
||||
const entries = await fetchIndex(`${server.url}index.json`);
|
||||
expect(entries.map((e) => e.name).sort()).toEqual(['migration', 'no-secrets']);
|
||||
} finally {
|
||||
server.stop(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('an index that returns an error status is reported with the status', async () => {
|
||||
const server = serve({ '/index.json': { body: 'nope', status: 503 } });
|
||||
try {
|
||||
expect(fetchIndex(`${server.url}index.json`)).rejects.toThrow(/503/);
|
||||
} finally {
|
||||
server.stop(true);
|
||||
}
|
||||
});
|
||||
|
||||
const SKILL_BODY = `---
|
||||
name: migration
|
||||
description: Write a database migration
|
||||
---
|
||||
|
||||
Migrations live in db/migrations and are never renumbered.`;
|
||||
|
||||
test('stage returns what would be written without writing it', async () => {
|
||||
const server = serve({ '/m.md': { body: SKILL_BODY } });
|
||||
try {
|
||||
const entry: RegistryEntry = {
|
||||
kind: 'skill',
|
||||
name: 'migration',
|
||||
description: 'Write a database migration',
|
||||
url: `${server.url}m.md`,
|
||||
};
|
||||
|
||||
const staged = await stage(entry);
|
||||
expect(staged.path).toBe(join(skillsDir(), 'migration.md'));
|
||||
expect(staged.preview).toContain('never renumbered');
|
||||
expect(await Bun.file(staged.path).exists()).toBe(false);
|
||||
} finally {
|
||||
server.stop(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('install writes the skill where the loader will find it', async () => {
|
||||
const server = serve({ '/m.md': { body: SKILL_BODY } });
|
||||
try {
|
||||
const installed = await install({
|
||||
kind: 'skill',
|
||||
name: 'migration',
|
||||
description: 'd',
|
||||
url: `${server.url}m.md`,
|
||||
});
|
||||
|
||||
expect(installed.path).toBe(join(skillsDir(), 'migration.md'));
|
||||
expect(await Bun.file(installed.path).text()).toContain('name: migration');
|
||||
|
||||
const { loadSkills } = await import('../src/skills');
|
||||
const loaded = await loadSkills(home);
|
||||
const skill = loaded.find((s) => s.name === 'migration');
|
||||
expect(skill?.origin).toBe('registry');
|
||||
} finally {
|
||||
server.stop(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('a body whose name disagrees with the index is refused', async () => {
|
||||
const server = serve({ '/m.md': { body: SKILL_BODY } });
|
||||
try {
|
||||
expect(
|
||||
stage({ kind: 'skill', name: 'something-else', description: 'd', url: `${server.url}m.md` }),
|
||||
).rejects.toThrow(/calls itself "migration"/);
|
||||
} finally {
|
||||
server.stop(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('a skill with no frontmatter is refused rather than installed as prose', async () => {
|
||||
const server = serve({ '/m.md': { body: 'just some text, no frontmatter' } });
|
||||
try {
|
||||
expect(stage({ kind: 'skill', name: 'migration', description: 'd', url: `${server.url}m.md` })).rejects.toThrow(
|
||||
/frontmatter/,
|
||||
);
|
||||
} finally {
|
||||
server.stop(true);
|
||||
}
|
||||
});
|
||||
|
||||
test('a plugin manifest is validated before it can be staged', async () => {
|
||||
const server = serve({
|
||||
'/good.json': { body: manifest() },
|
||||
'/bad.json': { body: manifest({ deny: [{ tools: ['bash'], commandPattern: '([', reason: 'r' }] }) },
|
||||
});
|
||||
try {
|
||||
const staged = await stage({
|
||||
kind: 'plugin',
|
||||
name: 'no-secrets',
|
||||
description: 'd',
|
||||
url: `${server.url}good.json`,
|
||||
});
|
||||
expect(staged.path).toBe(join(pluginsDir(), 'no-secrets.json'));
|
||||
expect(staged.preview).toContain('denies write_file');
|
||||
|
||||
expect(
|
||||
stage({ kind: 'plugin', name: 'no-secrets', description: 'd', url: `${server.url}bad.json` }),
|
||||
).rejects.toThrow(/invalid pattern/);
|
||||
} finally {
|
||||
server.stop(true);
|
||||
}
|
||||
});
|
||||
@@ -106,6 +106,40 @@ test('the status bar reports model, agent, thinking, context, and spend', () =>
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('a context limit turns the raw token count into a percentage', () => {
|
||||
const app = render(
|
||||
<StatusBar
|
||||
model="gpt-5"
|
||||
agent="default"
|
||||
thinking="medium"
|
||||
contextTokens={60_000}
|
||||
contextLimit={120_000}
|
||||
cost="$0.10"
|
||||
toolCount={14}
|
||||
/>,
|
||||
);
|
||||
const frame = app.lastFrame() ?? '';
|
||||
expect(frame).toContain('50% ctx');
|
||||
expect(frame).not.toContain('60000');
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('the context percentage is capped at 100 rather than running over', () => {
|
||||
const app = render(
|
||||
<StatusBar
|
||||
model="m"
|
||||
agent="a"
|
||||
thinking="t"
|
||||
contextTokens={300_000}
|
||||
contextLimit={120_000}
|
||||
cost="$1"
|
||||
toolCount={1}
|
||||
/>,
|
||||
);
|
||||
expect(app.lastFrame()).toContain('100% ctx');
|
||||
app.unmount();
|
||||
});
|
||||
|
||||
test('the info panel renders a markdown body', () => {
|
||||
const app = render(<InfoPanel title="tools" hint="7 offered" lines={'- `read_file`\n- `bash`'} />);
|
||||
const frame = app.lastFrame() ?? '';
|
||||
|
||||
Reference in New Issue
Block a user