Agentic coding CLI on Bun, Ink, and the AI SDK. Core: streamText loop with SDK-level tool approval so a denied call provably never executes; endpoint fallback for OpenAI reasoning models; retry with backoff. Tools: read/write/edit/glob/grep/bash, path-jailed, gitignore-aware, ripgrep with a JS fallback, binary rejection, live bash streaming. Agents: five variants crossing thinking level with tool restriction; plan and review withhold mutating tools from the model. Extensibility: frontmatter skills with on-demand bodies, plugin host with blocking hooks, MCP stdio and HTTP, read-only subagents. State: durable per-project memory, session task lists, session persistence, compaction that repairs provider-item dependencies. Distribution: five-platform cross-compiled binaries with checksums, install scripts, CI on three operating systems. 404 tests, typecheck clean.
4.2 KiB
4.2 KiB
TODO
Next up. One item, one outcome, verifiable when done.
Longer-term direction lives in ROADMAP.md.
Now
Show reasoning in the transcript
src/session.ts already yields { type: 'reasoning' }; src/ui/App.tsx ignores it.
- Accumulate reasoning deltas into their own buffer, separate from
text - Render as a dim collapsed panel:
thinking... 412 tokens, expandable with a key - Drop it from the transcript when the turn ends — reasoning is not part of the answer
- Test: a model emitting
reasoning-deltaputs text on screen before anytext-delta
Show the file being touched
tool-call carries the path but the transcript only shows a summary after the call returns.
- Render an active-tool line while a call is in flight:
read src/session.ts - Clear it on
tool-resultortool-error - Test: a slow tool leaves its line on screen for the duration
Queue prompts typed during a turn
- Keep
PromptInputmounted whilebusy, alongside the spinner - Submitting while busy appends to a queue and shows
queued: 2 - Drain the queue in order when the turn ends
escclears the queue as well as aborting- Test: two prompts typed during a turn run in order afterwards
Next
activeTools gating per tool set
Needed before the tool count grows. Measured at 553 chars of schema per tool per request.
toolSetsin config: which sets are live- Sets:
core,git,edit-plus,net prepareStepnarrowsactiveToolsto the enabled sets/toolsshows which set each tool came from- Test: a disabled set's tools reach neither the wire nor the prompt
multi_edit
- Several
{ oldString, newString }edits against one file - Atomic: any failing match aborts the whole call, file untouched
- Each edit applied to the result of the previous one
- Approval prompt shows one combined diff
- Test: a failing second edit leaves the file exactly as it was
list_dir
- Tree view honouring
.gitignore, depth-limited, entry-capped - Marks directories and shows file sizes
- Test: respects ignore rules, stops at the depth limit
Git read-only tools
All approval-free, since none can mutate.
git_status,git_diff,git_log,git_show,git_blame- Structured output, not raw porcelain
- Fail clearly outside a repo instead of returning git's error text
- Test: each returns something usable in a temp repo, and a clean error outside one
@file completion
@in the input opens a path picker fed by the ignore-aware walker- Tab completes, continued typing narrows
- Completed path inserted as a plain relative path
- Test:
@src/narrows to files undersrc/
Maintenance
- Pricing table needs a source note and a date; rates drift and ours are hand-entered
estimateTokensdivides JSON length by four. Good enough for a compaction threshold, wrong enough to mislead in/cost. Either label it an estimate everywhere or use a real tokenizer- The subagent shares the parent's model. A cheaper model for search would cut cost
substantially on
exploreruns - No spend ceiling. A headless run that loops costs real money with nothing to stop it
Known rough edges
Not bugs exactly, but things that will bite someone.
/clearwipes the terminal scrollback.<Static>output is already committed, so clearing React state alone leaves it on screen. The escape sequence works but takes the user's earlier terminal history with it.- Compaction is lossy in a way the model cannot see. It is told the history was pruned, but not what was in the pruned part. A summary of the discarded span would be better than a count.
- Memory has no conflict resolution. Two contradictory notes both persist and both get
injected.
/memorymay merge them, or may keep both. bashcannot be interrupted independently.escaborts the whole turn, killing the command. There is no way to stop a runaway command and keep the turn.- Windows
cmd /cdiffers frombash -lc. A command the model writes for one shell may fail on the other. The prompt states the platform; it does not translate.