From f62ac3f688a096089eb3adedb2a03e59e11c696c Mon Sep 17 00:00:00 2001 From: asepharyana Date: Thu, 16 Jul 2026 03:22:34 +0700 Subject: [PATCH] docs(specs): tambah desain rombak context & compaction Rencana pengganti shortsend.rs dengan modul context/ (dedup, squash, shaping, tokens, window) plus mode ringkas opsional, disusun dari studi teknik rtk-ai/rtk, headroomlabs-ai/headroom, dan JuliusBrussee/caveman (ide saja, tanpa menyalin kode). Co-Authored-By: Claude Sonnet 5 --- ...7-16-context-compaction-overhaul-design.md | 264 ++++++++++++++++++ 1 file changed, 264 insertions(+) create mode 100644 docs/superpowers/specs/2026-07-16-context-compaction-overhaul-design.md diff --git a/docs/superpowers/specs/2026-07-16-context-compaction-overhaul-design.md b/docs/superpowers/specs/2026-07-16-context-compaction-overhaul-design.md new file mode 100644 index 0000000..3a859cb --- /dev/null +++ b/docs/superpowers/specs/2026-07-16-context-compaction-overhaul-design.md @@ -0,0 +1,264 @@ +# Context & Compaction Overhaul — Design + +**Status:** Approved, pending implementation plan +**Date:** 2026-07-16 +**Scope:** replaces `src/app/runtime/shortsend.rs`; touches `src/app/runtime/actions/mod.rs`, +`src/view/status.rs`, `src/model/settings.rs`, `src/app/subagent/division.rs`, `Cargo.toml` + +## Context + +The existing conversation-compaction system (`shortsend.rs`, 129 lines) only acts once the +context is already close to the model's window limit, and has accumulated inconsistencies +found during a codebase audit: + +1. Three different token-count heuristics for the same job: `/3` inside + `shortsend::shape_messages`, `/4` in the auto-compact loop + (`actions/mod.rs` ~line 1146), `/4` again in the live status bar (`view/status.rs:68`). +2. Manual `/compact` (`Action::Compact`, `actions/mod.rs:547-563`) passes `client: None` + because `apply_action` is synchronous, so it never gets LLM summarization — it always + falls back to the bare `"[prior conversation compacted]"` placeholder, unlike automatic + mid-turn compaction (`Some(&tc.client)`, line 1160). Undocumented asymmetry between the + two trigger paths. +3. `context_window` resolution (`model_roles.values().find(...).and_then(...).unwrap_or(...)`) + duplicated three times (`Action::Compact`, `spawn_turn`, `view/status.rs` twice). +4. No repeated-tool-call dedup: reading the same file (or running the same grep) twice in a + session keeps both full copies in context forever, until compaction eventually drops the + older one wholesale along with everything else from that period. +5. No per-result compression: a single large tool output (a big `bash` log, a large `grep` + result) is stored verbatim even when most of it is redundant or low-value. +6. Zero test coverage on `shortsend.rs`. + +Separately, research into three real, permissively-licensed open-source projects +(`rtk-ai/rtk`, Apache-2.0; `headroomlabs-ai/headroom`, Apache-2.0; `JuliusBrussee/caveman`, +MIT — verified via `gh api` for authenticity/license, and by cloning and reading source, not +taken from marketing blog posts) surfaced techniques worth reimplementing natively: + +- **rtk**: generic line-scan compression (strip comment/blank runs, brace-depth collapse of + function bodies, importance-ranked truncation ending in an unambiguous `[N more lines]` + marker — their own regression tests show a comment-shaped marker confuses the LLM into + retry-looping) plus structured per-toolchain parsing (e.g. `cargo --message-format=json` + bucketed into errors/warnings, boilerplate lines dropped). +- **headroom**: per-content-type compressors — logs (classify lines by level/stack-trace/ + summary, score, keep highest-value lines + surrounding context, adaptive cap), grep + results (group by file, score matches, cap globally and per-file), JSON (keep all + structural tokens — keys, brackets, colons — drop or shrink long low-entropy string + values, keep short values and UUID/hash-shaped high-entropy ones). +- **caveman**: a pure prompt/persona instruction (no algorithm) that tells the model to + write tersely — drop articles/filler/hedging, keep code/commands/errors verbatim — with an + explicit carve-out that disables terseness for destructive-op confirmations and security + warnings. This compresses *output* tokens, a different axis from everything else in this + design, which compresses *input* context. + +This is a from-scratch reimplementation of the underlying ideas, not a port — no code is +copied from any of the three projects. + +## Goals + +- One unified, always-on pipeline that keeps context lean from turn 1, not just once near + the limit. +- Deduplicate repeated tool calls: an older copy of a tool result superseded by an identical + later call (same tool name + same arguments) is replaced with a placeholder, for read-only + tools only. +- Compress large individual tool results (logs, JSON, generic text) at capture time, above a + size floor. +- Fix the three known inconsistencies (token heuristic, manual/auto asymmetry, + `context_window` duplication). +- Optional, off-by-default "concise mode" system-prompt toggle for terser model output. +- Full inline test coverage per repo convention. + +## Non-goals + +- Not adding a runtime dependency on `rtk`, `headroom`, or `caveman` themselves (as a binary, + proxy, or crate) — everything is implemented natively in Rust inside zesdex. +- Not building rtk's per-toolchain structured parsers (`cargo --message-format=json` + re-invocation, etc.) — too invasive for a general-purpose `bash` tool that runs arbitrary + commands zesdex doesn't control the flags of. Only the generic line-scan/log/JSON layer is + built. +- Not switching to an exact per-provider tokenizer — `tiktoken-rs` (BPE, cl100k_base/ + o200k_base) is an approximation good enough for the 85%/95% budget thresholds; it is not + used for billing-accurate counts. +- `caveman-compress`-style memory-file rewriting (the LLM-round-trip variant of caveman) is + out of scope — only the pure-prompt persona mechanism is adopted. + +## Architecture + +Replace `src/app/runtime/shortsend.rs` with `src/app/runtime/context/`: + +``` +context/ + mod.rs — prepare() facade: dedup → count → maybe-shape + tokens.rs — unified token counting (tiktoken-rs) + dedup.rs — cross-call tool-result deduplication + squash.rs — per-result compression (log/json/generic), applied at + tool-result construction time, upstream of prepare() + shaping.rs — budget-based drop + LLM summarize (renamed shortsend logic) + window.rs — shared context_window resolution +``` + +### `tokens.rs` + +```rust +pub fn count_tokens(text: &str) -> usize +pub fn count_message_tokens(msg: &ChatMessage) -> usize +``` + +Backed by `tiktoken-rs` (new dependency, pure Rust, embedded BPE vocab, no network calls at +runtime), using `o200k_base`. Replaces all three existing heuristic call sites: `shortsend`'s +internal `/3`, the auto-loop's `/4` (`actions/mod.rs` ~1146), and `status.rs:68`'s `/4`. + +### `dedup.rs` + +```rust +pub fn collapse(messages: &[ChatMessage]) -> Vec +``` + +Flow: walk messages, pair each `Role::Tool` message to its originating `ToolCall` via +`tool_call_id`. Key = `(function.name, sha256(canonical_json(function.arguments)))` (`sha2` +is already a dependency). Track the last index seen per key. For any earlier occurrence of a +key whose tool name is in the read-only set, replace that earlier `Tool` message's `content` +with a short placeholder (`"[duplicate result — superseded by a later identical call, see +below]"`); the assistant's tool-call entry (name + arguments) is left untouched, so the +action/audit trail stays intact. Mutating tools are never touched, even with identical +arguments, because call order and repetition can be semantically meaningful (e.g. retrying a +flaky `bash` command). + +Read-only classification reuses `subagent::division::tool_scope::READ_TOOLS` +(`src/app/subagent/division.rs:21`) rather than a new list — that `const` is made `pub` for +this purpose. It already enumerates exactly the read-only tool set (`read`, `grep`, `glob`, +`search`, `seqthink`, `recall`, `lsp_*`, `read_findings`). + +Runs every turn, unconditionally, before token counting — not gated on `should_shape`. + +### `squash.rs` + +```rust +pub fn apply(tool_name: &str, output: &str) -> String +``` + +Size floor: outputs under 1500 bytes pass through unchanged (compression only pays off on +large output, and touching small results risks losing detail with no token benefit). +Above the floor, dispatch by content shape: + +- `squash_json(&str) -> String` — hand-rolled JSON tokenizer; structural tokens (keys, + brackets, colons, commas, booleans, null) always kept; string values kept if ≤20 chars or + high-entropy (Shannon entropy ≥0.85 bits/char, catches UUIDs/hashes/paths — same threshold + headroom uses), otherwise replaced with `"…"` in place; array elements past the first 3 + compressed harder (values elided regardless of length/entropy). Applied when + `serde_json::from_str` on the output succeeds. +- `squash_log(&str) -> String` — line classifier (error/fail/warn/info/debug/trace by + keyword + stack-trace-frame detection) → score + (`level_score {1.0 error/fail, 0.5 warn, 0.1 info, 0.05 debug/trace} + 0.3 if + stack-trace-frame + 0.4 if summary-shaped line`) → keep up to 20 highest-scored error + lines, up to 10 highest-scored warning lines, all summary lines, plus a ±2-line context + window around each kept line → single `[N lines omitted]` marker for drops (not + comment-shaped, per rtk's own finding on LLM confusion). Applied when the output isn't + valid JSON and has ≥3 lines matching error/warn/stack-trace patterns. +- `squash_generic(&str, budget) -> String` — importance-ranked truncation: keeps the first 10 + and last 10 lines plus any line matching a small "looks important" heuristic (non-blank, + not a byte-for-byte repeat of the immediately preceding line), single `[N lines omitted]` + marker for the rest, capped to `budget` bytes overall (`budget` = the 1500-byte squash + floor doubled, i.e. 3000 bytes, chosen so the fallback path still yields a real reduction + on anything that triggered it). Fallback for anything that isn't JSON or log-shaped. + +Called once, at the single tool-result construction site +(`actions/mod.rs:1420`, `let tool_msg = ChatMessage::tool_result(tool_call.id.clone(), +output);`) — `output` is passed through `squash::apply(&tool_name, &output)` before being +wrapped. Runs before the result is ever archived or pushed into `msgs`, so compression is +permanent and applies uniformly whether or not compaction ever triggers. + +### `shaping.rs` + +Unchanged behavior from today's `shortsend.rs` (hysteresis `should_shape`, 70%-budget +newest-first retention, LLM summarization of dropped messages), moved as-is into this file +and updated to source token counts from `tokens.rs` instead of its own heuristic. + +### `window.rs` + +```rust +pub fn resolve(app_config: &AppConfig, settings: &Settings) -> usize +``` + +Replaces the three duplicated `model_roles.values().find(...).and_then(...).unwrap_or(...)` +blocks in `Action::Compact`, `spawn_turn`, and `view/status.rs` (×2). + +### `mod.rs` + +```rust +pub fn prepare( + messages: &[ChatMessage], + max_wire_tokens: usize, + force: bool, + client: Option<&LlmClient>, +) -> Vec +``` + +Orchestrates: `dedup::collapse` → per-message `squash::apply` is already baked into messages +by the time they arrive here (applied earlier, at construction) → `tokens::count_message_tokens` +sum → `shaping::should_shape` → conditionally `shaping::shape`. This is the single entry point +both `run_agent_turn`'s auto-loop and `Action::Compact` call. + +## Data flow (per turn) + +1. Tool executes → raw `output: String`. +2. `squash::apply(tool_name, &output)` — compress if over the size floor. +3. Wrapped into `ChatMessage::tool_result(...)`, archived, pushed to `msgs`. +4. Once per loop iteration: `context::prepare(&msgs, tc.context_window, force=false, + Some(&tc.client))` — dedup pass always runs; shaping runs only if `should_shape` trips. +5. Result pushed as `TurnEvent::Compacted` if shaping actually changed anything (same as + today), consumed on the main thread to update `SessionRuntime.messages`. + +## Fixing the manual/auto asymmetry + +`Action::Compact` (`actions/mod.rs:547`) currently runs synchronously inside `apply_action` +and can't block on an LLM call. Fix: make it spawn a background `std::thread::spawn` — the +same pattern `spawn_turn` already uses (`actions/mod.rs:694`) — that calls +`context::prepare(&rt.messages, max_wire_tokens, force=true, Some(&client))` and reports back +via `TurnEvent::Compacted`, identical to the automatic path. The toast sequence becomes +"Compacting…" immediately (optimistic, non-blocking) then "History compacted" when the +`TurnEvent` arrives. This gives manual `/compact` real LLM summarization instead of always +falling back to the placeholder. + +## Concise mode (separate from the `context/` module) + +- `Settings` (`src/model/settings.rs`) gains `pub concise_output: bool`, default `false`, + with `#[serde(default)]` for backward-compatible deserialization of existing + `settings.json` files (matching the existing `hive_mind_node_timeout_ms` precedent in the + same file). +- When `true`, `run_agent_turn`'s system-prompt assembly (`actions/mod.rs:930-936`) appends a + fourth section to `system_text`: a terse-writing instruction (persona-prompt only, no + algorithm — drop articles/filler/hedging/pleasantries, keep code/commands/error text + byte-exact) with an explicit carve-out disabling terseness for destructive-operation + confirmations and security-relevant warnings, mirroring caveman's own "Auto-Clarity" + safety exception. +- No UI toggle is in scope for this pass — confirmed no such mechanism exists today for any + boolean `Settings` field (`review_enabled`, `session_archive_enabled`, + `lsp_auto_provision` are all hand-edited in `settings.json`, same as this one will be). + +## New dependency + +`tiktoken-rs` — pure Rust, embedded BPE vocab (`cl100k_base`/`o200k_base`), no network calls +at runtime, MIT/Apache-2.0 dual-licensed. Added to `Cargo.toml`. + +## Testing + +Inline `#[cfg(test)] mod tests` per repo convention, one per new file: + +- `dedup.rs`: same tool+args → older result replaced; different args → no-op; mutating tool + with identical args → both kept in full; unmatched `tool_call_id` (malformed history) → + no panic, treated as unpaired. +- `squash.rs`: JSON input under/over the size floor; JSON with long low-entropy string values + gets them elided while short/UUID-shaped values survive; log input with error/warn lines + keeps highest-scored lines and emits exactly one `[N lines omitted]` marker; generic text + keeps first/last N lines. +- `tokens.rs`: known-string token counts against fixed expected values; empty string → 0. +- `shaping.rs`: port the behavioral cases implied by today's hysteresis logic (85% trigger + when not previously shaped, 95% once shaped) plus budget-drop ordering. +- `window.rs`: role match resolves to the role's `context_window`; no match falls back to + `default_context_window`. + +## Migration + +- Delete `src/app/runtime/shortsend.rs`; all three call sites (`actions/mod.rs` auto-loop, + `Action::Compact`, and the module path itself) updated to `context::`. +- `view/status.rs` switches its live token display to `tokens::count_tokens`, so the status + bar finally matches what compaction measures internally.