fix(plan): perbaiki bug entropy gate dan fixture test squash.rs

Ditemukan implementer Task 4 sebelum commit apapun (BLOCKED, bukan
kode salah): entropi Shannon mentah per-karakter tidak membedakan
prosa dari identifier acak — prosa berulang skor ~3.89 bit/char,
lebih tinggi dari UUID (~3.39). Tambah syarat "tanpa spasi" sebelum
cek entropi (meniru pre-filter headroom sendiri), turunkan ambang ke
3.0 pada skala mentah. Fixture test array JSON juga diperbesar
(repeat 5 -> 8) karena sebelumnya tidak pernah melewati SQUASH_FLOOR_BYTES
yang diasumsikan test itu sendiri.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
asepharyana
2026-07-16 07:49:43 +07:00
co-authored by Claude Sonnet 5
parent 12a03fd3d1
commit ceb84790bb
2 changed files with 39 additions and 14 deletions
@@ -647,7 +647,7 @@ mod tests {
#[test] #[test]
fn json_array_elements_past_third_are_squashed_harder() { fn json_array_elements_past_third_are_squashed_harder() {
let long_str = "the quick brown fox jumps over the lazy dog again and again ".repeat(5); let long_str = "the quick brown fox jumps over the lazy dog again and again ".repeat(8);
let value = serde_json::json!({ let value = serde_json::json!({
"items": [long_str.clone(), long_str.clone(), long_str.clone(), long_str.clone()], "items": [long_str.clone(), long_str.clone(), long_str.clone(), long_str.clone()],
}); });
@@ -756,9 +756,9 @@ pub fn apply(tool_name: &str, output: &str) -> String {
/// Compress a JSON tool result by keeping all structural content (keys, /// Compress a JSON tool result by keeping all structural content (keys,
/// array/object shape) and eliding long, low-entropy string *values*, /// array/object shape) and eliding long, low-entropy string *values*,
/// while keeping short values (<=20 chars) and high-entropy ones (UUIDs, /// while keeping short values (<=20 chars) and high-entropy single-token
/// hashes, paths) intact. Array elements past the first 3 are elided /// ones (UUIDs, hashes, paths) intact. Array elements past the first 3
/// regardless of length/entropy. /// are elided regardless of length/entropy.
/// ///
/// Why walk a parsed `Value` instead of hand-rolling a JSON tokenizer: /// Why walk a parsed `Value` instead of hand-rolling a JSON tokenizer:
/// `serde_json` already handles escaping/nesting correctly (this /// `serde_json` already handles escaping/nesting correctly (this
@@ -778,10 +778,25 @@ fn squash_json(text: &str) -> String {
/// Recursively elide long, low-entropy string values in place. /// Recursively elide long, low-entropy string values in place.
/// `in_late_array` is true once past the first 3 elements of an /// `in_late_array` is true once past the first 3 elements of an
/// enclosing array, tightening the elision rule for the rest of it. /// enclosing array, tightening the elision rule for the rest of it.
///
/// Why the `!s.contains(' ')` gate before the entropy check: raw
/// per-character Shannon entropy alone does NOT separate "meaningful
/// prose" from "random-looking identifier" — verified empirically,
/// repeated English prose scores ~3.89 bits/char, *higher* than a UUID's
/// ~3.39 or a SHA-256 hex digest's ~3.66, because prose draws from a
/// wide, fairly-balanced character set too. What actually distinguishes
/// identifiers from prose is that identifiers are a single unbroken
/// token — this mirrors headroom's own approach (its entropy check is
/// "cheaply pre-filtered by 'no spaces'" before scoring). Multi-word
/// values never reach the entropy branch at all; only whitespace-free
/// tokens do, where entropy correctly separates "abc123" or "aaaaaaaa"
/// (low, elided if long) from a UUID/hash/API-key-shaped string (high,
/// kept).
fn squash_json_value(value: &mut serde_json::Value, in_late_array: bool) { fn squash_json_value(value: &mut serde_json::Value, in_late_array: bool) {
match value { match value {
serde_json::Value::String(s) => { serde_json::Value::String(s) => {
let keep = !in_late_array && (s.len() <= 20 || shannon_entropy(s) >= 0.85); let looks_like_identifier = !s.contains(' ') && shannon_entropy(s) >= 3.0;
let keep = !in_late_array && (s.len() <= 20 || looks_like_identifier);
if !keep { if !keep {
*s = "".to_string(); *s = "".to_string();
} }
@@ -800,9 +815,13 @@ fn squash_json_value(value: &mut serde_json::Value, in_late_array: bool) {
} }
} }
/// Shannon entropy in bits per character — used to distinguish /// Shannon entropy in bits per character — used, after the `squash_json`
/// high-entropy strings (UUIDs, hashes, random IDs, worth keeping) from /// caller's own "no internal whitespace" pre-filter, to distinguish
/// low-entropy prose (safe to elide). /// high-entropy single-token strings (UUIDs, hashes, random IDs, worth
/// keeping) from low-entropy ones (e.g. `"aaaaaaaaaa"`, safe to elide).
/// 3.0 sits comfortably below a UUID's ~3.39 and a SHA-256 hex digest's
/// ~3.66 (both empirically measured with this exact formula) while
/// staying well above a degenerate repeated-character string's 0.0.
fn shannon_entropy(s: &str) -> f64 { fn shannon_entropy(s: &str) -> f64 {
if s.is_empty() { if s.is_empty() {
return 0.0; return 0.0;
@@ -148,12 +148,18 @@ bytes pass through unchanged (compression only pays off on large output, and tou
results risks losing detail with no token benefit). Above the floor, dispatch by content results risks losing detail with no token benefit). Above the floor, dispatch by content
shape: shape:
- `squash_json(&str) -> String`hand-rolled JSON tokenizer; structural tokens (keys, - `squash_json(&str) -> String`walks a parsed `serde_json::Value` (not a hand-rolled
brackets, colons, commas, booleans, null) always kept; string values kept if ≤20 chars or tokenizer — `serde_json` already handles escaping/nesting correctly, reusing it is simpler
high-entropy (Shannon entropy ≥0.85 bits/char, catches UUIDs/hashes/paths — same threshold and more robust); structural tokens (keys, brackets, colons, commas, booleans, null) always
headroom uses), otherwise replaced with `"…"` in place; array elements past the first 3 kept; string values kept if ≤20 chars or "identifier-shaped" (no internal whitespace *and*
compressed harder (values elided regardless of length/entropy). Applied when Shannon entropy ≥3.0 bits/char — catches UUIDs/hashes/paths), otherwise replaced with `"…"`
`serde_json::from_str` on the output succeeds. in place; array elements past the first 3 compressed harder (values elided regardless of
length/entropy). Applied when `serde_json::from_str` on the output succeeds. The
no-whitespace pre-filter matters: raw per-character entropy alone doesn't separate prose
from identifiers — repeated English prose measures ~3.89 bits/char, higher than a UUID's
~3.39 — because prose also draws from a wide character set. headroom's own entropy gate is
"cheaply pre-filtered by 'no spaces'" before scoring for the same reason; multi-word values
never reach the entropy check at all under this rule.
- `squash_log(&str) -> String` — line classifier (error/fail/warn/info/debug/trace by - `squash_log(&str) -> String` — line classifier (error/fail/warn/info/debug/trace by
keyword + stack-trace-frame detection) → score keyword + stack-trace-frame detection) → score
(`level_score {1.0 error/fail, 0.5 warn, 0.1 info, 0.05 debug/trace} + 0.3 if (`level_score {1.0 error/fail, 0.5 warn, 0.1 info, 0.05 debug/trace} + 0.3 if