Collapsed-history images now carry who-said-what AND when, instead of leaving
the model to reconstruct it from a flattened, role-ambiguous blob (which led it
to resurface the opening turn as if it were the live request).
- Turn-index recency anchor: each collapsed turn serializes as <user t="N"> /
<assistant t="N"> with an ABSOLUTE turn index (larger N = more recent), so the
model can tell turn 1 from turn 60. Absolute (never "N ago"/"i of total") so
frozen chunks stay byte-identical and keep hitting the prompt cache. Mirrored on
the GPT path. (history.ts, openai-history.ts, openai.ts)
- colorByRole: role tags are tinted via a parallel "slot string" carried from
serialize time (structure-through), replacing the parse-back that miscolored a
body quoting a literal <tag>. (render.ts, history.ts)
- neutralizeSentinel: a pre-existing ↵ (U+21B5) in content is swapped to ⏎
(U+23CE) in render-prep so reflow packs newlines instead of bailing to a raw,
unpacked render — common when the content is about pxpipe itself. Render-only;
originals preserved. (render.ts, transform.ts, history.ts, openai-history.ts)
- Per-model GPT vision/render profiles, retunable via PXPIPE_GPT_PROFILES,
behavior-identical to the old hardcoded values. (gpt-model-profiles.ts)
Review fixes folded in:
- Banner intro/outro consolidated to a single source of truth; the GPT path now
aliases the Anthropic constant instead of byte-copying it (can't drift).
- slotCopyBody neutralizes literal slot markers to a width-equivalent control
char (U+0003), not a space the minifier would strip and misalign.
459 tests pass; tsc clean.
Collapse old conversation history into rendered PNG sections so the model
reads a compact image instead of re-billed text, while preserving prompt
caching and tool-call behavior. Measures real vs compressed token/cost.
Core:
- GPT history collapse (openai-history.ts): append-only, o200k token-length
sectioning. Sections seal only at a tool-closed boundary (open call-id set
empty), so a function_call and its function_call_output never split across
the collapse cut. Fixes the OpenAI 400 "No tool call found for function
call output with call_id ..." that hit long Responses-API sessions.
- Anthropic cache contract (history.ts): append-only per-chunk rendering;
cache_control markers are preserved/moved, never added; chunk boundaries
align with caller marker seams for byte-stable prefix caching.
- GPT image budget (openai.ts): detail:'original' for gpt-5.x, flagship
vision-multiplier fix, patch cap; schema-strip preserves real descriptions.
- Savings accounting (openai-savings.ts): cached_tokens + vision-token basis.
Model scope (applicability.ts):
- Default imaged scope = claude-fable-5 + gpt-5.6.
- gpt-5.5 and claude-opus-4-8 stay opt-in: same pipeline, but they degrade
reading dense imaged history (gist drift), so silently imaging them by
default is wrong. Promotion is gated on an OCR/recall threshold.
Dashboard: GPT + Anthropic rendering, per-family model toggles, persisted
metrics, thumbnail-expired session UI, reflow/newline handling.
Tests: cache-alignment (GPT + Anthropic), history sectioning + tool-boundary
invariants, savings, dashboard, sessions/restart-restore. 452 passing.
The history image sits after the slab in prefix order, but the single cache
breakpoint was anchored on the slab. So the largest block — the ~141k-token
history image — only cached when the caller's roaming downstream marker happened
to land after it. When it didn't, the whole history image re-created at the 1.25x
cache_create rate every turn instead of amortizing into 0.1x reads. On a warm
session that turned a real compression win into a net loss (the "411% more"
the dashboard correctly surfaced).
Fix is pure relocation: move pxpipe's single breakpoint from the slab onto the
LAST history image (which sits after the slab), so slab + history cache as one
stable prefix — created once, then read. Marker count is unchanged: pxpipe still
never ADDS a marker (Task #21), it only moves the caller's existing one. When no
slab imaged, the history path keeps its own relocation.
- transform.ts: relocateAnchorToHistoryImage on the collapse-success branch
- history.test.ts: two old "history carries no marker" invariants upgraded to
assert the relocation (last history image carries it; exactly one marker total)
replay() rebuilt RecentRow from the JSONL but never set
session_saved_so_far_delta, img_id, or contextHistory — so after a restart
old rows showed "-" in Saved and had no Details link, while live rows did.
Live update() does all three; replay() now matches.
- replay() sets session_saved_so_far_delta (baselineInputEff - actualInputEff,
same cache-weighted math as live) so the Saved column survives a restart.
- replay() rebuilds the Context Map breakdown from persisted fields
(baseline_tokens, bucket_chars, image_count, usage) with a synthetic
nextImageId, and sets img_id/img_ids so the Details link resolves.
- The PNG ring is in-memory and not restorable; restored entries carry
imageIds:[] and a `restored` flag, and the Details panel shows
"thumbnails expired when the proxy restarted" instead of dead <img>s.
- +2 replay tests (restart-restore.test.ts).
376 tests pass; typecheck + build clean.
Claude-Session: https://claude.ai/code/session_01RV8uiMYNhrAdTUzJ9fkpcp
The session hero computed (1 - sent / raw_count_tokens_baseline) — the same
cache-blind math just removed from the Details panel. It priced the all-text
baseline at full rate, ignoring that most of it would have been cheap
cache-reads, so a warm net-loss session showed a big "fewer tokens" win up top
while Details/Saved showed a loss. Two panels, one session, opposite signs.
- renderSessionSummaryFragment now reads baselineInputWeighted /
actualInputWeighted (the same cache-weighted pair as Details + the Saved
column) instead of rawActualTokens / rawBaselineTokens.
- Headline is input-only; dropped the output-on-both-sides lumping (output is
never "sent" and only dampened the %).
- Honest labels: "effective tokens … after normal cache discounts"; still no $
assumptions (0.1x/1.25x are token multipliers, not $/Mtok).
- types.ts: corrected stale field comments (raw fields are math-drawer only).
- +4 hero tests pinning direction to the weighted pair and that output can't
move the headline %.
374 tests pass; typecheck + build clean.
The Details panel divided the raw count_tokens baseline by sent tokens
(cache-blind), over-claiming "X% smaller" even when the cache-aware row
showed a net loss — the two panels could flatly contradict each other.
- dashboard.ts: feed baselineInputEff/actualInputEff/haveBaseline into
ContextMapData (same weighted pair the recent row uses); push after
the eff-tokens are computed so the breakdown matches.
- fragments.ts: headline divides weighted baseline by weighted actual,
says "bigger" on a net loss, with a clarifying sub-line; no savings
claim when the baseline probe did not resolve.
- +4 tests (context-map.test.ts) pinning headline direction to the row.
370 tests pass; typecheck + build clean.
- Generalize isPxpipeSupportedGptModel to the GPT-5 family (gpt-5, 5.x,
-mini/-nano), default-on; [variant] tags stripped.
- OpenAI render switches to the 768px portrait strip (GPT_STRIP_COLS=152)
so the API's mandatory shortest-side->768 downscale never shrinks our
5px glyphs; downscale-free in both tile and patch regimes.
- Add OpenAI vision-token cost model (resolveVisionCost/openAIVisionTokens)
and rewire the Chat Completions profitability gate to use it instead of
the Anthropic 750px/token model.
- Add transformOpenAIResponses for the /v1/responses (Responses API) shape
used by Codex: instructions + input items + flat tools -> rendered
input_image parts; wire it into the proxy (was passthrough before).
- Tests: +12 (gate family, vision-token math incl. downscale path, chat +
responses transforms). 366 passing; Anthropic path unchanged.
README intentionally not updated (GPT OCR fidelity unverified until tested).
- keepSharp: caller predicate pins exact-match-critical blocks (IDs, hashes,
paths) as text, overriding the chars/token heuristic. Reported via
info.keptSharpBlocks. Never forces imaging; throwing predicate = false.
- emitRecoverable: optional recovery map (block -> original text + provenance)
so a stateful caller can re-inject or re-fetch imaged content. The recover
half of the fidelity contract; documented mitigation for lossy recall.
- edge/Workers-safe: process.env read is typeof-guarded; @napi-rs/canvas moved
to devDependencies (atlas is build-time only, runtime PNG encoder is pure-JS).
- comments: trimmed ~1,640 lines of essay/derivation/changelog narration across
src/ (transform.ts -73%), keeping one-line rationale; deep math points to docs/.
- New exports/types in index.ts + library.ts; +12 tests (keep-sharp, recoverable).
Build clean; 354/354 tests pass.
The real-corpus reflow check read whole transcript files; on a large ~/.claude/projects (multi-hundred-MB files) that blew the 30s timeout and aborted 'pnpm test' (hence prepublishOnly). Read only the first 8MB per file — enough for the per-file sample; a truncated trailing line just fails JSON.parse and is skipped. Corpus test now runs in ~5s.
- "How your context works" panel: per-request token flow (as-text -> real) plus
the exact-char bucket breakdown of what became images and a gallery of every
rendered page, opened via a "view" link on each recent-requests row.
- "compress models" chips are the union of a catalog (Fable 5, Opus 4.8/4.7,
Sonnet 4.6, Haiku 4.5), the PXPIPE_MODELS env scope, and the active scope, so
any env-enabled model stays toggleable (off <-> on); runtime-only override.
- Opus is OFF by default (Fable-5-only scope); opt-in via env or the chips.
Fixes from the code review:
- contextHistory cap aligned to the recent table (was 30 < 50, so older "view"
links silently showed the latest request's breakdown); an evicted or
unrecorded request now shows an explicit "no longer available" message.
- the session headline phrases a net loss as "N% more tokens" instead of
rendering a garbled "-7% fewer tokens".
Per-turn and per-session savings are the real baseline_eff - actual_eff with no
floor: a net-losing turn (cache_create-heavy image rewrite that costs more than
the text prefix it replaced) now reports the real loss instead of a fabricated
0. Updates baseline.ts and the caller comments that claimed a >=0 clamp the code
never applied -- the negative is intended and covered by tests/sessions.test.ts.
Fable 5 / Opus 4.8 accept up to 2576px / 4784 visual tokens per image, but a
request with >20 images (pxpipe always sends many) is held to the stricter
<=2000px/side rule, so the real ceiling is ~1932x1932 (1928x1928 = 69x69 = 4761
tokens). MAX_HEIGHT_PX 1568->1932; dense tool/history pages now render at
DENSE_CONTENT_COLS=384 / DENSE_CONTENT_CHARS_PER_IMAGE=92160 (1928x1928 full
page) -- fewer image blocks at the same OCR-validated 5x8 cell. The static slab
is unchanged (313 cols / 1573x1280).
Also fixes a regression the code review caught: the break-even gate,
truncateForBudget, and estimateImageCount predicted against the slab geometry
(313 cols / 159 rows) while the dense renderer emits 384 cols / 240 rows, so
large tool_results were truncated far earlier than the 10-image cap required --
silently dropping output that would have rendered. A per-page char budget is
threaded through the prediction chain (defaults to READABLE so the slab and unit
tests are byte-identical) plus a denseGateGeometry helper, so the gate prices the
same page the renderer actually produces.
- applicability: isPxpipeSupportedModel reads PXPIPE_MODELS (comma-sep,
default claude-fable-5 only) and strips bracket/variant tags like [1m];
Opus becomes opt-in with no code change. Default behavior unchanged.
- FINDINGS: Opus 4.8 re-test - arithmetic 98/100 (-2pp), verbatim 6/15
(was 93/100, 0/15): improved but still taxed -> default stays Fable-only.
- FINDINGS + eval/glyph-matrix/sweep: root cause of the read tax is per-glyph
RESOLUTION, not font shape (confusions scale with px and vanish by 10x16).
Opus needs ~4x glyph area to read reliably; the 1568px ceiling makes that
break-even, so there is no profitable Opus operating point.
- src/dashboard/fragments.ts: server-side HTML fragment renderers reusing the
same JSON payloads (HTML and JSON surfaces can't drift)
- node.ts/dashboard.ts: /fragments/<name> routes (header, session-summary,
recent, latest, sessions, stats, toggle) with htmx polling + Alpine for
local state; toggle is a POST-returning-fragment htmx swap
- vendored htmx 2.0.4 + Alpine 3.14.9 in src/dashboard/vendor.ts (offline,
single-file npm package, no CDN)
- deleted Svelte components/stores/bundle build step; dropped svelte/vite
devDeps; build is now just tsc + esbuild
- tests: +5 fragment contract tests (routing, toggle state, HTML escaping of
source text, 404) - 339 total
- PXPIPE_PROVIDER=cloudflare-ai-gateway + PXPIPE_GATEWAY_BASE_URL derive both
family routes ({base}/anthropic, {base}/openai) from one configurable base
- PXPIPE_GATEWAY_HEADERS injects auth headers (JSON or k=v;k2=v2) on every
upstream request incl. count_tokens probes
- OpenAI paths drop /v1 to match gateway shape; Anthropic paths unchanged;
unknown paths pass through on the family route; streaming untouched
- no hardcoded hostnames; tests use fake gateway URLs/tokens (14 new)
The prior commit captured only the rename (a stale pathspec aborted the
content staging). This applies the actual edits: retitle to FINDINGS,
'## Verdict' + superseded TL;DR, and the two cross-reference updates.
Widen the model gate to /^claude-opus-4-(?:[7-9]|[1-9]\d)(?:-|$)/ (Opus 4.7
and newer; previously 4.6/4.7). Wire isPixelpipeSupportedModel into the proxy
boundary in proxy.ts — the proxy previously gated on economics only and
ignored the model, so it compressed 4.8 traffic despite the docs claiming a
4.6/4.7 scope. Unsupported models now pass through with reason
'unsupported_model'.
Update public-api tests to the 4.7+ spec (320/320 pass).
Correct the docs to live measurement: reverse the POSTMORTEM "dead" verdict
(pixelpipe is a lossy gist-compressor saving ~68% on real dense Claude Code
traffic; the verbatim 0/15 needle finding stands as a caveat, not the
verdict), rewrite README Status/Limitations, and add a correction pointer to
the eval README. Original POSTMORTEM body preserved below the correction.
BREAKING CHANGE: Opus 4.6 and older are no longer supported.
Verbatim-risk guard (skip imaging unique IDs/hashes/exact values) still pending.
Pixelpipe now always uses the full 1568px canvas (cols=313) and packs up
to 50,000 chars per image. The multi-column code path and the
shrink-to-content path are no longer used by transform.ts - both were
sacrificing token savings on dense content to make sparse content marginally
more readable, and the readability gain wasn't real (the renderer's
output was unreadable in either layout once the content was actually dense).
Key changes:
- READABLE_CHARS_PER_IMAGE: 6_000 -> 50_000
- DEFAULT_COLS: 100 -> 313 (full 1568px / 5px cell)
- shrinkColsToContent() is now a no-op (returns cols unchanged)
- MinToolResultChars / MinReminderChars default to 50k (was 6k)
- transform.ts always renders single-column at full canvas
Practical impact: a 6 KB tool_result that previously rendered as 4 narrow
508x488 pages (1324 image tokens, 86% of text cost) now renders as 1
full-canvas 1568x488 page (~331 image tokens, 22% of text cost) - 4x
more savings on the same content with no quality change.
Rebuilt dist/. Tests: 315/315 pass.
After the new atlas geometry (CELL_W=5, CELL_H=8, LINES_PER_IMAGE=195,
MaxCharsPerImage(100)=19500) and the WIP shrinkWidth=true default in
transformRequest, a single 100-col image now holds ~19.5k chars of dense
ASCII at cpt=4 (vs ~15.6k previously). Combined with the gate hoisting
the width-shrink before prediction, the existing test fixtures (which
mostly used 'a'.repeat(N) with no newlines) now compress profitably at
every length above MIN_COMPRESS_CHARS=2000.
Per user direction ("use the actual gate function to compute new
expected values", "prefer rewrite over delete unless premise is
genuinely gone"):
- render.test.ts: 19 isCompressionProfitable() expectations rewritten.
Numeric-literal callers (which used a removed back-compat overload)
converted to 'a'.repeat(N) strings. Per-test comments updated to
describe the new break-even framing under MaxCPI(100)=19500. Real-shape
regression fixtures (PRODUCTION_SLAB_135H_DENSE, 169H_HEAVY) now
ACCEPT under the larger atlas; comments note the conservative-side
shift no longer rejects these dense shapes.
- history.test.ts: 5 expectations rewritten. The amortized-horizon and
prior-warm-tokens gates no longer reject at the historical fixture
sizes because image side cost has dropped well below text side cost
even at cpt=4.5. Reasons updated from 'not_profitable' to 'collapsed'
where the gate now collapses cleanly; amort() and cold() calls flipped
from .toBe(false) to .toBe(true) with rationale.
src/ untouched. Final count: 315 passed, 7 skipped, 0 failed (was 24 failing).
npx vitest run: 100% green.
Two fixes to the image-cost prediction that gates compression:
1. **Content-aware image cost.** The old gate used a fixed
`TOKENS_PER_IMAGE_SINGLE_COL = 2500` constant per image, modelling
every tool_result block as a full 1568px canvas. The renderer
actually produces images sized to the wrapped line count, so a
5,000-char tool_result becomes a ~80px-tall image costing ~54 tokens,
not 2,500. Result: the break-even for small blocks was over-estimated
by ~50x and the gate rejected anything below ~10k chars even though
image-form was cheaper.
Replace `estImages × effectiveTokensPerImage(numCols)` with a single
content-aware `imageTokensCost(text, cols, numCols, imageCountCap)`
that mirrors the renderer's height math:
`tokens = (rows + 1) × overhead + fullImageRows × perRow`, derived
from the cold-miss anchor and applied row-proportionally for partial
final images. Calibrated against N=33 production cold-miss
measurements (median multiplier 1.05, mean 1.32 — high-image case is
~1.05x of the documented `width × height / 750` rate).
2. **Width-shrinking on per-block paths.** For per-block call sites
(tool_result, reminder) the renderer now narrows the canvas width to
the longest wrapped line via `shrinkColsToContent`. Multi-col
layouts already pack tightly; this captures the single-col savings
that were previously paying for whitespace. The gate threads
`shrinkWidth` through to `evalCompressionProfitability` /
`isCompressionProfitable` / `isCompressionProfitableAmortized` so
the prediction matches what the renderer will actually produce.
System-slab and history-collapse paths pass `shrinkWidth=false` —
they fill the configured canvas deliberately.
Knobs removed: `MinToolResultChars` and `MinReminderChars` now
default to 0 (let the math decide). The TransformOptions fields are kept
for env-override observability but no longer set a floor in production.
`isCompressionProfitable` / `evalCompressionProfitability` /
`isCompressionProfitableAmortized` are now string-only at the public
API; the numeric-length fallback is gone (it under-counted images and
silently leaked net-losing compressions through).
Tests: gate-flap and history.test updated for the new physics, but 24
tests in render.test.ts still assert the OLD over-estimation as truth
(e.g. "5,000 chars is not profitable", "7000 chars stays as text").
They need to be deleted or have their expected verdicts inverted — left
for the next session to finish the rewrite cleanly. Typecheck passes.
Measured impact: the production tail's $1.43 flap is gone and per-block
tool_result/reminder content is now genuinely compressed when it would
save tokens, not just when it crosses the (wrong) 10k-char threshold.
Production data (ocproxy Opus 4.7, 2026-05-22 23:15–23:44) showed the
slab break-even gate ping-ponging between text and image modes within a
single session, paying cache_create on whichever side it just flipped
away from. Two flips in 30 minutes cost ~$0.71 + ~$0.40 of avoidable
cache_create on a single conversation.
Root cause: priorWarmTokens was asymmetric. It penalised compressing
when the un-rewritten text path was warm, but did not penalise
decompressing when the rewritten image path was warm. Per-turn cost
favored flipping; the flip forced cache_create on the new side; next
turn flipped back.
This change adds the symmetric counterpart:
text→image flip: warm text invalidated → burn applied to IMAGE side
(existing behavior, priorWarmTokens)
image→text flip: warm image invalidated → burn applied to TEXT side
(NEW, priorWarmImageTokens)
Once a session commits to a mode, staying in that mode is cheaper than
flipping by the burn cost. Mode changes only happen when the per-turn
delta exceeds the burn, which is exactly when the flip is actually
worth the cache_create. Foundational fix — no thresholds.
Observability:
- New exported `evalCompressionProfitability()` decomposes the slab
gate's evaluation into imageTokens, textTokens, burnImageSide,
burnTextSide, and verdict. Same constants and helper as the gate
itself, no drift risk.
- New `TransformInfo.gateEval` field stamps that decomposition into
every applied or not_profitable event so hosts can persist the verdict
margin and prove flap-prevention efficacy from telemetry.
Back-compat:
- New params default to 0; single-arg callers behave identically.
- 315 existing tests unchanged; +7 new tests pin the new semantics
(back-compat, asymmetric ⇒ symmetric, both-sides-warm cancellation,
eval verdict matches gate verdict, burn term math).
Pack text into a continuous sentinel-delimited stream (↵ = U+21B5) so
wrapLines fills every row to `cols` instead of leaving dead right-margin.
Adds reflow/dereflow, renderTextToPngsReflow{,MultiCol} variants, a full
test suite, and an A/B eval harness with L1 OCR + L2 session results.
Snap the collapse cutoff to a `collapseChunk`-sized grid (default 50) so
the rendered history image stays byte-identical for a full window of
turns. That lets Anthropic's prompt cache `cache_read` the history
prefix (0.1x) instead of re-billing `cache_create` (1.25x) every turn.
The cutoff floors at `minCollapsePrefix` (clamped to the raw cutoff) so
short conversations still collapse without the boundary drifting.
Also log `history_image_sha8` per request -- a sha8 of the concatenated
history-image base64. An unchanged hash across consecutive collapsed
events is ground-truth proof the prompt cache is being hit; a hash that
moves every turn signals the cache-key drift regression is back.
Production data (ocproxy 7d, Opus 4.7) shows pixelpipe HURTS short
conversations and helps long ones:
input.count <10 : +$0.17 (HURTS) over 14 rows
input.count 10-19: +$0.04 (HURTS) over 22 rows
input.count 20-39: -$0.56 (helps) over 29 rows
input.count 40+ : -$3.21 (helps) over 237 rows
Diagnosis: warm-cache share split.
bucket actual_warm_share baseline_warm_share
<10 60% 100%
10-19 84% 97%
20-39 88% 96%
40+ 92% 58%
For short conversations the un-rewritten path is already nearly fully
warm-cached. Rewriting the prefix bytes (slab compression / image
substitution / history collapse) gives the new prefix a different
cache key, so Anthropic charges cache_create (1.25×) on the new
prefix's first turn — destroying the prior warm cache. The
conversation often ends before the new cache amortizes.
Fix: add a 'priorWarmTokens' option to TransformOptions. When >0, the
break-even gate penalises the image side by
burn = priorWarmTokens × (CACHE_CREATE_RATE − CACHE_READ_RATE)
Per-turn gate eats the full burn; amortized gate spreads textLifetime
across N turns so longer horizons re-accept once savings outweigh burn.
Default 0 preserves byte-identical behavior for callers that don't
plumb it through (cold-start safe).
Tests: 3 new (NaN clamp, per-turn flip, horizon flip). 302 total pass.
count_tokens validates structural pairing and rejects truncated prefix
bodies that contain tool_use blocks whose tool_result was in the
dropped tail, with: 'messages.<N>: tool_use ids were found without
tool_result blocks'. That's been firing on ~15% of Opus 4.7 rows
(75/503 last 7d) and excluding them from baseline_probe_status='ok'.
Append a synthetic user message with minimal tool_result blocks for
any orphan ids. Adds ~5-10 tokens per orphan (within ~1% of truth)
and converts those rows from 'partial' to 'ok' so they contribute
to honest savings rollups.
Source: ocproxy daemon.log under OCPROXY_PIXELPIPE_DEBUG=1.
N=391 Opus 4.7 measured rows (last 7d) show real chars-per-token = 1.91
(avg_outgoing_text_chars=231,925 / avg_real_input_tokens=115,893). The
older 2.5 constant was calibrated on Opus 4.5/4.6 and over-estimated cpt
by ~30%, under-counting text-token cost and causing the break-even gate
to reject most history-collapse opportunities (283/391 = 72% of measured
ok rows came back history_reason='not_profitable').
Resize one render.test.ts fixture (6060 → 4848 chars) so the 'low cpt
unlocks a previously-rejected slab' assertion still exercises the
intended before/after edge under the new default.
Adds isCompressionProfitableAmortized() and a new
historyAmortizationHorizon option (default 1 = current per-turn
behaviour). When the host opts in to horizon ≥ 2, the history-collapse
gate evaluates expected lifetime cost over N future turns instead of
the cold per-turn comparison:
I × (CC + CR×(N-1)) < T × CR × N
where CC = 1.25 (cache_create), CR = 0.10 (cache_read)
At N=5 this accepts collapses with I/T < 0.30; at N=10 with I/T < 0.47.
The worst-case-for-image / best-case-for-text framing is on purpose —
we only collapse when the math wins under pessimistic warm-cache
assumptions, so a single-turn session can never pay a tax.
This is the 'try-then-decide-with-bounded-horizon' design from the
README's new 'unsolved part: multi-turn amortization' section. Picked
over JIT-style session state (rejected: introduces durable state),
always-collapse (rejected: dishonest per-request), and cache-bust-driven
(rejected: signal arrives one turn late).
- New option threaded through DEFAULTS, TransformOptions, PixelpipeOptions
- Both history call sites (early-exit helper + main success path) use the
amortized gate when horizon > 1, fall back to the cold gate otherwise
- 4 new unit tests covering fallback, accept-on-long-history, reject-on-
tiny-text, and the subset-relationship between cold and amortized gates
- 298 tests pass
Adds the data collection half of the adaptive chars-per-token plan
without touching the gate. Every isCompressionProfitable call site
now credits its chars to a named bucket on TransformInfo, and those
counts flow through toTrackEvent into the persisted JSONL.
transform.ts:
- New BucketName union: 'static_slab' | 'reminder' |
'tool_result_{structured,log,prose}' | 'history'
- bumpBucket() helper lazily allocates info.bucketChars
- toolResultBucket() routes by classifyContent shape so the post-
compaction text we credit matches the shape the gate evaluated
- Bump sites: static slab (combinedRaw), reminder (reminderRaw),
tool_result single-string (innerRaw), tool_result multi-block
per TextBlock (innerTextRaw), history (collapsedChars)
- Surfaces historyTextChars separately so history's no-collapse
rejections still contribute a denominator
tracker.ts:
- TrackEvent gains bucket_chars?: Partial<Record<BucketName, number>>
and history_text_chars?: number
- toTrackEvent copies them when populated, drops empty bucket maps
so happy-path events stay small
tests/tracker.test.ts:
- Asserts snake_case field names and nested shape land correctly
- Asserts absent buckets and pass-through events stay clean
No gate logic changed. CHARS_PER_TOKEN / SLAB_CHARS_PER_TOKEN /
HISTORY_CHARS_PER_TOKEN still drive isCompressionProfitable. Phase 2
will fit per-bucket slopes from the captured events and wire a
bucket-aware override.
docs/ADAPTIVE_CPT_PLAN.md: full plan + empirical basis (slab-only
regression on 87 events recovered cpt≈1.50 vs the baked 2.5
constant — 40% on non-slab compressions ship today).
Verified: tsc --noEmit clean; 290/290 tests pass (+2 new);
build emits dist/node.js.
Adds a third tee branch in createProxy that walks the response body and
counts Unicode code units across text_delta / thinking_delta /
input_json_delta / tool_use blocks, plus a redacted_thinking block count.
Same shape on streaming (SSE) and non-stream (content[]) responses.
Numbers are independent of Anthropic's usage.output_tokens - they give a
real ruler against the redacted_thinking-inflated bill that surfaced in
the May-2026 weekly-meter audit (#22).
Wiring:
OutputMeasurement type in proxy.ts (textChars, thinkingChars,
toolUseChars, redactedBlockCount)
teeForUsage now returns measurementPromise alongside usagePromise +
errorBodyPromise; Promise.all in handle() awaits all three
ProxyEvent gains optional measurement: OutputMeasurement
TrackEvent gains text_chars_measured / thinking_chars_measured /
tool_use_chars_measured / redacted_block_count_measured (optional,
non-breaking JSONL shape)
toTrackEvent() copies them through
Tests (6 new in proxy-usage.test.ts):
text_delta count across multi-event SSE
message_delta output_tokens overrides message_start placeholder
thinking_delta + redacted_thinking block count
tool_use input_json_delta reassembly
non-stream content[] walks same shape
5xx no-body path leaves measurement undefined
288 tests pass, tsc --noEmit clean, dist/node.js builds.
Dashboard surfacing lands in a follow-up so the wire-up stays reviewable.
Refs #22
Anthropic's /v1/messages usage block returns more fields than we were
capturing. The flat 4-field view (input_tokens, output_tokens,
cache_creation_input_tokens, cache_read_input_tokens) historically dropped:
- cache_creation.ephemeral_5m_input_tokens (1.25x rate price tier)
- cache_creation.ephemeral_1h_input_tokens (2x rate price tier)
- server_tool_use.web_search_requests (billed per-request, not per-token)
Extends:
- Usage interface (types.ts): nested cache_creation + server_tool_use
- TrackEvent (tracker.ts): cache_create_5m_tokens, cache_create_1h_tokens,
web_search_requests as optional fields
- toTrackEvent copy logic: passes through when defined, omits when absent
(a real 0 stays 0 - the dashboard needs to tell "missing" from "zero")
Also adds .gitignore entries for draumode MCP scratch files
(*.draumode.ts, *.excalidraw) so exploratory diagrams don't leak into
the repo.
Pure data collection. No behavior change. Sets up the dashboard fixes
(tasks #20 and #21) where the data will actually get rendered.
282 tests pass, typecheck clean, build clean.
Fragility #3 from end-of-session audit. The previous gate tests pinned
isCompressionProfitable(slab, 100, 2, 2.5) === true against 'A'.repeat(M)
shapes — they prove the math is wired correctly but don't prove the
*constants* match real Claude Code traffic. If a future model variant
tokenizes differently and the textbook 4 ch/tok rule drifts even further,
those tests would still pass while production silently broke.
Anonymized 6 production shapes from events.jsonl (2026-05-19..2026-05-20):
- PRODUCTION_SLAB_161K (multi-col, 52 chars/row, 11 ings/141 lines/2 cols)
- PRODUCTION_SLAB_135K_DENSE (neuline-heavy, 19 chars/row, 24 ings)
- PRODUCTION_SLAB_169K_HEAVY (very dense, 16 chars/row — actually rejected)
- BELOW_MIN_CHARS_TINY (142 chars, pre-filter)
- BELOW_MIN_CHARS_BORDERLINE (1123 chars, pre-filter)
- HISTORY_COLLAPSED_LONG_SESSION (collapsed path captured)
Each fixture stores its real captured decision plus the shape parameters
(origChars, numCols, approxCharsPerRow). synthesizeText() pads 'A'.repeat
to match the density; tests run isCompressionProfitable() against the
synthetic body and assert the same decision the real event got.
\#\# Notable: 135K_DENSE divergence
The real 135K neuline-heavy event was ACCEPTED (compressed=true) but the
synthetic 'A'.repeat(19) reconstruction REJECTS even at cpt=2.5. Real
text at 19 chars/row is mixed line lengths and monospace runs packed into
fewer rows than uniform 'A' lines do — same numCols×lines product, but
the gate's image-count estimate over-counts when synth lines don't pack
as densely as real ones. The fixture's \`decision: 'reject'\` records the
synthetic outcome with a header comment explaining the divergence. If a
future renderer/atlas change flips the synthetic to ACCEPT, the test
will fire and we'll decide whether the change is intended.
\#\# Export changes
SLAB_CHARS_PER_TOKEN and HISTORY_CHARS_PER_TOKEN exported from transform.ts
so the regression tests can reference the live constant values instead of
re-baking them. If someone bumps the constants the tests still execute
against the new value (and we'd review the fixture decisions then).
\#\# Tests
8 files / 280 tests pass, typecheck clean, build clean. 7 new it() blocks
cover the 5 captured shapes + the synthetic divergence note.
The override-gate previously used `o.charsPerToken !== CHARS_PER_TOKEN`
to detect "host pinning". After CHARS_PER_TOKEN=4 was replaced with
slab/history-specific 2.5 constants, a host that deliberately passes the
literal `4` (the same as DEFAULTS) was silently swapped to 2.5 - the
gate couldn't distinguish "host passed nothing" from "host passed 4".
Fix: check the raw opts (\`opts.charsPerToken !== undefined\`) instead
of comparing the merged value against a sentinel. Host-supplied `4` now
flows through as `4`, even though it matches the default.
Tests pin both directions:
- slab path (render.test.ts): explicit cpt=4 keeps the 161k slab
REJECTED at cpt=4 (would be ACCEPTED at the built-in 2.5)
- history path (history.test.ts): explicit cpt=4 keeps a 14-turn closed
conversation REJECTED at cpt=4 (would collapse at the built-in 2.5),
with companion cpt=2.5 test proving the fixture actually straddles
275 tests pass, typecheck clean, build clean.
Production telemetry (N=10 events with history_reason="not_profitable",
2026-05-18..05-20) shows body-level chars-per-token on rejected history
events: median 1.09, max 1.10. The gate at cpt=4 was estimating text as
3.7× cheaper than reality on every history-collapse decision in the
sampled traffic — same shape of bug as the slab path fixed in c36ec3f.
Why this matters more than the slab fix:
• The slab is cached (cache_create×1.25 once, cache_read×0.10 after).
Slab compression saves bytes mostly on cold misses + 2k tok/warm hit.
• History content is NOT cached. Once turns leave the cache breakpoint
budget (4 max per request), every subsequent request pays 100% for
them. A long agentic session re-bills the same closed-prefix turns
on every turn — exactly the workload Variant C was designed to
compress, and exactly the path the gate was silently rejecting.
Apply same fix as c36ec3f: HISTORY_CHARS_PER_TOKEN=2.5 (upper bound of
empirical data, identical safety property to SLAB_CHARS_PER_TOKEN since
history content shape ⊆ slab content shape in tool-heavy sessions).
Host can still override via TransformOptions.charsPerToken. The
override-gate `o.charsPerToken !== CHARS_PER_TOKEN` means unspecified
requests (DEFAULTS.charsPerToken=4) get the empirical cpt while explicit
non-default host values win — same intentional pattern as the slab path.
Test: regression fixture pinning a 14-turn × 1500 chars/turn conversation
where:
• cpt=4 path → 'not_profitable' rejection (image cost 5k > text 4k)
• cpt=2.5 path → 'collapsed' with 10 turns merged (text 6.4k > image 5k)
Counterfactual half uses cpt=4.5 to bypass the override-gate (passing
exactly 4 collides with the DEFAULTS sentinel).
273 tests pass, typecheck clean, build clean.