The text-based Tool Reference triggered Anthropic's model-cloning
safeguard (stop_reason: refusal -> silent Opus fallback) because the
per-tool stub said docs lived in "the system prompt" — the exact
phrasing the 169521c retrip identified as the trigger for the imaged
slab, now reintroduced in text form.
- Stub now reads: 'ⓘ Full docs: see "## Tool: <name>" in the Tool
Reference section.' — no "system prompt"/"authoritative" framing.
- Tool Reference header is provenance-framed as first-party ("pxpipe
(this user's local proxy) moved the full tool documentation...").
- tracker.ts now maps tool_docs_chars into events.jsonl so refusals
are attributable to the text-reference path.
- Regression test: stub + reference header must never match
/system prompt|authoritative/i and must carry provenance framing
(scoped to the header; quoted third-party tool docs below it may
legitimately contain the phrase).
Verified: 614/614 tests, tsc --noEmit clean, and a live cold-start
through the rebuilt proxy on claude-fable-5 with the text reference
active — no refusal event, no fallback.
The in-image banner announced 'SYSTEM PROMPT + TOOL DOCS ... treat as
authoritative system instructions' inside a user turn. Anthropic's
model-cloning classifier read that as a replayed/extracted prompt and
refused (apiRefusalCategory: reasoning_extraction), forcing Claude Code
to fall back to claude-opus-4-8 — outside compress scope, so the whole
session ran passthrough with zero savings.
Reworded to first-party provenance framing (pxpipe rendered this
session's own configuration; no 'system prompt'/'authoritative').
Two headless probes through :47824 now serve on claude-fable-5[1m]
with the imaged slab (img:4, ~94k chars) and zero refusal events;
identical shape was refused pre-fix (req_011CccRnutLzRJNVZ7bRKG8Q).
Anthropic's safety classifier OCRs rendered images. A *legible* image of the
system prompt + tool docs (bash / permissions / credential / file-write
vocabulary) trips its prompt-injection heuristic and forces a model downgrade
to Opus, which costs far more than imaging the (already prompt-cached) slab
ever saves. Empirically the caption wording is NOT the trigger — the OCR'd
content is — so no caption rewrite fixes it; the slab must stay text.
Gate at the slab profitability check with PXPIPE_TEXT_SLAB=1. Default is
unchanged (legacy imaged slab) so the cache-align / cache-stability suites
stay valid. Both demos now launch the pxpipe proxy with the toggle set.
Verified: identical ~300KB-slab request images 8 PNGs by default, 0 with the
toggle (reason=slab_text). Full suite: 614 passed.
Both demos' a.sh/b.sh now request claude-fable-5[1m] for the default `fable`
arm, consistent with the opus/sonnet arms which already use [1m].
setup.sh PXPIPE_MODELS stays the BASE claude-fable-5 on purpose: the proxy
strips [variant] tags from the incoming model before matching against the
allow-list verbatim (src/core/applicability.ts), so base claude-fable-5
already covers claude-fable-5[1m]. Writing [1m] into the allow-list would
make the stripped incoming base no longer match -> pxpipe silently stops
compressing. Added a comment in each setup.sh pinning that.
Measured 2026-07-01 (count_tokens sweep, claude-sonnet-4-5; see
docs/LEGIBILITY-AUDIT-2026-07-01.md): the API downscales every image to fit
BOTH long-edge ≤1568 AND ~1.15 MP (≈1,143,750 px), then bills ≈px/750
(~1525 tok/img cap). The old 1932×1932 page hit the cap but was resampled
0.555× before the vision encoder, so 5×8 glyphs arrived at ~2.8×4.4 px — the
root cause of the legibility misses this session.
New page shape 1568×728 = 1,141,504 px fits both bounds → glyphs reach the
encoder WYSIWYG (and still ≤2000 px/side for >20-image requests).
render.ts geometry:
MAX_HEIGHT_PX 1932 -> 728
MAX_WIDTH_PX 1932 -> 1568
DEFAULT_COLS 313 -> 312 (1568 px exactly; 313=1573 would 0.997× blur)
DENSE_CONTENT_COLS 384 -> 312
READABLE/DENSE chars 50000/92160 -> 28080 (312×90, tracks real pagination)
Decouple the GPT path: OpenAI resizes differently (2048 bbox, 768 short side),
so a 768-wide strip up to 2048 tall survives un-resampled. Add
GPT_MAX_HEIGHT_PX = 1932 (gpt-model-profiles.ts) where the built-in cost
numbers were calibrated; openai-history.ts now uses it instead of the
now-Anthropic-specific MAX_HEIGHT_PX. Doc-only touch-ups in library.ts/export.ts.
tests/render + tests/paging updated to the new geometry (171 pass).
Holds the glyph cell at prod 5x8 and varies ONLY render style
(prod/onebit/color/grid/cgrid) at identical pixel dims => identical
image-token cost. Answers what the size sweep could not: at fixed cost,
does any style beat the ~10% exact-read baseline? Same content, grader,
and claude-CLI read method as eval/glyph-matrix/sweep.
The function-form transform intentionally returns {} on the active
path; the gate runs on static DEFAULTS (charsPerToken=4, priorWarm*=0)
and there is no live-alpha feedback loop from the dashboard. Record the
telemetry that justified this (2026-06: 897 sessions / 21,347 measured
rows, 5 mode flips ever, losses 0.8% of wins, all cache-create
amortization) so nobody "fixes" it without re-running that
reconciliation.
Negative Saved/lost cells get a small "create" pill (with explanatory
tooltip) when repricing the newly written prefix at the read rate turns
the loss positive: saved + cache_create x (1.25 - 0.1) > 0. Those rows
are the one-time cache-write premium being amortized, not gate
failures. Rates/token counts interpolate from CACHE_CREATE_RATE /
CACHE_READ_RATE so the copy cannot drift from the math.
Verified against ~/.pxpipe/events.jsonl replay: 3/3 negative cells in
the live 50-row window marked; warm-turn losses (cc=0) correctly bare.
- extractFactSheetEntries() returns {token, count}; the kept-token set
and order are byte-identical to the old extractFactSheetTokens(),
which now delegates. Counts render as a ×N suffix, so tally questions
over imaged content are answerable from the text sheet instead of by
counting 5x8 px glyph rows.
- Same-token spans matched by overlapping patterns dedupe by offset
(token + match index) so nothing double-counts.
- New tier-0 shape: uppercase hyphenated codes with a digit
(PROJ-1482, CVE-2024-30078); digit lookahead is bounded.
- export.ts consumes entries across all pages; 14 new/extended tests.
a.sh/b.sh -> claude-sonnet-5[1m] (was the stale claude-sonnet-4-6),
matching the opus alias's [1m] pattern. setup.sh -> claude-sonnet-5
(bare id, feeds PXPIPE_MODELS). claude-sonnet-5 was already in the dashboard
model catalog (fragments.ts); old 4.6 still reachable via the verbatim-id
fallback. No src/ changes.
A fresh same-session prior within the TTL is no longer sufficient to price
the text counterfactual warm: deriveBaselineWarmth now also requires the
static-prefix hash (system_sha8) to match. When opencode rotates the system
prompt / tool docs mid-session, the cacheable prefix changes, so a text-only
client would hit a new provider cache key too — pricing it warm against a
cold actual fabricated a huge phantom "loss" (the dashboard's 800%-worse
report). cr>0 still rescues a genuine warm read with no in-memory prior.
Wired through update(), replay(), and aggregateSessions via system_sha8.
Also surface losses honestly in the recent table (Saved/lost: negative
deltas in red instead of hidden as "—") and clarify headings as
billing-equivalent input tokens. Docs updated to match.
Drop the duplicate CHARS_PER_TOKEN=3.7 from export.ts and import REPORT_CHARS_PER_TOKEN from transform.ts so the reporting-estimate constant has one source of truth. Clarify the dashboard baseline-warmth invariant comment. No behavior change.
The savings accounting priced the text baseline as cold whenever pxpipe's own
image cache missed this turn (cr==0), fabricating a 1.25x create the text path
would never have paid and inflating reported savings ~2x. Decouple them: the
text prefix reads warm whenever a fresh same-session prior exists within the
300s TTL (wall-clock), unioned with an observed read (cr>0) for the post-restart
case. Centralised in deriveBaselineWarmth; used by all three call sites
(sessions, dashboard, fragments).
Empirical replay over ~/.pxpipe/events.jsonl (13,402 message rows): honest
headline 504.6M tokens saved vs the old cr-gated 1.05B; 1,409/12,305 fresh-prior
rows (11.4%) read warm despite a busted image cache. Dashboard now narrates the
busted-image case explicitly. Adds baseline/sessions/context-map regression
tests. Measurement only — no change to proxy behavior.
typecheck + 600 tests + build all green.
The magic-word @pxpipe directive fired unreliably and added per-request body
scanning to the hot path, so drop it entirely: removed the directive module +
its test and unwired it from transform.ts / proxy.ts / node.ts / index.ts.
transformRequest now leaves message bodies byte-identical — no per-request
directive parsing.
In its place, `pxpipe export <paths>` renders code/text to dense PNG pages via
the same SDK primitive the proxy uses, so exported pages are byte-identical to
what ships to the model. Width is locked to proxy density (--cols rejected as
an unknown option, guarded by tests) to prevent drift. Added --open to preview
the rendered pages plus a result footer (artifact summary + token-savings line).
Tests: export.test.ts 64 passed (incl. --open parse + default coverage); full
suite 589 passed / 29 files; typecheck clean.
Clears the Node 20 runtime deprecation warning on the runner.
package.json already sets packageManager: pnpm@10.21.0, so action-setup@v6
resolves the pnpm version with no action input - same as the prior bare usage.
Intermittent prefix-cache busts on long sessions (#11): the single cache
breakpoint relocated onto the LAST history image — the still-growing chunk
whose bytes change every window advance — so the slab+history prefix
re-created at 1.25x instead of being read.
- relocateAnchorToHistoryImage pins collapseHistory's carry-over ordinal
(the last byte-stable chunk) instead of the last image.
- demoteProtectedHeadText keeps pxpipe's slab scaffolding (images +
fact-sheet + the [End of rendered context.] boundary) verbatim and only
quarantines the user's stale opening turn, so the relocator can still
find the slab anchor (#14 demotion had clobbered the boundary it keys on).
- cachePrefixDigest joins parts with \x00 so block-boundary differences
can't hash-collide into false prefix-stability.
520/520 green, tsc clean.
Read-only digest of the exact pinned prefix (tools+system+messages through
the imaged history/slab boundary, live tail excluded), computed after final
marker placement and emitted per-turn via tracker. Attributes a prompt-cache
bust as pxpipe-side (sha8 drifts turn-over-turn) vs upstream eviction (sha8
stable while cache_create spikes). Behavior-neutral: byte-identical requests,
518/518 green.
Top-right toggle flips between the warm-light theme and a new dark palette (same flame identity, inverted neutrals). Pure client-side: a <head> script sets the theme before first paint (no flash) from localStorage, defaulting to prefers-color-scheme; no new route or server round-trip. Themed via the existing :root CSS variables plus dark fix-ups for the few hard-coded light spots.
The 64-token budget was filled longest-first, so long low-risk doc-URLs evicted short zero-redundancy tokens (git SHAs, ports, CLI flags, CONST_IDS) — the strings a coding agent must reproduce verbatim. Allocate by priority tier (ids/flags/numbers > paths/versions > URLs, capped); preserve substring-collapse and byte-stable, cache-safe output. Adds an eviction regression test.
writeWebResponse now respects write() backpressure via a drain/close race,
cancels the upstream body reader when the client hangs up, and treats socket
aborts (UND_ERR_SOCKET/ECONNRESET/EPIPE/AbortError, "terminated",
"other side closed") as clean disconnects instead of throwing. Releases the
reader lock and removes the close listener in finally.
Recovers precision-critical strings (paths/ids/versions) from the imaged
slab as cache-stable text so they survive OCR. Byte-identical across turns,
validated on both Claude and gpt-5.5 paths.
Adds the 2026-06-23 entry: pxpipe's right comparison is /compact (itself lossy),
not perfect recall. compact-vs-gist analysis (same job, opposite loss profiles —
pxpipe wins on Fable, peer-on-Opus with abstention, recoverability as the
structural edge); code-grounded notes on the open-source harnesses (codex,
opencode); ties to the agent-memory (reducer/channel-routing) direction. n=1
behavioral parts flagged; eval numbers cross-checked against existing entries.
Docs only — not shipped in the npm package.
Anthropic-path recency anchor for collapsed history (commit 9eb66f5): banner now
warns against reopening low-N turns, plus a bounded ~300-char text pointer to the
most-recent collapsed user turn after the images. Patch — a guardrail fix on the
visible path; cache-safe (pointer sits after the cache_control breakpoint).
Adversarially reviewed: 8 findings, all low/nit — 6 are positive confirmations
(cache-safe, matcher in sync via shared constant, relocation robust to the trailing
block, GPT path unaffected, no off-by-one); 2 benign (tool-result-only user falls
back to an earlier user turn with a correct t-index; mid-word preview truncation).
474 tests pass; tsc + build clean.
GPT autonomous agents (OpenCode/gpt-5.5) no longer lose the live request to
history imaging — the most-recent user turn is pinned as legible text between
before/after history images, with the guard echoing it. Fixes snap/confabulation
/off-task drift; validated on gpt-5.5 (stays on task, no off-task edits).
Patch, not minor: GPT is opt-in/WIP and hidden in the dashboard, and the visible
Anthropic path is byte-identical (pin commit 08a2b0a touched only the GPT path).
Also: clarifying comment on the minCollapseTokens gate (counts imageable work
only; sub-floor histories stay fully text by design — request stays legible).
473 tests pass; tsc + build clean.
Autonomous GPT agents (OpenCode/gpt-5.5) send ONE human request then a long
run of tool turns. The lone request is the OLDEST turn, so history collapse
imaged it first and the model lost it — "I wonder what the user actually asked"
→ off-task drift (observed: agent edited compaction.ts instead of answering a
compare question). The live-request guard's "the request is the trailing user
message" heuristic assumes interactive chat and points at nothing here.
Fix: keep the most-recent user turn OVERALL as legible TEXT, spliced between
before-pin and after-pin history images inside the synthetic user message;
older user turns stay imaged (must not look live — snap-to-first-prompt guard).
The guard now echoes the request verbatim (capped). History is imaged on both
sides of the pin, so compression barely changes.
Cache safety (adversarially reviewed):
- Pin ONLY when the latest user turn is INSIDE the collapse range. If it's in
the kept tail (ordinary interactive turn) it's already native text — pinning
an older in-range turn would migrate the pin across collapse-chunk boundaries
and re-image frozen history. Restricting to the latest turn fixes its position
until the next prompt, so before/after sections stay byte-stable across a run.
- Undersized before-pin remainder merges into the previous before-section rather
than emitting a sub-threshold (net-negative) image.
Both Chat Completions and Responses paths. 473 tests pass; tsc + build clean.
Not yet released — validate on the gpt-5.5 machine first.
b9418bf repriced cr>0-but-no-prior turns warm in dashboard.ts (update + replay)
but left the identical gate in sessions.ts unfixed — so the Sessions panel still
fabricated inflated savings after a restart/eviction or on the first tracked turn
(a cr>0 turn priced cold bills the text counterfactual a 1.25x create on a prefix
we KNOW was cached). Mirror the dashboard fix so "honest cache math" holds on every
panel.
- sessions.ts: warm follows the observed read (cr>0) alone; with no fresh prior,
assume the cacheable prefix was fully reused (prevCacheable = cacheable) rather
than mispricing a known-warm read.
- sessions.test.ts: the "real prefix compression" case had cr=100 on its first
turn (observably warm) yet asserted the cold-priced 22490 fabrication; corrected
to the honest 1790 (warm) + 95 = 1885. Now a regression guard for this fix.
- dashboard.ts: drop the stale "no per-session state" replay comment — replay now
mirrors update()'s per-session warmth.
Found by adversarial review of b9418bf. 462 tests pass; tsc clean.
Collapsed-history images now carry who-said-what AND when, instead of leaving
the model to reconstruct it from a flattened, role-ambiguous blob (which led it
to resurface the opening turn as if it were the live request).
- Turn-index recency anchor: each collapsed turn serializes as <user t="N"> /
<assistant t="N"> with an ABSOLUTE turn index (larger N = more recent), so the
model can tell turn 1 from turn 60. Absolute (never "N ago"/"i of total") so
frozen chunks stay byte-identical and keep hitting the prompt cache. Mirrored on
the GPT path. (history.ts, openai-history.ts, openai.ts)
- colorByRole: role tags are tinted via a parallel "slot string" carried from
serialize time (structure-through), replacing the parse-back that miscolored a
body quoting a literal <tag>. (render.ts, history.ts)
- neutralizeSentinel: a pre-existing ↵ (U+21B5) in content is swapped to ⏎
(U+23CE) in render-prep so reflow packs newlines instead of bailing to a raw,
unpacked render — common when the content is about pxpipe itself. Render-only;
originals preserved. (render.ts, transform.ts, history.ts, openai-history.ts)
- Per-model GPT vision/render profiles, retunable via PXPIPE_GPT_PROFILES,
behavior-identical to the old hardcoded values. (gpt-model-profiles.ts)
Review fixes folded in:
- Banner intro/outro consolidated to a single source of truth; the GPT path now
aliases the Anthropic constant instead of byte-copying it (can't drift).
- slotCopyBody neutralizes literal slot markers to a width-equivalent control
char (U+0003), not a space the minifier would strip and misalign.
459 tests pass; tsc clean.
Collapse old conversation history into rendered PNG sections so the model
reads a compact image instead of re-billed text, while preserving prompt
caching and tool-call behavior. Measures real vs compressed token/cost.
Core:
- GPT history collapse (openai-history.ts): append-only, o200k token-length
sectioning. Sections seal only at a tool-closed boundary (open call-id set
empty), so a function_call and its function_call_output never split across
the collapse cut. Fixes the OpenAI 400 "No tool call found for function
call output with call_id ..." that hit long Responses-API sessions.
- Anthropic cache contract (history.ts): append-only per-chunk rendering;
cache_control markers are preserved/moved, never added; chunk boundaries
align with caller marker seams for byte-stable prefix caching.
- GPT image budget (openai.ts): detail:'original' for gpt-5.x, flagship
vision-multiplier fix, patch cap; schema-strip preserves real descriptions.
- Savings accounting (openai-savings.ts): cached_tokens + vision-token basis.
Model scope (applicability.ts):
- Default imaged scope = claude-fable-5 + gpt-5.6.
- gpt-5.5 and claude-opus-4-8 stay opt-in: same pipeline, but they degrade
reading dense imaged history (gist drift), so silently imaging them by
default is wrong. Promotion is gated on an OCR/recall threshold.
Dashboard: GPT + Anthropic rendering, per-family model toggles, persisted
metrics, thumbnail-expired session UI, reflow/newline handling.
Tests: cache-alignment (GPT + Anthropic), history sectioning + tool-boundary
invariants, savings, dashboard, sessions/restart-restore. 452 passing.