57 Commits

Author SHA1 Message Date
teamchong 69eb96d17e docs: README scope wording - imaging covers system prompt + tool docs, not just old history 2026-06-12 05:40:57 -04:00
teamchong 5704c84237 docs: add real-world failure FAQ; scrub usernames/home paths from eval corpus 2026-06-11 17:17:19 -04:00
teamchong 62790e7371 docs: add FAQ - end-to-end bill math vs touched-only, measurement method, compression scope 2026-06-11 16:53:18 -04:00
teamchong d990faf646 eval(verbatim): expand to n=15 dense hex recall on Fable 5 - 13/15 vs Opus 0/15, both misses single-glyph confabulations 2026-06-11 16:48:12 -04:00
teamchong 7d62a4c715 eval(swe-bench-pro): expand to 19 pairs — ON 14/19 vs OFF 15/19, agree 18/19; navidrome 3/3 replication clears the ON loss 2026-06-11 15:47:54 -04:00
teamchong 156945977e eval(swe-bench-pro): 10-pair bench - ON 6/9 vs OFF 7/9, -61% per-request, receipts committed 2026-06-11 10:43:02 -04:00
teamchong b8e1e855bc feat(eval/swe-bench-pro): resumable 10-pair runner on bench-dedicated proxies
- ON 47823 / OFF 47824 with own PXPIPE_LOG files: operator session can
  never pollute bench measurement (separation by construction)
- resume = skip existing patch files; clean stop on quota errors
- Docker grading documented as proxy-free (no model calls in containers)
- README: mention Pro pilot alongside Lite
2026-06-11 08:05:06 -04:00
teamchong 8e8cb09223 docs: SWE-bench tables lead with per-request -65%, demote confounded run totals 2026-06-10 23:49:38 -04:00
teamchong 83764a18a0 docs: surface SWE-bench parity in honest-part section, link eval dir 2026-06-10 23:31:55 -04:00
teamchong 76542b4a3c eval(swe-bench): Lite pilot 10 paired instances — 10/10 resolved both arms, ON $27.27 vs OFF $53.61 (-49%)
- run_pilot.py harness (paired generation via proxy ON 47821 / OFF 47822)
- official swebench Docker grading via colima, reports + patches committed
- README.md results table, FINDINGS.md task-completion parity update
- repo README benchmark table gains SWE-bench row
2026-06-10 23:25:24 -04:00
teamchong 5274e74bd4 docs: lead with whole-bill savings (59%, all 13.7k requests) over touched-only 72% 2026-06-10 19:08:32 -04:00
teamchong c6a0d61f00 eval(gist-recall): A/B imaged vs text history — 98/98 both arms, 0 confabulation
- 3 tiers: facts-at-depth w/ distractors (50), hard near-miss distractors
  on 45k-char sessions (30), 3x-mutated state tracking (18); 16 unanswerable
  confabulation guards. claude-fable-5, real production renderer, exact-match
  grading, no LLM grader.
- text 98/98, image 98/98, 0 wrong, 0 confabulated UNKNOWNs in either arm
- README benchmark table + FINDINGS dated update; PNGs gitignored,
  probes/results.jsonl committed as receipts
2026-06-10 18:54:03 -04:00
teamchong 1685935759 docs: remove em dashes from README 2026-06-10 07:28:20 -04:00
teamchong e57cc07c1b chore(publish): rename npm package to pxpipe-proxy (registry typo-squat filter blocked pxpipe); CLI command stays pxpipe 2026-06-10 07:06:22 -04:00
teamchong 63c1eab83e chore(publish): absolute README image URL, repo/homepage/bugs fields, Fable-scoped description+keywords, drop pnpm engine 2026-06-10 06:59:02 -04:00
teamchong 40a3eb5cdd docs: caption matches the real render measurements 2026-06-09 21:58:37 -04:00
teamchong e520749b84 docs: example render is now real transformRequest output (reflow + instruction banner) 2026-06-09 21:57:50 -04:00
teamchong 7a80bd1799 docs: fix image caption to match new dense example render 2026-06-09 21:45:46 -04:00
teamchong 898be465a3 docs: README example image now matches real dashboard density (full dense page, synthetic content) 2026-06-09 21:45:12 -04:00
teamchong 4ffb75e665 docs: drop GPT 5.5 route mention, Fable 5 only 2026-06-09 20:36:16 -04:00
teamchong c35b3ab2d4 docs: rewrite README for readability — hook first, real render example, honest-caveats section 2026-06-09 19:10:49 -04:00
teamchong 65cd8aab71 docs: hide unproven Cloudflare Workers quick start 2026-06-09 19:08:49 -04:00
teamchong 5eb7ec6156 docs: demote unmeasured OpenAI route to a footnote, fix stale test count and ocproxy mention 2026-06-09 19:08:33 -04:00
teamchong b5504d166e docs: drop redundant (npm: pxpipe) from README title 2026-06-09 19:05:56 -04:00
teamchong 99bb931014 chore: rename all pixelpipe references to pxpipe
- code, tests, docs, README, wrangler.toml, help text
- data dir ~/.pixelpipe -> ~/.pxpipe (PXPIPE_LOG env var)
- rebuilt dist; 323/323 tests
2026-06-09 19:04:09 -04:00
teamchong a1c51b3543 feat: publish as pxpipe on npm
- rename package to pxpipe (pixelpipe is an npm security placeholder), bin: pxpipe
- move @napi-rs/canvas to dependencies (runtime dep for npx consumers)
- drop only-allow pnpm preinstall (blocked npm/npx installs)
- add prepublishOnly guard (typecheck + test + build)
- exclude sourcemaps from tarball; README quick start: npx pxpipe
2026-06-09 19:01:22 -04:00
teamchong 0e3c95a079 docs: add CLI proxy quick start, fix stale constants table and test count 2026-06-09 18:33:34 -04:00
teamchong 3d0f9ef08a scope: narrow model gate to fable-5, drop opus; switch dense render to bare 5x8 cell
- gate regex: fable-5 only (opus 4.7/4.8 read tax gone on fable: 100/100 novel-arithmetic, identical image billing)
- remove dead opus-4.6 cpt fork in transform.ts
- DENSE_RENDER_STYLE: 7x10 padded -> bare 5x8 cell (~42% fewer image tokens/page, recall flat in A/B)
- dashboard: compression defaults ON
- README/FINDINGS/TRANSFORM_INFO updated with 2026-06-09 measurements
2026-06-09 18:33:34 -04:00
Steven Chong 613ca3bc71 feat(proxy): support gpt 5.5 chat completions 2026-06-05 00:35:26 -04:00
teamchong f6150ba2bd docs(readme): lead benchmark with the clean 93%, demote contaminated GSM8K to a footnote
One headline number (93% novel reading), 0/15 as the boundary, GSM8K 96%
moved to a flagged footnote so the result isn't ambiguous.
2026-05-31 18:50:31 -04:00
teamchong 05a5ac98dc docs(eval): reproducible reading-fidelity benchmark + honest README table
Does the model actually read pixelpipe's render, or coast on memorized answers?
- novel random-number problems (un-memorizable): text 100% vs image 93% (-7pp real reading tax)
- GSM8K (in training data): 97% vs 96% -- inflated ~3pp by recall
- verbatim recall from a dense render: 0/15

README 'Benchmarks' section leads with the clean 93% and reports the failure
rows too (no cherry-pick). Harness in eval/gsm8k/: gen_novel.py, render_cfg.mjs
(pixelpipe's real renderTextToPngs), bench.py -- claude -p, exact-match grading.
2026-05-31 18:48:19 -04:00
teamchong 378dd7f6c4 feat(applicability)!: scope to opus 4.7+ and enforce model gate at proxy
Widen the model gate to /^claude-opus-4-(?:[7-9]|[1-9]\d)(?:-|$)/ (Opus 4.7
and newer; previously 4.6/4.7). Wire isPixelpipeSupportedModel into the proxy
boundary in proxy.ts — the proxy previously gated on economics only and
ignored the model, so it compressed 4.8 traffic despite the docs claiming a
4.6/4.7 scope. Unsupported models now pass through with reason
'unsupported_model'.

Update public-api tests to the 4.7+ spec (320/320 pass).

Correct the docs to live measurement: reverse the POSTMORTEM "dead" verdict
(pixelpipe is a lossy gist-compressor saving ~68% on real dense Claude Code
traffic; the verbatim 0/15 needle finding stands as a caveat, not the
verdict), rewrite README Status/Limitations, and add a correction pointer to
the eval README. Original POSTMORTEM body preserved below the correction.

BREAKING CHANGE: Opus 4.6 and older are no longer supported.
Verbatim-risk guard (skip imaging unique IDs/hashes/exact values) still pending.
2026-05-29 22:34:55 -04:00
Steven Chong 2f78ecaf9f fix(transform): use conservative cpt defaults for opus 4.6 2026-05-26 15:46:19 -04:00
Steven Chong d3992c6bad feat(applicability): enable pixelpipe for opus 4.6 2026-05-26 15:35:12 -04:00
Steven Chong be545b7ccd feat(render): full-canvas single-column rendering, 50k chars/page
Pixelpipe now always uses the full 1568px canvas (cols=313) and packs up
to 50,000 chars per image. The multi-column code path and the
shrink-to-content path are no longer used by transform.ts - both were
sacrificing token savings on dense content to make sparse content marginally
more readable, and the readability gain wasn't real (the renderer's
output was unreadable in either layout once the content was actually dense).

Key changes:
- READABLE_CHARS_PER_IMAGE: 6_000 -> 50_000
- DEFAULT_COLS: 100 -> 313 (full 1568px / 5px cell)
- shrinkColsToContent() is now a no-op (returns cols unchanged)
- MinToolResultChars / MinReminderChars default to 50k (was 6k)
- transform.ts always renders single-column at full canvas

Practical impact: a 6 KB tool_result that previously rendered as 4 narrow
508x488 pages (1324 image tokens, 86% of text cost) now renders as 1
full-canvas 1568x488 page (~331 image tokens, 22% of text cost) - 4x
more savings on the same content with no quality change.

Rebuilt dist/. Tests: 315/315 pass.
2026-05-25 14:20:43 -04:00
Steven Chong 393f72ee7c docs(readme): reframe tagline — pixels instead of a transcript, context as UI 2026-05-22 23:38:20 -04:00
teamchong a8d2f62d63 docs(cell): update stale 7×10 references to current 5×8 production cell
The reflow-inimage finding (commit bb9c231) made 5×8 viable, so
DEFAULT_CELL_W_BONUS / DEFAULT_CELL_H_BONUS reverted to 0 — but many
comments and the README still said 7×10 was production.

Touched:
- README.md             — eval table, density bullet, sweep narrative,
                          cell-pitch section, measurement header,
                          Limitations section
- src/core/render.ts    — RenderStyle docstring + cell-bonus docstrings
                          + multicol intro comment
- src/core/transform.ts — TransformOptions.reflow docstring, DEFAULTS
                          comments (minReminderChars floor rationale,
                          maxImagesPerToolResult cap, reflow gate
                          rationale, anchor scaling formula, LINES_PER_IMAGE
                          worked example)

No code behavior changed — every constant still resolves through CELL_W /
CELL_H, and the eval harness still overrides the cell via cellWBonus /
cellHBonus when comparing variants. The 5×8 production path is what was
already shipped on 81550f6 (reflow + grayscale + L1/L2 harness) and
bb9c231 (reflow-inimage).
2026-05-22 22:59:20 -04:00
teamchong 4774aeb14c feat(eval): in-image instruction variant (reflow-inimage) — +1.04pp vs baseline
Co-render the OCR instruction into the same PNG as the content,
delimited by '===…===' bands, with the API system field dropped.

L1 OCR fidelity, Opus 4.7, 20 production blocks, 7×10 cell:
- baseline (text-only):       97.91% mean / 96.25% min
- reflow (separate system):   91.99% mean / 82.59% min  (-5.93pp)
- reflow-inimage:             98.95% mean / 96.42% min  (+1.04pp)

reflow-inimage wins on all 20/20 blocks vs reflow, and beats the
text-only baseline on 17/20 blocks. Three blocks hit 100%. The
-5.93pp reflow regression that the cell-pitch sweep partially
recovered disappears entirely when the instruction is co-rendered.

Mechanism: when system carries the instruction and the image
carries the content, the model does cross-modal binding to figure
out what the image is for. Co-rendering reduces it to a
single-modal task with an unambiguous parse rule.

Files:
- eval/eval-l1-ocr.mjs:        add reflow-inimage variant + prompt
- README.md:                   new section before history compression
- eval/EXPERIMENT_LOG.md:      attempt #2 writeup
- eval/results/{l1-report.md, l1-results.json}: regenerated
2026-05-22 22:04:10 -04:00
teamchong 81550f690e feat(render): pack reflow across newlines + grayscale atlas + L1/L2 eval harness
- wrapLines packs continuously; NL_SENTINEL stays as an inline marker, never breaks a row
- atlas-gray.ts: anti-aliased grayscale variant (gated via RenderStyle.aa, unused in production)
- gen-atlas.ts: optional ATLAS_GRAY=1 emission
- eval/: standalone Anthropic-client L1 OCR fidelity + L2 session replay harness, claude -p backed
- README: reflect packed reflow + measured Opus 4.7 numbers
2026-05-22 20:48:40 -04:00
Steven Chong d87aecec95 feat(render): use 5x8 code-font atlas 2026-05-21 17:41:30 -04:00
Steven Chong 6b508c367d docs(readme): summarize VLM text-density research for renderer tuning
Capture the practical conclusions from recent VLM/OCR papers:
ReadBench (arXiv:2505.19091), typographic perturbation work
(arXiv:2604.12371 / 2604.25102), and typography-gap work
(arXiv:2603.08497).

The main guidance for pixelpipe is: don't increase DPI; margin tweaks
are tiny; the real foundational experiment is a denser-but-readable
bitmap atlas (e.g. 5x11 -> 4x8/4x7) with exact retrieval tests before
shipping.
2026-05-21 16:39:21 -04:00
Steven Chong 3d5d27d0db fix: Opus 4.7 pricing is $5/$25/MTok, not $2.50/$12.50 — revert prior "fix"
Earlier commit a0f282f "corrected" Opus 4.7 rates downward to $2.50
input / $12.50 output based on a misread of the pricing source. Per
docs.claude.com/en/docs/about-claude/pricing (current), Opus 4.7 is:

  Base input:       $5    / MTok   (same as Opus 4.5, 4.6)
  Output:           $25   / MTok
  5m cache write:   $6.25 / MTok
  Cache hits/reads: $0.50 / MTok

Same dollar pricing as Opus 4.5 / 4.6 — what changed in 4.7 is the
tokenizer, not the per-token rate. The README now flags this so future
contributors don't assume hardcoded chars-per-token / image-token
constants tuned on earlier models are still safe on 4.7; only
count_tokens against the target model is honest.

Touches:
  - src/dashboard.ts: ASSUMED_INPUT_USD_PER_MTOK back to 5.0, with a
    long comment block citing the source and warning about tokenizer
    differences. Updated card subtext on the headline accordingly.
  - README.md: pricing table back to $5/$25/$6.25/$0.50, with a
    callout that the tokenizer changed in 4.7.
2026-05-21 12:38:10 -04:00
Steven Chong 09e4f635ea docs(readme): stop claiming '$ saved'; report token deltas only
The README headlined a '~76% fewer tokens' framing, a worked example
with a 'savings' column, and an Opus pricing table that used the old
Opus-4-base rates ($5/$25/MTok input/output) instead of Opus 4.7's
actual rates ($2.50/$12.50). Combined, it read like an aggregate
$-savings claim that the data does not yet support.

What's actually true and stays:
  - Pixelpipe ships fewer input tokens on cold-miss requests, measured
    by Anthropic's own count_tokens endpoint. The 173,783 -> 41,321
    anecdote is one real request, not a session average.
  - Image-tokens-vs-text is a real per-request encoding-density
    observation.

What was misleading and is removed/softened:
  - Opening tagline no longer claims '~76% fewer tokens, 100% quality'
    as if it generalizes. Replaced with an experimental-status note.
  - Worked-example 'savings' column renamed 'delta' and contextualized.
  - 'How we report numbers' section added that explains the
    baseline_probe_status gate, says we don't currently make an
    aggregate $ saved claim, and warns the reader to treat any host
    dashboard that surfaces $ saved without the gate as marketing.
  - Pricing table corrected to Opus 4.7 rates.
2026-05-21 12:26:26 -04:00
Steven Chong 3d22a04356 docs(readme): document multi-turn amortization problem and design space
The break-even gate in collapseHistory is per-turn — it answers 'is the
image cheaper than the text on this single request?' That question has
the wrong answer once Anthropic has cached the prior text-based prefix,
because text-at-10% beats image-at-40%-cold. But the prefix's lifetime
cost is what matters, not the per-turn cost, and lifetime cost favours
the image once the cache eventually expires.

This is the same shape of decision a JIT compiler, a DB optimiser, or
ZFS block compression makes. Documenting the analogues, the four
credible designs we considered (try-then-decide / session-state /
always-collapse / cache-bust-driven), and why we picked try-then-decide
as v1.

The Why-it's-hard section gets a new paragraph framing the asymmetric
cache invalidation. The history-compression section gets a new
'The unsolved part: multi-turn amortization' subsection with the full
design space and the decision criteria.

No code changes — this is design documentation for future contributors
so they don't reinvent the analysis.
2026-05-21 10:13:00 -04:00
Steven Chong 0aa2cce88f Expose pixelpipe library API 2026-05-20 23:47:32 -04:00
teamchong 51c8f10012 docs(README): drop duplicate blockquote, change headline to "pixels instead of text"
- Headline: "pixels instead of tokens" → "pixels instead of text" (images
  bill as tokens too — the honest framing is pixels vs text)
- Removed the redundant "Why this works in 2026" blockquote since the
  Opus 4.7 explanation already lives in the intro paragraph
2026-05-20 13:38:49 -04:00
teamchong 5325effb65 docs(README): reframe as context-as-UI, not cost arbitrage
The intro now leads with the hypothesis (pixels pack more semantic info
per unit of context than serialized text) and treats the token reduction
as measurable evidence rather than the headline.

Adds a "What this is NOT" section to address the ToS / billing-loophole
read directly:
- Not ToS evasion (uses the documented vision API as documented)
- Not a billing loophole (savings hold even if Anthropic re-prices)
- Not cost arbitrage (cost is a side effect of density)

Same numbers, same code, different stance: defensive about the framing
("I'm not paying for churn I didn't cause") rather than offensive about
the pricing.
2026-05-20 13:33:47 -04:00
teamchong 3c759ed662 docs(README): cite Opus 4.7 vision upgrade as what makes this possible
Pre-4.7 vision couldn't reliably OCR dense monospace glyphs - errors
would corrupt the prompt before the model read it. Opus 4.7 bumped
the long-edge image cap 1568->2576px (3.3x pixels) and document-OCR
benchmarks (DocVQA 87->94%, ChartQA 80->88%) per Anthropic's release
notes.

Pixelpipe still renders at 1568x1568 - this is a model fidelity
change, not a renderer change. But it's why this approach works now
when it didn't a year ago.

Cites anthropic.com/news/claude-opus-4-7.
2026-05-20 10:10:11 -04:00
teamchong 3173e8ad92 docs(README): drop Opus 4 column from rate table
Proxy targets Opus 4.7. Input pricing has been flat across Opus 4.5/4.6/4.7
($5/$25 per MTok), so the historical Opus 4 column ($15/$75) was framing
noise unrelated to actual cost.
2026-05-20 09:44:27 -04:00
teamchong 699ce3023f docs(README): explain history compression (Variant C)
~80-line section covering the 4-breakpoint cache cliff that history
collapse addresses, the closed-prefix walk, the 4 honesty gates, and
today's HISTORY_CHARS_PER_TOKEN=2.5 fix wired in b563751.

Links to the live data point in events.jsonl (12:30:01 event,
collapsed_turns=175, collapsed_chars=180,684) and the relevant
transform.ts / core/history.ts call sites.
2026-05-20 09:40:38 -04:00