docs(readme): lead benchmark with the clean 93%, demote contaminated GSM8K to a footnote

One headline number (93% novel reading), 0/15 as the boundary, GSM8K 96%
moved to a flagged footnote so the result isn't ambiguous.
This commit is contained in:
teamchong
2026-05-31 18:50:31 -04:00
parent 05a5ac98dc
commit f6150ba2bd
+16 -28
View File
@@ -41,42 +41,30 @@ families are not enabled.
## Benchmarks (reproducible)
Does imaging preserve answers? We measured it three ways on `claude-opus-4-8`,
and — unlike a marketing table — we report where it breaks.
**One number: on short, readable content the model reads pixelpipe's render
~93% of the time, at ~38% fewer tokens.** Measured clean — with novel
random-number problems it cannot have memorized, on `claude-opus-4-8`:
**1. Reading fidelity, clean — novel random-number problems (zero memorization possible):**
| test | N | baseline (text) | pixelpipe (image) | delta |
|---|---:|---:|---:|---:|
| novel arithmetic, random numbers | 100 | 100% | **93%** | **7pp** |
This is the honest reading number. With the answer un-memorizable, the model
reads pixelpipe's render correctly ~93% of the time on short content; the other
~7% are misreads (`10200``9400`, `7873``7793`) or come back unreadable.
**2. GSM8K — the standard suite, but it's in training data:**
| benchmark | N | baseline (text) | pixelpipe (image) | tokens/problem |
| test | N | text | pixelpipe (image) | tokens |
|---|---:|---:|---:|---|
| GSM8K (math) | 100 | 97% | 96% | 61 → 38 (**38%**) |
| novel arithmetic (un-memorizable) | 100 | 100% | **93%** | **38%** |
GSM8K *looks* near-lossless (1pp), but that's inflated — the model recognizes
memorized problems and recalls the answer even when it misreads the image. The
novel test strips that out: the real reading cost is ~7pp, not ~1.
The ~7% gap is real misreads (`10200``9400`, `7873``7793`), not noise — imaging
is lossy even where it mostly works.
**3. Where it fails completely — dense / verbatim:**
**The boundary** — push to dense, exact-recall content and it collapses:
| test | baseline (text) | pixelpipe (image) |
| test | text | pixelpipe (image) |
|---|---:|---:|
| verbatim recall — exact 12-char hex from a *dense* render | 15/15 (100%) | **0/15 (0%)** |
| verbatim recall — 12-char hex from a *dense* render | 15/15 | **0/15** |
So pixelpipe reads **short, readable** content at ~93% (a real ~7% misread tax)
and **silently fails** on dense walls and exact recall. The number you get depends
entirely on what you feed it.
~93% on short / readable, **0%** on dense / exact-recall. There is no free lunch —
full analysis in [`FINDINGS.md`](FINDINGS.md).
Reproduce: [`eval/gsm8k/`](eval/gsm8k/) (GSM8K + the novel reading test) and
[`eval/needle-haystack/`](eval/needle-haystack/) (the verbatim failure). Full
analysis: [`FINDINGS.md`](FINDINGS.md).
<sub>We also ran the standard **GSM8K** suite: 96% imaged. But GSM8K is in training
data, so the model recalls memorized answers through its own misreads — inflating
the score ~3pp over the clean novel number above, which is why we don't lead with
it. Reproduce: [`eval/gsm8k/`](eval/gsm8k/) · [`eval/needle-haystack/`](eval/needle-haystack/).</sub>
---