mirror of
https://github.com/teamchong/pxpipe.git
synced 2026-07-22 02:02:51 +02:00
docs(readme): lead benchmark with the clean 93%, demote contaminated GSM8K to a footnote
One headline number (93% novel reading), 0/15 as the boundary, GSM8K 96% moved to a flagged footnote so the result isn't ambiguous.
This commit is contained in:
@@ -41,42 +41,30 @@ families are not enabled.
|
||||
|
||||
## Benchmarks (reproducible)
|
||||
|
||||
Does imaging preserve answers? We measured it three ways on `claude-opus-4-8`,
|
||||
and — unlike a marketing table — we report where it breaks.
|
||||
**One number: on short, readable content the model reads pixelpipe's render
|
||||
~93% of the time, at ~38% fewer tokens.** Measured clean — with novel
|
||||
random-number problems it cannot have memorized, on `claude-opus-4-8`:
|
||||
|
||||
**1. Reading fidelity, clean — novel random-number problems (zero memorization possible):**
|
||||
|
||||
| test | N | baseline (text) | pixelpipe (image) | delta |
|
||||
|---|---:|---:|---:|---:|
|
||||
| novel arithmetic, random numbers | 100 | 100% | **93%** | **−7pp** |
|
||||
|
||||
This is the honest reading number. With the answer un-memorizable, the model
|
||||
reads pixelpipe's render correctly ~93% of the time on short content; the other
|
||||
~7% are misreads (`10200`→`9400`, `7873`→`7793`) or come back unreadable.
|
||||
|
||||
**2. GSM8K — the standard suite, but it's in training data:**
|
||||
|
||||
| benchmark | N | baseline (text) | pixelpipe (image) | tokens/problem |
|
||||
| test | N | text | pixelpipe (image) | tokens |
|
||||
|---|---:|---:|---:|---|
|
||||
| GSM8K (math) | 100 | 97% | 96% | 61 → 38 (**−38%**) |
|
||||
| novel arithmetic (un-memorizable) | 100 | 100% | **93%** | **−38%** |
|
||||
|
||||
GSM8K *looks* near-lossless (−1pp), but that's inflated — the model recognizes
|
||||
memorized problems and recalls the answer even when it misreads the image. The
|
||||
novel test strips that out: the real reading cost is ~7pp, not ~1.
|
||||
The ~7% gap is real misreads (`10200`→`9400`, `7873`→`7793`), not noise — imaging
|
||||
is lossy even where it mostly works.
|
||||
|
||||
**3. Where it fails completely — dense / verbatim:**
|
||||
**The boundary** — push to dense, exact-recall content and it collapses:
|
||||
|
||||
| test | baseline (text) | pixelpipe (image) |
|
||||
| test | text | pixelpipe (image) |
|
||||
|---|---:|---:|
|
||||
| verbatim recall — exact 12-char hex from a *dense* render | 15/15 (100%) | **0/15 (0%)** |
|
||||
| verbatim recall — 12-char hex from a *dense* render | 15/15 | **0/15** |
|
||||
|
||||
So pixelpipe reads **short, readable** content at ~93% (a real ~7% misread tax)
|
||||
and **silently fails** on dense walls and exact recall. The number you get depends
|
||||
entirely on what you feed it.
|
||||
~93% on short / readable, **0%** on dense / exact-recall. There is no free lunch —
|
||||
full analysis in [`FINDINGS.md`](FINDINGS.md).
|
||||
|
||||
Reproduce: [`eval/gsm8k/`](eval/gsm8k/) (GSM8K + the novel reading test) and
|
||||
[`eval/needle-haystack/`](eval/needle-haystack/) (the verbatim failure). Full
|
||||
analysis: [`FINDINGS.md`](FINDINGS.md).
|
||||
<sub>We also ran the standard **GSM8K** suite: 96% imaged. But GSM8K is in training
|
||||
data, so the model recalls memorized answers through its own misreads — inflating
|
||||
the score ~3pp over the clean novel number above, which is why we don't lead with
|
||||
it. Reproduce: [`eval/gsm8k/`](eval/gsm8k/) · [`eval/needle-haystack/`](eval/needle-haystack/).</sub>
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user