mirror of
https://github.com/teamchong/pxpipe.git
synced 2026-07-22 02:02:51 +02:00
455ab636a2
Probe 1: both fable-5 and sonnet-5 bill vision on a 28x28 patch grid (image_tokens = 3 + ceil(W/28)*ceil(H/28)); snap rows≡6 (mod 7), cols≡4 (mod 28) to avoid stranding paid patch area. Probe 2: patch-boundary straddling does NOT affect OCR accuracy (cols: 4.47% vs 4.60%; rows: z=-0.45, 7-offset paired sweep with line fixed effects). Real misread drivers: high-entropy runs (bimodal derailment on base64 blobs), 5x8 confusables (w/W, 8/0), line wraps. Harnesses: count-tokens-sweep, accuracy-phase-probe (line-DP-aligned scoring), rescore-sweep (offline paired analysis).
4.4 KiB
4.4 KiB
Patch-grid probe findings
Probe 1 — billing staircase (count_tokens, free) ✅
- Both
claude-fable-5andclaude-sonnet-5bill vision on a 28×28 px patch grid. The "Sonnet uses 32×32" rumor is false: image_tokens steps occur exactly when W or H crosses a multiple of 28, on both axes, both models (CSVs in this dir). - Formula:
image_tokens = 3 + ceil(W/28) * ceil(H/28)(fixed +3 per image). Spot check at H=28: W=53..56 → 5 tokens (2 patches), W=57..60 → 6 (3 patches). - The docs'
(W×H)/750is just an approximation of 784 px²/patch (28²) plus the constant. - No server-side resample at ≤1568 px long edge: a 1.73 MP image bills the full grid, so there is no ~1.15 MP cap kicking in — what we render is exactly what the model sees.
Geometry consequences for pxpipe (CELL 5×8, PAD_X = PAD_Y = 4)
- Height = 8 + 8·rows → hits a 28-multiple iff rows ≡ 6 (mod 7)
(every 7 rows = 56 px = exactly 2 patch rows).
MAX_HEIGHT_PX = 728at 90 rows = 26 patch rows, a perfect fit. - Width = 8 + 5·cols (+ atlas slack) → hits a 28-multiple iff cols ≡ 4 (mod 28) (every 28 cols = 140 px = 5 patch cols). 312 cols → 1568 px = 56 patch cols, also a perfect fit at the width cap.
- Full-cap page: 312 × 90 = 28,080 chars for 3 + 56·26 = 1459 tokens ≈ 19.2 chars/token — the density ceiling for this geometry.
- Snapping rule: pick rows ≡ 6 (mod 7) and cols ≡ 4 (mod 28); anything else strands already-paid patch area (up to 27 px per axis, ~1–3% of the bill).
- Phase structure: gcd(5,28)=1 → all 28 horizontal glyph↔patch phases occur in every image; gcd(8,28)=4 → only 7 distinct vertical phases.
Probe 2 — accuracy vs phase ✅ (verdict: phase alignment does NOT matter)
- Question: do glyphs straddling a patch boundary misread more? A 5 px glyph
straddles when
x mod 28 ≥ 24; an 8 px glyph row straddles wheny mod 28 ≥ 21. If straddle phases dominate errors, phase-locked pitch (fractional 5.6 / 9.33 px advances, ≈24% density cost) could pay for itself; if flat, keep packed 5×8 and close the alignment theory. - Method notes (pitfalls that produced false signals first):
- Pure-random char grids trip the safety layer ("looks like credentials") — use real repo source as content.
temperatureis rejected by fable/sonnet-5 → sampling is stochastic; single runs are NOT repeatable (same image scored 169 vs 237 errs).- Models wrap/merge lines (sonnet emitted 101 lines for 90); positional line pairing turns one slip into a phase-flat ~27% error smear. The harness now does banded line-level DP alignment before char scoring (sonnet went 72.54% → 99.87% on the identical response).
- Within one image, row phase aliases content line-type every 7 lines
(8·7 = 56 = 2 patches) → row buckets are confounded. Controlled sweep:
prepend k = 0..6 blank lines (shifts content by 8k px through all 7 row
phases, content identical), then score with per-line fixed effects
(
rescore-sweep.mjs, offline, from dumped responses).
- Results (claude-fable-5, 312×90 production geometry):
- Columns (within-image, content-controlled by construction): straddle 4.47% vs aligned 4.60% on 3,691 chars — null, per-phase table flat.
- Rows (7-offset paired sweep, 439 line×run cells): straddle excess −0.16%±0.28 vs aligned +0.03%±0.29, z = −0.45 — null.
- Real code at full page: fable 99.96% (1/2,276), sonnet 99.87% (3/2,276). 28,080 chars for 1,459 image tokens ≈ 19.2 chars/token with ~0.1% CER.
- What actually causes misreads (in error-yield order):
- Long high-entropy runs (base64/hex blobs): bimodal derailment — the same line on the same pixels scored 4% and 73% error across runs. The decoder loses lock mid-run with no language prior to recover; this, not geometry, is the production misread mechanism.
- 5×8 confusables:
w→W(11× in one page),s→S,c→C,K→H,M→N,8→0,(→O,:→.— legibility floor, context-corrected in real code. - Line wraps on long lines — harmless after alignment, but consumers that trust exact line numbers must re-align.
- Recommendations: keep packed 5×8 (fractional-advance idea rejected); snap
dims per Probe 1 for billing only; route/flag high-entropy lines (e.g.
64 chars of base64-ish content) as literal text instead of pixels.