mirror of
https://github.com/teamchong/pxpipe.git
synced 2026-07-22 02:02:51 +02:00
028f6a7cd8
Anthropic bills images by a 28-px patch grid (⌈w/28⌉×⌈h/28⌉ after tier downscale), not (w·h)/750 — the old fit overcharged large images ~4–5%. - new src/core/anthropic-vision.ts: exact patch + downscale model matching Anthropic's reference (incl. extreme aspect ratios); wired into the unified visionTokensForModel router. - gate counts real 28-px patches; 10% bias renamed ANTHROPIC_GATE_MARGIN, gate-only. - removed ANTHROPIC_PIXELS_PER_TOKEN / IMAGE_COST_SAFETY_MARGIN; cols 313 → 312 (slab = 1568 px cap). - docs (README / RENDER_SIZING / TRANSFORM_INFO / audit) synced to shipped geometry (1568×728, 312 cols) and the patch billing model.
101 lines
4.4 KiB
Markdown
101 lines
4.4 KiB
Markdown
# How pxpipe sizes rendered images
|
||
|
||
Image geometry is selected by the exact model id. `/v1/responses` is only a
|
||
wire protocol: Claude, GPT, and Grok on that path do not share a renderer.
|
||
|
||
## Current behavior
|
||
|
||
A page is profile-width and variable-height. It is never padded to a square.
|
||
|
||
```text
|
||
cellWidth = atlasWidth + cellWBonus
|
||
cellHeight = atlasHeight + cellHBonus
|
||
width = 2 * PAD_X + columns * cellWidth
|
||
height = 2 * PAD_Y + usedRows * cellHeight
|
||
```
|
||
|
||
Width may shrink to the widest rendered line, capped by the profile. Height
|
||
grows with content until the profile cap, then overflow becomes another page.
|
||
|
||
| model rule | atlas / effective cell | columns | full width | max height |
|
||
|---|---|---:|---:|---:|
|
||
| Claude / Anthropic | Spleen + Unifont, 5×8 | 312 | 1568 px | 728 px |
|
||
| opt-in `gpt-5.6-sol` | Spleen + Unifont, 5×8 | 152 | 768 px | 1932 px |
|
||
| opt-in Grok 4.5 | Spleen + Unifont, 5×8 white AA + IDS + factsheet (no grid) | 152 | 768 px | 512 px |
|
||
| other OpenAI fallback | Spleen + Unifont, 5×8 | 152 | 768 px | 1932 px |
|
||
|
||
Complete profile values and evidence are in
|
||
[`MODEL_RENDER_PROFILES.md`](MODEL_RENDER_PROFILES.md).
|
||
|
||
The Sol row is the **currently configured** geometry, not a successful recall
|
||
claim. Paid raw-image calls tested both the earlier 6×11/126-column geometry and
|
||
this 5×8/152-column profile: each returned 0/4 exact identifiers with four
|
||
confabulations; 5×8 also missed gist. Sol is consequently opt-in. The effective 9×12/84-column Sol
|
||
candidate is local-only pending a model call. See
|
||
[`eval/sol-profile/RESULTS.md`](../eval/sol-profile/RESULTS.md).
|
||
|
||
## Provider boundaries
|
||
|
||
Anthropic resizes images to fit both a 1568px long edge and roughly 1.15MP.
|
||
The 1568×728 Claude page stays inside both bounds, so rasterized glyphs reach
|
||
the vision encoder without the old 0.555× resample.
|
||
|
||
OpenAI-shaped profiles use portrait strips at or below the 768px short-side
|
||
floor. GPT 5.6 Sol and Grok both use 152×5+8, a 768px strip (1932 px and 512 px
|
||
tall respectively).
|
||
|
||
## Font and Unicode
|
||
|
||
Atlases are generated at build time; runtime never loads font files. The compact
|
||
JetBrains atlas contains common code/Latin/symbol ranges. Missing codepoints
|
||
fall back to the full Spleen/Unifont atlas, preserving multilingual rendering.
|
||
|
||
Spacing, anti-aliasing, grid, color cycle, and newline-marker style are all part
|
||
of the model profile. The same style object is used for static slabs and history.
|
||
|
||
## Gate parity
|
||
|
||
The profitability gate resolves the same profile as the renderer and derives:
|
||
|
||
- cell width and height
|
||
- content-width shrink
|
||
- rows per page
|
||
- page count
|
||
- per-image token cost
|
||
|
||
Changing cell spacing or font therefore cannot leave cost prediction on a
|
||
hardcoded 5×8 canvas. Claude uses the Anthropic 28-px patch formula (with a
|
||
conservative margin at the gate), GPT uses model-specific tile/patch pricing, and
|
||
Grok uses the measured tokens-per-megapixel rate.
|
||
|
||
## Why not square pages
|
||
|
||
Billing follows pixel area or patches/tiles, not visual symmetry. Padding a
|
||
portrait or short final page to a square adds cost without adding information.
|
||
The target is readable characters per billed pixel under each provider's resize
|
||
contract, not a preferred aspect ratio.
|
||
|
||
## Measurement history
|
||
|
||
| date | change | finding |
|
||
|---|---|---|
|
||
| 2026-05-21 | R3 reflow | recovered line-end dead space by marking hard newlines with `↵` |
|
||
| 2026-05-22 | grayscale atlas + L1/L2 harness | measured style and instruction-band effects |
|
||
| 2026-06-09 | Fable-only 5×8 scope | Fable read dense cells better than Opus |
|
||
| 2026-06-17 | ~1932px experiment | fewer pages, later shown to be server-resampled on Anthropic |
|
||
| 2026-07-01 | Claude 1568×728 clamp | made billed pixels survive Anthropic preprocessing |
|
||
| 2026-07-09 | exact-model profiles | separated Claude, GPT 5.6 Sol, and Grok geometry/style |
|
||
| 2026-07-11 | Grok profile | Shared 5×8 + IDS + factsheet production stack; Grok opt-in only |
|
||
| 2026-07-11 | default scope | Fable only; Sol and Grok remain opt-in after broader quality tests |
|
||
|
||
Older benchmark receipts retain their historical geometry; do not rewrite those
|
||
numbers as if they were generated by the current profile.
|
||
|
||
## Changing a profile
|
||
|
||
1. Add or re-run a reader-specific recall/style harness.
|
||
2. Measure the provider's real image-token accounting.
|
||
3. Keep the strip inside its no-resize boundary.
|
||
4. Test slab, history, gate, and cache determinism together.
|
||
5. Do not generalize one variant's result to sibling model ids.
|