mirror of
https://github.com/teamchong/pxpipe.git
synced 2026-07-22 02:02:51 +02:00
5eb80a4461
Several families share /v1/responses. Cost, geometry, cache rates, and whether imaging is even allowed must follow the model that serves the request. Otherwise Claude and Grok inherit GPT defaults, the gate lies, and the dashboard reports nonsense. Claude on Responses: - Bill images by Anthropic pixel area; cache 0.1x, output 5x. - Use Anthropic page geometry (312 cols x 728 px). GPT height overstated image cost ~2.6x and flipped every Opus slab to not_profitable, so enabled Claude stayed text-only with blank As text / Saved. History profitability gate: - Reflow packs hard newlines with an inline ↵ glyph. countVisualRows still treated ↵ as a row break, so estimateImageCount overstated pages ~6x on reflowed history and collapsed Claude/Grok history never fired. Match the renderer: only hard newlines start a visual row. Grok: - Keep production 5x8 packing and a verbatim fact-sheet for OCR-hard tokens (paths/hex/ports/camelCase). Measured ~1000 image tok/MPix; cache 0.25x, output 3x. - Leave Grok out of DEFAULT_MODEL_BASES (opt-in only, same bar as Opus). Pure-image exact OCR fails at 5x8; do not image it silently. Also thread optional per-model render style through the Responses path and allow eval-only atlas overrides in gen-atlas.