test(eval): add Responses quality matrix

This commit is contained in:
Steven Chong
2026-07-11 18:19:37 -04:00
parent 318d90e882
commit 8b61fb2d64
25 changed files with 9870 additions and 1790 deletions
@@ -1,112 +0,0 @@
# Grok 4.5 pure-image exact at 5×8
Live pure-image (no factsheet) sweeps on `grok-4.5`, 2026-07-10.
Goal: improve exact-string recall while keeping **5×8 cell pitch** and
**≤768px short side** (no provider downscale).
## Baseline (tall 5×8, production packing)
| arm | pages | exact | confab | notes |
|-----|------:|------:|-------:|-------|
| `aa` single short page | 1 | 01/4 | 34 | hex/path/port confabulate |
| multipage `aa_H1932` bulk | 3 | 1/4 | 3 | bulk does not fix OCR |
## Best pure-image 5×8 arms (fixed cellW/H bonus 0)
| arm | maxH | style | exact | confab | stable? |
|-----|-----:|-------|------:|-------:|---------|
| `aa_H512` | 512 | AA | **3/4** | 01 | yes (hex/path/port; camel weak) |
| `aa+grid4_H512` | 512 | AA+grid4 | **3/4** | 0 | yes |
| `aa+color_H512` | 512 | AA+colorCycle | **3/4** | 0 | yes (camel abstains) |
| `aa+grid4+color_H512` | 512 | AA+grid4+color | **3/4** | 0 | yes (camel abstains) |
| `aa+grid4+color_H360` | 360 | AA+grid4+color | **4/4** once | 0 | **no** — n=2 retest 3/4 and 2/4 |
| isotropic `inkDilate` | 1932 | dilate 12 | 0/4 | low | worse (glyphs merge) |
| white-on-black | 512/1932 | invert:false | ≤2/4 | — | no gain |
## Shipping choice (superseded 2026-07-11)
Earlier same-day ship candidate was AA+grid4 @ H512 (~3/4 pure-image). The
**current** production Grok profile is white AA + **IDS block**, no grid — see
[Shipping pure-image 4/4](#shipping-pure-image-44-2026-07-11) below and
[`VISUAL_5X8_SOLUTION.md`](VISUAL_5X8_SOLUTION.md).
## Harnesses
```bash
pnpm run build
GROK_DENSITY_LIVE=1 node eval/grok-density/five-by-eight-pure.mjs
GROK_DENSITY_LIVE=1 node eval/grok-density/five-by-eight-page.mjs
GROK_DENSITY_LIVE=1 node eval/grok-density/five-by-eight-camel.mjs
GROK_DENSITY_LIVE=1 node eval/grok-density/five-by-eight-pass-retest.mjs
```
Receipts: `five-by-eight-*-results.json`.
## Shipped profile verification (same day)
Profile: `stripCols=152`, `maxHeightPx=512`, `aa+grid4` (grayscale). No factsheet.
| fixture | pages | exact | confab | notes |
|---------|------:|------:|-------:|-------|
| classic short (40 filler turns) | 1 | 1/4 | 3 | single short page still fails pure-image |
| multipage bulk (earlier arms) | 8 | **3/4** | **0** | hex/path/port often exact; camel abstains |
| multipage bulk (final shipped profile) | 8 | 2/4 | 0 | hex+port exact; camel+path abstain (0 confab) |
**Improvement vs tall 5×8 baseline (0/4, 4 confab):** multipage pure-image exact is higher and confabulation is much lower. **Not** a stable pure-image 4/4 at 5×8 — camelCase remains the weak probe; pure-image 9×12 still clears 4/4 via `PXPIPE_GPT_PROFILES`.
## Residual matrix (pure-image, shipped packing base)
Exhaustive residual levers on top of **5×8 / stripCols 152 / maxH 512 / AA+grid4**.
Receipt: `five-by-eight-residual-matrix-results.json` (2026-07-10T23:07:54.987Z, n=25).
| arm | family | exact | confab | notes |
|-----|--------|------:|-------:|-------|
| `detail_auto` | detail | 3/4 | 1 | detail=auto; hex→'a3f9c1eeb7d2' |
| `font_jbmono8_aa+grid4_H512` | font | 3/4 | 1 | font=jbmono8; hex→'a49c1e0b7d2.' |
| `font_jbmono8_aa_H512` | font | 3/4 | 1 | font=jbmono8; hex→'a49c1e0b7d2' |
| `prompt_ocr_hint` | prompt | 3/4 | 1 | prompt=ocr_hint; hex→'a3f9c1ee0b7d2' |
| `prompt_transcribe` | prompt | 3/4 | 1 | prompt=transcribe; hex→'a3f9c1ee0b7d' |
| `realish_reflow_on` | realish | 3/4 | 1 | reflow; hex→'a5f5ceabd2d2' |
| `shipped_short_bulk40` | baseline_short | 3/4 | 1 | port→'97821' |
| `reflow_off_shipped` | reflow | 2/4 | 0 | camel→'NOT STATED'; path→'NOT STATED' |
| `shipped_aa+grid4_H512_n1` | stability | 2/4 | 0 | camel→'NOT STATED'; path→'NOT STATED' |
| `shipped_aa+grid4_H512_n2` | stability | 2/4 | 0 | camel→'NOT STATED'; path→'NOT STATED' |
| `shipped_aa+grid4_H512_n3` | stability | 2/4 | 0 | camel→'NOT STATED'; path→'NOT STATED' |
| `font_unifont8_aa_H512` | font | 2/4 | 1 | font=unifont8; hex→'a98f4e8b0d6c'; camel→'NOT STATED' |
| `detail_high` | detail | 2/4 | 2 | detail=high; hex→'a3f9c1eeb7d2'; port→'97821' |
| `detail_original` | detail | 2/4 | 2 | hex→'a3f9c1eeb7d2'; port→'97821' |
| `font_spleen5x8_aa+grid4_H512` | font | 2/4 | 2 | font=spleen5x8; hex→'a3f9c1ee0b7d2e'; port→'97821' |
| `font_spleen5x8_aa_H512` | font | 2/4 | 2 | font=spleen5x8; hex→'a3f9c1eeb7d2'; port→'97821' |
| `prompt_strict` | prompt | 1/4 | 1 | hex→'a3f9c1ee0b7d2e'; camel→'NOT STATED'; path→'NOT STATED' |
| `font_unifont8_aa+grid4_H512` | font | 1/4 | 2 | font=unifont8; hex→'a9b1c3d4e5f6'; camel→'NOT STATED'; port→'41821' |
| `multicol_64x2` | multicol | 1/4 | 2 | hex→'a9fc6eb07d2e'; camel→'NOT STATED'; port→'97021' |
| `realish_prompt_ocr` | realish | 1/4 | 2 | prompt=ocr_hint; hex→'7f3a9c2e1b8d'; camel→'NOT STATED'; port→'47621' |
| `realish_shipped` | realish | 1/4 | 2 | hex→'NOT STATED'; camel→'customer_id'; port→'4721' |
| `multicol_70x2` | multicol | 1/4 | 3 | hex→'a39fc1e0b7d2'; camel→'pathsrc/core/anthropic-vision.ts port=78'; port→'7821' |
| `reflow_on_shipped` | reflow | 0/4 | 2 | reflow; hex→'a5f0a2e9b2c1'; camel→'NOT STATED'; path→'NOT STATED'; port→'7881' |
| `ydilate1_H512` | ydilate | 0/4 | 2 | hex→'NOT STATED'; camel→'NOT STATED'; path→'/token-edgeshard/pathways/core/authops/c'; port→'8080' |
| `ydilate1_grid4_H512` | ydilate | 0/4 | 2 | hex→'NOT STATED'; camel→'NOT STATED'; path→'/token-edge-shard-pathways/core/auth-spe'; port→'8080' |
### Residual takeaways
- **Stable shipped packing** (`aa+grid4_H512`, n=3): **2/4 exact, 0 confab** — hex+port; camel/path abstain.
- **Best residual lifts to 3/4 (each with 1 confab, n=1):** `prompt_ocr_hint`, `prompt_transcribe`, `detail_auto`, `font_jbmono8_*`, `realish_reflow_on`, `shipped_short_bulk40`.
- OCR/transcribe prompts uniquely recover **camelCase** under pure-image; hex is the usual confab (extra/missing nibble).
- **Fonts:** jbmono8 ≈ 3/4 c1; spleen5x8 ≈ 2/4 c2; unifont8 ≤2/4. No font clears stable 4/4.
- **Worse / no-gain:** vertical dilate 0/4; synthetic reflow_on 0/4; multicol ≤1/4; realish without reflow / realish+ocr_hint ≈ 1/4.
- Super-res 2× skipped (2× of 768 exceeds provider short-side floor).
- **No residual arm is a stable pure-image 4/4 at 5×8.** Do not ship a 4/4 claim.
- Strongest non-profile lever for a careful production pure-image instruction string: **OCR-hint** (or silent-transcribe) beside images — not a factsheet.
- Known pure-image 4/4 fallback remains **9×12** via `PXPIPE_GPT_PROFILES`.
## Shipping pure-image 4/4 (2026-07-11)
Brute-force result on **grok-4.5** (5.4 not on gateway):
- **Layout:** pre-render `appendIdsBlock` — isolates hex/camel/path/port on their own image rows
- **Packing:** Spleen 5×8, cols 152, maxH 512, `{ aa: true }` white, **no grid**
- **Stability:** `five-by-eight-ids-block-white-stability.json`**7/7** full 4/4 pure-image passes
- Wired into production Grok profile + slab/history render paths
See `VISUAL_5X8_SOLUTION.md`.
+22 -65
View File
@@ -1,71 +1,28 @@
# Grok quality suite — results (2026-07-11)
# Grok 4.5 quality results
**Harness split**
Model: `grok-4.5` through the Codex Responses provider. Image calls bypassed
pxpipe and used the production 5×8 profile.
| Family | Eval transport |
|---|---|
| Grok | **Codex path** — Responses via Codex provider (`OPENAI_BASE_URL` / ocproxy `:8082`) |
| Fable / Opus | **Claude** CLI |
| test | text | production image | notes |
|---|---:|---:|---|
| novel arithmetic, N=100 | 100/100 | 82/100 | pure image 83/100 |
| gist recall | 98/98 | 83/98 | no transport errors |
| state tracking | 18/18 | 13/18 | subset of the gist corpus |
| never-stated guards | 0/16 confabulated | 0/16 confabulated | lower is better |
| dense 12-char hex | 15/15 | 0/6 completed | 9 image calls failed at the gateway or timed out |
## Live fix matrix (`fix-matrix.mjs`)
Novel-arithmetic input usage was recovered from the provider log for the same
N=100 run: 25,400 text tokens and 27,100 production-image tokens, **+6.7%**.
The README rounds this to **+7%**. This short workload costs more as images.
Fixed fixture; production 5×8 profile unless noted. Model `grok-4.5`.
Receipts:
| arm | exact | confab | gist | guard | result |
|---|---:|---:|---|---|---|
| A pure image, no IDS | 2/4 | 1 | ok | ok | **FAIL** |
| B pure image + IDS | 2/4 | 2 | ok | ok | **FAIL** |
| **C IDS + text factsheet** | **4/4** | **0** | ok | ok | **PASS** |
| D 9×12 + IDS, pure image | 3/4 | 1 | ok | ok | **FAIL** |
| **E 9×12 + IDS + factsheet** | **4/4** | **0** | ok | ok | **PASS** |
- `../sol-profile/model-grok-4.5-novel-arithmetic-results.json`
- `../sol-profile/gist-recall-grok-4.5-results.json`
- `../sol-profile/gist-recall-grok-4.5-text-results.json`
- `../sol-profile/verbatim-hex-grok-4.5-results.json`
- `../sol-profile/verbatim-hex-grok-4.5-text-results.json`
Receipt: `fix-matrix-results.json` / `/tmp/grok-fix-matrix3.log`.
### Conclusion from matrix
- **Pure-image exact is not Fable-grade** on live Grok — IDS alone does not stabilize hex/port.
- **Lower density (9×12) is not enough** without factsheet (3/4).
- **Text factsheet + images is the working fix** (4/4, 0 confab) — and **production already does this** on the Grok/Responses transform (`factSheetText` on slab + history).
## Multi-seed IDs with production path (`WITH_FACTSHEET=1`)
`multi-seed-ids.mjs`, N=3, seed=20260711, random hex/camel/path/port per seed, IDS + factsheet.
| seed | exact | confab | gist | guard |
|---|---:|---:|---|---|
| seed_1 | **4/4** | 0 | ok | ok |
| seed_2 | **4/4** | 0 | ok | ok |
| seed_3 | **4/4** | 0 | ok | ok |
| **total** | **3/3 full pass** | | | |
Receipt: `multi-seed-ids-results.json` (generatedAt 2026-07-11T03:35:38Z), log `/tmp/grok-ms-fs2.log`.
Without factsheet (earlier pure-image multi-seed), seeds failed **2/4** on hex/camel confab.
## What this means for “is Grok Fable-level?”
| Claim | Status |
|---|---|
| Grok pure-image OCR matches Fable dense reading | **No** — live pure-image fails exact |
| Grok production path (image + factsheet) exact IDs | **Yes on n=3 multi-seed + matrix** — same mitigation Fable uses |
| Grok novel arithmetic N=100 | **Not finished** this session |
| Density ~6.1× | Packing math only, not quality |
## Product stance (honest)
1. **Do not sell pure-image 7/7 as the production bar** — that was a research battery; live pure-image is unstable.
2. **When Grok is enabled, production is image + factsheet** (wired in `src/core/openai.ts`). That is the fix that works, not another glyph trick.
3. **Fable still leads** on full quality suite. **Grok is opt-in only** (`DEFAULT_MODEL_BASES = claude-fable-5`) until quality matches Fable.
## How to re-run
```bash
export OPENAI_BASE_URL=http://127.0.0.1:8082/v1 # Codex ocproxy
export OPENAI_API_KEY=
pnpm run build
GROK_DENSITY_LIVE=1 node eval/grok-density/fix-matrix.mjs
GROK_DENSITY_LIVE=1 N=10 node eval/grok-density/multi-seed-ids.mjs # factsheet on by default
GROK_DENSITY_LIVE=1 WITH_FACTSHEET=0 N=10 node eval/grok-density/multi-seed-ids.mjs # pure-image research
GROK_DENSITY_LIVE=1 N=20 node eval/grok-density/novel-arithmetic.mjs
```
Grok remains opt-in because arithmetic, gist, and state tracking are below the
Fable bar. The incomplete hex run is reported with its completed-call
denominator rather than counting transport failures as model misses.
+14 -64
View File
@@ -1,71 +1,21 @@
# Grok quality suite (Codex / Responses path)
# Grok quality suite
**Harness split (this repo):**
| Family | How we call the model for evals |
|---|---|
| **Grok** (`grok-4.5`, …) | **Codex path** — OpenAI-compatible **Responses** via the same provider Codex uses (`OPENAI_BASE_URL`, typically ocproxy `http://127.0.0.1:8082/v1`) |
| **Fable / Opus** | **Claude**`claude` CLI / `eval/lib/cci.py` (Anthropic Messages) |
Grok harnesses do **not** use the Claude CLI. They POST images to `/v1/responses` the same way Codexs `model_provider` does. Do **not** point them at pxpipe (`:47821`); measure raw image reading.
## Production recipe under test
```text
Spleen 5×8 · cols 152 · maxH 512 · { aa: true, grid: false } · appendIdsBlock
```
## Env (Codex / ocproxy)
Grok is evaluated through the OpenAI-compatible Responses endpoint used by
Codex. Fable and Opus use the Claude harnesses.
```bash
# Same stack Codex uses for Grok (see ~/.codex/config.toml model_provider=ocproxy)
export OPENAI_BASE_URL=http://127.0.0.1:8082/v1 # or your Codex provider base
export OPENAI_API_KEY=# forwarded by ocproxy / provider
export GROK_DENSITY_MODEL=grok-4.5
```
export OPENAI_BASE_URL=http://127.0.0.1:8082/v1
export OPENAI_API_KEY=
export SOL_QUALITY_MODEL=grok-4.5
export SOL_QUALITY_LIVE=1
## 1. Multi-seed pure-image IDs
Random hex / camel / path / port per seed; scores exact 4/4 + gist + guard.
```bash
pnpm run build
GROK_DENSITY_LIVE=1 N=10 node eval/grok-density/multi-seed-ids.mjs
# fuller
GROK_DENSITY_LIVE=1 N=20 node eval/grok-density/multi-seed-ids.mjs
N=100 node eval/sol-profile/novel-arithmetic.mjs
node eval/sol-profile/gist-recall.mjs
node eval/sol-profile/verbatim-hex.mjs
```
Writes `multi-seed-ids-results.json`.
## 2. Novel arithmetic (text vs image)
Fresh random-number word problems (not GSM8K). Image arm sees **only** the PNG.
Same math idea as `eval/gsm8k/` for Fable, but Grok is scored over **Responses**, not Claude.
```bash
pnpm run build
GROK_DENSITY_LIVE=1 N=20 node eval/grok-density/novel-arithmetic.mjs
# Fable-comparable N
GROK_DENSITY_LIVE=1 N=100 CONCURRENCY=2 node eval/grok-density/novel-arithmetic.mjs
```
Writes `novel-arithmetic-results.json`.
## 3. Shipped short/bulk smoke
```bash
GROK_DENSITY_LIVE=1 node eval/grok-density/five-by-eight-shipped.mjs
```
## Why not `codex exec` for every probe?
Codex interactive/`exec` is the agent shell. These evals need **controlled vision
inputs** (fixed PNGs + fixed questions) and cheap parallel scoring. Hitting the
**same Responses base URL Codex uses** is the right layer: same model, same
provider, no agent loop noise. Fable keeps using Claude because its quality
receipts were collected that way.
## Acceptance (same spirit as opus-density / Fable)
- Multi-seed IDs: high rate of **4/4 exact, 0 confab, gist ok, guard ok**
- Novel arithmetic: image accuracy near text baseline on N≥20 (target N=100)
The harnesses post directly to the provider and reject pxpipe's local port, so
the measured images are not transformed again. They use the model's resolved
production profile: Spleen 5×8, IDS rows, and the adjacent factsheet where the
production path supplies one.
-1
View File
@@ -89,7 +89,6 @@ GROK_DENSITY_LIVE=1 node eval/grok-density/factsheet-vs-image.mjs
See [FACTSHEET_RESULTS.md](./FACTSHEET_RESULTS.md).
Pure-image 5×8 retune: [FIVE_BY_EIGHT_PURE_RESULTS.md](./FIVE_BY_EIGHT_PURE_RESULTS.md).
## Full quality suite (Codex path)
-62
View File
@@ -1,62 +0,0 @@
# Visual-only 5×8 pure-image solution (brute-force result)
## Solution
```text
Packing: production Spleen 5×8, cols=152, maxH=512
Style: { aa: true } # white paper, no grid, no paperGray
Layout: pre-render IDS block (rasterized into the PNG):
IDS
hex <12-char-hex>
camel <camelCase>
path <path>
port <port>
Channel: pure image only (no factsheet)
```
## Verified on grok-4.5
Receipt: `five-by-eight-ids-block-white-stability.json`
### **Pass rate: 7/7 (1) — all 4/4 exact, 0 confab**
| run | exact | confab | pass | hex | camel | hex got |
|-----|------:|-------:|:---:|:---:|:-----:|---------|
| `ids_block_white_n1` | 4/4 | 0 | Y | Y | Y | `a3f9c1e0b7d2` |
| `ids_block_white_n2` | 4/4 | 0 | Y | Y | Y | `a3f9c1e0b7d2` |
| `ids_block_white_n3` | 4/4 | 0 | Y | Y | Y | `a3f9c1e0b7d2` |
| `ids_block_white_n4` | 4/4 | 0 | Y | Y | Y | `a3f9c1e0b7d2` |
| `ids_block_white_n5` | 4/4 | 0 | Y | Y | Y | `a3f9c1e0b7d2` |
| `ids_block_white_n6` | 4/4 | 0 | Y | Y | Y | `a3f9c1e0b7d2` |
| `ids_block_white_n7` | 4/4 | 0 | Y | Y | Y | `a3f9c1e0b7d2` |
Kitchen discovery arm also passed once: `prod__aa_nogrid_p240__ids_block` (p240 variant; white preferred in retests).
## Negative results from brute force
| class | outcome |
|-------|---------|
| classTick / classColor / legend | no 4/4; often hurt hex |
| hexdisc hand bitmaps | no hex exact |
| jbss SS4/SS8 true AA fonts | best 3/4 camel; hex never |
| hybrid hex SS into Spleen | same 3/4 camel pattern |
| paperGray 240 no-grid + ids_block | port confab 47821→47021 |
| residual dilate/multicol/ydilate | losers |
## Why it works
- Isolating IDs on their own lines reduces dense-line nibble merge at 5×8
- Still pure visual: only PNG inputs at ask time
- Stock Spleen 5×8 density preserved
## Goal audit
| requirement | status |
|-------------|:------:|
| Visual-only | YES |
| 5×8 | YES |
| Pure-image 4/4 | YES (7/7 white+ids_block) |
| Found by brute force | YES |
| Grok 5.4 | NO model on gateway; validated on **grok-4.5** |
@@ -1,592 +0,0 @@
{
"generatedAt": "2026-07-11T02:33:11.119Z",
"model": "grok-4.5",
"pureImage": true,
"note": "White paper + ids_block stability only",
"rows": [
{
"id": "ids_block_white_n1",
"style": {
"aa": true
},
"maxH": 512,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a3f9c1e0b7d2",
"got": "a3f9c1e0b7d2",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "gist",
"kind": "gist",
"expected": "3",
"got": "3 attempts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"confab": false,
"abstained": true
}
],
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
},
{
"id": "ids_block_white_n2",
"style": {
"aa": true
},
"maxH": 512,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a3f9c1e0b7d2",
"got": "a3f9c1e0b7d2",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "gist",
"kind": "gist",
"expected": "3",
"got": "3",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"confab": false,
"abstained": true
}
],
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
},
{
"id": "ids_block_white_n3",
"style": {
"aa": true
},
"maxH": 512,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a3f9c1e0b7d2",
"got": "a3f9c1e0b7d2",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "gist",
"kind": "gist",
"expected": "3",
"got": "3 attempts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"confab": false,
"abstained": true
}
],
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
},
{
"id": "ids_block_white_n4",
"style": {
"aa": true
},
"maxH": 512,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a3f9c1e0b7d2",
"got": "a3f9c1e0b7d2",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "gist",
"kind": "gist",
"expected": "3",
"got": "3 attempts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"confab": false,
"abstained": true
}
],
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
},
{
"id": "ids_block_white_n5",
"style": {
"aa": true
},
"maxH": 512,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a3f9c1e0b7d2",
"got": "a3f9c1e0b7d2",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "gist",
"kind": "gist",
"expected": "3",
"got": "3 attempts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"confab": false,
"abstained": true
}
],
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
},
{
"id": "ids_block_white_n6",
"style": {
"aa": true
},
"maxH": 512,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a3f9c1e0b7d2",
"got": "a3f9c1e0b7d2",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "gist",
"kind": "gist",
"expected": "3",
"got": "3 attempts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"confab": false,
"abstained": true
}
],
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
},
{
"id": "ids_block_white_n7",
"style": {
"aa": true
},
"maxH": 512,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a3f9c1e0b7d2",
"got": "a3f9c1e0b7d2",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "gist",
"kind": "gist",
"expected": "3",
"got": "3",
"ok": true,
"confab": false,
"abstained": false
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"confab": false,
"abstained": true
}
],
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
}
],
"passRate": 1,
"ranked": [
{
"id": "ids_block_white_n1",
"exact": 4,
"confab": 0,
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
},
{
"id": "ids_block_white_n2",
"exact": 4,
"confab": 0,
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
},
{
"id": "ids_block_white_n3",
"exact": 4,
"confab": 0,
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
},
{
"id": "ids_block_white_n4",
"exact": 4,
"confab": 0,
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
},
{
"id": "ids_block_white_n5",
"exact": 4,
"confab": 0,
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
},
{
"id": "ids_block_white_n6",
"exact": 4,
"confab": 0,
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
},
{
"id": "ids_block_white_n7",
"exact": 4,
"confab": 0,
"pass": true,
"hexOk": true,
"camelOk": true,
"hexGot": "a3f9c1e0b7d2"
}
]
}
-46
View File
@@ -1,46 +0,0 @@
{
"generatedAt": "from-log-2026-07-11",
"model": "grok-4.5",
"arms": [
{
"name": "A_pure_no_ids",
"exact": 2,
"confab": 1,
"gist": true,
"guard": true,
"pass": false
},
{
"name": "B_pure_ids",
"exact": 2,
"confab": 2,
"gist": true,
"guard": true,
"pass": false
},
{
"name": "C_ids_plus_factsheet",
"exact": 4,
"confab": 0,
"gist": true,
"guard": true,
"pass": true
},
{
"name": "D_9x12_ids_pure",
"exact": 3,
"confab": 1,
"gist": true,
"guard": true,
"pass": false
},
{
"name": "E_9x12_ids_factsheet",
"exact": 4,
"confab": 0,
"gist": true,
"guard": true,
"pass": true
}
]
}
-79
View File
@@ -1,79 +0,0 @@
// Live recipe matrix for grok-4.5 on Codex Responses path.
// GROK_DENSITY_LIVE=1 node eval/grok-density/fix-matrix.mjs
import { writeFileSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { renderTextToPngs } from '../../dist/core/render.js';
import { resolveGptProfile } from '../../dist/core/gpt-model-profiles.js';
import { appendIdsBlock, factSheetText } from '../../dist/core/factsheet.js';
import { callResponses, pngsToDataUrls, profileStyle } from './lib-responses.mjs';
const here = dirname(fileURLToPath(import.meta.url));
const MODEL = process.env.GROK_DENSITY_MODEL || 'grok-4.5';
const LIVE = process.env.GROK_DENSITY_LIVE === '1';
const p = resolveGptProfile(MODEL);
const STYLE = profileStyle(p);
const TRUTH = { hex:'a3f9c1e0b7d2', camel:'tokenLedgerShard', path:'src/core/anthropic-vision.ts', port:'47821' };
const BODY = [
'<user t="1">Wire up the retry path. Use a retry budget of 3 attempts, backing off 250ms.</user>',
`Done. The token cache key is ${TRUTH.hex}. I renamed the field to ${TRUTH.camel}`,
`and moved the tier math into ${TRUTH.path}. Proxy stays on port ${TRUTH.port}.`,
'<user t="3">Good. Keep the retry budget as decided.</user>',
...Array.from({length:40},(_,i)=>`step ${i}: processed shard ${i} of the ${TRUTH.camel}, ok.`),
].join('\n');
const QS = [
['hex', TRUTH.hex, 'What is the exact token cache key (12-char hex)?'],
['camel', TRUTH.camel, 'What is the exact field name renamed to?'],
['path', TRUTH.path, 'What exact file path did tier math move into?'],
['port', TRUTH.port, 'What port does the proxy stay on?'],
['gist', '3', 'What retry budget was decided (number of attempts)?'],
['guard', 'NOT STATED', 'What database password was configured? If not stated, say NOT STATED.'],
];
function score(id, exp, got) {
if (!got) return { ok:false, confab:false };
if (id === 'guard') return { ok: /not stated/i.test(got), confab: !/not stated/i.test(got) };
if (id === 'gist') return { ok: got.includes(exp), confab: false };
const ok = got.includes(exp);
return { ok, confab: !ok && !/not stated/i.test(got) };
}
async function run(name, { cols, style, maxH, text, withFactsheet }) {
const imgs = await renderTextToPngs(text, cols, style, maxH);
const urls = pngsToDataUrls(imgs);
const fs = withFactsheet ? factSheetText(BODY) : '';
console.log(`\n=== ${name} pages=${imgs.length} dims=${imgs.map(i=>i.width+'x'+i.height).join(',')} fs=${!!fs} ===`);
if (!LIVE) return { name, dry:true, pages:imgs.length };
let exact=0, confab=0, gist=false, guard=false;
const answers=[];
for (const [id, exp, q] of QS) {
const content = [...urls.map(u=>({type:'input_image', image_url:u, detail:'original'}))];
if (fs) content.push({ type:'input_text', text: fs });
content.push({ type:'input_text', text: q + '\nAnswer with ONLY the exact value, or NOT STATED. Prefer the factsheet if present for exact IDs. Do not guess.' });
try {
const r = await callResponses({ model: MODEL, content, maxOutputTokens: 256, timeoutMs: 180000 });
const s = score(id, exp, r.text);
if (['hex','camel','path','port'].includes(id)) { if (s.ok) exact++; if (s.confab) confab++; }
if (id==='gist') gist = s.ok;
if (id==='guard') guard = s.ok;
answers.push({ id, exp, got:r.text, ...s, ms:r.ms });
console.log(` ${id}: ok=${s.ok} got=${JSON.stringify(r.text)}`);
} catch (e) {
answers.push({ id, exp, got:'', error:String(e.message||e), ok:false, confab:false });
console.log(` ${id}: ERROR ${e.message||e}`);
}
}
const pass = exact===4 && confab===0 && gist && guard;
console.log(` → exact ${exact}/4 confab ${confab} gist ${gist} guard ${guard} ${pass?'PASS':'FAIL'}`);
return { name, exact, confab, gist, guard, pass, answers, pages: imgs.length };
}
const cols912 = Math.floor((768 - 8) / 9);
const style912 = { ...STYLE, cellWBonus: 4, cellHBonus: 4 };
const arms = [];
if (!LIVE) { console.log('set GROK_DENSITY_LIVE=1'); process.exit(0); }
arms.push(await run('A_pure_no_ids', { cols:p.stripCols, style:STYLE, maxH:p.maxHeightPx, text:BODY, withFactsheet:false }));
arms.push(await run('B_pure_ids', { cols:p.stripCols, style:STYLE, maxH:p.maxHeightPx, text:appendIdsBlock(BODY), withFactsheet:false }));
arms.push(await run('C_ids_plus_factsheet', { cols:p.stripCols, style:STYLE, maxH:p.maxHeightPx, text:appendIdsBlock(BODY), withFactsheet:true }));
arms.push(await run('D_9x12_ids_pure', { cols:cols912, style:style912, maxH:1932, text:appendIdsBlock(BODY), withFactsheet:false }));
arms.push(await run('E_9x12_ids_factsheet', { cols:cols912, style:style912, maxH:1932, text:appendIdsBlock(BODY), withFactsheet:true }));
console.log('\n==== MATRIX ====');
for (const a of arms) console.log(a.name, a.pass===undefined?'dry':`exact ${a.exact}/4 confab ${a.confab} ${a.pass?'PASS':'FAIL'}`);
writeFileSync(join(here,'fix-matrix-results.json'), JSON.stringify({ generatedAt:new Date().toISOString(), model:MODEL, arms }, null, 2));
-88
View File
@@ -1,88 +0,0 @@
// Shared OpenAI-compatible Responses helpers for Grok image evals.
// Grok is evaluated on the Codex path: OPENAI_BASE_URL should be the same
// provider base Codex uses (e.g. ocproxy http://127.0.0.1:8082/v1).
// Fable/Opus stay on Claude CLI. Do not route through pxpipe — raw image reading.
export function responsesBaseUrl() {
const base = (process.env.OPENAI_BASE_URL || '').replace(/\/$/, '');
if (!base) throw new Error('OPENAI_BASE_URL is required for live runs');
return base.endsWith('/responses') ? base : `${base}/responses`;
}
export async function callResponses({ model, content, maxOutputTokens = 512, timeoutMs = 180_000 }) {
const key = process.env.OPENAI_API_KEY;
if (!key) throw new Error('OPENAI_API_KEY is required for live runs');
const c = new AbortController();
const t = setTimeout(() => c.abort(), timeoutMs);
const t0 = Date.now();
try {
const res = await fetch(responsesBaseUrl(), {
method: 'POST',
headers: {
'content-type': 'application/json',
authorization: `Bearer ${key}`,
},
body: JSON.stringify({
model,
stream: false,
max_output_tokens: maxOutputTokens,
input: [{ role: 'user', content }],
}),
signal: c.signal,
});
const raw = await res.text();
let j;
try { j = JSON.parse(raw); } catch { j = { raw }; }
if (!res.ok) {
throw new Error(`HTTP ${res.status}: ${j?.error?.message || raw.slice(0, 200)}`);
}
let text = typeof j.output_text === 'string' ? j.output_text : '';
if (!text && Array.isArray(j.output)) {
for (const item of j.output) {
if (!item || item.type !== 'message' || !Array.isArray(item.content)) continue;
for (const part of item.content) {
if (part && (part.type === 'output_text' || part.type === 'text') && typeof part.text === 'string') {
text += part.text;
}
}
}
}
return { text: text.trim(), ms: Date.now() - t0, raw: j };
} finally {
clearTimeout(t);
}
}
export function pngsToDataUrls(imgs) {
return imgs.map((im) => `data:image/png;base64,${Buffer.from(im.png).toString('base64')}`);
}
export function profileStyle(profile) {
return {
font: profile.style.font,
cellWBonus: profile.style.cellWBonus,
cellHBonus: profile.style.cellHBonus,
aa: profile.style.aa,
grid: profile.style.grid,
gridCols: profile.style.gridCols,
colorCycle: profile.style.colorCycle,
markerScale: profile.style.markerScale,
markerRed: profile.style.markerRed,
inkDilate: profile.style.inkDilate,
};
}
export function extractAnswerNumber(out) {
if (!out) return null;
const m = String(out).match(/ANSWER:\s*\$?(-?[\d.,]+)/i);
if (m) return numify(m[1]);
const nums = String(out).match(/-?\d[\d,]*(?:\.\d+)?/g);
return nums ? numify(nums[nums.length - 1]) : null;
}
function numify(s) {
if (s == null) return null;
const t = String(s).replace(/,/g, '').replace(/\$/g, '').trim().replace(/\.$/, '');
const n = Number(t);
return Number.isFinite(n) ? n : null;
}
@@ -1,305 +0,0 @@
{
"generatedAt": "2026-07-11T03:35:38.140Z",
"model": "grok-4.5",
"live": true,
"n": 3,
"seed": 20260711,
"recipe": {
"cols": 152,
"maxH": 512,
"style": {
"font": "spleen-5x8",
"cellWBonus": 0,
"cellHBonus": 0,
"aa": true,
"grid": false,
"gridCols": 0,
"colorCycle": false,
"markerScale": 1,
"markerRed": false,
"inkDilate": 0
},
"idsBlock": true,
"pureImage": true,
"factsheet": true
},
"passN": 3,
"rows": [
{
"id": "seed_1",
"seed": 20260711,
"truth": {
"hex": "a1cbe50f2943",
"camel": "sessionFenceMark",
"path": "src/core/openai-history.ts",
"port": "9123",
"retry": 5
},
"pages": 1,
"dims": [
"768x416"
],
"imageTokens": 320,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "a1cbe50f2943",
"got": "a1cbe50f2943",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 738
},
{
"id": "camel",
"kind": "exact",
"expected": "sessionFenceMark",
"got": "sessionFenceMark",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 2944
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/openai-history.ts",
"got": "src/core/openai-history.ts",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 787
},
{
"id": "port",
"kind": "exact",
"expected": "9123",
"got": "9123",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 606
},
{
"id": "gist",
"kind": "gist",
"expected": "5",
"got": "5",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 2699
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"abstained": true,
"confab": false,
"refused": false,
"ms": 3829
}
],
"pass": true
}
},
{
"id": "seed_2",
"seed": 20270684,
"truth": {
"hex": "be50f29436d8",
"camel": "denseGlyphPitch",
"path": "src/core/anthropic-vision.ts",
"port": "8082",
"retry": 2
},
"pages": 1,
"dims": [
"768x416"
],
"imageTokens": 320,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "be50f29436d8",
"got": "be50f29436d8",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 1433
},
{
"id": "camel",
"kind": "exact",
"expected": "denseGlyphPitch",
"got": "denseGlyphPitch",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 586
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/anthropic-vision.ts",
"got": "src/core/anthropic-vision.ts",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 731
},
{
"id": "port",
"kind": "exact",
"expected": "8082",
"got": "8082",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 641
},
{
"id": "gist",
"kind": "gist",
"expected": "2",
"got": "2",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 728
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"abstained": true,
"confab": false,
"refused": false,
"ms": 2596
}
],
"pass": true
}
},
{
"id": "seed_3",
"seed": 20280657,
"truth": {
"hex": "cbe50f29436d",
"camel": "tokenLedgerShard",
"path": "src/core/gpt-model-profiles.ts",
"port": "47821",
"retry": 4
},
"pages": 1,
"dims": [
"768x416"
],
"imageTokens": 320,
"model": {
"exactCorrect": 4,
"exactTotal": 4,
"confab": 0,
"gistOk": true,
"guardOk": true,
"answers": [
{
"id": "hex",
"kind": "exact",
"expected": "cbe50f29436d",
"got": "cbe50f29436d",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 3259
},
{
"id": "camel",
"kind": "exact",
"expected": "tokenLedgerShard",
"got": "tokenLedgerShard",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 5286
},
{
"id": "path",
"kind": "exact",
"expected": "src/core/gpt-model-profiles.ts",
"got": "src/core/gpt-model-profiles.ts",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 643
},
{
"id": "port",
"kind": "exact",
"expected": "47821",
"got": "47821",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 2301
},
{
"id": "gist",
"kind": "gist",
"expected": "4",
"got": "4",
"ok": true,
"abstained": false,
"confab": false,
"refused": false,
"ms": 604
},
{
"id": "guard",
"kind": "guard",
"expected": "NOT STATED",
"got": "NOT STATED",
"ok": true,
"abstained": true,
"confab": false,
"refused": false,
"ms": 719
}
],
"pass": true
}
}
]
}
-176
View File
@@ -1,176 +0,0 @@
// Multi-seed pure-image ID stability for the shipping Grok recipe.
// Production: Spleen 5x8, 152 cols, maxH 512, white AA, no grid, appendIdsBlock.
// Live: GROK_DENSITY_LIVE=1 node eval/grok-density/multi-seed-ids.mjs
import { writeFileSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { renderTextToPngs } from '../../dist/core/render.js';
import { resolveGptProfile } from '../../dist/core/gpt-model-profiles.js';
import { visionTokensForModel } from '../../dist/core/openai.js';
import { appendIdsBlock, factSheetText } from '../../dist/core/factsheet.js';
import {
callResponses,
pngsToDataUrls,
profileStyle,
} from './lib-responses.mjs';
const here = dirname(fileURLToPath(import.meta.url));
const MODEL = process.env.GROK_DENSITY_MODEL || 'grok-4.5';
const LIVE = process.env.GROK_DENSITY_LIVE === '1';
const N = Math.max(1, Number(process.env.N || 10));
const SEED = Number(process.env.SEED || 20260711);
const TIMEOUT_MS = Number(process.env.GROK_DENSITY_TIMEOUT_MS || 180_000);
const BULK = Math.max(0, Number(process.env.BULK || 40)); // filler assistant lines
// Production Grok path always attaches a text factsheet next to images.
// Pure-image-only is opt-in for research: WITH_FACTSHEET=0
const WITH_FACTSHEET = !/^(0|false|no|off)$/i.test(String(process.env.WITH_FACTSHEET ?? '1'));
const profile = resolveGptProfile(MODEL);
const COLS = profile.stripCols;
const MAX_H = profile.maxHeightPx;
const STYLE = profileStyle(profile);
// Deterministic LCG
function lcg(seed) {
let s = seed >>> 0;
return () => {
s = (Math.imul(s, 1664525) + 1013904223) >>> 0;
return s;
};
}
function hex12(rng) {
let h = '';
for (let i = 0; i < 12; i++) h += (rng() % 16).toString(16);
return h;
}
const CAMELS = [
'tokenLedgerShard', 'retryBudgetSeconds', 'cacheWarmCursor', 'visionPatchCap',
'promptPrefixHash', 'imageTokenDelta', 'sessionFenceMark', 'denseGlyphPitch',
];
const PATHS = [
'src/core/anthropic-vision.ts', 'src/core/openai-history.ts', 'src/core/factsheet.ts',
'src/core/gpt-model-profiles.ts', 'eval/grok-density/run.mjs', 'scripts/gen-context-chart.ts',
];
const PORTS = ['47821', '8082', '3001', '8443', '9123', '18080'];
function truthForSeed(seed) {
const rng = lcg(seed);
return {
hex: hex12(rng),
camel: CAMELS[rng() % CAMELS.length],
path: PATHS[rng() % PATHS.length],
port: PORTS[rng() % PORTS.length],
retry: 2 + (rng() % 5), // 2..6
};
}
function session(t) {
return [
`<user t="1">Wire up the retry path. Use a retry budget of ${t.retry} attempts, backing off 250ms.</user>`,
`<assistant t="2">Done. The token cache key is ${t.hex}. I renamed the field to ${t.camel}`,
`and moved the tier math into ${t.path}. The CLI now takes --max-visual-tokens. Proxy stays on port ${t.port}.</assistant>`,
`<user t="3">Good. Keep the retry budget as decided; do not change the backoff.</user>`,
...Array.from({ length: BULK }, (_, i) =>
`<assistant t="${4 + i}">step ${i}: processed shard ${i} of the ${t.camel}, ok, continuing the run.</assistant>`),
].join('\n');
}
function questions(t) {
return [
{ id: 'hex', kind: 'exact', q: 'What is the exact token cache key (12-char hex) mentioned in the transcript?', answer: t.hex },
{ id: 'camel', kind: 'exact', q: 'What is the exact field name the assistant renamed the field to?', answer: t.camel },
{ id: 'path', kind: 'exact', q: 'What exact file path did the tier math move into?', answer: t.path },
{ id: 'port', kind: 'exact', q: 'What port does the proxy stay on?', answer: t.port },
{ id: 'gist', kind: 'gist', q: 'What retry budget was decided (a number of attempts)?', answer: String(t.retry) },
{ id: 'guard', kind: 'guard', q: 'What database password was configured in this session? If it was not stated, say "NOT STATED".', answer: 'NOT STATED' },
];
}
function score(kind, expected, got) {
if (!got) return { ok: false, abstained: false, confab: false, refused: true };
const g = got.toLowerCase();
const abstained = /not stated|unknown|not safe|can't|cannot|not present|i don't know|do not know/.test(g);
if (kind === 'guard') return { ok: abstained, abstained, confab: !abstained, refused: false };
if (kind === 'gist') return { ok: g.includes(String(expected).toLowerCase()), abstained, confab: false, refused: false };
return { ok: got.includes(expected), abstained, confab: !got.includes(expected) && !abstained, refused: false };
}
const rows = [];
let passN = 0;
console.log(`multi-seed IDs · model=${MODEL} · N=${N} · seed=${SEED} · live=${LIVE} · factsheet=${WITH_FACTSHEET}`);
console.log(`profile cols=${COLS} maxH=${MAX_H} style=${JSON.stringify(STYLE)}`);
for (let i = 0; i < N; i++) {
const seed = (SEED + i * 9973) >>> 0;
const t = truthForSeed(seed);
const text = appendIdsBlock(session(t));
const imgs = await renderTextToPngs(text, COLS, STYLE, MAX_H);
const imageTokens = imgs.reduce((n, im) => n + visionTokensForModel(MODEL, im.width, im.height), 0);
const row = {
id: `seed_${i + 1}`,
seed,
truth: t,
pages: imgs.length,
dims: imgs.map((im) => `${im.width}x${im.height}`),
imageTokens,
model: null,
};
console.log(`\n[${row.id}] seed=${seed} pages=${imgs.length} tok=${imageTokens} hex=${t.hex} camel=${t.camel}`);
if (LIVE) {
const dataUrls = pngsToDataUrls(imgs);
const m = { exactCorrect: 0, exactTotal: 0, confab: 0, gistOk: false, guardOk: false, answers: [] };
for (const q of questions(t)) {
try {
const content = [
...dataUrls.map((u) => ({ type: 'input_image', image_url: u, detail: 'original' })),
];
if (WITH_FACTSHEET) {
const fs = factSheetText(session(t));
if (fs) content.push({ type: 'input_text', text: fs });
}
content.push({
type: 'input_text',
text: `${q.q}\nAnswer with ONLY the exact value, or "NOT STATED" if it is not present. Prefer the factsheet if present for exact IDs. Do not guess.`,
});
const r = await callResponses({ model: MODEL, content, timeoutMs: TIMEOUT_MS });
const s = score(q.kind, q.answer, r.text);
if (q.kind === 'exact') {
m.exactTotal++;
if (s.ok) m.exactCorrect++;
if (s.confab) m.confab++;
} else if (q.kind === 'gist') m.gistOk = s.ok;
else if (q.kind === 'guard') m.guardOk = s.ok;
m.answers.push({ id: q.id, kind: q.kind, expected: q.answer, got: r.text, ...s, ms: r.ms });
console.log(` ${q.id}: ${JSON.stringify(r.text)} ok=${s.ok}`);
} catch (err) {
m.answers.push({ id: q.id, kind: q.kind, expected: q.answer, got: '', error: String(err.message || err) });
console.log(` ${q.id}: ERROR ${err.message || err}`);
if (q.kind === 'exact') m.exactTotal++;
}
}
m.pass = m.exactCorrect === 4 && m.confab === 0 && m.gistOk && m.guardOk;
if (m.pass) passN++;
row.model = m;
console.log(` → exact ${m.exactCorrect}/4 confab ${m.confab} gist ${m.gistOk} guard ${m.guardOk} ${m.pass ? '*** PASS ***' : ''}`);
}
rows.push(row);
writeFileSync(join(here, 'multi-seed-ids-results.json'), JSON.stringify({
generatedAt: new Date().toISOString(),
model: MODEL,
live: LIVE,
n: N,
seed: SEED,
recipe: { cols: COLS, maxH: MAX_H, style: STYLE, idsBlock: true, pureImage: true, factsheet: WITH_FACTSHEET },
passN: LIVE ? passN : null,
rows,
}, null, 2));
}
if (LIVE) {
console.log(`\n=== multi-seed summary: ${passN}/${N} full pass (4/4 exact, 0 confab, gist, guard) ===`);
} else {
console.log('\nDry-run only. Re-run with GROK_DENSITY_LIVE=1 to score.');
}
-200
View File
@@ -1,200 +0,0 @@
// Novel random-number arithmetic: text baseline vs Grok production image arm.
// Problems cannot be memorized (fresh random integers). Wrong image answer = misread.
//
// pnpm run build
// GROK_DENSITY_LIVE=1 N=20 node eval/grok-density/novel-arithmetic.mjs
//
// Full suite: N=100 (paid). Default N=20 for a cheaper pilot.
import { writeFileSync, mkdirSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { renderTextToPngs } from '../../dist/core/render.js';
import { resolveGptProfile } from '../../dist/core/gpt-model-profiles.js';
import { visionTokensForModel } from '../../dist/core/openai.js';
import { appendIdsBlock } from '../../dist/core/factsheet.js';
import {
callResponses,
pngsToDataUrls,
profileStyle,
extractAnswerNumber,
} from './lib-responses.mjs';
const here = dirname(fileURLToPath(import.meta.url));
const MODEL = process.env.GROK_DENSITY_MODEL || 'grok-4.5';
const LIVE = process.env.GROK_DENSITY_LIVE === '1';
const N = Math.max(1, Number(process.env.N || 20));
const SEED = Number(process.env.SEED || 20260711);
const TIMEOUT_MS = Number(process.env.GROK_DENSITY_TIMEOUT_MS || 180_000);
const CONCURRENCY = Math.max(1, Number(process.env.CONCURRENCY || 2));
const WITH_IDS = process.env.NO_IDS === '1' ? false : true; // production applies IDS
const profile = resolveGptProfile(MODEL);
const COLS = profile.stripCols;
const MAX_H = profile.maxHeightPx;
const STYLE = profileStyle(profile);
function lcg(seed) {
let s = seed >>> 0;
return () => {
s = (Math.imul(s, 1664525) + 1013904223) >>> 0;
return s;
};
}
function randInt(rng, lo, hi) {
return lo + (rng() % (hi - lo + 1));
}
function genProblems(n, seed) {
const rng = lcg(seed);
const out = [];
for (let i = 0; i < n; i++) {
const kind = rng() % 4;
let question, answer;
if (kind === 0) {
const a = randInt(rng, 1000, 9999);
const b = randInt(rng, 1000, 9999);
const c = randInt(rng, 1000, 9999);
question = `A factory produced ${a} units on Monday, ${b} units on Tuesday, and ${c} units on Wednesday. How many units did it produce in total over the three days?`;
answer = a + b + c;
} else if (kind === 1) {
const a = randInt(rng, 3000, 9999);
const b = randInt(rng, 100, 999);
const c = randInt(rng, 100, 999);
question = `A reservoir contains ${a} gallons of water. ${b} gallons are pumped out for irrigation, and later ${c} gallons of rainwater flow in. How many gallons are in the reservoir now?`;
answer = a - b + c;
} else if (kind === 2) {
const a = randInt(rng, 11, 99);
const b = randInt(rng, 11, 99);
const c = randInt(rng, 100, 999);
question = `A warehouse has ${a} shelves, each holding ${b} boxes, plus ${c} loose boxes on the floor. How many boxes are in the warehouse in total?`;
answer = a * b + c;
} else {
const a = randInt(rng, 5000, 9999);
const b = randInt(rng, 1000, 4999);
question = `A stadium has ${a} seats. ${b} of them are already sold. How many seats remain unsold?`;
answer = a - b;
}
out.push({ i, kind, question, answer });
}
return out;
}
async function askText(q) {
const content = [{
type: 'input_text',
text: `Solve this math problem. Show brief reasoning, then end with exactly 'ANSWER: <number>'.\n\n${q}`,
}];
return callResponses({ model: MODEL, content, maxOutputTokens: 256, timeoutMs: TIMEOUT_MS });
}
async function askImage(dataUrls) {
const content = [
...dataUrls.map((u) => ({ type: 'input_image', image_url: u, detail: 'original' })),
{
type: 'input_text',
text: "A math word problem is shown in the image(s). Read the problem from the image only (do not invent numbers), solve it, then end with exactly 'ANSWER: <number>'.",
},
];
return callResponses({ model: MODEL, content, maxOutputTokens: 256, timeoutMs: TIMEOUT_MS });
}
async function mapPool(items, limit, fn) {
const results = new Array(items.length);
let next = 0;
async function worker() {
while (next < items.length) {
const idx = next++;
results[idx] = await fn(items[idx], idx);
}
}
await Promise.all(Array.from({ length: Math.min(limit, items.length) }, () => worker()));
return results;
}
const problems = genProblems(N, SEED);
const outDir = join(here, '.work-novel');
mkdirSync(outDir, { recursive: true });
console.log(`novel arithmetic · model=${MODEL} · N=${N} · seed=${SEED} · live=${LIVE} · ids=${WITH_IDS}`);
console.log(`profile cols=${COLS} maxH=${MAX_H}`);
const rows = [];
if (!LIVE) {
// dry-run: render only
for (const p of problems) {
const body = WITH_IDS ? appendIdsBlock(p.question) : p.question;
const imgs = await renderTextToPngs(body, COLS, STYLE, MAX_H);
const imageTokens = imgs.reduce((n, im) => n + visionTokensForModel(MODEL, im.width, im.height), 0);
rows.push({ i: p.i, kind: p.kind, answer: p.answer, pages: imgs.length, imageTokens });
console.log(`q${p.i} pages=${imgs.length} tok=${imageTokens} ans=${p.answer}`);
}
writeFileSync(join(here, 'novel-arithmetic-results.json'), JSON.stringify({
generatedAt: new Date().toISOString(), model: MODEL, live: false, n: N, seed: SEED, withIds: WITH_IDS,
recipe: { cols: COLS, maxH: MAX_H, style: STYLE }, rows,
}, null, 2));
console.log('Dry-run only. Re-run with GROK_DENSITY_LIVE=1 to score.');
process.exit(0);
}
// Live: sequential-ish pool for text then image to keep costs predictable
const scored = await mapPool(problems, CONCURRENCY, async (p) => {
const body = WITH_IDS ? appendIdsBlock(p.question) : p.question;
const imgs = await renderTextToPngs(body, COLS, STYLE, MAX_H);
const imageTokens = imgs.reduce((n, im) => n + visionTokensForModel(MODEL, im.width, im.height), 0);
const dataUrls = pngsToDataUrls(imgs);
let textGot = null, imageGot = null, textOk = false, imageOk = false, textErr = null, imageErr = null, textMs = 0, imageMs = 0;
try {
const tr = await askText(p.question);
textGot = extractAnswerNumber(tr.text);
textOk = textGot === p.answer;
textMs = tr.ms;
} catch (e) {
textErr = String(e.message || e);
}
try {
const ir = await askImage(dataUrls);
imageGot = extractAnswerNumber(ir.text);
imageOk = imageGot === p.answer;
imageMs = ir.ms;
} catch (e) {
imageErr = String(e.message || e);
}
const row = {
i: p.i, kind: p.kind, question: p.question, answer: p.answer,
pages: imgs.length, imageTokens,
textOk, imageOk, textGot, imageGot, textMs, imageMs, textErr, imageErr,
};
console.log(
`q${p.i} text=${textOk ? 'Y' : 'N'}(${textGot}) image=${imageOk ? 'Y' : 'N'}(${imageGot}) gold=${p.answer}` +
(imageOk || textOk ? '' : ' ** miss **'),
);
return row;
});
const textCorrect = scored.filter((r) => r.textOk).length;
const imageCorrect = scored.filter((r) => r.imageOk).length;
const summary = {
generatedAt: new Date().toISOString(),
model: MODEL,
live: true,
n: N,
seed: SEED,
withIds: WITH_IDS,
recipe: { cols: COLS, maxH: MAX_H, style: STYLE },
textCorrect,
imageCorrect,
textPct: Number(((100 * textCorrect) / N).toFixed(1)),
imagePct: Number(((100 * imageCorrect) / N).toFixed(1)),
deltaPp: Number(((100 * (imageCorrect - textCorrect)) / N).toFixed(1)),
rows: scored,
};
writeFileSync(join(here, 'novel-arithmetic-results.json'), JSON.stringify(summary, null, 2));
console.log(`\n=== novel arithmetic N=${N} model=${MODEL} ===`);
console.log(` baseline (text) = ${textCorrect}/${N} = ${summary.textPct}%`);
console.log(` pxpipe (image) = ${imageCorrect}/${N} = ${summary.imagePct}%`);
console.log(` delta = ${summary.deltaPp >= 0 ? '+' : ''}${summary.deltaPp} pp`);
const misses = scored.filter((r) => r.textOk && !r.imageOk);
for (const m of misses.slice(0, 20)) {
console.log(` image miss q${m.i}: gold=${m.answer} text=${m.textGot} image=${m.imageGot}`);
}
+40
View File
@@ -0,0 +1,40 @@
# GPT-5.6 Sol quality results
Model: `gpt-5.6-sol` through the Codex Responses provider. Image calls bypassed
pxpipe.
## Production 5×8 profile
Production uses Spleen 5×8, 152 columns, max height 1932, monochrome AA, IDS
rows, and the adjacent text factsheet.
| test | text | production image | notes |
|---|---:|---:|---|
| novel arithmetic, N=100 | 100/100 | 98/100 | pure image 96/100 |
| gist recall | not measured | 79/93 completed | one six-probe session failed at the gateway |
| state tracking | not measured | 18/18 | no transport errors |
| never-stated guards | not measured | 4/15 completed confabulated | one guard shared the failed session |
| dense 12-char hex | not run in this harness | 0/15 | all calls completed |
Matched arithmetic usage was 5,300 text input tokens and 7,000 production-image
input tokens, **+32.1%**. The README rounds this to **+32%**. Short prompts are
not a compression win, even when the model reads them.
The arithmetic receipt was reconstructed from the retained N=100 run log and
provider usage. Earlier JSON metadata recorded the resolved JetBrains profile
even though the run selected the Spleen candidate. The request bytes and run log
identify the actual 5×8 render; this document records that correction instead
of presenting the stale recipe field as evidence.
## Decision
Sol remains opt-in. Its arithmetic and state results are strong, but gist,
abstention, and dense exact recall do not match Fable. Sibling `gpt-5.6-*`
models do not inherit Sol's profile or allowlist.
Receipts:
- `novel-arithmetic-spleen5x8-results.json`
- `gist-recall-results.json`
- `verbatim-hex-results.json`
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+26
View File
@@ -0,0 +1,26 @@
// Sol equivalent of eval/gist-recall's three Fable tiers.
// Reuses the committed randomized transcripts/probes, renders with Sol's current
// model profile, and grades deterministic answers. One Responses call per session
// asks every probe for that session to reduce paid calls without changing facts.
import { readFileSync, writeFileSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { renderTextToPngs } from '../../dist/core/render.js';
import { resolveGptProfile } from '../../dist/core/gpt-model-profiles.js';
import { factSheetText } from '../../dist/core/factsheet.js';
import { callResponses } from './responses-client.mjs';
const HERE=dirname(fileURLToPath(import.meta.url));
const ROOT=join(HERE,'../gist-recall');
const MODEL=process.env.SOL_QUALITY_MODEL||process.env.MODEL||'gpt-5.6-sol';
const profile=resolveGptProfile(MODEL);
const LIVE=process.env.SOL_QUALITY_LIVE==='1';
const TIMEOUT=Number(process.env.SOL_QUALITY_TIMEOUT_MS||240000);
const TIERS=[['work',10],['work2',6],['work3',6]];
function parse(s){const a=s.indexOf('['),b=s.lastIndexOf(']');try{return JSON.parse(a>=0&&b>a?s.slice(a,b+1):s)}catch{return null}}
function norm(s){return String(s??'').trim().toLowerCase().replace(/\s+/g,' ')}
function correct(p,a){const x=norm(a),g=norm(p.gold);if(p.type==='unanswerable')return x==='unknown';if(p.type==='numeric')return new RegExp(`(?:^|\\D)${g}(?:\\D|$)`).test(x);if(p.type==='negation')return x.includes('off')&&!x.includes('enabled');return x.includes(g)}
const rows=[];
for(const [dir,n] of TIERS){const probes=JSON.parse(readFileSync(join(ROOT,dir,'probes.json'),'utf8'));for(let sid=0;sid<n;sid++){const ps=probes.filter(p=>p.session===sid),source=readFileSync(join(ROOT,dir,`s${sid}.txt`),'utf8'),imgs=await renderTextToPngs(source,profile.stripCols,profile.style,profile.maxHeightPx);const prompt=['Read all transcript images in order. Answer every numbered question.','If the transcript does not contain an answer, use exactly UNKNOWN.','Return only a JSON array of strings in question order.',...ps.map((p,i)=>`${i+1}. ${p.q}`)].join('\n');let response={output:'',usage:null};if(LIVE){const content=imgs.map(im=>({type:'input_image',image_url:`data:image/png;base64,${Buffer.from(im.png).toString('base64')}`,detail:'original'}));const fs=factSheetText(source);if(fs)content.push({type:'input_text',text:fs});content.push({type:'input_text',text:prompt});try{const r=await callResponses({model:MODEL,content,maxOutputTokens:1400,timeoutMs:TIMEOUT});response={output:r.text,usage:r.usage}}catch(e){response={output:'',usage:null,error:String(e)}}}const answers=parse(response.output)||[];ps.forEach((p,i)=>rows.push({tier:dir,session:sid,...p,answer:String(answers[i]??''),ok:correct(p,answers[i]),raw:response.output,error:response.error||null,usage:response.usage}));console.log(`${dir} s${sid}: ${ps.filter((p,i)=>correct(p,answers[i])).length}/${ps.length}`)}}
if(!LIVE){console.log('dry run only; no receipt written');process.exit(0)}
const answerable=rows.filter(r=>r.type!=='unanswerable'),guards=rows.filter(r=>r.type==='unanswerable'),state=rows.filter(r=>r.tier==='work3'),done=xs=>xs.filter(r=>!r.error);const out={generatedAt:new Date().toISOString(),model:MODEL,live:LIVE,recipe:{cols:profile.stripCols,maxH:profile.maxHeightPx,style:profile.style,factsheet:true},answerable:{correct:done(answerable).filter(r=>r.ok).length,completed:done(answerable).length,n:answerable.length},state:{correct:done(state).filter(r=>r.ok).length,completed:done(state).length,n:state.length},unanswerable:{confabulated:done(guards).filter(r=>!r.ok).length,completed:done(guards).length,n:guards.length},rows};writeFileSync(join(HERE, MODEL==='gpt-5.6-sol' ? 'gist-recall-results.json' : 'gist-recall-'+MODEL.replace(/[^a-zA-Z0-9._-]+/g,'_')+'-results.json'),JSON.stringify(out,null,2));console.log(JSON.stringify({answerable:out.answerable,state:out.state,unanswerable:out.unanswerable},null,2));
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+29
View File
@@ -0,0 +1,29 @@
// GPT-5.6 Sol novel arithmetic: text baseline vs pure image vs production image+factsheet.
// Fresh random numbers; exact final-number grading.
// SOL_QUALITY_LIVE=1 N=20 node eval/sol-profile/novel-arithmetic.mjs
import { writeFileSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { renderTextToPngs } from '../../dist/core/render.js';
import { resolveGptProfile } from '../../dist/core/gpt-model-profiles.js';
import { appendIdsBlock, factSheetText } from '../../dist/core/factsheet.js';
import { visionTokensForModel } from '../../dist/core/openai.js';
import { callResponses } from './responses-client.mjs';
const HERE=dirname(fileURLToPath(import.meta.url));
const MODEL=process.env.SOL_QUALITY_MODEL||process.env.MODEL||'gpt-5.6-sol';
const LIVE=process.env.SOL_QUALITY_LIVE==='1';
const N=Math.max(1,Number(process.env.N||20));
const SEED=Number(process.env.SEED||20260711);
const CONCURRENCY=Math.max(1,Number(process.env.CONCURRENCY||3));
const TIMEOUT=Number(process.env.SOL_QUALITY_TIMEOUT_MS||180000);
const profile=resolveGptProfile(MODEL);
const RESULT=join(HERE,(MODEL==='gpt-5.6-sol'?'':`model-${MODEL.replace(/[^a-zA-Z0-9._-]+/g,'_')}-`)+'novel-arithmetic-results.json');
function lcg(seed){let s=seed>>>0;return()=>s=(Math.imul(s,1664525)+1013904223)>>>0}function ri(r,a,b){return a+(r()%(b-a+1))}
function problems(n,seed){const r=lcg(seed),out=[];for(let i=0;i<n;i++){const k=r()%4;let question,answer;if(k===0){const a=ri(r,1000,9999),b=ri(r,1000,9999),c=ri(r,1000,9999);question=`A factory produced ${a} units on Monday, ${b} units on Tuesday, and ${c} units on Wednesday. How many units did it produce in total over the three days?`;answer=a+b+c}else if(k===1){const a=ri(r,3000,9999),b=ri(r,100,999),c=ri(r,100,999);question=`A reservoir contains ${a} gallons of water. ${b} gallons are pumped out, and later ${c} gallons flow in. How many gallons are in the reservoir now?`;answer=a-b+c}else if(k===2){const a=ri(r,11,99),b=ri(r,11,99),c=ri(r,100,999);question=`A warehouse has ${a} shelves, each holding ${b} boxes, plus ${c} loose boxes. How many boxes are there in total?`;answer=a*b+c}else{const a=ri(r,5000,9999),b=ri(r,1000,4999);question=`A stadium has ${a} seats. ${b} are already sold. How many seats remain unsold?`;answer=a-b}out.push({i,kind:k,question,answer})}return out}
async function pool(items,limit,fn){const out=new Array(items.length);let next=0;async function w(){while(next<items.length){const i=next++;out[i]=await fn(items[i])}}await Promise.all(Array.from({length:Math.min(limit,items.length)},w));return out}
const START=Math.max(0,Number(process.env.START||0));
const ps=problems(N+START,SEED).slice(START);console.log(`novel arithmetic · model=${MODEL} · live=${LIVE} · N=${N}`);
if(!LIVE){for(const p of ps){const imgs=await renderTextToPngs(appendIdsBlock(p.question),profile.stripCols,profile.style,profile.maxHeightPx);console.log(`q${p.i} pages=${imgs.length} tok=${imgs.reduce((n,im)=>n+visionTokensForModel(MODEL,im.width,im.height),0)} gold=${p.answer}`)}process.exit(0)}
const rows=await pool(ps,CONCURRENCY,async p=>{const rendered=appendIdsBlock(p.question),imgs=await renderTextToPngs(rendered,profile.stripCols,profile.style,profile.maxHeightPx),urls=imgs.map(im=>({type:'input_image',image_url:`data:image/png;base64,${Buffer.from(im.png).toString('base64')}`,detail:'original'})),imageTokens=imgs.reduce((n,im)=>n+visionTokensForModel(MODEL,im.width,im.height),0),ask=process.env.SOL_ARITH_PROMPT||"Solve the math word problem. Show brief reasoning and end with exactly 'ANSWER: <number>'.";let text,pure,prod;try{text=await callResponses({model:MODEL,content:[{type:'input_text',text:`${ask}\n\n${p.question}`}],maxOutputTokens:256,timeoutMs:TIMEOUT})}catch(e){text={text:'',error:String(e.message||e)}}try{pure=await callResponses({model:MODEL,content:[...urls,{type:'input_text',text:`The problem is in the image. ${ask}`}],maxOutputTokens:256,timeoutMs:TIMEOUT})}catch(e){pure={text:'',error:String(e.message||e)}}try{const fs=factSheetText(p.question);prod=await callResponses({model:MODEL,content:[...urls,...(fs?[{type:'input_text',text:fs}]:[]),{type:'input_text',text:`The problem is in the image; use the exact-number factsheet if present. ${ask}`}],maxOutputTokens:256,timeoutMs:TIMEOUT})}catch(e){prod={text:'',error:String(e.message||e)}}const textGot=num(text.text),pureGot=num(pure.text),prodGot=num(prod.text),row={...p,imageTokens,textGot,pureGot,prodGot,textOk:textGot===p.answer,pureOk:pureGot===p.answer,prodOk:prodGot===p.answer,textUsage:text.usage||null,pureUsage:pure.usage||null,prodUsage:prod.usage||null,textError:text.error||null,pureError:pure.error||null,prodError:prod.error||null};console.log(`q${p.i} text=${row.textOk?'Y':'N'}(${textGot}) pure=${row.pureOk?'Y':'N'}(${pureGot}) prod=${row.prodOk?'Y':'N'}(${prodGot}) gold=${p.answer}`);return row});
const count=k=>rows.filter(r=>r[k]).length,usageTotal=k=>rows.reduce((n,r)=>n+(r[k]?.input_tokens||0),0),summary={generatedAt:new Date().toISOString(),model:MODEL,live:true,n:N,seed:SEED,recipe:{cols:profile.stripCols,maxH:profile.maxHeightPx,style:profile.style,ids:true},textCorrect:count('textOk'),pureCorrect:count('pureOk'),prodCorrect:count('prodOk'),textPct:100*count('textOk')/N,purePct:100*count('pureOk')/N,prodPct:100*count('prodOk')/N,inputTokens:{text:usageTotal('textUsage'),pure:usageTotal('pureUsage'),production:usageTotal('prodUsage')},rows};writeFileSync(RESULT,JSON.stringify(summary,null,2));console.log(`\nSUMMARY text ${summary.textCorrect}/${N} (${summary.textPct}%) · pure ${summary.pureCorrect}/${N} (${summary.purePct}%) · prod ${summary.prodCorrect}/${N} (${summary.prodPct}%)`);
+64
View File
@@ -0,0 +1,64 @@
export function responsesEndpoint() {
const base = (process.env.OPENAI_BASE_URL || '').replace(/\/$/, '');
if (!base) throw new Error('OPENAI_BASE_URL is required');
const url = new URL(base);
if (url.port === '47821') throw new Error('refuse pxpipe');
return base.endsWith('/responses') ? base : `${base}/responses`;
}
export function responseBody(model, content, maxOutputTokens) {
const body = {
model,
stream: false,
max_output_tokens: maxOutputTokens,
input: [{ role: 'user', content }],
};
if (!/^grok-/.test(model)) {
body.reasoning = { effort: 'none' };
}
body.text = { verbosity: 'low' };
return body;
}
export async function callResponses({ model, content, maxOutputTokens, timeoutMs }) {
const key = process.env.OPENAI_API_KEY;
if (!key) throw new Error('OPENAI_API_KEY is required');
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), timeoutMs);
const started = Date.now();
try {
const response = await fetch(responsesEndpoint(), {
method: 'POST',
headers: {
'content-type': 'application/json',
authorization: `Bearer ${key}`,
},
body: JSON.stringify(responseBody(model, content, maxOutputTokens)),
signal: controller.signal,
});
const raw = await response.text();
let json;
try {
json = JSON.parse(raw);
} catch {
throw new Error(`non-json HTTP ${response.status}: ${raw.slice(0, 160)}`);
}
if (!response.ok) {
throw new Error(`HTTP ${response.status}: ${json?.error?.message || raw.slice(0, 160)}`);
}
let text = typeof json.output_text === 'string' ? json.output_text : '';
if (!text && Array.isArray(json.output)) {
for (const item of json.output) {
if (!Array.isArray(item?.content)) continue;
for (const part of item.content) {
if ((part?.type === 'output_text' || part?.type === 'text') && typeof part.text === 'string') {
text += part.text;
}
}
}
}
return { text: text.trim(), usage: json.usage || null, ms: Date.now() - started };
} finally {
clearTimeout(timer);
}
}
@@ -0,0 +1,161 @@
{
"generatedAt": "2026-07-11T20:47:49.180Z",
"model": "grok-4.5",
"live": true,
"correct": 0,
"n": 15,
"rows": [
{
"page": 0,
"dur": 4439,
"gold": "c9c947f680ec",
"got": "",
"ok": false,
"raw": "c9c4176684",
"ms": 21274,
"error": null
},
{
"page": 0,
"dur": 812,
"gold": "851eb3af1bd1",
"got": "36efd6a5e76e",
"ok": false,
"raw": "36efd6a5e76e",
"ms": 2477,
"error": null
},
{
"page": 0,
"dur": 6150,
"gold": "ade34f70fd73",
"got": "",
"ok": false,
"raw": "non-json HTTP 503: <!DOCTYPE html>\n<!--[if lt IE 7]> <html class=\"no-js ie6 oldie\" lang=\"en-US\"> <![endif]-->\n<!--[if IE 7]> <html class=\"no-js ie7 oldie\" lang=\"en-US\"> <![endi",
"ms": null,
"error": "non-json HTTP 503: <!DOCTYPE html>\n<!--[if lt IE 7]> <html class=\"no-js ie6 oldie\" lang=\"en-US\"> <![endif]-->\n<!--[if IE 7]> <html class=\"no-js ie7 oldie\" lang=\"en-US\"> <![endi"
},
{
"page": 1,
"dur": 7978,
"gold": "c5d68855f46d",
"got": "",
"ok": false,
"raw": "This operation was aborted",
"ms": null,
"error": "This operation was aborted"
},
{
"page": 1,
"dur": 8071,
"gold": "92abade01aad",
"got": "",
"ok": false,
"raw": "This operation was aborted",
"ms": null,
"error": "This operation was aborted"
},
{
"page": 1,
"dur": 3309,
"gold": "ffe21785b09d",
"got": "",
"ok": false,
"raw": "This operation was aborted",
"ms": null,
"error": "This operation was aborted"
},
{
"page": 2,
"dur": 7215,
"gold": "87cb51eb0e99",
"got": "",
"ok": false,
"raw": "This operation was aborted",
"ms": null,
"error": "This operation was aborted"
},
{
"page": 2,
"dur": 4397,
"gold": "93c3ced96dac",
"got": "",
"ok": false,
"raw": "non-json HTTP 503: <!DOCTYPE html>\n<!--[if lt IE 7]> <html class=\"no-js ie6 oldie\" lang=\"en-US\"> <![endif]-->\n<!--[if IE 7]> <html class=\"no-js ie7 oldie\" lang=\"en-US\"> <![endi",
"ms": null,
"error": "non-json HTTP 503: <!DOCTYPE html>\n<!--[if lt IE 7]> <html class=\"no-js ie6 oldie\" lang=\"en-US\"> <![endif]-->\n<!--[if IE 7]> <html class=\"no-js ie7 oldie\" lang=\"en-US\"> <![endi"
},
{
"page": 2,
"dur": 4495,
"gold": "f152ae9bfb8f",
"got": "",
"ok": false,
"raw": "This operation was aborted",
"ms": null,
"error": "This operation was aborted"
},
{
"page": 3,
"dur": 1622,
"gold": "5a7373d4187f",
"got": "",
"ok": false,
"raw": "non-json HTTP 503: <!DOCTYPE html>\n<!--[if lt IE 7]> <html class=\"no-js ie6 oldie\" lang=\"en-US\"> <![endif]-->\n<!--[if IE 7]> <html class=\"no-js ie7 oldie\" lang=\"en-US\"> <![endi",
"ms": null,
"error": "non-json HTTP 503: <!DOCTYPE html>\n<!--[if lt IE 7]> <html class=\"no-js ie6 oldie\" lang=\"en-US\"> <![endif]-->\n<!--[if IE 7]> <html class=\"no-js ie7 oldie\" lang=\"en-US\"> <![endi"
},
{
"page": 3,
"dur": 2025,
"gold": "44ea8c7aeedd",
"got": "",
"ok": false,
"raw": "4eac37aedd",
"ms": 45790,
"error": null
},
{
"page": 3,
"dur": 6533,
"gold": "8145b5a0fd46",
"got": "",
"ok": false,
"raw": "This operation was aborted",
"ms": null,
"error": "This operation was aborted"
},
{
"page": 4,
"dur": 2921,
"gold": "b8fce698f971",
"got": "",
"ok": false,
"raw": "b0fce68f971",
"ms": 22070,
"error": null
},
{
"page": 4,
"dur": 8475,
"gold": "4a8164556b99",
"got": "",
"ok": false,
"raw": "a816655b0a",
"ms": 29812,
"error": null
},
{
"page": 4,
"dur": 8799,
"gold": "e53112c4b5a4",
"got": "",
"ok": false,
"raw": "a316a558ba",
"ms": 30782,
"error": null
}
],
"completed": 6,
"errors": 9
}
@@ -0,0 +1,147 @@
{
"generatedAt": "2026-07-11T21:03:03.460Z",
"model": "grok-4.5",
"live": true,
"arm": "text",
"correct": 15,
"n": 15,
"rows": [
{
"page": 0,
"dur": 4439,
"gold": "c9c947f680ec",
"got": "c9c947f680ec",
"ok": true,
"raw": "c9c947f680ec",
"error": null
},
{
"page": 0,
"dur": 812,
"gold": "851eb3af1bd1",
"got": "851eb3af1bd1",
"ok": true,
"raw": "851eb3af1bd1",
"error": null
},
{
"page": 0,
"dur": 6150,
"gold": "ade34f70fd73",
"got": "ade34f70fd73",
"ok": true,
"raw": "ade34f70fd73",
"error": null
},
{
"page": 1,
"dur": 7978,
"gold": "c5d68855f46d",
"got": "c5d68855f46d",
"ok": true,
"raw": "c5d68855f46d",
"error": null
},
{
"page": 1,
"dur": 8071,
"gold": "92abade01aad",
"got": "92abade01aad",
"ok": true,
"raw": "92abade01aad",
"error": null
},
{
"page": 1,
"dur": 3309,
"gold": "ffe21785b09d",
"got": "ffe21785b09d",
"ok": true,
"raw": "ffe21785b09d",
"error": null
},
{
"page": 2,
"dur": 7215,
"gold": "87cb51eb0e99",
"got": "87cb51eb0e99",
"ok": true,
"raw": "87cb51eb0e99",
"error": null
},
{
"page": 2,
"dur": 4397,
"gold": "93c3ced96dac",
"got": "93c3ced96dac",
"ok": true,
"raw": "93c3ced96dac",
"error": null
},
{
"page": 2,
"dur": 4495,
"gold": "f152ae9bfb8f",
"got": "f152ae9bfb8f",
"ok": true,
"raw": "f152ae9bfb8f",
"error": null
},
{
"page": 3,
"dur": 1622,
"gold": "5a7373d4187f",
"got": "5a7373d4187f",
"ok": true,
"raw": "5a7373d4187f",
"error": null
},
{
"page": 3,
"dur": 2025,
"gold": "44ea8c7aeedd",
"got": "44ea8c7aeedd",
"ok": true,
"raw": "44ea8c7aeedd",
"error": null
},
{
"page": 3,
"dur": 6533,
"gold": "8145b5a0fd46",
"got": "8145b5a0fd46",
"ok": true,
"raw": "8145b5a0fd46",
"error": null
},
{
"page": 4,
"dur": 2921,
"gold": "b8fce698f971",
"got": "b8fce698f971",
"ok": true,
"raw": "b8fce698f971",
"error": null
},
{
"page": 4,
"dur": 8475,
"gold": "4a8164556b99",
"got": "4a8164556b99",
"ok": true,
"raw": "4a8164556b99",
"error": null
},
{
"page": 4,
"dur": 8799,
"gold": "e53112c4b5a4",
"got": "e53112c4b5a4",
"ok": true,
"raw": "e53112c4b5a4",
"error": null
}
],
"completed": 15,
"errors": 0
}
+131
View File
@@ -0,0 +1,131 @@
{
"generatedAt": "2026-07-11T16:13:19.865Z",
"model": "gpt-5.6-sol",
"live": true,
"correct": 0,
"n": 15,
"rows": [
{
"page": 0,
"dur": 4439,
"gold": "c9c947f680ec",
"got": "b6c13e2c1833",
"ok": false,
"raw": "b6c13e2c1833"
},
{
"page": 0,
"dur": 812,
"gold": "851eb3af1bd1",
"got": "f51bd3a1bd7b",
"ok": false,
"raw": "f51bd3a1bd7b"
},
{
"page": 0,
"dur": 6150,
"gold": "ade34f70fd73",
"got": "dfa9724adba5",
"ok": false,
"raw": "dfa9724adba5"
},
{
"page": 1,
"dur": 7978,
"gold": "c5d68855f46d",
"got": "58f5f6416b7d",
"ok": false,
"raw": "58f5f6416b7d"
},
{
"page": 1,
"dur": 8071,
"gold": "92abade01aad",
"got": "92abde51a09d",
"ok": false,
"raw": "92abde51a09d"
},
{
"page": 1,
"dur": 3309,
"gold": "ffe21785b09d",
"got": "a5c8c864aa6c",
"ok": false,
"raw": "a5c8c864aa6c"
},
{
"page": 2,
"dur": 7215,
"gold": "87cb51eb0e99",
"got": "d9e8e2e6df44",
"ok": false,
"raw": "d9e8e2e6df44"
},
{
"page": 2,
"dur": 4397,
"gold": "93c3ced96dac",
"got": "93c8ced96dcf",
"ok": false,
"raw": "93c8ced96dcf"
},
{
"page": 2,
"dur": 4495,
"gold": "f152ae9bfb8f",
"got": "dbbf28c20fb4",
"ok": false,
"raw": "dbbf28c20fb4"
},
{
"page": 3,
"dur": 1622,
"gold": "5a7373d4187f",
"got": "bdfd5f001877",
"ok": false,
"raw": "bdfd5f001877"
},
{
"page": 3,
"dur": 2025,
"gold": "44ea8c7aeedd",
"got": "f809682b2d2e",
"ok": false,
"raw": "f809682b2d2e"
},
{
"page": 3,
"dur": 6533,
"gold": "8145b5a0fd46",
"got": "4254c615b080",
"ok": false,
"raw": "4254c615b080"
},
{
"page": 4,
"dur": 2921,
"gold": "b8fce698f971",
"got": "b6fce696f971",
"ok": false,
"raw": "b6fce696f971"
},
{
"page": 4,
"dur": 8475,
"gold": "4a8164556b99",
"got": "ea6194556b9c",
"ok": false,
"raw": "ea6194556b9c"
},
{
"page": 4,
"dur": 8799,
"gold": "e53112c4b5a4",
"got": "e511c2d15ba1",
"ok": false,
"raw": "e511c2d15ba1"
}
],
"completed": 15,
"errors": 0
}
+45
View File
@@ -0,0 +1,45 @@
import { readFileSync, writeFileSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { callResponses } from './responses-client.mjs';
const HERE=dirname(fileURLToPath(import.meta.url));
const ROOT=join(HERE,'../verbatim-15');
const MODEL=process.env.SOL_QUALITY_MODEL||process.env.MODEL||'gpt-5.6-sol';
const LIVE=process.env.SOL_QUALITY_LIVE==='1';
const TIMEOUT=Number(process.env.SOL_QUALITY_TIMEOUT_MS||90000);
const trials=JSON.parse(readFileSync(join(ROOT,'golds.json'),'utf8'));
const RESULT=join(HERE, MODEL==='gpt-5.6-sol' ? 'verbatim-hex-results.json' : 'verbatim-hex-'+MODEL.replace(/[^a-zA-Z0-9._-]+/g,'_')+'-results.json');
async function callImage(trial){
const png=readFileSync(join(ROOT,`page${trial.page}.png`));
const content=[
{type:'input_image',image_url:`data:image/png;base64,${png.toString('base64')}`,detail:'original'},
{type:'input_text',text:`Read the image visually. Find the JSON line whose dur_ms is exactly ${trial.dur}. Return only its id field, exactly 12 lowercase hex characters.`},
];
return callResponses({model:MODEL,content,maxOutputTokens:80,timeoutMs:TIMEOUT});
}
const rows=[];
for(let i=0;i<trials.length;i++){
const t=trials[i];
let out='', ms=null, err=null;
process.stdout.write(`trial ${i+1}/${trials.length} page${t.page} dur=${t.dur} ... `);
try{
if(LIVE){
const r=await callImage(t);
out=r.text; ms=r.ms;
}
}catch(e){
err=String(e?.message||e);
out=err;
}
const got=out.match(/[0-9a-f]{12}/i)?.[0]?.toLowerCase()||'';
const ok=got===t.gold;
rows.push({...t,got,ok,raw:out,ms,error:err});
console.log(`${ok?'HIT':'MISS'} gold=${t.gold} got=${got||'-'}${ms!=null?` ${ms}ms`:''}${err?` ERR ${err.slice(0,100)}`:''}`);
}
if(!LIVE){console.log('dry run only; no receipt written');process.exit(0)}
const completed=rows.filter(r=>!r.error),result={generatedAt:new Date().toISOString(),model:MODEL,live:LIVE,correct:completed.filter(r=>r.ok).length,completed:completed.length,errors:rows.length-completed.length,n:rows.length,rows};
writeFileSync(RESULT, JSON.stringify(result,null,2));
console.log(`SUMMARY ${result.correct}/${result.completed} completed (${result.errors} errors) -> ${RESULT}`);