eval: complete Kimi K3 quality suite

This commit is contained in:
Steven Chong
2026-07-17 20:02:32 -04:00
parent efef8cf0d8
commit f3fb2a8a21
4 changed files with 2169 additions and 0 deletions
+8
View File
@@ -177,16 +177,20 @@ used for these novel-arithmetic rows.
| gist recall A/B (decisions, values, paths, names, negations; distractors; 15k45k char sessions) | Fable 5 | 98/arm | **98/98** |
| same gist corpus, production images + factsheet | `gpt-5.6-sol` | 98 | **83/98** |
| same gist corpus, production images + factsheet | `grok-4.5` | 98 | **83/98** |
| same gist corpus, production images + factsheet | `moonshotai/kimi-k3` | 98 | **84/98** |
| state tracking (value mutated 3×, final/first/count) | Fable 5 | 18/arm | **18/18** |
| same state-tracking corpus | `gpt-5.6-sol` | 18 | **17/18** |
| same state-tracking corpus | `grok-4.5` | 18 | **13/18** |
| same state-tracking corpus | `moonshotai/kimi-k3` | 18 | **15/18** |
| confabulation on never-stated facts (lower is better) | Fable 5 | 16/arm | **0/16** |
| same never-stated probes (lower is better) | `gpt-5.6-sol` | 16 | **4/16** |
| same never-stated probes (lower is better) | `grok-4.5` | 16 | **0/16** |
| same never-stated probes (lower is better) | `moonshotai/kimi-k3` | 16 | **1/16** |
| verbatim 12-char hex, dense render | Opus | 15 | **0/15** |
| verbatim 12-char hex, dense render | Fable 5 | 15 | **13/15** |
| verbatim 12-char hex, same dense pages | `gpt-5.6-sol` | 15 | **0/15** |
| verbatim 12-char hex, same dense pages | `grok-4.5` | 15 | **0/15** |
| verbatim 12-char hex, same dense pages | `moonshotai/kimi-k3` | 15 | **0/15** |
**Harness split:** Fable/Opus quality and SWE-bench rows use **Claude**; Sol and Grok quality use
**Codexs Responses provider** (`OPENAI_BASE_URL`). Kimi K3
@@ -195,6 +199,10 @@ Cloudflare Messages bridge — see the
[`K3 receipt`](eval/sol-profile/model-moonshotai_kimi-k3-novel-arithmetic-results.json) and
[`eval/grok-density/QUALITY_SUITE.md`](eval/grok-density/QUALITY_SUITE.md).
K3 semantic and exact-recall receipts:
[`gist/state/guards`](eval/sol-profile/gist-recall-moonshotai_kimi-k3-results.json) and
[`dense hex`](eval/sol-profile/verbatim-hex-moonshotai_kimi-k3-results.json).
Sol receipts: [`eval/sol-profile/QUALITY_RESULTS.md`](eval/sol-profile/QUALITY_RESULTS.md).
Grok receipts: [`eval/grok-density/QUALITY_RESULTS.md`](eval/grok-density/QUALITY_RESULTS.md).
SWE-bench is not copied to Sol: its runner is Claude Code/Fable-specific
@@ -0,0 +1,23 @@
# Kimi K3 quality results
Model: `moonshotai/kimi-k3` through pxpipe's Anthropic Messages to Cloudflare
Chat Completions bridge. The run used the generic GPT production profile:
Spleen 5x8, 152 columns, max height 1932, and the adjacent text factsheet.
| test | production image | notes |
|---|---:|---|
| novel arithmetic, N=100 | 79/100 | all calls completed |
| gist recall | 84/98 | all sessions completed |
| state tracking | 15/18 | subset of the gist corpus |
| never-stated guards | 1/16 confabulated | lower is better |
| dense 12-char hex | 0/15 | all calls completed after transient retries |
The semantic and exact-recall runs executed on the remote K3-configured proxy;
the existing process was not restarted or modified. Dense hex used a 16,000
token output cap because K3 performs mandatory reasoning.
Receipts:
- `model-moonshotai_kimi-k3-novel-arithmetic-results.json`
- `gist-recall-moonshotai_kimi-k3-results.json`
- `verbatim-hex-moonshotai_kimi-k3-results.json`
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,161 @@
{
"generatedAt": "2026-07-18T00:01:42.799Z",
"model": "moonshotai/kimi-k3",
"live": true,
"correct": 0,
"completed": 15,
"errors": 0,
"n": 15,
"rows": [
{
"page": 0,
"dur": 4439,
"gold": "c9c947f680ec",
"got": "",
"ok": false,
"raw": "",
"ms": 300469,
"error": null
},
{
"page": 0,
"dur": 812,
"gold": "851eb3af1bd1",
"got": "",
"ok": false,
"raw": "",
"ms": 38209,
"error": null
},
{
"page": 0,
"dur": 6150,
"gold": "ade34f70fd73",
"got": "a3f9c0e6b1d8",
"ok": false,
"raw": "a3f9c0e6b1d8",
"ms": 282380,
"error": null
},
{
"page": 1,
"dur": 7978,
"gold": "c5d68855f46d",
"got": "",
"ok": false,
"raw": "",
"ms": 300445,
"error": null
},
{
"page": 1,
"dur": 8071,
"gold": "92abade01aad",
"got": "",
"ok": false,
"raw": "",
"ms": 187706,
"error": null
},
{
"page": 1,
"dur": 3309,
"gold": "ffe21785b09d",
"got": "3d9c7e1a5b28",
"ok": false,
"raw": "3d9c7e1a5b28",
"ms": 57104,
"error": null
},
{
"page": 2,
"dur": 7215,
"gold": "87cb51eb0e99",
"got": "",
"ok": false,
"raw": "",
"ms": 300247,
"error": null
},
{
"page": 2,
"dur": 4397,
"gold": "93c3ced96dac",
"got": "92e9dc498842",
"ok": false,
"raw": "92e9dc498842",
"ms": 14112,
"error": null
},
{
"page": 2,
"dur": 4495,
"gold": "f152ae9bfb8f",
"got": "",
"ok": false,
"raw": "",
"ms": 300300,
"error": null
},
{
"page": 3,
"dur": 1622,
"gold": "5a7373d4187f",
"got": "e9a32a7a7980",
"ok": false,
"raw": "e9a32a7a7980",
"ms": 132304,
"error": null
},
{
"page": 3,
"dur": 2025,
"gold": "44ea8c7aeedd",
"got": "",
"ok": false,
"raw": "",
"ms": 283222,
"error": null
},
{
"page": 3,
"dur": 6533,
"gold": "8145b5a0fd46",
"got": "",
"ok": false,
"raw": "",
"ms": 300748,
"error": null
},
{
"page": 4,
"dur": 2921,
"gold": "b8fce698f971",
"got": "7a9f4d2c8b16",
"ok": false,
"raw": "7a9f4d2c8b16",
"ms": 228213,
"error": null
},
{
"page": 4,
"dur": 8475,
"gold": "4a8164556b99",
"got": "bb9adff44ffe",
"ok": false,
"raw": "bb9adff44ffe",
"ms": 262684,
"error": null
},
{
"page": 4,
"dur": 8799,
"gold": "e53112c4b5a4",
"got": "",
"ok": false,
"raw": "",
"ms": 300454,
"error": null
}
]
}