Submit gpt-5.6-terra and gpt-5.6-luna sweeps (20 runs, p001/p003) - #220
Conversation
📝 WalkthroughWalkthroughAdded 20 SenseBench Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Comment |
Full lexen-v1 runs for both new OpenAI mid- and low-tier models across all five
reasoning efforts the API accepts (none, low, medium, high, xhigh) on prompts
p001 and p003.
gpt-5.6-terra p001: 0.9120 / 0.9350 / 0.9368 / 0.9418 / 0.9422
gpt-5.6-terra p003: 0.9105 / 0.9323 / 0.9253 / 0.9309 / 0.9395
gpt-5.6-luna p001: 0.8772 / 0.8834 / 0.9025 / 0.9161 / 0.9220
gpt-5.6-luna p003: 0.8780 / 0.8883 / 0.8980 / 0.9111 / 0.9126
The reasoning-effort value "max" is omitted: the OpenAI API rejects it for both
models ("Supported values are: 'none', 'low', 'medium', 'high', and 'xhigh'"),
even though the docs model page lists it for the gpt-5.6 family.
Repeated runs of the same configuration vary by roughly +/-0.7pp (terra low
p003 measured 0.9323 and 0.9251 on two identical runs), so sub-1pp differences
within an effort ladder should not be read as real.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 20 new terra and luna runs move the site totals, which the machine-checked answer block states verbatim: 198 -> 218 verified runs, 63 -> 65 models, and the latest run date 25 -> 31 July 2026. Every other figure was recomputed and is unchanged. GPT-5.5 at xhigh on p001 still leads at 95.60% (95% CI 95.00-96.17, 4,647/4,861), and the McNemar tests were re-run rather than carried over: Claude Fable 5 remains the only model indistinguishable from it (p = 0.20), while GPT-5.6 Sol (p = 0.019), Gemini 3.1 Pro (p = 0.031) and Claude Opus 5 (p = 0.004) stay significantly below. The best of the new runs, GPT-5.6 Terra at 94.22%, is well below the leader (p < 0.001). Names GPT-5.6 Sol explicitly where the block previously said "GPT-5.6": the family now has three models on the board, so the bare version is ambiguous. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
bc57cb4 to
16ae2bc
Compare
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 7bafad0f-cde9-4ae5-ab1a-89e81dc5514e
⛔ Files ignored due to path filters (20)
results/gpt-5.6-luna-high-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-high-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-low-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-low-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-medium-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-medium-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-none-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-none-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-xhigh-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-luna-xhigh-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-high-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-high-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-low-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-low-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-medium-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-medium-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-none-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-none-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-xhigh-reasoning-p001-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gzresults/gpt-5.6-terra-xhigh-reasoning-p003-lexen-v1-20260731/calls.jsonl.gzis excluded by!**/*.gz
📒 Files selected for processing (41)
results/gpt-5.6-luna-high-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-high-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-high-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-high-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-low-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-low-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-low-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-low-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-medium-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-medium-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-medium-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-medium-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-none-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-none-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-none-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-none-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-xhigh-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-xhigh-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-luna-xhigh-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-luna-xhigh-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-high-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-high-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-high-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-high-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-low-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-low-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-low-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-low-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-medium-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-medium-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-medium-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-medium-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-none-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-none-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-none-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-none-reasoning-p003-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-xhigh-reasoning-p001-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-xhigh-reasoning-p001-lexen-v1-20260731/run.jsonresults/gpt-5.6-terra-xhigh-reasoning-p003-lexen-v1-20260731/predictions.jsonlresults/gpt-5.6-terra-xhigh-reasoning-p003-lexen-v1-20260731/run.jsonsrc/sensebench/site/templates/index.html.j2
📜 Review details
⏰ Context from checks skipped due to timeout. (3)
- GitHub Check: build
- GitHub Check: checks
- GitHub Check: links
🧰 Additional context used
🪛 OpenGrep (1.26.0)
results/gpt-5.6-terra-high-reasoning-p003-lexen-v1-20260731/run.json
[ERROR] 61-61: Possible credit card number (PAN) detected in source code. Credit card numbers should never be hardcoded or stored in source files. Use a secrets manager or tokenization service instead.
(coderabbit.pii.credit-card-number)
results/gpt-5.6-terra-high-reasoning-p001-lexen-v1-20260731/run.json
[ERROR] 80-80: Possible credit card number (PAN) detected in source code. Credit card numbers should never be hardcoded or stored in source files. Use a secrets manager or tokenization service instead.
(coderabbit.pii.credit-card-number)
results/gpt-5.6-terra-low-reasoning-p001-lexen-v1-20260731/run.json
[ERROR] 80-80: Possible credit card number (PAN) detected in source code. Credit card numbers should never be hardcoded or stored in source files. Use a secrets manager or tokenization service instead.
(coderabbit.pii.credit-card-number)
results/gpt-5.6-luna-low-reasoning-p003-lexen-v1-20260731/run.json
[ERROR] 61-61: Possible credit card number (PAN) detected in source code. Credit card numbers should never be hardcoded or stored in source files. Use a secrets manager or tokenization service instead.
(coderabbit.pii.credit-card-number)
results/gpt-5.6-terra-medium-reasoning-p001-lexen-v1-20260731/run.json
[ERROR] 61-61: Possible credit card number (PAN) detected in source code. Credit card numbers should never be hardcoded or stored in source files. Use a secrets manager or tokenization service instead.
(coderabbit.pii.credit-card-number)
🔇 Additional comments (16)
src/sensebench/site/templates/index.html.j2 (1)
45-51: LGTM!Also applies to: 60-60
results/gpt-5.6-luna-medium-reasoning-p001-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-medium-reasoning-p003-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-none-reasoning-p001-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-none-reasoning-p003-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-xhigh-reasoning-p001-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-xhigh-reasoning-p003-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-terra-medium-reasoning-p003-lexen-v1-20260731/run.json (1)
1-5: LGTM!Also applies to: 11-16, 18-36, 39-52, 54-62, 64-88
results/gpt-5.6-terra-none-reasoning-p001-lexen-v1-20260731/run.json (1)
1-5: LGTM!Also applies to: 11-16, 18-36, 39-52, 54-62, 64-88
results/gpt-5.6-terra-none-reasoning-p003-lexen-v1-20260731/run.json (1)
1-5: LGTM!Also applies to: 11-16, 18-36, 39-52, 54-62, 64-88
results/gpt-5.6-terra-xhigh-reasoning-p001-lexen-v1-20260731/run.json (1)
1-5: LGTM!Also applies to: 11-16, 18-36, 39-52, 54-62, 64-88
results/gpt-5.6-terra-xhigh-reasoning-p003-lexen-v1-20260731/run.json (1)
1-5: LGTM!Also applies to: 11-16, 18-36, 39-52, 54-62, 64-88
results/gpt-5.6-luna-high-reasoning-p001-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-high-reasoning-p003-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-low-reasoning-p001-lexen-v1-20260731/run.json (1)
1-89: LGTM!results/gpt-5.6-luna-low-reasoning-p003-lexen-v1-20260731/run.json (1)
1-89: LGTM!
| { | ||
| "schema_version": "sensebench-run-v2", | ||
| "run_id": "gpt-5.6-terra-high-reasoning-p001-lexen-v1-20260731", |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift
Add the required call-record artifacts.
The PR stack adds no calls.jsonl.gz files for the 20 new run directories. src/sensebench/runs/loaders.py:53-71 loads that file without a fallback. Loading and verification will fail before leaderboard ingestion.
results/gpt-5.6-terra-high-reasoning-p001-lexen-v1-20260731/run.json#L1-L3: Add the completecalls.jsonl.gzartifact.results/gpt-5.6-terra-high-reasoning-p003-lexen-v1-20260731/run.json#L1-L3: Add the completecalls.jsonl.gzartifact.results/gpt-5.6-terra-low-reasoning-p001-lexen-v1-20260731/run.json#L1-L3: Add the completecalls.jsonl.gzartifact.results/gpt-5.6-terra-low-reasoning-p003-lexen-v1-20260731/run.json#L1-L3: Add the completecalls.jsonl.gzartifact.results/gpt-5.6-terra-medium-reasoning-p001-lexen-v1-20260731/run.json#L1-L3: Add the completecalls.jsonl.gzartifact.
📍 Affects 5 files
results/gpt-5.6-terra-high-reasoning-p001-lexen-v1-20260731/run.json#L1-L3(this comment)results/gpt-5.6-terra-high-reasoning-p003-lexen-v1-20260731/run.json#L1-L3results/gpt-5.6-terra-low-reasoning-p001-lexen-v1-20260731/run.json#L1-L3results/gpt-5.6-terra-low-reasoning-p003-lexen-v1-20260731/run.json#L1-L3results/gpt-5.6-terra-medium-reasoning-p001-lexen-v1-20260731/run.json#L1-L3
Full
lexen-v1runs for OpenAI's two new mid- and low-tier models, across every reasoning effort the API accepts, on promptsp001andp003. 20 runs, all 4,861 items each, all passingsensebench verify.Accuracy
Cost per full run ranges $4.39–$10.35 for terra and $0.44–$1.23 for luna.
Notes
maxeffort is absent by necessity. The API rejects it for both models —Unsupported value: 'reasoning_effort' does not support 'max' with this model. Supported values are: 'none', 'low', 'medium', 'high', and 'xhigh'— although the docs model page listsmaxfor the gpt-5.6 family. This appears to apply togpt-5.6-solonly.Run-to-run variance is about ±0.7pp. Two identical runs of terra
lowon p003 produced 0.9323 and 0.9251. Differences under ~1pp within an effort ladder (e.g. terra p001 high 0.9418 vs xhigh 0.9422) should not be read as real effects. The larger moves — terranone→lowon p001 at +2.3pp, and luna's full climb at +4.5pp — are well outside that band.Luna is a notable price/quality point. Against the existing cheap OpenAI runs in
results/, luna occupies the whole upper Pareto frontier: its best (0.9220 at $1.23) beatsgpt-5-minihigh (0.9072 at $3.90) by +1.5pp at 3.2× lower cost, and lunamedium(0.9025 at $0.85) matchesgpt-5-minimedium (0.9054 at $2.33) for 2.7× less. Worth noting thatgpt-5-nanohas cheaper tokens than luna ($0.05/$0.40 vs $0.20/$1.20) yet costs more at comparable accuracy, because it emits 11.5M output tokens at high effort against luna's 534K at xhigh.Runner:
github:vassiliphilippov. Sampling:--max-tokens 32768, default temperature, unseeded.