ProgramBench task detail

psampaz__go-mod-outdated.bb79367

Task-level results for this Codex /goal scaffold. ProgramBench baseline context is cached from the official task page; Codex scored tests are after active-branch and ignored-test filtering.

285generated behavioral tests
98.2%official best score
83.9%Codex best score
2Codex result rows

Codex Results by Model

#ModelRunAgentModeScoreEvalTestsEst. costCallsWallEvidence
1 GPT 5.5 (xhigh) 20260517T094700Z Codex /goal Mini-SWE-compatible no internet 83.9% ok 239/285 $2.13 45 0.14h manifest · eval json · eval summary · usage audit
2 GPT 5.5 (xhigh) 20260518T232006Z Codex /goal Paper prompt + /goal no internet 70.2% ok 200/285 n/a 26 0.11h not exported

Why Scores Differ

Official ProgramBench rows are the public mini-SWE-agent submissions. GoalBench rows are separate Codex /goal submissions. Same model label does not mean the same agent, prompt, tool loop, or generated implementation. A GoalBench row is resolved only when the ProgramBench-scored pass rate is exactly 100%.

GPT 5.5 (xhigh) · 20260517T094700Z

83.9% from 239/285 ProgramBench-scored tests. Raw public eval statuses: failure: 49, passed: 288, skipped: 5.

Agent trace summary unavailable.

Likely miss class: behavioral mismatch, exact version/output mismatch.

  • 27501e100bd6 tests.test_complex_scenarios.test_local_replace (failure)
  • 27501e100bd6 tests.test_error_handling.test_unexpected_eof (failure)
  • 27501e100bd6 tests.test_error_handling.test_invalid_json_syntax (failure)
  • 27501e100bd6 tests.test_replace_handling.test_replace_uses_replaced_version (failure)
  • 27501e100bd6 tests.test_replace_handling.test_replace_without_update (failure)
  • 27501e100bd6 tests.test_replace_handling.test_replace_with_update (failure)

manifest · eval json · eval summary · usage audit

GPT 5.5 (xhigh) · 20260518T232006Z

Public eval evidence has not been exported for this row yet.

Official ProgramBench Results by Model

#ModelProviderScoreCostCalls
1 Claude Opus 4.6 Anthropic 98.2% $1.18 85
2 GPT 5.5 (high) OpenAI 90.5% $1.93 28
3 Claude Opus 4.7 (xhigh) Anthropic 87.0% $3.05 96
4 Gemini 3 Flash Google 86.7% $0.14 55
5 GPT 5.5 (xhigh) OpenAI 82.5% $2.08 32
6 GPT 5.5 OpenAI 82.5% $0.85 15
7 GPT 5.4 OpenAI 80.4% $0.15 10
8 Claude Opus 4.7 Anthropic 76.1% $0.77 43
9 Gemini 3.1 Pro Google 73.7% $0.77 48
10 Claude Sonnet 4.6 Anthropic 71.6% $1.13 100
11 Claude Haiku 4.5 Anthropic 54.7% $0.53 108
12 GPT 5.4 mini OpenAI 46.7% $0.16 255
13 GPT 5 mini OpenAI 1.4% $0.02 16