ProgramBench task detail

ismaelgv__rnr.fc0733b

Task-level results for this Codex /goal scaffold. ProgramBench baseline context is cached from the official task page; Codex scored tests are after active-branch and ignored-test filtering.

683generated behavioral tests
94.0%official best score
90.8%Codex best score
3Codex result rows

Codex Results by Model

#ModelRunAgentModeScoreEvalTestsEst. costCallsWallEvidence
1 GPT 5.5 (xhigh) 20260518T232006Z Codex /goal Paper prompt + /goal no internet 90.8% ok 620/683 n/a 85 0.19h not exported
2 GPT 5.5 (xhigh) 20260517T094700Z Codex /goal Mini-SWE-compatible no internet 89.9% ok 614/683 $5.91 112 0.27h manifest · eval json · eval summary · usage audit
3 GPT 5.5 (xhigh) 20260519T215046Z Codex /goal Paper prompt + Goal contract no internet 86.4% ok 590/683 n/a 62 0.19h not exported

Why Scores Differ

Official ProgramBench rows are the public mini-SWE-agent submissions. GoalBench rows are separate Codex /goal submissions. Same model label does not mean the same agent, prompt, tool loop, or generated implementation. A GoalBench row is resolved only when the ProgramBench-scored pass rate is exactly 100%.

GPT 5.5 (xhigh) · 20260518T232006Z

Public eval evidence has not been exported for this row yet.

GPT 5.5 (xhigh) · 20260517T094700Z

89.9% from 614/683 ProgramBench-scored tests. Raw public eval statuses: failure: 70, passed: 670.

Agent trace summary unavailable.

Likely miss class: behavioral mismatch, exact version/output mismatch.

  • 8210eaf4f82e tests.test_edge_cases.test_whitespace_in_replacement (failure)
  • 8210eaf4f82e tests.test_edge_cases.test_replace_with_period (failure)
  • 8210eaf4f82e tests.test_edge_cases.test_empty_replacement (failure)
  • 7877cb75132e tests.test_errors.test_invalid_regex_unclosed_bracket (failure)
  • 7877cb75132e tests.test_advanced_regex.test_backreference_in_pattern (failure)
  • 7877cb75132e tests.test_advanced_regex.test_replace_limit_unlimited (failure)

manifest · eval json · eval summary · usage audit

GPT 5.5 (xhigh) · 20260519T215046Z

Public eval evidence has not been exported for this row yet.

Official ProgramBench Results by Model

#ModelProviderScoreCostCalls
1 GPT 5.5 (xhigh) OpenAI 94.0% $6.29 73
2 GPT 5.5 (high) OpenAI 92.7% $2.75 50
3 Claude Opus 4.7 (xhigh) Anthropic 89.6% $50.30 594
4 Claude Opus 4.7 Anthropic 82.1% $4.44 154
5 GPT 5.5 OpenAI 81.7% $1.10 19
6 Gemini 3.1 Pro Google 79.8% $1.95 184
7 Claude Sonnet 4.6 Anthropic 78.0% $20.12 374
8 GPT 5.4 OpenAI 67.9% $0.37 14
9 Claude Opus 4.6 Anthropic 66.6% $11.86 257
10 GPT 5.4 mini OpenAI 25.9% $0.04 18
11 Claude Haiku 4.5 Anthropic 21.4% $1.33 154
12 Gemini 3 Flash Google 3.8% $0.26 107
13 GPT 5 mini OpenAI 3.8% $0.03 24