ProgramBench task detail

blake3-team__blake3.15e83a5

Task-level results for this Codex /goal scaffold. ProgramBench baseline context is cached from the official task page; Codex scored tests are after active-branch and ignored-test filtering.

647generated behavioral tests
97.8%official best score
96.3%Codex best score
3Codex result rows

Codex Results by Model

#ModelRunAgentModeScoreEvalTestsEst. costCallsWallEvidence
1 GPT 5.5 (xhigh) 20260517T094700Z Codex /goal Mini-SWE-compatible no internet 96.3% ok 623/647 $5.75 97 0.25h manifest · eval json · eval summary · usage audit
2 GPT 5.5 (xhigh) 20260518T232006Z Codex /goal Paper prompt + /goal no internet 65.4% ok 423/647 $3.43 58 0.18h manifest · eval json · eval summary · agent summary · usage audit
3 GPT 5.5 (xhigh) 20260519T215046Z Codex /goal Paper prompt + Goal contract no internet 64.6% ok 418/647 $2.26 37 0.14h not exported

Why Scores Differ

Official ProgramBench rows are the public mini-SWE-agent submissions. GoalBench rows are separate Codex /goal submissions. Same model label does not mean the same agent, prompt, tool loop, or generated implementation. A GoalBench row is resolved only when the ProgramBench-scored pass rate is exactly 100%.

GPT 5.5 (xhigh) · 20260517T094700Z

96.3% from 623/647 ProgramBench-scored tests. Raw public eval statuses: failure: 24, passed: 660, skipped: 3.

Agent trace summary unavailable.

Likely miss class: behavioral mismatch.

  • 06dabfabaea7 tests.test_check.test_check_unicode_replacement_rejected (failure)
  • 06dabfabaea7 tests.test_check.test_check_conflicts_with_derive_key (failure)
  • 06dabfabaea7 tests.test_check.test_check_windows_path_on_linux (failure)
  • 06dabfabaea7 tests.test_errors.test_keyed_mode_cannot_use_stdin_dash (failure)
  • 06dabfabaea7 tests.test_check.test_check_conflicts_with_raw (failure)
  • 06dabfabaea7 tests.test_check.test_check_malformed_lines_reported (failure)

manifest · eval json · eval summary · usage audit

GPT 5.5 (xhigh) · 20260518T232006Z

65.4% from 423/647 ProgramBench-scored tests. Raw public eval statuses: failure: 229, passed: 455, skipped: 3.

Agent trace: 163 shell command(s), 132 executable-observation command(s), 8 build command(s), 2 blocked attempt(s). Raw Codex logs are not published.

Likely miss class: behavioral mismatch, interactive terminal behavior mismatch.

  • 06dabfabaea7 tests.test_basic.test_stdin_large_input (failure)
  • 06dabfabaea7 tests.test_edge_cases_final.test_keyed_mode_with_escaped_filename (failure)
  • 06dabfabaea7 tests.test_check.test_check_windows_path_on_linux (failure)
  • 06dabfabaea7 tests.test_check.test_check_unicode_replacement_rejected (failure)
  • 06dabfabaea7 tests.test_basic.test_no_args_reads_stdin (failure)
  • 06dabfabaea7 tests.test_advanced.test_seek_and_length_with_stdin (failure)

manifest · eval json · eval summary · agent summary · usage audit

GPT 5.5 (xhigh) · 20260519T215046Z

Public eval evidence has not been exported for this row yet.

Official ProgramBench Results by Model

#ModelProviderScoreCostCalls
1 GPT 5.5 (xhigh) OpenAI 97.8% $5.97 49
2 Claude Opus 4.7 Anthropic 97.5% $3.88 84
3 Claude Opus 4.7 (xhigh) Anthropic 96.8% $6.88 119
4 GPT 5.5 (high) OpenAI 95.1% $2.65 34
5 Claude Opus 4.6 Anthropic 94.9% $5.68 171
6 GPT 5.5 OpenAI 92.4% $1.06 15
7 Claude Sonnet 4.6 Anthropic 92.3% $4.51 213
8 Gemini 3.1 Pro Google 90.6% $2.22 103
9 Gemini 3 Flash Google 51.3% $0.47 88
10 GPT 5.4 OpenAI 11.9% $0.23 10
11 GPT 5 mini OpenAI 9.1% $0.03 18
12 Claude Haiku 4.5 Anthropic 8.3% $0.99 148
13 GPT 5.4 mini OpenAI 0.2% $0.05 16