ProgramBench task detail

alecthomas__chroma.8d04def

Task-level results for this Codex /goal scaffold. ProgramBench baseline context is cached from the official task page; Codex scored tests are after active-branch and ignored-test filtering.

503generated behavioral tests
41.7%official best score
0.0%Codex best score
3Codex result rows

Codex Results by Model

#ModelRunAgentModeScoreEvalTestsEst. costCallsWallEvidence
1 GPT 5.5 (xhigh) 20260517T094700Z Codex /goal Mini-SWE-compatible no internet 0.0% ok 0/503 $4.76 74 0.18h manifest · eval json · eval summary · usage audit
2 GPT 5.5 (xhigh) 20260518T232006Z Codex /goal Paper prompt + /goal no internet 0.0% ok 0/503 n/a 57 0.21h not exported
3 GPT 5.5 (xhigh) 20260519T215046Z Codex /goal Paper prompt + Goal contract no internet 0.0% ok 0/503 n/a 41 0.19h not exported

Why Scores Differ

Official ProgramBench rows are the public mini-SWE-agent submissions. GoalBench rows are separate Codex /goal submissions. Same model label does not mean the same agent, prompt, tool loop, or generated implementation. A GoalBench row is resolved only when the ProgramBench-scored pass rate is exactly 100%.

GPT 5.5 (xhigh) · 20260517T094700Z

0.0% from 0/503 ProgramBench-scored tests. Raw public eval statuses: not_run: 531.

Agent trace summary unavailable.

Likely miss class: behavioral mismatch.

  • 06dabfabaea7 tests.test_advanced_flags.test_profile_creates_valid_pprof_file (not_run)
  • 06dabfabaea7 tests.test_advanced_flags.test_profile_different_formatters (not_run)
  • 06dabfabaea7 tests.test_advanced_flags.test_profile_multiple_files (not_run)
  • 06dabfabaea7 tests.test_advanced_flags.test_profile_with_check_mode (not_run)
  • 06dabfabaea7 tests.test_advanced_flags.test_profile_with_file_input (not_run)
  • 06dabfabaea7 tests.test_advanced_flags.test_profile_with_larger_input_captures_data (not_run)

manifest · eval json · eval summary · usage audit

GPT 5.5 (xhigh) · 20260518T232006Z

Public eval evidence has not been exported for this row yet.

GPT 5.5 (xhigh) · 20260519T215046Z

Public eval evidence has not been exported for this row yet.

Official ProgramBench Results by Model

#ModelProviderScoreCostCalls
1 GPT 5.5 (high) OpenAI 41.7% $4.27 54
2 GPT 5.5 OpenAI 27.2% $1.51 22
3 Claude Sonnet 4.6 Anthropic 16.1% $37.36 565
4 GPT 5.5 (xhigh) OpenAI 13.1% $9.46 80
5 GPT 5.4 OpenAI 9.3% $0.33 9
6 Claude Opus 4.7 (xhigh) Anthropic 8.7% $5.07 126
7 Gemini 3 Flash Google 7.6% $0.24 64
8 Claude Haiku 4.5 Anthropic 6.2% $0.65 114
9 Claude Opus 4.7 Anthropic 4.5% $0.48 20
10 Gemini 3.1 Pro Google 4.4% $1.90 141
11 GPT 5.4 mini OpenAI 2.2% $0.03 8
12 GPT 5 mini OpenAI 1.2% $0.02 16
13 Claude Opus 4.6 Anthropic 0.6% $16.38 297