ProgramBench task detail

tarka__xcp.5e5b448

Task-level results for this Codex /goal scaffold. ProgramBench baseline context is cached from the official task page; Codex scored tests are after active-branch and ignored-test filtering.

1184generated behavioral tests
95.8%official best score
93.8%Codex best score
3Codex result rows

Codex Results by Model

#ModelRunAgentModeScoreEvalTestsEst. costCallsWallEvidence
1 GPT 5.5 (xhigh) 20260517T094700Z Codex /goal Mini-SWE-compatible no internet 93.8% ok 1111/1184 $5.28 102 0.24h manifest · eval json · eval summary · usage audit
2 GPT 5.5 (xhigh) 20260519T215046Z Codex /goal Paper prompt + Goal contract no internet 93.7% ok 1109/1184 $3.59 65 0.21h not exported
3 GPT 5.5 (xhigh) 20260518T232006Z Codex /goal Paper prompt + /goal no internet 18.9% ok 224/1184 $3.24 54 0.21h manifest · eval json · eval summary · agent summary · usage audit

Why Scores Differ

Official ProgramBench rows are the public mini-SWE-agent submissions. GoalBench rows are separate Codex /goal submissions. Same model label does not mean the same agent, prompt, tool loop, or generated implementation. A GoalBench row is resolved only when the ProgramBench-scored pass rate is exactly 100%.

GPT 5.5 (xhigh) · 20260517T094700Z

93.8% from 1111/1184 ProgramBench-scored tests. Raw public eval statuses: failure: 72, passed: 1143, skipped: 21.

Agent trace summary unavailable.

Likely miss class: behavioral mismatch.

  • 49179779960b tests.test_advanced.test_reflink_never_warning_message (failure)
  • 49179779960b tests.test_attributes.test_timestamp_preservation_default (failure)
  • 49179779960b tests.test_config_additional.test_reflink_always_on_non_reflink_filesystem (failure)
  • 49179779960b tests.test_cli_flags.test_multiple_verbose_flags_increase_verbosity_fixed (failure)
  • 49179779960b tests.test_attributes.test_timestamp_nanosecond_precision (failure)
  • 49179779960b tests.test_attributes.test_timestamp_preservation_across_dst_changes (failure)

manifest · eval json · eval summary · usage audit

GPT 5.5 (xhigh) · 20260519T215046Z

Public eval evidence has not been exported for this row yet.

GPT 5.5 (xhigh) · 20260518T232006Z

18.9% from 224/1184 ProgramBench-scored tests. Raw public eval statuses: failure: 972, passed: 243, skipped: 21.

Agent trace: 135 shell command(s), 103 executable-observation command(s), 4 build command(s), 0 blocked attempt(s). Raw Codex logs are not published.

Likely miss class: behavioral mismatch.

  • 0ed2ee2b4c94 eval.tests.test_advanced_options.test_no_progress_flag (failure)
  • 0ed2ee2b4c94 eval.tests.test_comprehensive_coverage.test_parblock_very_small_block_size (failure)
  • 0ed2ee2b4c94 eval.tests.test_backup_modes.test_backup_numbered_increments (failure)
  • 0ed2ee2b4c94 eval.tests.test_advanced_options.test_parfile_driver (failure)
  • 0ed2ee2b4c94 eval.tests.test_coverage_improvements.test_fsync_with_file (failure)
  • 0ed2ee2b4c94 eval.tests.test_comprehensive_coverage.test_dereference_with_parblock (failure)

manifest · eval json · eval summary · agent summary · usage audit

Official ProgramBench Results by Model

#ModelProviderScoreCostCalls
1 GPT 5.5 (xhigh) OpenAI 95.8% $6.37 62
2 Claude Opus 4.6 Anthropic 92.6% $5.01 138
3 Claude Sonnet 4.6 Anthropic 91.7% $8.99 330
4 Claude Opus 4.7 (xhigh) Anthropic 91.5% $13.71 263
5 Claude Opus 4.7 Anthropic 84.7% $2.57 85
6 GPT 5.4 OpenAI 84.5% $0.28 10
7 Gemini 3.1 Pro Google 73.7% $0.72 42
8 GPT 5.5 OpenAI 38.7% $1.33 18
9 GPT 5.4 mini OpenAI 22.5% $0.04 18
10 Gemini 3 Flash Google 18.6% $0.24 57
11 GPT 5 mini OpenAI 4.3% $0.02 16
12 Claude Haiku 4.5 Anthropic 0.0% $0.73 107
13 GPT 5.5 (high) OpenAI n/a $1.95 25