ProgramBench task detail

jonas__tig.8334123

Task-level results for this Codex /goal scaffold. ProgramBench baseline context is cached from the official task page; Codex scored tests are after active-branch and ignored-test filtering.

1586generated behavioral tests
85.2%official best score
85.1%Codex best score
3Codex result rows

Codex Results by Model

#ModelRunAgentModeScoreEvalTestsEst. costCallsWallEvidence
1 GPT 5.5 (xhigh) 20260518T232006Z Codex /goal Paper prompt + /goal no internet 85.1% ok 1350/1586 n/a 60 0.16h not exported
2 GPT 5.5 (xhigh) 20260517T094700Z Codex /goal Mini-SWE-compatible no internet 84.1% ok 1334/1586 $3.08 62 0.17h manifest · eval json · eval summary · usage audit
3 GPT 5.5 (xhigh) 20260519T215046Z Codex /goal Paper prompt + Goal contract no internet 83.9% ok 1330/1586 n/a 45 0.12h not exported

Why Scores Differ

Official ProgramBench rows are the public mini-SWE-agent submissions. GoalBench rows are separate Codex /goal submissions. Same model label does not mean the same agent, prompt, tool loop, or generated implementation. A GoalBench row is resolved only when the ProgramBench-scored pass rate is exactly 100%.

GPT 5.5 (xhigh) · 20260518T232006Z

Public eval evidence has not been exported for this row yet.

GPT 5.5 (xhigh) · 20260517T094700Z

84.1% from 1334/1586 ProgramBench-scored tests. Raw public eval statuses: error: 1, failure: 304, passed: 1924, skipped: 8.

Agent trace summary unavailable.

Likely miss class: behavioral mismatch, interactive terminal behavior mismatch, terminal rendering mismatch.

  • 8ec0cc6e2379 tests.test_tmux_fixed.test_tmux_log_view (failure)
  • 8ec0cc6e2379 tests.test_tmux_fixed.test_tmux_show_view (failure)
  • f276bd1d6c4e eval.tests.test_env_and_config.test_tigrc_user_custom_location (failure)
  • f276bd1d6c4e eval.tests.test_env_and_config.test_tigrc_user_nonexistent (failure)
  • f276bd1d6c4e eval.tests.test_env_and_config.test_tigrc_system_custom_location (failure)
  • f276bd1d6c4e eval.tests.test_env_and_config.test_tig_diff_opts (failure)

manifest · eval json · eval summary · usage audit

GPT 5.5 (xhigh) · 20260519T215046Z

Public eval evidence has not been exported for this row yet.

Official ProgramBench Results by Model

#ModelProviderScoreCostCalls
1 GPT 5.5 (xhigh) OpenAI 85.2% $4.82 39
2 GPT 5.5 (high) OpenAI 85.0% $2.06 32
3 GPT 5.5 OpenAI 84.2% $0.82 16
4 Gemini 3.1 Pro Google 83.9% $1.02 72
5 Claude Opus 4.7 (xhigh) Anthropic 83.8% $1.57 65
6 Claude Opus 4.7 Anthropic 83.6% $0.84 44
7 Claude Haiku 4.5 Anthropic 82.9% $0.22 80
8 Claude Opus 4.6 Anthropic 79.6% $10.78 281
9 GPT 5 mini OpenAI 76.9% $0.01 12
10 Gemini 3 Flash Google 71.2% $0.19 54
11 GPT 5.4 mini OpenAI 39.0% $0.03 8
12 GPT 5.4 OpenAI 38.1% $0.12 7
13 Claude Sonnet 4.6 Anthropic 0.0% $14.93 382