ProgramBench task detail

bellard__quickjs.d7ae12a

Task-level results for this Codex /goal scaffold. ProgramBench baseline context is cached from the official task page; Codex scored tests are after active-branch and ignored-test filtering.

3034generated behavioral tests
31.0%official best score
14.5%Codex best score
3Codex result rows

Codex Results by Model

#ModelRunAgentModeScoreEvalTestsEst. costCallsWallEvidence
1 GPT 5.5 (xhigh) 20260519T215046Z Codex /goal Paper prompt + Goal contract no internet 14.5% ok 440/3034 n/a 90 0.35h not exported
2 GPT 5.5 (xhigh) 20260517T094700Z Codex /goal Mini-SWE-compatible no internet 0.0% ok 0/3034 $6.76 118 0.27h manifest · eval json · eval summary · usage audit
3 GPT 5.5 (xhigh) 20260518T232006Z Codex /goal Paper prompt + /goal no internet 0.0% ok 0/3034 n/a 76 0.22h not exported

Why Scores Differ

Official ProgramBench rows are the public mini-SWE-agent submissions. GoalBench rows are separate Codex /goal submissions. Same model label does not mean the same agent, prompt, tool loop, or generated implementation. A GoalBench row is resolved only when the ProgramBench-scored pass rate is exactly 100%.

GPT 5.5 (xhigh) · 20260519T215046Z

Public eval evidence has not been exported for this row yet.

GPT 5.5 (xhigh) · 20260517T094700Z

0.0% from 0/3034 ProgramBench-scored tests. Raw public eval statuses: failure: 3038, skipped: 6.

Agent trace summary unavailable.

Likely miss class: behavioral mismatch, exact version/output mismatch, interactive terminal behavior mismatch.

  • e1d7a4e20e53 tests.test_advanced_features.test_proxy_preventextensions_trap (failure)
  • e1d7a4e20e53 tests.test_conversion_coercion_final_gaps.test_tonumber_string_formats (failure)
  • e1d7a4e20e53 tests.test_examples.test_circular_import (failure)
  • e1d7a4e20e53 tests.test_final_easy_wins.test_string_charat_null_this (failure)
  • e1d7a4e20e53 tests.test_error_path_sweep.test_date_invalid_string (failure)
  • e1d7a4e20e53 tests.test_import_attributes.test_basic_json_import_with_type_attribute (failure)

manifest · eval json · eval summary · usage audit

GPT 5.5 (xhigh) · 20260518T232006Z

Public eval evidence has not been exported for this row yet.

Official ProgramBench Results by Model

#ModelProviderScoreCostCalls
1 Claude Opus 4.7 (xhigh) Anthropic 31.0% $31.65 261
2 GPT 5.5 (xhigh) OpenAI 21.3% $7.83 76
3 GPT 5.5 (high) OpenAI 10.9% $3.51 30
4 GPT 5.5 OpenAI 4.0% $1.67 26
5 Claude Opus 4.6 Anthropic 3.6% $16.39 248
6 GPT 5.4 OpenAI 1.2% $0.21 13
7 Claude Opus 4.7 Anthropic 1.2% $0.52 18
8 Gemini 3.1 Pro Google 1.0% $1.30 119
9 Claude Sonnet 4.6 Anthropic 0.0% $82.98 995
10 Claude Haiku 4.5 Anthropic 0.0% $0.55 124
11 Gemini 3 Flash Google 0.0% $0.44 96
12 GPT 5.4 mini OpenAI 0.0% $0.01 8
13 GPT 5 mini OpenAI 0.0% $0.01 10