Task Details

Task pages mirror ProgramBench's per-task view for this Codex /goal scaffold: scored behavioral tests, best score, results by model/mode, and links to sanitized evidence.

Pending rows are full-run targets waiting for Codex results. Once a task is evaluated, its page shows the GoalBench score beside cached official ProgramBench task context, plus links to the official ProgramBench task page for baseline comparison.

What Each Task Page Shows

Official context. Generated test count, official best score, and official model rows are cached from ProgramBench public task pages.
GoalBench results. Each Codex /goal result is shown by run version, model, inference mode, score, evaluated tests, estimated cost, calls, and wall-clock time.
Why scores differ. When public evidence exists, task pages list failed test names and a compact miss-class summary so near-solves like cmatrix are explainable without opening raw JSON.
Evidence links. Public artifacts include sanitized eval summaries, public eval JSON, usage audit, and manifest. Raw Codex session logs and submission tarballs remain local by default.

Metric Contract

Resolved means the ProgramBench behavioral test pass rate is exactly 100%. Almost resolved follows ProgramBench's public threshold of at least 95%. Scores are computed with ProgramBench's own evaluation summary logic after active-branch and ignored-test filtering.

Scope

GoalBench is not the official mini-SWE-agent leaderboard. It is a scaffold measurement for Codex /goal on the same ProgramBench task family, with modes and compliance labels shown explicitly.