Skip to content
← All news

News1 min read

0% and 0.5% alone, 3.5% together: a review loop on ProgramBench

Two models that solve 0% and 0.5% of ProgramBench on their own solve 3.5% as a builder and a reviewer in a loop — and the curve was still rising when the four-hour clock ran out.

Ongoing

Runs are still going and the numbers below are a checkpoint, not a final result.

ProgramBench is hard enough that single-shot numbers are close to the floor, which makes it an unusually clean instrument for measuring a loop.

Solved
Opus-4.8, on its own0%
GPT-5.5, on its own0.5%
Humanize 2 — Opus-4.8 building, GPT-5.5 reviewing3.5%
Reported state of the art (Opus-5)4.5%

Two models that solve almost nothing separately solve seven times the better one's share when they are arranged as a builder and an independent reviewer inside a Ralph loop. Across tasks, the loop averages a 14.2% improvement over the single-shot baseline.

The part that bothers us ​

Performance was still rising at roughly 1.2% per round in the later rounds when the runs were cut off.

That is not a result, it is a missing one. The runs are constrained by a four-hour time budget rather than by having converged, so what we have measured is partly the budget rather than the method. The honest statement is: this is the number at four hours, the slope at four hours is positive, and we do not yet know where it goes.

Which is the sort of thing FlowBench exists to stop us from guessing at.

Read next

NewsNo certificate, no win: QCode Discovery's search now fails closed17,520 definitions across 21 lattices, and no candidate counts as a win without an exact construction, an independent verification and a replay against the IBM baselines.Read it BlogModel level, tool level, flow levelThe same model, measured three ways — raw API, the vendor's CLI, and the CLI inside a flow — across PutnamBench, Physics Cup, SuperChem and an HLE subset. The flow is worth more than the gap between model generations.Read it News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it