Skip to content
← All news

News1 min read

No certificate, no winQCode Discovery's search now fails closed

17,520 definitions across 21 lattices, and no candidate counts as a win without an exact construction, an independent verification and a replay against the IBM baselines.

Search over quantum codes has an unpleasant property: the cheap metric that ranks candidates and the expensive certificate that proves one is genuinely better are not the same thing, and the gap between them is exactly where a long unattended run will settle if you let it.

The QCode Discovery pipeline has been moved onto the other side of that gap. It is now fail-closed: verified certification, not heuristic score, is what makes a candidate a result.

  • 17,520 definitions processed across 21 lattices;
  • 5,178 winner-capable definitions persisted for further work.

What a win now requires ​

No candidate is reported as a win without all three of:

  1. an exact construction, not a score that implies one exists;
  2. an independent verification of that construction; and
  3. a strict final replay that revalidates the established IBM baselines it is being compared against.

The third is the one that gets skipped in practice, and it is the one that catches the failure mode we care about — a baseline that drifted, or was measured under conditions that no longer hold, making a comparison look like a win when the only thing that moved was the reference.

This is the same discipline as everything else here: the run is allowed to be wrong, and it is not allowed to be wrong quietly.

Read next

News0% and 0.5% alone, 3.5% together: a review loop on ProgramBenchTwo models that solve 0% and 0.5% of ProgramBench on their own solve 3.5% as a builder and a reviewer in a loop — and the curve was still rising when the four-hour clock ran out.Read it BlogModel level, tool level, flow levelThe same model, measured three ways — raw API, the vendor's CLI, and the CLI inside a flow — across PutnamBench, Physics Cup, SuperChem and an HLE subset. The flow is worth more than the gap between model generations.Read it News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it