Skip to content
← All news

News2 min read

14 of 19 Kaggle competitions in the top 5%, five inside the top 1%

Ten agent workflows on live Kaggle competitions: of nineteen completed, fourteen finished in the top 5% — and our first read is that disagreement between agents, not a better model, did it.

Still running

Fourteen of these competitions are open. Public ranks move, and the private board is the only one that counts. Everything below is a snapshot with a date on it.

Kaggle is the setting we keep coming back to, because it is one of the few places an agent loop can be wrong in public: thousands of humans on the same problem, a deadline that does not negotiate, and a private leaderboard that arrives after every decision has already been made.

What was run ​

Built on Humanize 2, we designed roughly ten distinct agent workflows and pointed all of them at live competitions. They differ in the things we think actually matter on a long run:

  • how context is managed across hours of work,
  • how several skills are orchestrated inside one run,
  • how much of an end-to-end machine-learning engineering pipeline the agents own,
  • whether a second agent reviews the first one's work, and
  • whether attempts run in parallel or in sequence.

Where they landed ​

Completed — nineteen competitions. Fourteen finished in the top 5%, including five results at or inside the top 1%.

Ongoing — fourteen official competitions. Six currently sit in the top 5%, four of them in the top 3%, and the numbers are still moving.

For comparison, Codex on GPT-5.5 at xhigh reasoning effort generally peaked at around the top-5% level across most of the same competitions. Same models, different arrangement, and the arrangement is worth the tail of the distribution.

What we think is doing the work ​

Our preliminary read is that the gains come primarily from the diversity of reasoning strategies that multi-agent collaboration makes available. A single strong agent narrows early and commits; several agents arranged to disagree keep more than one line of attack alive long enough to be tested against a validation split that has not been quietly bent to fit.

That is a hypothesis with a number attached, not a conclusion. Further analysis is ongoing.

What these figures are, and are not ​

A finish and an estimate are two different claims, and we do not report them as one. A finish is an exact position on a closed leaderboard. An estimate is a score, submitted after the competition closed and placed against the frozen final board — real as a score, and not a rank, a medal, or evidence of having competed. Both appear above; the audit says which is which, one row at a time.

The audit · the live leaderboard · HMA

Read next

NewsBiohub closed at #188 of 3,947: top 5%, not the #8 the public board showedIn August an agent workflow sat 8th on the Biohub cell-tracking public leaderboard. The private board has now spoken: 188th of 3,947, top 4.8%. A real finish, a smaller one, and what the Kaggle tally looks like when only final ranks count.Read it News78.2% medal rate on MLE-bench in six hours, by making two agents take turnsTwo coding agents alternating over one workspace, each starting fresh, beat both of them working alone on 75 MLE-bench tasks: 78.2% any-medal against 72.4% and 68.0%. Self-reported, and not on the official leaderboard.Read it News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it