Skip to content
← All news

News2 min read

Biohub closed at #188 of 3,947: top 5%, not the #8 the public board showed

In August an agent workflow sat 8th on the Biohub cell-tracking public leaderboard. The private board has now spoken: 188th of 3,947, top 4.8%. A real finish, a smaller one, and what the Kaggle tally looks like when only final ranks count.

On 15 August, one of our agent workflows on Kaggle was 8th of 2,378 on the public leaderboard of Biohub – Cell Tracking During Development. That was the top 0.34%, and it was among the ongoing results we reported then, under a note that the numbers were still moving.

The competition has closed. The authenticated final private rank is 188th of 3,947, the top 4.76%. That is a genuine top-5% finish on a closed board. It is also 188th, not 8th.

We do not know how much of the gap is private-board shake-up and how much is everything that happened after 15 August: 1,569 more teams entered, and the rest of the field kept improving. The point stands either way. A public rank in the middle of a competition is a weather report. We said so in August, and this is what it looks like.

The other boards that closed ​

CompetitionBest tracked, 15 August, public boardBest tracked, final private board
Biohub – Cell Tracking8 / 2,378 · top 0.34%188 / 3,947 · top 4.76%
PTCG AI Battle955 / 6,829 · top 14.0%654 / 6,807 · top 9.61%
CUHK-X Small Model Track31 / 199 · top 15.6%56 / 326 · top 17.2%
Hyperspectral Object Detection10 / 40 · top 25%77 / 309 · top 24.9%

One more closed with a striking number we cannot yet count. Predicting Smartphone Addiction shows an account at 2nd of 3,531, but on the public board as of the snapshot. That account's final private rank has not been recorded. The best authenticated private rank we have there is 564th. Until the private rank is in, rank 2 is not a finish, and we do not report it as one.

The tally, counting only final ranks ​

The AgentKaggle Team Radar (snapshot 2026-10-05 09:18 UTC) tracks 39 completed competitions. Taking the best tracked result in each, 16 are in the top 5%. Most of those are not finishes:

  • 3 are authenticated final private ranks: Predicting Student Health Risk (top 2.00%), ROGII Wellbore Geology (top 2.24%) and Biohub (top 4.76%);
  • 1 is a public rank at snapshot, the Smartphone Addiction result above; and
  • 12 are late-submission estimates: real scores placed against a frozen final board, but not ranks, medals or evidence of having competed.

The audit's charts are drawn with a fixed list of 12 competitions left out, so a chart rendered from today's data shows 27, not 39. All 12 are outside the top 15%, so leaving them out flatters the picture. The counts above use all 39. Of the 13 competitions still open, 3 sit in the top 5% of their public boards today. One of those 3 is Spaceship Titanic, which is a getting-started practice competition.

The set of tracked accounts has also changed since August, with some added and some dropped. That is why we compare competitions one by one here, and do not compare August's totals with today's.

The audit · Team Radar · HMA

Read next

News78.2% medal rate on MLE-bench in six hours, by making two agents take turnsTwo coding agents alternating over one workspace, each starting fresh, beat both of them working alone on 75 MLE-bench tasks: 78.2% any-medal against 72.4% and 68.0%. Self-reported, and not on the official leaderboard.Read it News14 of 19 Kaggle competitions in the top 5%, five inside the top 1%Ten agent workflows on live Kaggle competitions: of nineteen completed, fourteen finished in the top 5% — and our first read is that disagreement between agents, not a better model, did it.Read it News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it