Skip to content
← All posts

Blog2 min read

Flow Sciencedifferent tasks need different flows

Flow-level ablations on MLE-bench, HLE, ProgramBench and Anthropic's performance take-home, with the models held fixed. The best flow differs from one benchmark to the next.

Part one argued that the flow is a dimension of its own: which agent goes next, what it is asked, what it remembers, and when it stops. It also argued that one flow cannot solve every problem, because building a project is a constraint satisfaction problem and making a kernel fast is an optimization problem.

If that is right, it should show up in a measurement. So we ran flows against each other with the models held fixed. The best flow differs from one benchmark to the next.

The figure is rebuilt from the charts in our October 2026 "Humanize Intro" talk. Pick a benchmark from the tabs, hover over the chart, tap it or use the arrow keys to read values, and click a legend entry to hide that series. These are our own runs, Humanfia-reported, 2026-10-05, unless a figure links to something published.

Four benchmarks, one set of models ​

Fig. 1

Flow-level ablations on four benchmarks

0%20%40%60%80%100%0h1h2h3h4h5h6hActive time per task

MLE-bench, 75 tasks, six hours: Flame Chase and Claude Opus 5 /goal both end at 52 of 75 (69.3%), Parallel Flame Chase at 50 (66.7%), GPT-5.6 Sol /goal at 48 (64.0%). Flame Chase leads for most of the run.

75 tasks (22 lite, 38 medium, 15 hard), max effort. Each valid submission replaces the task’s previous status, so a later medal loss subtracts a point. This is the deck’s latest-submission metric; HMA’s published best-of-run figure is 78.2% any-medal. Read off the “Humanize Intro” deck, slides 8 and 22; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.

  • MLE-bench. Flame Chase leads from about the first hour and ties Claude Opus 5's /goal at 52 of 75 tasks after six hours. The published HMA result uses the same alternation and scores each task's best submission: 78.2% any-medal.
  • HLE. /goal is fastest at first. Flame Chase catches up by about an hour and ends highest. The Ralph loop and RLCR are slow until a late jump.
  • ProgramBench. Flame Chase ends highest (90.2), ahead of Opus 5 /goal (87.8), RLAR (86.4) and GPT-5.6 Sol /goal (79.8).
  • Anthropic's performance take-home. Every single-model flow stops between about 1,045 and 1,080 cycles. Only the two-model Flame Chase keeps going.

What the ablations say ​

No row of that list has the same shape as another. On one benchmark /goal is fastest at first, on another it ties at the end, and on the take-home every single-model flow stops while the two-model Flame Chase keeps going. The flow is a choice to make per task, which is why Humanize is a framework for writing flows, not a single loop.

Next ​

Running these ablations surfaced four things we did not expect: about diversity, about collaboration, about how the goal is posed, and about how models change over a long run. They are in part three, "Four findings, and a science of flows".

Read next

BlogFlow Science: four findings, and engineering or a science?What the flow-level ablations found that we did not expect: higher diversity, higher score; better collaboration, better performance; a progressive goal keeps agents on track; and Claude degrades across turns where GPT does not. Then: is a flow just engineering?Read it BlogFlow Science: flow is the new scaling dimensionModels scale with parameters and data, and agents scale with tools. The next dimension is the flow. Why a single /goal session is not enough on long, hard tasks, and why one flow cannot solve every problem.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it