Part one argued that the flow is a dimension of its own: which agent goes next, what it is asked, what it remembers, and when it stops. It also argued that one flow cannot solve every problem, because building a project is a constraint satisfaction problem and making a kernel fast is an optimization problem.
If that is right, it should show up in a measurement. So we ran flows against each other with the models held fixed. The best flow differs from one benchmark to the next.
The figure is rebuilt from the charts in our October 2026 "Humanize Intro" talk. Pick a benchmark from the tabs, hover over the chart, tap it or use the arrow keys to read values, and click a legend entry to hide that series. These are our own runs, Humanfia-reported, 2026-10-05, unless a figure links to something published.
Four benchmarks, one set of models
Flow-level ablations on four benchmarks
MLE-bench, 75 tasks, six hours: Flame Chase and Claude Opus 5 /goal both end at 52 of 75 (69.3%), Parallel Flame Chase at 50 (66.7%), GPT-5.6 Sol /goal at 48 (64.0%). Flame Chase leads for most of the run.
75 tasks (22 lite, 38 medium, 15 hard), max effort. Each valid submission replaces the task’s previous status, so a later medal loss subtracts a point. This is the deck’s latest-submission metric; HMA’s published best-of-run figure is 78.2% any-medal. Read off the “Humanize Intro” deck, slides 8 and 22; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.
- MLE-bench. Flame Chase leads from about the first hour and ties Claude Opus 5's
/goalat 52 of 75 tasks after six hours. The published HMA result uses the same alternation and scores each task's best submission: 78.2% any-medal. - HLE.
/goalis fastest at first. Flame Chase catches up by about an hour and ends highest. The Ralph loop and RLCR are slow until a late jump. - ProgramBench. Flame Chase ends highest (90.2), ahead of Opus 5
/goal(87.8), RLAR (86.4) and GPT-5.6 Sol/goal(79.8). - Anthropic's performance take-home. Every single-model flow stops between about 1,045 and 1,080 cycles. Only the two-model Flame Chase keeps going.
What the ablations say
No row of that list has the same shape as another. On one benchmark /goal is fastest at first, on another it ties at the end, and on the take-home every single-model flow stops while the two-model Flame Chase keeps going. The flow is a choice to make per task, which is why Humanize is a framework for writing flows, not a single loop.
Next
Running these ablations surfaced four things we did not expect: about diversity, about collaboration, about how the goal is posed, and about how models change over a long run. They are in part three, "Four findings, and a science of flows".