Part one made the case that the flow is the next dimension of scaling agents, and that one flow cannot solve every problem. Part two ran flows against each other on four benchmarks with the models held fixed, and the best flow differed from one benchmark to the next.
This last part is what those runs taught us beyond the leaderboard: four findings we did not expect, and the question they raise. Every figure is rebuilt from the charts in our October 2026 "Humanize Intro" talk. Hover over a figure, tap it or use the arrow keys to read values, and click a legend entry to hide that series. These are our own runs, Humanfia-reported, 2026-10-05, unless a figure links to something published.
Higher diversity, higher score
Two different models taking turns beat either model working alone, whether you measure tokens, dollars or hours. The Flame Chase of Claude Fable 5 and GPT-5.6 Sol goes below either model's own Ralph loop and keeps going after both have stopped.
Two models alternating against each model alone
Take-home, two models alternating against each model alone: the Flame Chase of Claude Fable 5 and GPT-5.6 Sol ends near 1,000 cycles; each model in a Ralph loop on its own stops around 1,035 to 1,065, on any of tokens, cost or time.
Seed mean; the source shades the seed range, which is not redrawn. Same task as the take-home ablation. Read off the “Humanize Intro” deck, slide 23; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.
Better collaboration, better performance
Parallelism helps only as much as the lanes can work together. Three Flame Chase lanes that share reports beat one chase. The same three lanes with a coordinator merging their work through Git pull requests beat both, on wall-clock time and on the tokens along the critical path.
Flame Chase, Parallel Flame Chase and its Git/PR variant, 12 hours
Take-home over 12 hours, three ways of collaborating: Flame Chase ends near 1,013 cycles, Parallel Flame Chase near 1,004, and Parallel Flame Chase coordinated through Git pull requests near 986 — the best, on time and on critical-path tokens alike.
Three-seed mean. Tokens count from the first valid evaluation. Critical-path tokens are the coordinator’s plus the busiest of the three lanes; for plain Flame Chase it equals the total. Read off the “Humanize Intro” deck, slide 24; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.
A progressive goal keeps agents on track
In a Ralph loop, every round starts from a fresh session with the same goal. Agents lose their way: a round redoes work or polishes the wrong thing. If the goal advances as milestones are met, the same model on the same task gets below 1,000 cycles. With a fixed goal it levels off near 1,060.
Ralph loop, one fixed goal against a progressive goal
Take-home, GPT-5.6 Sol in a Ralph loop: given a progressive goal it gets below 1,000 cycles; given the one fixed goal it levels off near 1,060, on any axis.
Seed mean; the source shades the seed range, which is not redrawn. Read off the “Humanize Intro” deck, slide 25; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.
Claude degrades across turns, GPT does not, and Flame Chase helps
Measured by output tokens per response over a 24-hour run, GPT-5.6 Sol works at the same length from the first hour to the last in every flow. Claude Fable 5 starts with very long responses. In the flows that keep one session going (Always continue, Always prompt, ARAR), it falls to under a hundred tokens per response within five to ten hours. At that point it is barely working. The Flame Chase, where a fresh Fable session alternates with GPT, keeps it near a thousand. On the take-home, that is also where the cycles keep falling.
Output tokens per response over a 24-hour run
GPT-5.6 Sol: every flow holds between roughly 300 and 1,000 output tokens per response for the whole 24 hours.
Smoothed per-flow trend lines, max effort; the source’s per-response scatter is not redrawn. Read off the “Humanize Intro” deck, slides 26 and 27; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.
The talk also showed this measurement for a Ralph loop at every effort level on ten models, five from each vendor. The GPT models stay flat. Some of the Claude models decay over the run, most clearly Fable 5 at both effort levels and Opus 4.8 at xhigh.
Engineering, or a science?
Is a flow just one more kind of prompt, context, skill or loop engineering? We think it can be more than that.
- XXX engineeringone prompt, context, skill or loop · one task
Prompt, context, skill and loop engineering: each fix is tuned to the task in front of it, and the dirty work does not carry over to the next one.
- Flow engineeringone flow · one task
A whole method written as a flow and pointed at one kind of work: HOA at Lean, KDA at kernels, HMA at MLE-bench. It works, and it is still one flow per job.
- Flow scienceone flow family · every task that reduces to one problem
A family — the Ralph loop, the Flame Chase, the progressive goal — generalises to every task that reduces to the same problem class, and can be ablated and analysed like any other object of study.
A flow family can be studied the way a model architecture is: hold everything else fixed, vary one thing and measure the result. These are the axes along which we are ablating the Flame Chase:
Varies: Claude Fable 5 ↔ GPT-5.6 Sol · Opus 5 ↔ GPT-5.6 Sol · one model alone
Already on this site: Fig. 3 compares two models alternating with each model alone.
The work needs a runtime that can express any of these variations as a short piece of Python, run it for a day on real hardware and record every turn on one clock. That is what Humanize is. The engine underneath it is described in Deep Tech.
This closes the series. Start again from part one, or go to part two for the ablations these findings came out of.