Skip to content
← All posts

Blog3 min read

Flow Sciencefour findings, and engineering or a science?

What the flow-level ablations found that we did not expect: higher diversity, higher score; better collaboration, better performance; a progressive goal keeps agents on track; and Claude degrades across turns where GPT does not. Then: is a flow just engineering?

Part one made the case that the flow is the next dimension of scaling agents, and that one flow cannot solve every problem. Part two ran flows against each other on four benchmarks with the models held fixed, and the best flow differed from one benchmark to the next.

This last part is what those runs taught us beyond the leaderboard: four findings we did not expect, and the question they raise. Every figure is rebuilt from the charts in our October 2026 "Humanize Intro" talk. Hover over a figure, tap it or use the arrow keys to read values, and click a legend entry to hide that series. These are our own runs, Humanfia-reported, 2026-10-05, unless a figure links to something published.

Higher diversity, higher score ​

Two different models taking turns beat either model working alone, whether you measure tokens, dollars or hours. The Flame Chase of Claude Fable 5 and GPT-5.6 Sol goes below either model's own Ralph loop and keeps going after both have stopped.

Fig. 1

Two models alternating against each model alone

10001100120013001400150001M2M3M4MOutput tokens spent

Take-home, two models alternating against each model alone: the Flame Chase of Claude Fable 5 and GPT-5.6 Sol ends near 1,000 cycles; each model in a Ralph loop on its own stops around 1,035 to 1,065, on any of tokens, cost or time.

Seed mean; the source shades the seed range, which is not redrawn. Same task as the take-home ablation. Read off the “Humanize Intro” deck, slide 23; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.

Better collaboration, better performance ​

Parallelism helps only as much as the lanes can work together. Three Flame Chase lanes that share reports beat one chase. The same three lanes with a coordinator merging their work through Git pull requests beat both, on wall-clock time and on the tokens along the critical path.

Fig. 2

Flame Chase, Parallel Flame Chase and its Git/PR variant, 12 hours

950100010501100115012000h2h4h6h8h10h12hActive elapsed time

Take-home over 12 hours, three ways of collaborating: Flame Chase ends near 1,013 cycles, Parallel Flame Chase near 1,004, and Parallel Flame Chase coordinated through Git pull requests near 986 — the best, on time and on critical-path tokens alike.

Three-seed mean. Tokens count from the first valid evaluation. Critical-path tokens are the coordinator’s plus the busiest of the three lanes; for plain Flame Chase it equals the total. Read off the “Humanize Intro” deck, slide 24; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.

A progressive goal keeps agents on track ​

In a Ralph loop, every round starts from a fresh session with the same goal. Agents lose their way: a round redoes work or polishes the wrong thing. If the goal advances as milestones are met, the same model on the same task gets below 1,000 cycles. With a fixed goal it levels off near 1,060.

Fig. 3

Ralph loop, one fixed goal against a progressive goal

10001100120013001400150001M2M3MOutput tokens spent

Take-home, GPT-5.6 Sol in a Ralph loop: given a progressive goal it gets below 1,000 cycles; given the one fixed goal it levels off near 1,060, on any axis.

Seed mean; the source shades the seed range, which is not redrawn. Read off the “Humanize Intro” deck, slide 25; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.

Claude degrades across turns, GPT does not, and Flame Chase helps ​

Measured by output tokens per response over a 24-hour run, GPT-5.6 Sol works at the same length from the first hour to the last in every flow. Claude Fable 5 starts with very long responses. In the flows that keep one session going (Always continue, Always prompt, ARAR), it falls to under a hundred tokens per response within five to ten hours. At that point it is barely working. The Flame Chase, where a fresh Fable session alternates with GPT, keeps it near a thousand. On the take-home, that is also where the cycles keep falling.

Fig. 4

Output tokens per response over a 24-hour run

101001k10k100k0h4h8h12h16h20h24hActive elapsed time in a single run

GPT-5.6 Sol: every flow holds between roughly 300 and 1,000 output tokens per response for the whole 24 hours.

Smoothed per-flow trend lines, max effort; the source’s per-response scatter is not redrawn. Read off the “Humanize Intro” deck, slides 26 and 27; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.

The talk also showed this measurement for a Ralph loop at every effort level on ten models, five from each vendor. The GPT models stay flat. Some of the Claude models decay over the run, most clearly Fable 5 at both effort levels and Opus 4.8 at xhigh.

Engineering, or a science? ​

Is a flow just one more kind of prompt, context, skill or loop engineering? We think it can be more than that.

  1. XXX engineeringone prompt, context, skill or loop · one task

    Prompt, context, skill and loop engineering: each fix is tuned to the task in front of it, and the dirty work does not carry over to the next one.

  2. Flow engineeringone flow · one task

    A whole method written as a flow and pointed at one kind of work: HOA at Lean, KDA at kernels, HMA at MLE-bench. It works, and it is still one flow per job.

  3. Flow scienceone flow family · every task that reduces to one problem

    A family — the Ralph loop, the Flame Chase, the progressive goal — generalises to every task that reduces to the same problem class, and can be ablated and analysed like any other object of study.

A flow family can be studied the way a model architecture is: hold everything else fixed, vary one thing and measure the result. These are the axes along which we are ablating the Flame Chase:

Varies: Claude Fable 5 ↔ GPT-5.6 Sol · Opus 5 ↔ GPT-5.6 Sol · one model alone

Already on this site: Fig. 3 compares two models alternating with each model alone.

The work needs a runtime that can express any of these variations as a short piece of Python, run it for a day on real hardware and record every turn on one clock. That is what Humanize is. The engine underneath it is described in Deep Tech.

This closes the series. Start again from part one, or go to part two for the ablations these findings came out of.

Read next

BlogFlow Science: flow is the new scaling dimensionModels scale with parameters and data, and agents scale with tools. The next dimension is the flow. Why a single /goal session is not enough on long, hard tasks, and why one flow cannot solve every problem.Read it BlogFlow Science: different tasks need different flowsFlow-level ablations on MLE-bench, HLE, ProgramBench and Anthropic's performance take-home, with the models held fixed. The best flow differs from one benchmark to the next.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it