Skip to content
← All posts

Blog3 min read

Flow Scienceflow is the new scaling dimension

Models scale with parameters and data, and agents scale with tools. The next dimension is the flow. Why a single /goal session is not enough on long, hard tasks, and why one flow cannot solve every problem.

Flow is the new dimension of scaling agents. Models scale with parameters and data, and agents scale with tools. The next dimension is the flow: which agent goes next, what it is asked, what it remembers, and when it stops. Change only the flow and the same models land somewhere else.

This is the first of three posts on what we have measured so far. This one makes the case: a single session is not enough on long, hard tasks, and no one flow can serve every task. The second runs flows against each other on four benchmarks with the models held fixed. The third is what those runs found that we did not expect, and why we think flows are a science rather than one more kind of engineering.

Every figure is rebuilt from the charts in our October 2026 "Humanize Intro" talk. Hover over a figure, tap it or use the arrow keys to read values, and click a legend entry to hide that series. These are our own runs, Humanfia-reported, 2026-10-05, unless a figure links to something published.

From session to flow ​

A coding agent's own /goal is one session pursuing one objective until it says it is done. That works for small tasks. On long, hard ones it is not enough. On Anthropic's original performance take-home, /goal with GPT-5.6 Sol spends 2.6 million output tokens to reach about 1,080 clock cycles. A Ralph loop on the same model is below 1,100 within half a million. A Flame Chase alternating Claude Fable 5 and GPT-5.6 Sol keeps improving and finishes near 1,010.

Fig. 1

A simple /goal is not enough for complex tasks

1000110012001300140015000500k1M1.5M2M2.5M3MOutput tokens spent

The performance take-home, best cycle count against output tokens: /goal reaches about 1,080 cycles after 2.6 million tokens; the Ralph loop is below 1,100 within half a million; Flame Chase ends near 1,010.

The Always-continue arm is in the source but too faint there to trace, so it is left out here. Read off the “Humanize Intro” deck, slide 2; Humanfia-reported, 2026-10-05. Hover, tap or use the arrow keys to read values; click a legend entry to hide it.

The talk summed this up as "80 cycles better, 80% tokens saved" against /goal. Humanfia-reported, 2026-10-05.

No free lunch: one flow cannot solve every problem ​

Our first flow, RLCR Flow (Ralph loop with Codex review), is good at building a project from scratch. Coding of that kind is a constraint satisfaction problem. Every requested feature must be implemented, the tests must pass and the lints must be clean, and we do not care about the exact shape of the code. One feasible solution is enough.

A CUDA kernel is a different problem. A kernel that merely runs is not the goal: we want the fastest one that is still correct. That makes it an optimization problem, and real runs already guarantee that a candidate is feasible. For this we built the Flame-Chase loop: two fresh agents take turns on one shared repository. In our runs, Claude Code with Claude Fable 5 tends toward large refactors, which sometimes make things slower. Codex with GPT-5.6 Sol tends to fine-tune and gets stuck in a local optimum. Alternating the two lets the run jump from one local optimum to a better one.

Coding · constraint satisfaction

feasible
  • features implemented
  • tests pass
  • lints clean
One feasible solution is the whole job. → RLCR: propose, check, repeat until every check passes.

Kernels · optimisation

fine-tuning alone stops herecost ↓ is better · one axis of a huge design space
  • small step · fine-tune (the model that polishes)
  • big jump · refactor (the model that rewrites)
The best one is the job. → Flame-Chase: alternate the two, and the run hops from one local optimum to a better one.
Problem classWhat "done" meansFlow family
Constraint satisfactionany solution that passes every checkRLCR, the Ralph loop
Optimizationthe best solution foundFlame Chase, Parallel Flame Chase
Decisiona yes or no, with a certificateopen
Predictiona calibrated estimateopen
Game-theoretica strategy that holds against othersopen
Countingan exact numberopen

Theorem proving, hyperparameter tuning and many other tasks need flows of their own. That is why Humanize is a framework for writing flows, not a single loop.

Next ​

If the flow matters this much, the best flow should change from one task to the next, and it should show when the models are held fixed. That is the experiment in part two, "Different tasks need different flows".

Read next

BlogFlow Science: four findings, and engineering or a science?What the flow-level ablations found that we did not expect: higher diversity, higher score; better collaboration, better performance; a progressive goal keeps agents on track; and Claude degrades across turns where GPT does not. Then: is a flow just engineering?Read it BlogFlow Science: different tasks need different flowsFlow-level ablations on MLE-bench, HLE, ProgramBench and Anthropic's performance take-home, with the models held fixed. The best flow differs from one benchmark to the next.Read it BlogKDA²: Kernel Design Agents (KDA) optimize Kimi Delta Attention (KDA)Our agents wrote Kimi Delta Attention kernels that run up to 2.96× faster than FlashKDA on B300 with a tenth of its state error. Here is how, and how the agents tried to cheat along the way.Read it