Skip to content

HumanfiaWe build the flow around the agents

HumanfiaHumanfia

blog

What came back

Every number we have published, and what it takes to check it.

HOA

Lean-Eval: first place, and 172 research-level proofs

A fully agentic run took first place on the Lean-Eval leaderboard with 172 research-level mathematics problems — every accepted proof sorry-free and independently re-verified.

Zhengyang Zhang · Hongzhou LinRead it

Method

The review is the next prompt

Read it

Architecture

Four layers and a referee

Read it

HKA

Ten workflows, nineteen competitions

Read it

KDA

The MSA indexer, 3.3× faster in production

KDA 1.5 found the hardware underutilization in the MSA prefill and decode indexers on B300 — a 6.5× geometric mean on prefill, 3.3× on long-context decode, and bitwise-identical output throughout.

Jiaming TangRead it

Humanize 2

ProgramBench: 3.5%, against 0.5% and 0%

The same two models that solve 0.5% and 0% of ProgramBench on their own solve 3.5% of it arranged as a builder and a reviewer in a loop — and the curve had not flattened when the clock ran out.

Read it

KDA

A quality tier, so the fast kernel can ship at all

Most profitable diffusion fusions change the output at the level of bf16 rounding order — enough to fail a bitwise CI, not enough for a human to see. So we gave the request a quality tier instead of arguing about the diff.

Read it

KDA

KDA 1.5 goes past human SOTA on every contest track

A rerun of the MLSys 2026 FlashInfer kernel contest with CuteDSL support, a profiling skill and a better loop — 1.25× to 1.39× past the best human entries on all three tracks.

Dongyun ZouRead it

Humanize 2

QCode Discovery, made fail-closed

Read it

HOA

Physics and quantum, formalized end to end

Read it

HOA

Lean-Eval: 149 of 219, and second place

A general agent loop on a general model, second on a leaderboard led by a purpose-built prover — and six problems behind it.

Read it

HOA

Six of six at IMO 2026, on two different backends

A fully agentic, YOLO-style run solved every problem of the 2026 International Mathematical Olympiad, machine-checked in Lean 4 — 3.2× faster than the previously reported agentic result, with both backends closing all six.

Read it

All posts · RSS

ecosystem

The projects, and the loop between them

drivesis pointed atis scored by THE FLOWS WE WRITE · WHAT IT RUNS RLARFlame ChaseHumanize 1Ralph LoopGoalHumanize 2THE RUNTIMEWHAT IT DRIVESThe agentsclaude · codex · dsh · agyand every other CLI youalready log intoAPPLICATIONKDAKernel Design AgentsFaster, or it is notAPPLICATIONHOAHumanize Olympic AgentsLean accepts it, or it does notAPPLICATIONHKAHumanize Kaggle AgentKaggle says so, or it does notTHE REFEREEFlowBenchScores the flows againsteach other. Ours included.MEASURESSELECTSSHIPS IT BACK
Scroll the diagram sideways to see all of it. Humanize 2 runs the flows and drives the agents you already log into; a flow is pointed at one of the three applications, which is where it is found out; and FlowBench scores what came back and sends the winner into the next flow. Hover any piece to hold it.
ApplicationsWhere a flow is found out

HOA, KDA and HKA. Chosen because somebody else keeps the scoreboard.

FlowsThe method, as code

RLAR, Flame Chase, Humanize 1, the Ralph loop — directories of Python anyone can read, fork or beat.

RuntimeHumanize 2: Agent Flow System

Opens and resumes sessions, takes the turns a flow asks for, puts work in a container or on another machine, and writes the run down as a timeline.

BackendsThe agents you already log into

claude · codex · dsh · agy · grok · kimi · qwen · pi · opencode · mimo. We hold no API key.

The refereeFlowBench

Scores the flows against each other on long-horizon work, and the winner becomes the next default. Ours included.

measures · selects · ships it back into the flows

Everything points downward: a flow asks for turns, the runtime takes them on a CLI you already pay for, and an application is where the result stops being our opinion. The one arrow that runs the other way is the benchmark's.