Blog
Results, post-mortems and the occasional strong opinion about long-horizon agent work. Subscribe by RSS.
Four layers and a referee
How Humanfia is put together — a runtime, a flowverse, a benchmark and two applications — and why the arrow from the benchmark back into the flows is the only part that matters.
Read itThe review is the next prompt
Two small decisions inside RLAR — the reviewer's words go to the actor verbatim, and "done" is read off a field rather than a sentence — that changed how long-horizon runs behave.
Read itHumanfia: from automated idea factory to realization
Introducing an open-source agent workflow for serious L3–L4 development — and the argument that the workflow, not the model, is the thing worth building.
Read it