Skip to content
← All news

News1 min read

567 files, one engineeragents port gem5's build from SCons to CMake

gem5's whole build system, ported from SCons to CMake by one engineer directing agents for weeks — the least glamorous long-horizon task there is, and the one that taught us the most.

gem5's build system was migrated from SCons to CMake in full: 567 files changed, by one engineer and a set of agents running for weeks.

This is the oldest result on this page and, for what we were trying to learn, one of the most useful — because a build-system migration is the least glamorous long-horizon task there is, and it has almost none of the properties that make a benchmark flattering.

There is no clever insight to have. There is no moment where a good idea collapses the search. There is a very large number of mechanical changes, each of which is easy, and a build that either works or does not, and a long stretch in the middle where it does not work for a reason that has nothing to do with the change you just made.

Work of that shape is exactly where a session that starts fresh every turn falls apart, where a run that keeps everything in context drowns, and where an agent left unsupervised will eventually declare victory on a build that compiles a subset. It is the reason we started caring about the loop rather than the model, and it is the kind of task FlowBench is being assembled out of.

One engineer, some weeks, 567 files, and a build that works.

The pull request, gem5/gem5#2969 · FlowBench

Read next

News672 of 672: we closed PutnamBench's last two, and three others got there tooGPT-5.6 in a worker–reviewer loop proved all 672 PutnamBench statements in Lean, at $44.50 a problem. The top line is now a four-way tie, and what separates the four is cost.Read it NewsFull marks on five science exams, and the fine print on every oneIPhO 2026 theory 30 of 30, IChO 2026 formalized 68 of 68, IOI 2026 six of six, IBO 2024 theory 100 of 100, and quantum information theory 40 of 40. What each score is, who graded it, and what it does not show.Read it News30 first solves on Lean-Eval v1, the most of anyone — and second by one on totalLean-Eval froze its first 128-problem release and moved everything older to an archive. On v1 our GPT-5.6 flow has the most first solves and is one problem behind the total leader; on the archive it has 170 of 171.Read it