Skip to content

fixed_interrupt_flame_chase ​

flame_chase on a clock. Two agents take turns in one workspace, each turn a fresh session, but a turn does not end when the agent says so: it ends after k accepted experiments, as a trusted evaluator counts them. At the end, the agent that did not write the latest accepted candidate reviews them all blind, and may pick an older one. This is the HMA paper's fixed-k alternation, as a flow.

flowverse v0.1.1

In hmz, open /flow → Flowverses → official and install fixed_interrupt_flame_chase. Or run the release without installing anything:

hmz exec -f 'git+https://github.com/humanfia/fixed-interrupt-flame-chase-flow@v0.1.1#fixed_interrupt_flame_chase' …
release
v0.1.1 · commit a43abbe
earlier
v0.1.0
licence
Apache-2.0

Read from humanfia/flowverse when this site was built.

sh
hmz exec -f fixed_interrupt_flame_chase \
    -a first_chaser=claude/claude-opus-5:max -a second_chaser=codex/gpt-5.6-sol:max \
    -p budget.duration=6h "$(cat TASK.md)"
sh
hmz exec -f fixed_interrupt_flame_chase \
    -a first_chaser=claude/claude-opus-5:max -a second_chaser=codex/gpt-5.6-sol:max \
    -p budget.duration=6h \
    -p gate_config=/ABS/trusted/gate.json -p run_dir=/ABS/results/new-task-run \
    "$(cat TASK.md)"
fixed_interrupt_flame_chasesimulated

The whole run, at rest. Step through it with the buttons, or drag the bar.

Each turn is a fresh session, closed the moment the evaluator has counted k accepted experiments — a number the agents are never told. At the end the agent that did not write the latest accepted candidate reviews them all blind, and may pick an older one.

The run, turn by turn
  1. first_chaser — experiments; a session opened for this turn; hands evaluator in the tree: k accepted
  2. evaluator — closes the turn; no turn of a model; hands second_chaser in the tree: the workspace
  3. second_chaser — experiments; a session opened for this turn; hands evaluator in the tree: k accepted
  4. evaluator — closes the turn; no turn of a model; hands first_chaser in the tree: the review reserve
  5. first_chaser — blind review; a session opened for this turn; hands the finish as words: the final candidate

Then round again: alternating, until only the review reserve is left.

It ends when the review settles the final candidate, a provider or the evaluator fails, or the budget runs out.

When to use it ​

When a command measures the work, a turn that runs until the agent feels finished wastes the measurement: an agent can polish one idea for an hour that a score already rejected. Here the evaluator, not the model, decides when the other agent gets the tree, so neither one holds it for long without results. The final review guards the other way, against the last accepted candidate winning only because it was last.

Two admission backends do the counting, chosen by whether gate_config is given:

FlowBench (gate_config unset)native MLE (gate_config set)
Countsrecords the cell's evaluator scored without an errorthe native evaluator's receipts
An actor submits withsubmit.py submit (built with the workspace's submit.sh)submit.py submit PATH.csv
Admission serviceinside the flow processevaluator.py, which you start

The native backend needs a trusted evaluator started beside the run, with control files outside the workspace. Its README walks through it.

Roles and params ​

RoleWhat it isHow it is filled
first_chaseragent, required-a first_chaser=…Takes the odd turns, each in a fresh session, and may be the reviewer.
second_chaseragent, required-a second_chaser=…Takes the even turns, each in a fresh session, and may be the reviewer.
workspaceenvironment, localthe directory you start in; no -eThe task's shared workspace, which both chasers work in. Files carry over; transcripts do not.

Each agent role takes one -a role=CLI[@PROVIDER]/MODEL[:EFFORT]; several roles may share one -a, comma-separated. There is no -e to give: workspace is a local environment, the directory you start the run in, and an -e naming it is refused. See Command-line specs.

Every param has a default, so FlowBench, which passes none, can run it:

ParamDefault
gate_configunsetThe trusted native evaluator's connection file, an absolute path. Unset for a FlowBench task.
run_dirunsetA new absolute output directory outside the workspace. Unset makes one under ~/.fixed_interrupt_flame_chase/.
evaluator_urlhttp://evaluatorThe FlowBench evaluator, used when there is no gate_config.
max_valid_submissions_per_session5k: accepted experiments before a turn closes.
active_time_limit_secondsunsetThe whole run's wall clock, review included. Unset takes the budget's duration, else six hours.
review_reserve_seconds900Wall time kept back at the end for review and finalization.
review_turn_seconds600The longest the reviewer may run.
rest_seconds1Seconds between one turn and the next, at most 60.

The rest, finalize_reserve_seconds, cleanup_reserve_seconds, build_timeout_seconds and poll_seconds, are in the README.

What ends it ​

  • The wall clock. Turns alternate until the exploration deadline, which is the wall clock less the review reserve. A score that lands after it is not counted.
  • One final review. With at least two accepted candidates, the agent opposite the latest one's author sees them all, by identity and authorship, with no scores, and nominates one. Silence or an unknown nomination keeps the standing candidate. There is one ballot and no retry.
  • The budget. A budget tighter than the wall clock can end the run earlier, even before the review.
  • A failure. A provider or evaluator error during exploration fails the run; it is not treated as a handoff.

Picking it up ​

It cannot be picked up, by design: an admission ledger cannot be trusted twice. After a failed run, keep its evidence and start a new one, with a fresh run_dir and fresh evaluator state. run_dir holds contract.json, turns.json, review.json and result.json.

Local is not isolated

The paper's runner gives every turn a new container and home directory. A local workspace does not: sessions are fresh, but the home directory, installed packages and the machine are shared, and nothing outside the workspace is hidden from an agent that goes looking. Use the HMA reproduction's isolated runner for results that must match the paper's containment.

See also ​