rlar
Have every round of work reviewed, and stop when the reviewer agrees it is done. An actor works in one session that remembers; a fresh reviewer reads the repository after each round, and its review is the actor's next prompt, word for word.
Nothing to install: it is in hmz from the first run, with chat and the other loops, and its code is humanize's flows/builtin/rlar.
❯ $rlar add undo and redo to the editorhmz exec -f rlar \
-a actor=claude/claude-opus-5:high -a reviewer=codex/gpt-5.6-sol:high \
-p budget.duration=6h,budget.cost=60 "$(cat TASK.md)"The whole run, at rest. Step through it with the buttons, or drag the bar.
Every review answers two things: whether the task is done, and what the actor hears next.
The run, turn by turn
- actor — the task; a session opened for this turn; hands reviewer in the tree: a round of work
- reviewer — reviews the work; a session opened for this turn; hands actor as words: done: false, and its notes
- actor — the notes; another turn of the session it already had; hands reviewer in the tree: a round of work
- reviewer — reviews the work; a session opened for this turn; hands the finish as words: done: true
Then round again: while the review says there is more to do.
It ends when the reviewer says done, 3 failures in a row, or the budget runs out.
When to use it
When "is it done?" should not be answered by the agent that did the work. Every review comes from a session that has just started: it reads the repository itself and knows nothing of how the work was arrived at. Give both roles the same model if you like; that asymmetry is still the point.
Each review answers in a fixed shape, so the flow reads a field rather than hunting for a phrase:
| Field | What the reviewer says |
|---|---|
done | true only if everything asked for is implemented, works, and nothing was faked, stubbed or special-cased to pass |
notes | The review itself, written to the actor: what is wrong or missing and what to do next, citing files and lines |
The reviewer is told to be skeptical, and to treat reward hacking, such as weakened tests or stubbed-out work, as the thing it is there to catch. How it reads a round and writes its notes is the flow's own skill, review-notes, carried by every review session.
Roles and params
| Role | What it is | How it is filled | |
|---|---|---|---|
actor | agent, required | -a actor=… | Does the work, in one session held for the whole run. |
reviewer | agent, required | -a reviewer=… | Reads the repository after each round, in a fresh session every time. |
workspace | environment, local | the directory you start in; no -e | Where the actor works and the reviewer reads. |
Each agent role takes one -a role=CLI[@PROVIDER]/MODEL[:EFFORT]; several roles may share one -a, comma-separated. There is no -e to give: workspace is a local environment, the directory you start the run in, and an -e naming it is refused. See Command-line specs.
No params. The loop pauses 5 seconds between rounds.
What ends it
- The reviewer says
done. Its notes are the last thing the run prints. - The budget. The ceiling, not the usual end.
- Three failures in a row. A failed turn, or a review that does not fit the shape, is taken again next round; the third in a row ends the run with that failure.
Picking it up
--resume carries on the round count and the last review nobody has acted on. The actor's session is not picked up, so the resumed actor is sent both: the task, and under it that review, marked as a reading of work this session did not do.
That review may even be of another task: --resume picks up the newest run of rlar in this directory, whatever it was asked. A run the reviewer agreed with keeps nothing. See Picking a run up.
See also
- humanize1: the same actor-and-reviewer idea, with a plan agreed first
- flame_chase: two agents both working, rather than one reviewing