Useful tools · Running a team

A loop with no finish line is a money fire

An agent left running will keep running until something stops it, and the thing that usually stops it is the bill. Two things separate a system that works overnight from a money fire: a goal with an exit the machine can actually reach, and a checker that is not the same agent that did the work.

An agent left running does not eventually get there. A loop with a vague goal and no checker runs until something stops it, and the thing that stops it is usually the bill.

The two things that make the difference

A goal with an exit. "Improve my content" runs forever. "Read the five newest comments, draft one reply each, stop" finishes. The difference is not ambition, it is whether the instruction contains a condition the machine can evaluate as done.

A checker that is not the maker. A second pass that asks did this meet the goal before anything ships. Not the same agent reviewing its own work — every model is a soft grader on its own output. Ask it whether the code is good and it will tell you the code is beautiful.

Everything else in loop design is arrangement. Those two are load-bearing.

Which settles a question people argue about: unattended continuation is safe exactly when a checker exists, and reckless when one does not. "Always keep going" is a sensible rule right up until nothing is grading the output.

Runone round of real work
Scoremeasure the result
Checkerpass, or fail?
Fail → diagnose, change, restart, loop back to Run.
Pass → stop condition met. The loop ends.

A loop with no checker has no reason to stop — it will probe a dead shell as happily as it does real work, and just as expensively. The checker is what makes unattended repetition safe rather than reckless.

Where we are strong, and it is the checker

We run maker and checker as separate people who cannot do each other's jobs. Mason builds. Cassandra tries to break it using the wrong credential, to prove strangers stay out. Beck checks the owner can still get in, using the right one.

Beck's seat exists because a change once passed the first check and failed the second: a correct security fix that locked the owner out of his own dashboard. Two checkers with opposite failure modes catch things one checker never will, and the trust does not come from having a reviewer — it comes from the reviewer being structurally unable to build.

Where we are weak, and it is the second requirement

The playbook's second requirement is separate workspaces, so parallel agents do not collide. We do not have that. Everyone works in the same tree, with no isolated lanes — a real limitation of how we run rather than something we have solved.

There is a specific way that bites, and it is worth knowing whatever your setup looks like: a deploy can remove an edge function while every page still returns 200. The site looks perfect from outside. Every status check passes. The form quietly stops working, and the person who finds out is a customer who does not write in to tell you.

So do not test a deploy by loading pages. Send a real submission through your own contact form after every deploy and confirm it arrived. It takes a minute, and it is the only check that proves the thing you actually care about.

The order that keeps it cheap

Pick a boring, repeating task — the most repeated, not the hardest. Write the goal with an exit in it. Do the job once by hand and watch where it guesses wrong, because that tells you what context it is missing. Add the skill and the checker. Automate it last, and keep a human gate — draft, do not send — for the first week.

That last one is not a starter setting here, it is the design. Our mail script sends only on instruction. Our inbox specialist has no sending tool at all, on purpose.

What to do

Take your longest-running automation and check it has both: a condition it can reach, and a second pass it does not control. If it is missing either one, it is not a loop yet. It is just something running.

Want this running in your own practice? Let's talk.