Verdict. I’d only build a loop where the output is cheap to verify. Anything gated on taste stays a prompt.

what it is

Boris Cherny, who architected Claude Code, moved from manual prompting to loop engineering. He writes automated agent systems that decide what to build and when, instead of being told each step. His loops coordinate hundreds of agents watching his GitHub, Slack and Twitter. Anthropic engineers reportedly merge about 8 times more code per day. The leverage moved from prompt quality to system design.

His reference patterns are a CI-triage loop and dependency-bump automation (.claude/skills/dep-bump/SKILL.md plus a STATE.md and /goal). Addy Osmani and Peter Steinberger have made the same move.

what I found

Cherny’s test for whether a loop is worth building.

  1. The task repeats weekly or more, so the build pays for itself.
  2. The loop can verify its own output without me.
  3. There’s token budget to let it run.
  4. Agents can test their own work mid-run.

Most agent systems fail it. The hidden cost is comprehension debt, the gap between what the loop ships and my ability to reason about it.

I scored my own candidates against the four tests.

Candidate loopRepeatsSelf-verifiesVerdict
Pipeline health (did the content ingest run, did the schedule no-op)dailyyes, counts and timestampsStart here
Nightly flag-to-draft for postsdailypartlySecond, and only to the draft stage
Full auto-publishdailyno, it’s tasteNo

compared to