You type a prompt, read what comes back, spot the problem, type the next one. Somewhere in that cycle you are the part that decides whether the job is finished, so the work moves at the speed you check it. That’s the bottleneck, and it isn’t the model.
This is the fix. You write down what finished looks like as something a machine can check, and Claude keeps taking turns until a second model agrees it’s met. You stop being the loop.
This needs Claude Code, version 2.1.139 or later. Terminal, desktop app or headless, but not the chat app. Every command below is copy and paste.
The 3 pieces
The job. A completion condition, written so Claude’s own output can prove it. This is /goal. Setting one starts a turn straight away.
The check. After every turn, a separate small fast model reads your condition and the conversation so far, then returns yes, no, or impossible, with a short reason.
The retry. On a no, Claude starts another turn with that reason as guidance instead of handing control back to you. Nobody asks you anything.
The piece that decides everything is the first one. The evaluator does not run commands and does not read files. It only judges what Claude has already put in the conversation. So “the code is cleaner” can never pass, but “npm test exits 0, output shown” can, because Claude runs it and the result lands in the transcript where the evaluator can read it.
That’s the whole rule. If Claude can’t demonstrate it, you can’t loop it.
Before you start
What you need: Claude Code 2.1.139 or later, the workspace trust dialog accepted for the folder you’re in, and hooks not disabled in settings. /goal runs on the hooks system, so if any of those are missing the command tells you why rather than silently doing nothing.
How long: about 10 minutes to set up, and most of that is writing one good condition.
What it costs: the evaluator runs on the small fast model configured for your provider, Haiku by default on the Claude API, and those tokens are usually tiny next to the work itself. The turns are the real spend. A loop with no bound on it is a loop with no bound on your bill.
Works on: terminal, the desktop app, Remote Control, and non-interactive -p mode.
Step 1. Write the done condition
Time: 5 minutes. This is the step people skip, and it’s the one that decides whether the loop works. A condition that holds up over many turns has 4 parts.
- One measurable end state. A test result, an exit code, a file count, an empty queue.
- A stated check. How Claude proves it.
npm test exits 0,git status is clean,the script prints DONE. - Constraints that matter. Anything that must not change on the way there.
- A bound. A turn or time clause so it can’t run all night. Claude reports progress against it each turn.
The template:
/goal [end state]. Prove it by running [command] and showing me the output. Don't change [thing that must stay put]. Stop after [N] turns if it isn't met.
Conditions are capped at 4,000 characters, so there’s room to be specific.
Good and bad, side by side
Dead: /goal the auth module is refactored properly. Nothing in there a machine can read. The evaluator will guess, and it’ll guess yes.
Works: /goal every file in src/auth is under 300 lines and npm test exits 0. Prove it with wc -l and the test output. Don't touch any test file. Stop after 15 turns.
Dead: /goal fix the failing tests. No proof, no bound, no constraint.
Works: /goal npm test passes with 0 failures, shown in full output. Don't delete or skip any test to get there. Stop after 20 turns and tell me what's left.
That second constraint is not padding. Left alone, the fastest way to make tests pass is to delete them.
You’ll know it worked when you can read your own condition back and point at the exact command that proves it.
Step 2. Set it running
Time: 1 minute. Start Claude Code and paste your condition:
/goal all tests in test/auth pass and the lint step is clean
That starts a turn immediately, using the condition itself as the instruction. You don’t send a second prompt. A ◎ /goal active indicator shows how long it’s been going.
Every verdict the evaluator returns shows up in the transcript. Press Ctrl+O to see the reason behind it, which is the fastest way to find out your condition was vague.
You’ll know it worked when the indicator appears and Claude starts working without you prompting again.
Step 3. Stop it asking permission
Time: 1 minute. Here’s the part that catches people. A goal doesn’t change your permission mode. In manual mode Claude still stops and asks before any tool call your settings haven’t already allowed, so you’ve removed the per-turn tapping and kept the per-tool tapping. It looks like the loop is broken. It isn’t, it’s waiting on you.
Run it in auto mode. Auto mode sends each tool call past a classifier before it runs, which approves the routine stuff and blocks things like mass deletion or writes to credential files. The two do different jobs and you want both: auto mode removes the per-tool prompts, /goal removes the per-turn ones.
The other option, --dangerously-skip-permissions, turns the gate off entirely. Anthropic’s own guidance is that skipping permissions can be destructive and shouldn’t be used outside an isolated environment, so keep it for a container and use auto mode everywhere else.
You’ll know it worked when you walk away for 10 minutes and come back to turns that have run, not a prompt waiting for a y.
Step 4. Watch the first run, always
Run /goal with no arguments at any point to see where it’s at:
/goal
You get the condition, how long it’s been running, how many turns have been evaluated, the current token spend, and the evaluator’s most recent reason. The turn count and reason only appear after the first evaluation has run.
To stop it before it resolves:
/goal clear
stop, off, reset, none and cancel all work the same way. Starting a new conversation with /clear also removes it.
Never set a new condition and walk off before watching one full turn and one verdict. A bad condition costs money quietly.
Step 5. What actually stops it
The evaluator returns one of 3 verdicts.
- Not yet met. Claude keeps working, and the reason gets passed in as guidance for the next turn.
- Met. The goal clears and an achieved entry goes in the transcript.
- Impossible. The evaluator has decided the condition can never be satisfied. It clears itself and records the reason.
There’s also a stall guard. If Claude keeps answering the evaluator without using any tools for several turns in a row, Claude Code stops the loop, prints a warning and hands control back with the goal still set. Evaluation picks up again on your next prompt.
And 4 kinds of failure clear the goal outright: an authentication failure where Claude Code manages its own credentials, an exhausted credit balance, a context overflow that auto-compaction couldn’t clear, and a model that isn’t available. Anything else, including rate limits and overloaded servers, leaves the goal running.
Step 6. When the trigger is time, not a condition
Time: 2 minutes. /goal is for work with a finish line. When you’re watching for something to change instead, you want /loop, which re-runs a prompt on a clock inside your session.
/loop 5m check if the deployment finished and tell me what happened
Units are s, m, h and d. Leave the interval off and Claude picks its own gap each round, between 1 minute and an hour, based on whether anything is happening:
/loop check whether CI passed and address any review comments
Press Esc while it’s waiting to stop it. Three things worth knowing before you rely on it: tasks are session-scoped and die when the session does, recurring ones expire 7 days after creation, and a session holds a maximum of 50 at once.
Save your default loop
Drop a file at .claude/loop.md in the project, or ~/.claude/loop.md for everything, and a bare /loop runs that instead of the built-in maintenance prompt. Plain markdown, no required structure, written like you’d type the prompt. Edits take effect on the next iteration, so you can tune it while it runs. Keep it under 25,000 bytes.
Step 7. Run it while you’re asleep
/loop needs your session open, which makes it the wrong tool for overnight. There are 3 ways to schedule work and they’re not interchangeable.
| Cloud routines | Desktop tasks | /loop |
|
|---|---|---|---|
| Machine has to be on | No | Yes | Yes |
| Session has to be open | No | No | Yes |
| Access to local files | No, fresh clone | Yes | Yes |
| Permission prompts | None, runs autonomously | Set per task | Inherits from session |
| Shortest interval | 1 hour | 1 minute | 1 minute |
For anything that has to survive a closed laptop, use a cloud routine. For anything that needs your actual files, use a desktop scheduled task.
And /goal works headless, which is the version worth stealing:
claude -p "/goal CHANGELOG.md has an entry for every PR merged this week"
That runs the whole loop to completion in one invocation, no terminal session needed, so it drops straight into CI or a cron job. With the default text output nothing prints until it finishes, so a long run looks frozen. Add --output-format stream-json --verbose to watch it work. Ctrl+C to kill it.
3 conditions to steal
A test suite that has to go green
/goal npm test exits 0 with the full output shown. Don't delete, skip or weaken any existing test. Stop after 20 turns and summarise what's left.
A backlog that has to empty
/goal every open issue labelled "triage" has a priority label and an owner. Prove it by listing the remaining triage issues, which should be none. Don't close anything. Stop after 30 turns.
A pile of files that has to be processed
/goal every .csv in ./inbox has a matching .json in ./out. Prove it by listing both folders and showing the counts match. Don't modify anything in ./inbox. Stop after 25 turns.
Notice what all 3 have in common: the proof is a command whose output lands in the conversation, and there’s a bound on the end.
If it’s not working
/goal isn’t available. You’re below 2.1.139, the folder hasn’t been trusted, or hooks are turned off in settings. The command tells you which one, so read the message rather than guessing.
It said done and it isn’t. Your condition wasn’t provable from the transcript. The evaluator can’t run anything or open anything, it only reads what Claude surfaced. Rewrite it around a command and its output.
It runs forever. No bound. Add the turn clause. This is the expensive failure.
It stops after one turn. Usually manual mode, so it’s waiting on a permission prompt. Otherwise, if a subagent or a background command is still running when a turn ends, evaluation is skipped for that turn and picks up at the end of the next one that finishes clean.
It’s deleting things to pass. Your condition had no constraints. The check is the spec, so anything you didn’t forbid is on the table.
The bill was bigger than expected. The evaluator is cheap. The turns aren’t. Bound every goal.
The part nobody clips
Boris Cherny, who created Claude Code, is the reason this idea went everywhere. The line people quote is that his job now is to write loops. What gets left out is the constraint underneath it, and the constraint is the useful half: every loop he actually names has a success condition a machine can check for free. It’s the cost of checking, not the cleverness of the loop, that decides what you can automate.
So the honest test before you build one: how much does it cost to find out whether this is done? If checking is nearly free, loop it. If checking means you reading it, you’ve just built a machine that makes mistakes unattended.
One more, and it’s not optional. A met condition means your check passed. It does not mean the work is good, or safe, or the right approach. Read the diff.
The final check
- Claude Code is 2.1.139 or later and
/goalruns - My condition names a command that proves it
- My condition says what must not change
- My condition has a turn or time bound on it
- I’m in auto mode, not manual
- I watched one full turn and one verdict before walking away
- I ran
/goalmid-run and read the token spend - I read the diff after it said it was done
