Why your Claude Code routine reported success and did nothing

There is one sentence in Anthropic’s routines documentation that explains more confusion than anything else on the page: a green status in the run list means the session started and exited without an infrastructure error — it does not mean the task in your prompt succeeded.

That is the whole problem with unattended automation in one line. A routine runs with nobody watching. The run list is the only thing most people ever look at. And the run list is reporting on the wrong thing: it tells you the machine worked, not that the job got done.

This article covers the three ways a routine fails while reporting success, and the eight-point pass that catches all three before you schedule anything.

Three failures that look exactly like success

1. The source that could not be read

Your mail connector times out. The routine writes “no urgent messages”. You read that over coffee as good news.

It is not good news. It is no news — and the two are indistinguishable in the output. This is the most expensive failure mode in unattended work, because nothing about it looks wrong. A calendar connector that fails renders an empty schedule, and an empty schedule reads as a free day. You find out on Thursday, in the meeting you missed on Tuesday.

2. The check that never ran

A health-check routine that cannot reach your endpoint, and says nothing about it, reads exactly like a healthy system. Silence is ambiguous, and ambiguity always defaults to “fine” in the reader’s head.

Worth knowing the specific signature here: on the default cloud environment, network access is Trusted, which allows only Anthropic’s default allowlist. A request to a host outside that list fails with a 403 and the header x-deny-reason: host_not_allowed. That is a real, named, catchable error — but only if your prompt is written to catch it rather than to carry on.

3. The run that fired at the wrong time

Schedule a routine exactly on the hour and it can start several minutes late; Anthropic’s own recommendation is to pick 9:07 rather than 9:00. Local Desktop tasks are worse. A 7am brief missed because the laptop was shut can run as a catch-up when the machine wakes at 11pm — and still write “today” in the present tense, about a day that is nearly over.

None of these three produce an error. All three produce green.

The eight-point pass

Run any routine prompt — a template you accepted, or one you wrote yourself — through these eight checks before you schedule it. Each one closes a specific hole.

# Check The question it answers
1 Self-contained If a stranger ran this with no context, would it work?
2 Sources named Can every number in the output be traced to something read this run?
3 Verified Did the routine re-check its own work before reporting?
4 Idempotent If this runs twice, do I get one result or two?
5 Honest when blind Does an unreadable source look different from an empty one?
6 Least privilege Which connectors are attached, and which of them can write?
7 Fails loudly When it breaks, do I find out — and does it tell me what to do?
8 Time-aware If this fires nine hours late, is the output still true?

Check 1 — Self-contained

A routine session has no memory of your conversations. It cannot ask a follow-up question. Anything the prompt assumes you will supply is simply missing. Write it as instructions to a competent stranger who has never met you: name the files, name the windows, name the thresholds.

Check 2 — Sources named

Every figure in the output should be traceable to something the routine actually read during this run. The failure this prevents is subtle: a plausible number that came from nowhere. Require the prompt to list its sources and counts, and a fabricated figure has nowhere to hide.

Check 3 — Verified

Add a step where the routine re-opens its own output and checks it against what it read. Counts match the items listed. Times appear as read, not rounded. Conclusions trace to evidence. Make that step able to fail the run — a verification section that always passes is decoration.

Check 4 — Idempotent

Routines fire more than once. A GitHub trigger fires on every push to an open pull request, and Claude Code does not reuse sessions across events — five pushes means five independent runs. Key your state on something stable (the head commit SHA, not the PR number) so a repeat run writes nothing instead of writing a duplicate.

Check 5 — Honest when blind

This is the one almost every prompt fails, and the one that costs the most. The rule is short:

A source that could not be read must never render as a source with nothing in it.

The fix is a few lines, and it is the single highest-value edit you can make to any routine prompt:

For each source, record exactly one of two outcomes:
  read: <count> items for <explicit window>
  unavailable: <the exact error or refusal text>

Never infer absence from a failure. If a source is unavailable, the
section it feeds says "unavailable: <reason>". An empty section and
an unreadable section must never look the same.

Two outcomes, never one. That is the entire trick, and it converts the most dangerous failure in unattended automation into an obvious one.

Check 6 — Least privilege

The platform defaults against you here, and it is worth being blunt about it. When you create a routine, every connector currently attached to your account is included by default. And during a run, Claude can use every tool from an included connector — including writes — without stopping to ask for permission.

So a template you accept without editing arrives holding your entire connector list, write scopes and all, and will use any of it unprompted if the task seems to call for it. Strip that list down to what the routine actually needs before the first run, not after the first incident.

Check 7 — Fails loudly

Because green tells you nothing, the routine has to tell you itself. Give every run an explicit status — OK or PARTIAL — and make PARTIAL the result whenever a source was unavailable or a verification check failed. Append one line per run to a log file. Then a routine that has quietly been failing for a fortnight is a thing you can see, instead of a silence you have to notice.

Check 8 — Time-aware

Have the prompt read the clock first and compare it to the slot it was meant to run in. If it is hours late, say so in the first line of the output and report on the window the slot intended — not the last 24 hours from now. A brief written at 11pm should not contain a “today’s focus” list at all.

What this looks like finished

Applied together, these eight turn a prompt that describes a task into a prompt that describes a task and what to do when it doesn’t work — which is the only difference that matters once nobody is watching.

The companion piece to this article walks the whole pass through one real template end to end, with the complete hardened prompt you can paste in: How to harden the Briefing routine template.

Also useful: Claude Code routines — every limit, cap and gotcha, which covers the hourly caps, the dropped GitHub events, and the 72-hour GitHub expiry that switches a routine off entirely.

Cloud routines vs local scheduled tasks — which of the three scheduling options to use, and what each one does to you if you pick wrong.

The Routines Runbook is 32 ready-made routines built this way from the start, plus hardened replacements for all eight of Anthropic’s free templates and the fourteen guardrail patterns behind them. See what is in it, or take the free 5-routine starter.


Platform behaviour described here was checked against Anthropic’s Routines and Desktop scheduled tasks documentation on 5 October 2026. Routines are a research preview and the details move; this page carries the date it was verified so you can tell how stale it is. Not affiliated with Anthropic.