Assistants Analysis

Claude Code's auto mode catches most risky actions. Anthropic's numbers show what it misses

A 0.4% false positive rate sounds great. The 17% miss rate on real overeager actions is the number to plan around.

Anthropic published how Claude Code’s auto mode works, with measured error rates. That’s unusual for a security feature, and it makes the engineering post worth reading before you let an agent run unattended. The full write-up is on Anthropic’s engineering blog.

The starting point is approval fatigue. Anthropic says users approve 93% of permission prompts. When nearly every prompt is a yes, people stop reading them, and the prompt that matters gets the same reflexive click as the rest. The alternative many people reach for, the skip-permissions flag, removes the safety net entirely. Auto mode is the middle path: a classifier reviews actions, so you only get interrupted for things that look risky.

How it works

Two layers. For what Claude reads, a server-side probe screens tool outputs for prompt injection before they enter the agent’s context, adding a warning when it sees hostile content. For what Claude does, a classifier evaluates each action in two stages: a fast single-token yes or no, then chain-of-thought reasoning only when the first stage flags something. Anthropic says the classifier ran on Sonnet 4.6 at the time of writing.

One design choice stands out. The classifier sees only your messages and the agent’s tool calls. Claude’s own prose and tool results are stripped, which Anthropic calls reasoning-blind by design. That stops the agent talking the classifier into approving something.

The four threats it targets are overeager behavior, honest mistakes about scope, prompt injection, and misaligned models (not currently observed). Examples it should catch include deleting remote branches without explicit confirmation and hunting for alternative API tokens.

The numbers

Test setFinal miss rateFinal false positive rate
Real traffic (10,000 actions)n/a0.4%
Real overeager actions (52)17%n/a
Synthetic exfiltration (1,000)5.7%n/a

The first stage alone had an 8.5% false positive rate on real traffic. The second stage brings it down to 0.4%. That’s the price of a second pass, and it’s why the design has two.

Now the figure to plan around. On 52 real examples of overeager actions, the full system let 17% through. Anthropic is blunt about it: whether 17% is acceptable depends on what you’re comparing against, and auto mode is not a drop-in replacement for careful human review on high-stakes infrastructure. The sample is small, and the authors say the classifier tends to find approval-shaped evidence without checking whether it covers the blast radius of the action. Read that as: “yes, go ahead and clean up the branches” is not consent to delete the remote ones.

What that means for how you use it

Auto mode is the built-in starting mode for interactive terminal and VS Code sessions from Claude Code 2.1.283, according to Anthropic’s docs. So you may already be using it.

Match the mode to the environment. A scratch repo or a sandboxed container is a good fit, because the cost of a miss is a bad afternoon. A shell with production credentials is not.

Narrow the blast radius first. The best-practices guide pairs it with permission allowlists for commands you trust, and with sandboxing, which restricts filesystem and network access at the OS level. Those limit what a missed action can do, which is a different kind of protection from catching it.

Keep secrets out of reach. Of everything the classifier is asked to prevent, credential misuse is the one where a 17% miss is hardest to live with.

Next step: check which permission mode your sessions start in, and for anything that touches shared infrastructure, set it deliberately.