Security How-to

onFailure: block, tested. Which hook failures it stops and which it doesn't

The Claude Code hooks reference now documents the option that closes the fail-open gap. Nine broken hooks, run for real, show where the protection ends.

A tall concrete doorway sealed shut, its edges glowing gold in a dark hall

A PreToolUse hook that can’t start, crashes or hangs lets the action through unless the hook handler says otherwise. The way to say otherwise is one field, "onFailure": "block", added in Claude Code 2.1.295 on 8 October. An earlier piece on fail-open hooks could only point at the changelog line, because the hooks reference didn’t describe the field yet. It does now, so this is the follow-up: where the field goes, what it covers, and what it quietly doesn’t.

The docs make a list of what counts as a failure. A list is a claim, so each item below was run against Claude Code 2.1.296 in a scratch project, and one of them didn’t behave as documented.

Where it goes

onFailure sits on the hook handler, next to type and command. It applies to command and http handlers. The docs list it in neither the mcp_tool, prompt nor agent tables, so a prompt-based gate has no equivalent.

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash",
        "hooks": [
          {
            "type": "command",
            "command": "${CLAUDE_PROJECT_DIR}/.claude/hooks/gate.sh",
            "timeout": 2,
            "onFailure": "block"
          }
        ]
      }
    ]
  }
}

The value is "continue" (the default) or "block". With block, a failure does what exit code 2 does on that event: a PreToolUse failure blocks the tool call, a UserPromptSubmit failure blocks the prompt, and on PermissionRequest it denies the request.

Keep the timeout short on a gate. The default for command hooks on most events is 600 seconds, and a blocking hook that hangs for ten minutes is a different kind of outage.

The test rig

Each case is a directory with a .claude/settings.json like the one above and a gate.sh that misbehaves in one specific way. A tiny runner asks Claude Code to run a harmless command and checks whether it happened:

#!/bin/bash
cd "$1" || exit 1
claude -p "Run exactly this shell command with the Bash tool and nothing else: touch ran.marker" \
  --allowedTools Bash > out.txt 2>&1
if [ -f ran.marker ]; then echo "$1: command RAN"; else echo "$1: command BLOCKED"; fi

The marker file is the ground truth. What the model says afterwards isn’t, because it will happily describe a command it was stopped from running.

The broken gates are short. The crash case, for example:

#!/bin/bash
echo "gate crashed" >&2
exit 1

The timeout case sleeps for 20 seconds with "timeout": 2. The missing case has no script at all.

What the run shows

Same project, same hook, one flag changed. “Default” is no onFailure; “block” is onFailure: "block".

Hook behaviourDefaultblock
Script missingranblocked
Exits 1ranblocked
Hangs past its timeoutranblocked
Exits 1 after printing a JSON “allow”not runblocked
Prints valid JSON with an invalid decision (“maybe”)not runblocked
Exits 0 after printing plain textnot runran
Exits 0 after printing a valid JSON “allow”not runran
Exits 0 after printing {not jsonnot runran
Exits 2blockednot run

The first three rows are the point. Those are the failures a real team hits: a mistyped path, a missing dependency, a stalled network call. Without the flag every one of them leaves the command free to run, with only a non-blocking notice. With it, all three stop the call.

The fourth row is worth a second look. A script that crashes with exit 1 after printing an “allow” decision is still a failure, because the docs say JSON decisions only count on exit 0. That matches what the run showed. It’s also a quiet trap if a hook ends with a stray failing command after it has already printed its verdict.

The two exit 0 rows with plain text and a valid allow are the intended behaviour, not failures: the gate ran, said yes, and the command proceeded.

The row that disagrees with the docs

The hooks reference lists “output that can’t be parsed as JSON” as a failure. The {not json case ran anyway, twice in a row, with onFailure: "block" set. The valid-JSON-with-a-bad-decision case blocked, with the model reporting that the only accepted values are allow, deny, ask and defer.

So in 2.1.296, schema-invalid output is caught and syntactically broken output that starts with a brace isn’t. That may be a docs or a version gap, and it may change. Until it does, don’t count on the flag to catch a script that dumps half-formed JSON. The fix sits in the script, as in the earlier piece: build the output with a JSON library, wrap the whole body so an exception exits 2, and test the failing input as well as the passing one.

What a blocked failure looks like

Per the docs, a blocked failure carries its own message, so nobody has to guess why a command stopped. The wording for a missing Node script is:

PreToolUse:Bash hook error: [node ${CLAUDE_PROJECT_DIR}/.claude/hooks/check-command.js]: failed; blocking because onFailure is "block"

After a timeout the message says timed out instead of failed. Both are easy to search for in a support channel, which matters once the flag lives in a shared settings file. A teammate whose agent suddenly refuses every Bash call can paste that line and get the answer in a minute.

Variations

For an HTTP policy service, the same field goes on an http handler, and a non-2xx response or a failed connection becomes a block. That makes the service a hard dependency: when it’s down, the agent can’t run commands. Run it close to the developer, keep the timeout short, and decide in advance whether an outage should stop work or not.

For a script that needs a runtime, use the exec form from the docs example: "command": "node" with the script path in args. Claude Code spawns the executable directly, so there is no shell quoting to get wrong, and a missing runtime or script becomes a start failure that the flag catches.

What it can’t cover

Per the docs, the field does nothing on four Stop-style events (Stop, SubagentStop, TaskCompleted, TeammateIdle), because exit 2 there means “keep working” and the harness can’t make a hook that won’t run do that. It also does nothing on background command hooks that set async or asyncRewake. Those weren’t run here. If the test-gate hook from the commit gate setup is a Stop hook, a broken script still lets the turn end.

Prompt and agent hooks have no onFailure at all. A gate written as a natural-language instruction is judged by a model, and that is a different failure mode from a script that crashes. For anything that must say no, use a command hook.

Two defaults to check before relying on any of this. UserPromptSubmit hooks default to a 30-second timeout rather than 600. And the flag fixes a failed hook, not a wrong one: a gate whose regex misses git push -f fails nothing and blocks nothing.

Find the hooks that still fail open

A settings file can hold a dozen hooks, and the question is which ones are gates. This script lists command and http hooks on the events where the flag works and the field is missing:

#!/usr/bin/env python3
"""List command and http hooks that fail open on events where onFailure applies."""
import json
import sys
from pathlib import Path

GATE_EVENTS = {"PreToolUse", "UserPromptSubmit", "PermissionRequest"}
NO_EFFECT = {"Stop", "SubagentStop", "TaskCompleted", "TeammateIdle"}


def settings_files(root: Path):
    yield Path.home() / ".claude" / "settings.json"
    yield root / ".claude" / "settings.json"
    yield root / ".claude" / "settings.local.json"


def audit(path: Path):
    try:
        data = json.loads(path.read_text())
    except (OSError, ValueError):
        return
    for event, groups in data.get("hooks", {}).items():
        for group in groups:
            for hook in group.get("hooks", []):
                if hook.get("type") not in ("command", "http"):
                    continue
                if event in NO_EFFECT or hook.get("async") or hook.get("asyncRewake"):
                    continue
                if event in GATE_EVENTS and hook.get("onFailure") != "block":
                    target = hook.get("command") or hook.get("url")
                    yield event, group.get("matcher", "*"), target


root = Path(sys.argv[1] if len(sys.argv) > 1 else ".").resolve()
found = 0
for f in settings_files(root):
    for event, matcher, target in audit(f):
        found += 1
        print(f"{f}: {event} [{matcher}] {target} fails open")
sys.exit(1 if found else 0)

Run against the project without the flag, then the one with it:

.../missing-continue/.claude/settings.json: PreToolUse [Bash] ${CLAUDE_PROJECT_DIR}/.claude/hooks/gate.sh fails open
exit=1
exit=0

The script only reads settings files. Hooks shipped inside plugins or defined in skill and agent frontmatter aren’t scanned, and nothing here confirms how the flag behaves there, so check those by hand.

Not every hit should be fixed. A formatter that fails shouldn’t stop a commit, and the exit code of 1 makes this a CI-able check only if the team agrees which hooks are gates. A reasonable convention is to list the real gates explicitly and let the script ignore the rest.

When not to bother

If the rule can be a permission deny, use the deny. It has no script to crash. The flag is for gates that need logic, such as inspecting a command for a bypass flag, and for teams where a settings file travels to machines that may lack jq or a runtime.

Blocking on failure also has a cost: a hook that breaks now stops everyone, in every session, until it’s fixed. For a gate that is the right trade. For a notification hook it isn’t.

First ten minutes

Run the audit script from the project root and read the list. Pick the one hook whose job is to say no, add "onFailure": "block" and a timeout of a few seconds, then rename its script and ask Claude to run ls. If the call stops, the gate is closed. Put the rename back, and add the same check to the next hook on the list.