Moving to Sonnet 5.5? Four requests that now return a 400
Same price as Sonnet 5, faster, and better at agentic coding by Anthropic's numbers. But the API breaks in ways a model-ID swap won't catch.
4 min read
Everything we've published, newest first.
Same price as Sonnet 5, faster, and better at agentic coding by Anthropic's numbers. But the API breaks in ways a model-ID swap won't catch.
4 min read
Permission prompts depend on what mode you're in. A PreToolUse deny hook doesn't. Here's a small one that blocks force pushes, and where it falls short.
3 min read
Anthropic's harness splits planning, building and testing across agents. The cost jump was 20x, and the reasons it paid off are specific.
3 min read
Compile-time tool schemas, human-in-the-loop flows and resumable sessions, for teams that build agents in Kotlin.
2 min read
The agent that plans and delegates never writes the code. Cursor says it can keep a project's context for months.
3 min read
A reviewer that never saw the code being written catches different things. Here's the file, the invocation and the over-reporting trap.
3 min read
Anthropic's engineering team spent nearly $20,000 on it. The lessons about tests, locks and duplicate code transfer to much smaller jobs.
3 min read
The non-interactive mode turns Codex into a pipeline step. The defaults are conservative, and one warning matters a lot.
3 min read
A three-step workflow for features too big for one prompt, and what makes a spec worth handing over.
3 min read
Anthropic measured how much infrastructure settings move agentic coding scores. It's more than many leaderboard gaps.
3 min read
A 0.4% false positive rate sounds great. The 17% miss rate on real overeager actions is the number to plan around.
3 min read
GitHub's REST and GraphQL APIs can request a Copilot review, and the default review depth just changed.
3 min read
The numbers behind the release, what changes in practice, and why Anthropic's own benchmark caveat is worth reading.
3 min read
The --worktree flag gives each session its own checkout and branch. The setup details are where people get caught.
3 min read
A public preview feature that lets code run the steps and agents handle the judgment calls. Here's how to try it.
3 min read
It posts one comment per PR, skips drafts, and takes your own security rules. What's missing is any accuracy data.
3 min read
Anthropic's own advice is to prune ruthlessly. A one-question test, a before and after, and where each deleted rule should go instead.
3 min read
Global defaults, repo rules and per-service overrides stack in a fixed order. Know it, and a monorepo stops fighting your instructions.
3 min read
A bug prompt in three parts, and a rule for when to throw the whole session away.
4 min read
A Stop hook turns 'please run the tests' from a request into a rule. Here's a working setup and the trap to avoid.
3 min read
Annotate an OpenAPI spec, redeploy, and agents get tools without a separate MCP server to run.
3 min read
Most AI tooling stops at the pull request. This one keeps going until the change is live in each environment.
3 min read