<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Home on Tokenise</title><link>https://tokenise.rosvetic.com/</link><description>Recent content in Home on Tokenise</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 05 Oct 2026 23:18:00 +0000</lastBuildDate><atom:link href="https://tokenise.rosvetic.com/index.xml" rel="self" type="application/rss+xml"/><item><title>Moving to Sonnet 5.5? Four requests that now return a 400</title><link>https://tokenise.rosvetic.com/posts/sonnet-5-5-migration-breaking-changes/</link><pubDate>Mon, 05 Oct 2026 23:18:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/sonnet-5-5-migration-breaking-changes/</guid><description>&lt;p&gt;Claude Sonnet 5.5 shipped on September 28 at the same price as Sonnet 5: $2 per million input tokens and $10 per million output. Anthropic says it runs 30% or more faster. If you call the API directly, though, changing &lt;code&gt;claude-sonnet-5&lt;/code&gt; to &lt;code&gt;claude-sonnet-5-5&lt;/code&gt; isn&amp;rsquo;t the whole migration. Several requests that worked yesterday now fail.&lt;/p&gt;</description></item><item><title>A PreToolUse hook is the guardrail that survives bypass mode</title><link>https://tokenise.rosvetic.com/posts/claude-code-pretooluse-guardrail-hook/</link><pubDate>Mon, 05 Oct 2026 22:30:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-pretooluse-guardrail-hook/</guid><description>&lt;p&gt;Permission prompts are only as strong as the mode you&amp;rsquo;re running in. Switch to bypass mode, or launch with &lt;code&gt;--dangerously-skip-permissions&lt;/code&gt;, and the prompts are gone. A PreToolUse hook is different: according to the Claude Code &lt;a href="https://code.claude.com/docs/en/hooks-guide" target="_blank" rel="noopener noreferrer"&gt;hooks guide&lt;/a&gt;&#10;, it fires before any permission-mode check, and a &lt;code&gt;deny&lt;/code&gt; from it blocks the tool even in &lt;code&gt;bypassPermissions&lt;/code&gt; mode.&lt;/p&gt;</description></item><item><title>A separate evaluator agent is what makes long agent runs worth the money</title><link>https://tokenise.rosvetic.com/posts/planner-generator-evaluator-harness/</link><pubDate>Mon, 05 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/planner-generator-evaluator-harness/</guid><description>&lt;p&gt;Ask one agent to build a whole application and two things tend to go wrong. It loses the thread as its context fills, and it grades its own work generously. Anthropic&amp;rsquo;s engineering team wrote up a &lt;a href="https://www.anthropic.com/engineering/harness-design-long-running-apps" target="_blank" rel="noopener noreferrer"&gt;harness design for long-running application development&lt;/a&gt;&#10; that attacks both, and the numbers are honest about what it costs.&lt;/p&gt;</description></item><item><title>Google's Agent Development Kit hits 1.0 for Kotlin, with Android support</title><link>https://tokenise.rosvetic.com/posts/adk-for-kotlin-1-0/</link><pubDate>Mon, 05 Oct 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/adk-for-kotlin-1-0/</guid><description>&lt;p&gt;Google released version 1.0 of its Agent Development Kit for Kotlin on September 9, according to the &lt;a href="https://developers.googleblog.com/announcing-adk-for-kotlin-10-building-production-ready-ai-agents-in-kotlin-android-and-beyond/" target="_blank" rel="noopener noreferrer"&gt;Google developers blog&lt;/a&gt;&#10;. It&amp;rsquo;s a framework for building AI agents in idiomatic Kotlin, aimed at both server-side JVM apps and Android.&lt;/p&gt;</description></item><item><title>Cursor Projects puts a coordinator agent in charge of the work</title><link>https://tokenise.rosvetic.com/posts/cursor-projects-coordinator-agents/</link><pubDate>Mon, 05 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/cursor-projects-coordinator-agents/</guid><description>&lt;p&gt;Cursor&amp;rsquo;s September 10 release of Projects, in beta and rolling out to all users per its &lt;a href="https://cursor.com/changelog" target="_blank" rel="noopener noreferrer"&gt;changelog&lt;/a&gt;&#10;, is built on a specific idea: the agent in charge shouldn&amp;rsquo;t be the one writing the code.&lt;/p&gt;</description></item><item><title>Build a Claude Code subagent that reviews your diff with fresh eyes</title><link>https://tokenise.rosvetic.com/posts/subagent-fresh-context-reviewer/</link><pubDate>Sun, 04 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/subagent-fresh-context-reviewer/</guid><description>&lt;p&gt;An agent that just wrote a change is a poor reviewer of it. It remembers why each decision was made, so it reads its own code charitably. A reviewer that starts with nothing but the diff and a checklist has no such bias. Claude Code subagents are a convenient way to get one.&lt;/p&gt;</description></item><item><title>What 16 parallel agents building a C compiler teach about running agents at scale</title><link>https://tokenise.rosvetic.com/posts/parallel-claudes-c-compiler-lessons/</link><pubDate>Sun, 04 Oct 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/parallel-claudes-c-compiler-lessons/</guid><description>&lt;p&gt;Anthropic&amp;rsquo;s engineering team handed 16 Claude instances a goal, a shared repo and a test suite, and let them write a C compiler in Rust. The result, described in an &lt;a href="https://www.anthropic.com/engineering/building-c-compiler" target="_blank" rel="noopener noreferrer"&gt;engineering post&lt;/a&gt;&#10;, is a roughly 100,000-line compiler that can build Linux 6.9 on x86, ARM and RISC-V, plus projects like SQLite, postgres and redis. It took nearly two weeks, close to 2,000 Claude Code sessions, 2 billion input tokens and 140 million output tokens, for just under $20,000.&lt;/p&gt;</description></item><item><title>Run Codex from a script: codex exec, structured output and CI</title><link>https://tokenise.rosvetic.com/posts/codex-exec-in-scripts-and-ci/</link><pubDate>Sun, 04 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/codex-exec-in-scripts-and-ci/</guid><description>&lt;p&gt;Codex isn&amp;rsquo;t only a terminal app you sit in front of. &lt;code&gt;codex exec&lt;/code&gt; runs it without the interactive interface, so it can be a step in a script, a scheduled job or a CI pipeline. The &lt;a href="https://learn.chatgpt.com/docs/non-interactive-mode.md" target="_blank" rel="noopener noreferrer"&gt;non-interactive mode docs&lt;/a&gt;&#10; lay out the behavior, and the defaults tell you how OpenAI expects you to use it.&lt;/p&gt;</description></item><item><title>Let the agent interview you, then write the spec, then start fresh</title><link>https://tokenise.rosvetic.com/posts/agent-interview-spec-workflow/</link><pubDate>Sat, 03 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/agent-interview-spec-workflow/</guid><description>&lt;p&gt;Big features fail in a predictable way when you hand them to an agent in one prompt. The agent fills every gap in your description with a guess, builds confidently on the guesses, and you find out about the wrong ones at review. A better move is to make it ask first.&lt;/p&gt;</description></item><item><title>That 2-point benchmark lead might just be a bigger server</title><link>https://tokenise.rosvetic.com/posts/benchmark-gaps-infrastructure-noise/</link><pubDate>Sat, 03 Oct 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/benchmark-gaps-infrastructure-noise/</guid><description>&lt;p&gt;Every model launch comes with a table of coding benchmark scores, and the gaps are often a few points. Anthropic&amp;rsquo;s engineering team published a study that should make you read those gaps more carefully. In its &lt;a href="https://www.anthropic.com/engineering/infrastructure-noise" target="_blank" rel="noopener noreferrer"&gt;infrastructure noise analysis&lt;/a&gt;&#10;, it found that how much compute an eval runs on can move agentic coding scores by more than the leaderboard gap between top models.&lt;/p&gt;</description></item><item><title>Claude Code's auto mode catches most risky actions. Anthropic's numbers show what it misses</title><link>https://tokenise.rosvetic.com/posts/claude-code-auto-mode-numbers/</link><pubDate>Sat, 03 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-auto-mode-numbers/</guid><description>&lt;p&gt;Anthropic published how Claude Code&amp;rsquo;s auto mode works, with measured error rates. That&amp;rsquo;s unusual for a security feature, and it makes the engineering post worth reading before you let an agent run unattended. The &lt;a href="https://www.anthropic.com/engineering/claude-code-auto-mode" target="_blank" rel="noopener noreferrer"&gt;full write-up&lt;/a&gt;&#10; is on Anthropic&amp;rsquo;s engineering blog.&lt;/p&gt;</description></item><item><title>You can now ask Copilot for a code review from a script</title><link>https://tokenise.rosvetic.com/posts/copilot-code-review-api/</link><pubDate>Fri, 02 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/copilot-code-review-api/</guid><description>&lt;p&gt;GitHub says you can now request a Copilot code review through its REST and GraphQL APIs, and set the review effort on each request. The change landed in a &lt;a href="https://github.blog/changelog/2026-10-02-copilot-code-review-api-support-and-new-default-effort-level" target="_blank" rel="noopener noreferrer"&gt;changelog entry on October 2&lt;/a&gt;&#10;, and it&amp;rsquo;s available on Copilot Pro, Pro+, Max, Business and Enterprise.&lt;/p&gt;</description></item><item><title>Claude Opus 5.5 costs 20% less per token and Anthropic says it codes better</title><link>https://tokenise.rosvetic.com/posts/claude-opus-5-5-pricing-and-coding/</link><pubDate>Fri, 02 Oct 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-opus-5-5-pricing-and-coding/</guid><description>&lt;p&gt;Anthropic released Claude Opus 5.5 on September 22. The model id is &lt;code&gt;claude-opus-5-5&lt;/code&gt;, it&amp;rsquo;s on the Claude Platform, AWS, Google Cloud and Azure, and it&amp;rsquo;s cheaper than the model it replaces. The &lt;a href="https://www.anthropic.com/claude-opus-5-5" target="_blank" rel="noopener noreferrer"&gt;announcement&lt;/a&gt;&#10; says it does most work at the level of Claude Fable 5.1 while costing 40% less to run than Opus 5.&lt;/p&gt;</description></item><item><title>Run parallel Claude Code sessions without them stepping on each other</title><link>https://tokenise.rosvetic.com/posts/claude-code-worktrees-parallel-sessions/</link><pubDate>Fri, 02 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-worktrees-parallel-sessions/</guid><description>&lt;p&gt;Two Claude Code sessions in the same checkout will eventually edit the same file. Git worktrees are the standard answer: each session gets its own working directory and branch, backed by the same repository history. Claude Code has a flag for it, and the &lt;a href="https://code.claude.com/docs/en/worktrees" target="_blank" rel="noopener noreferrer"&gt;worktree docs&lt;/a&gt;&#10; cover more edge cases than the one-liner suggests.&lt;/p&gt;</description></item><item><title>Copilot CLI dynamic workflows: turn a repeatable chore into code with agents inside</title><link>https://tokenise.rosvetic.com/posts/copilot-cli-dynamic-workflows/</link><pubDate>Thu, 01 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/copilot-cli-dynamic-workflows/</guid><description>&lt;p&gt;GitHub added dynamic workflows to Copilot CLI and the Copilot app on October 1, in public preview and on every Copilot plan. The idea: you write down a multi-step job once, code runs the steps, and agents only step in where a judgment call is needed.&lt;/p&gt;</description></item><item><title>Cursor's Security Review bot reads your pull requests for exploitable bugs</title><link>https://tokenise.rosvetic.com/posts/cursor-security-review-bot/</link><pubDate>Thu, 01 Oct 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/cursor-security-review-bot/</guid><description>&lt;p&gt;Cursor launched a Security Review bot on September 23, alongside the Rollouts deploy monitor we covered in &lt;a href="https://tokenise.rosvetic.com/posts/cursor-rollouts-bot-watches-deploys/"&gt;an earlier post&lt;/a&gt;&#10;. It reads pull requests and reports vulnerabilities an attacker could actually exploit. Per the &lt;a href="https://cursor.com/changelog/rollouts-and-security-reviewer" target="_blank" rel="noopener noreferrer"&gt;Cursor changelog&lt;/a&gt;&#10;, it&amp;rsquo;s on Teams and Enterprise plans.&lt;/p&gt;</description></item><item><title>Your CLAUDE.md is probably too long. Here's how to cut it</title><link>https://tokenise.rosvetic.com/posts/prune-your-claude-md/</link><pubDate>Thu, 01 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/prune-your-claude-md/</guid><description>&lt;p&gt;A CLAUDE.md file starts life as five useful lines. Six months later it&amp;rsquo;s two hundred, half of which describe things Claude could work out by reading your code, and the instruction you actually care about is somewhere around line 140.&lt;/p&gt;</description></item><item><title>How Codex reads AGENTS.md, and how to layer it so the right rules win</title><link>https://tokenise.rosvetic.com/posts/codex-agents-md-layering/</link><pubDate>Wed, 30 Sep 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/codex-agents-md-layering/</guid><description>&lt;p&gt;Codex reads instructions from &lt;code&gt;AGENTS.md&lt;/code&gt; files before it starts work, and it doesn&amp;rsquo;t read just one. It builds a chain of them, in a documented order, and the order decides which rule wins when two disagree. If you&amp;rsquo;ve ever had Codex ignore a rule you were sure you&amp;rsquo;d written down, the layering is the first place to look.&lt;/p&gt;</description></item><item><title>Stop telling your AI agent to 'fix the bug'. Make it reproduce it first</title><link>https://tokenise.rosvetic.com/posts/reproduce-first-bug-prompt/</link><pubDate>Wed, 30 Sep 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/reproduce-first-bug-prompt/</guid><description>&lt;p&gt;&amp;ldquo;Fix the login bug&amp;rdquo; is a terrible prompt for a coding agent, and a frustrating one to debug when it goes wrong. The agent will read some code, find something plausible, change it, and tell you it&amp;rsquo;s fixed. Sometimes it is. Sometimes it has patched a bug you didn&amp;rsquo;t have.&lt;/p&gt;</description></item><item><title>Make Claude Code run your tests before it's allowed to say it's done</title><link>https://tokenise.rosvetic.com/posts/claude-code-stop-hook-tests/</link><pubDate>Wed, 30 Sep 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-stop-hook-tests/</guid><description>&lt;p&gt;Telling Claude Code to run the tests in CLAUDE.md is a request. It usually works, and sometimes it doesn&amp;rsquo;t. If you want it to work every time, Claude Code has a mechanism for that: a Stop hook, a script that runs when Claude thinks it&amp;rsquo;s finished and can refuse to let it stop.&lt;/p&gt;</description></item><item><title>Google Cloud API Gateway can now serve your REST API as MCP tools</title><link>https://tokenise.rosvetic.com/posts/api-gateway-rest-to-mcp/</link><pubDate>Tue, 29 Sep 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/api-gateway-rest-to-mcp/</guid><description>&lt;p&gt;Google Cloud API Gateway can now act as a remote Model Context Protocol server for your existing REST APIs. The feature is in public preview, announced on &lt;a href="https://developers.googleblog.com/turn-your-rest-apis-into-mcp-tools-with-google-cloud-api-gateway/" target="_blank" rel="noopener noreferrer"&gt;Google&amp;rsquo;s developer blog&lt;/a&gt;&#10; on September 24. You mark up an OpenAPI spec, deploy as usual, and the gateway translates MCP requests into REST calls to your backend.&lt;/p&gt;</description></item><item><title>Cursor's Rollouts bot watches your deploys after the merge</title><link>https://tokenise.rosvetic.com/posts/cursor-rollouts-bot-watches-deploys/</link><pubDate>Tue, 29 Sep 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/cursor-rollouts-bot-watches-deploys/</guid><description>&lt;p&gt;Cursor has added a bot called Rollouts that watches a change after it merges and reports whether it&amp;rsquo;s healthy in each environment. It was announced in the &lt;a href="https://cursor.com/changelog/rollouts-and-security-reviewer" target="_blank" rel="noopener noreferrer"&gt;Cursor changelog&lt;/a&gt;&#10; on September 23 and is available on Teams and Enterprise plans.&lt;/p&gt;</description></item><item><title>About</title><link>https://tokenise.rosvetic.com/about/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/about/</guid><description>&lt;p&gt;Tokenise covers AI for software development: coding assistants, AI in code review and pull requests, GitHub workflows, new developer tools, and what the frontier labs ship that actually changes how you work.&lt;/p&gt;</description></item></channel></rss>