<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Assistants on Tokenise</title><link>https://tokenise.rosvetic.com/categories/assistants/</link><description>Recent content in Assistants on Tokenise</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 05 Oct 2026 22:30:00 +0000</lastBuildDate><atom:link href="https://tokenise.rosvetic.com/categories/assistants/index.xml" rel="self" type="application/rss+xml"/><item><title>A PreToolUse hook is the guardrail that survives bypass mode</title><link>https://tokenise.rosvetic.com/posts/claude-code-pretooluse-guardrail-hook/</link><pubDate>Mon, 05 Oct 2026 22:30:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-pretooluse-guardrail-hook/</guid><description>&lt;p&gt;Permission prompts are only as strong as the mode you&amp;rsquo;re running in. Switch to bypass mode, or launch with &lt;code&gt;--dangerously-skip-permissions&lt;/code&gt;, and the prompts are gone. A PreToolUse hook is different: according to the Claude Code &lt;a href="https://code.claude.com/docs/en/hooks-guide" target="_blank" rel="noopener noreferrer"&gt;hooks guide&lt;/a&gt;&#10;, it fires before any permission-mode check, and a &lt;code&gt;deny&lt;/code&gt; from it blocks the tool even in &lt;code&gt;bypassPermissions&lt;/code&gt; mode.&lt;/p&gt;</description></item><item><title>A separate evaluator agent is what makes long agent runs worth the money</title><link>https://tokenise.rosvetic.com/posts/planner-generator-evaluator-harness/</link><pubDate>Mon, 05 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/planner-generator-evaluator-harness/</guid><description>&lt;p&gt;Ask one agent to build a whole application and two things tend to go wrong. It loses the thread as its context fills, and it grades its own work generously. Anthropic&amp;rsquo;s engineering team wrote up a &lt;a href="https://www.anthropic.com/engineering/harness-design-long-running-apps" target="_blank" rel="noopener noreferrer"&gt;harness design for long-running application development&lt;/a&gt;&#10; that attacks both, and the numbers are honest about what it costs.&lt;/p&gt;</description></item><item><title>Cursor Projects puts a coordinator agent in charge of the work</title><link>https://tokenise.rosvetic.com/posts/cursor-projects-coordinator-agents/</link><pubDate>Mon, 05 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/cursor-projects-coordinator-agents/</guid><description>&lt;p&gt;Cursor&amp;rsquo;s September 10 release of Projects, in beta and rolling out to all users per its &lt;a href="https://cursor.com/changelog" target="_blank" rel="noopener noreferrer"&gt;changelog&lt;/a&gt;&#10;, is built on a specific idea: the agent in charge shouldn&amp;rsquo;t be the one writing the code.&lt;/p&gt;</description></item><item><title>What 16 parallel agents building a C compiler teach about running agents at scale</title><link>https://tokenise.rosvetic.com/posts/parallel-claudes-c-compiler-lessons/</link><pubDate>Sun, 04 Oct 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/parallel-claudes-c-compiler-lessons/</guid><description>&lt;p&gt;Anthropic&amp;rsquo;s engineering team handed 16 Claude instances a goal, a shared repo and a test suite, and let them write a C compiler in Rust. The result, described in an &lt;a href="https://www.anthropic.com/engineering/building-c-compiler" target="_blank" rel="noopener noreferrer"&gt;engineering post&lt;/a&gt;&#10;, is a roughly 100,000-line compiler that can build Linux 6.9 on x86, ARM and RISC-V, plus projects like SQLite, postgres and redis. It took nearly two weeks, close to 2,000 Claude Code sessions, 2 billion input tokens and 140 million output tokens, for just under $20,000.&lt;/p&gt;</description></item><item><title>Let the agent interview you, then write the spec, then start fresh</title><link>https://tokenise.rosvetic.com/posts/agent-interview-spec-workflow/</link><pubDate>Sat, 03 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/agent-interview-spec-workflow/</guid><description>&lt;p&gt;Big features fail in a predictable way when you hand them to an agent in one prompt. The agent fills every gap in your description with a guess, builds confidently on the guesses, and you find out about the wrong ones at review. A better move is to make it ask first.&lt;/p&gt;</description></item><item><title>Claude Code's auto mode catches most risky actions. Anthropic's numbers show what it misses</title><link>https://tokenise.rosvetic.com/posts/claude-code-auto-mode-numbers/</link><pubDate>Sat, 03 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-auto-mode-numbers/</guid><description>&lt;p&gt;Anthropic published how Claude Code&amp;rsquo;s auto mode works, with measured error rates. That&amp;rsquo;s unusual for a security feature, and it makes the engineering post worth reading before you let an agent run unattended. The &lt;a href="https://www.anthropic.com/engineering/claude-code-auto-mode" target="_blank" rel="noopener noreferrer"&gt;full write-up&lt;/a&gt;&#10; is on Anthropic&amp;rsquo;s engineering blog.&lt;/p&gt;</description></item><item><title>Run parallel Claude Code sessions without them stepping on each other</title><link>https://tokenise.rosvetic.com/posts/claude-code-worktrees-parallel-sessions/</link><pubDate>Fri, 02 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-worktrees-parallel-sessions/</guid><description>&lt;p&gt;Two Claude Code sessions in the same checkout will eventually edit the same file. Git worktrees are the standard answer: each session gets its own working directory and branch, backed by the same repository history. Claude Code has a flag for it, and the &lt;a href="https://code.claude.com/docs/en/worktrees" target="_blank" rel="noopener noreferrer"&gt;worktree docs&lt;/a&gt;&#10; cover more edge cases than the one-liner suggests.&lt;/p&gt;</description></item><item><title>Your CLAUDE.md is probably too long. Here's how to cut it</title><link>https://tokenise.rosvetic.com/posts/prune-your-claude-md/</link><pubDate>Thu, 01 Oct 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/prune-your-claude-md/</guid><description>&lt;p&gt;A CLAUDE.md file starts life as five useful lines. Six months later it&amp;rsquo;s two hundred, half of which describe things Claude could work out by reading your code, and the instruction you actually care about is somewhere around line 140.&lt;/p&gt;</description></item><item><title>How Codex reads AGENTS.md, and how to layer it so the right rules win</title><link>https://tokenise.rosvetic.com/posts/codex-agents-md-layering/</link><pubDate>Wed, 30 Sep 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/codex-agents-md-layering/</guid><description>&lt;p&gt;Codex reads instructions from &lt;code&gt;AGENTS.md&lt;/code&gt; files before it starts work, and it doesn&amp;rsquo;t read just one. It builds a chain of them, in a documented order, and the order decides which rule wins when two disagree. If you&amp;rsquo;ve ever had Codex ignore a rule you were sure you&amp;rsquo;d written down, the layering is the first place to look.&lt;/p&gt;</description></item><item><title>Stop telling your AI agent to 'fix the bug'. Make it reproduce it first</title><link>https://tokenise.rosvetic.com/posts/reproduce-first-bug-prompt/</link><pubDate>Wed, 30 Sep 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/reproduce-first-bug-prompt/</guid><description>&lt;p&gt;&amp;ldquo;Fix the login bug&amp;rdquo; is a terrible prompt for a coding agent, and a frustrating one to debug when it goes wrong. The agent will read some code, find something plausible, change it, and tell you it&amp;rsquo;s fixed. Sometimes it is. Sometimes it has patched a bug you didn&amp;rsquo;t have.&lt;/p&gt;</description></item><item><title>Make Claude Code run your tests before it's allowed to say it's done</title><link>https://tokenise.rosvetic.com/posts/claude-code-stop-hook-tests/</link><pubDate>Wed, 30 Sep 2026 07:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/claude-code-stop-hook-tests/</guid><description>&lt;p&gt;Telling Claude Code to run the tests in CLAUDE.md is a request. It usually works, and sometimes it doesn&amp;rsquo;t. If you want it to work every time, Claude Code has a mechanism for that: a Stop hook, a script that runs when Claude thinks it&amp;rsquo;s finished and can refuse to let it stop.&lt;/p&gt;</description></item><item><title>Cursor's Rollouts bot watches your deploys after the merge</title><link>https://tokenise.rosvetic.com/posts/cursor-rollouts-bot-watches-deploys/</link><pubDate>Tue, 29 Sep 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/cursor-rollouts-bot-watches-deploys/</guid><description>&lt;p&gt;Cursor has added a bot called Rollouts that watches a change after it merges and reports whether it&amp;rsquo;s healthy in each environment. It was announced in the &lt;a href="https://cursor.com/changelog/rollouts-and-security-reviewer" target="_blank" rel="noopener noreferrer"&gt;Cursor changelog&lt;/a&gt;&#10; on September 23 and is available on Teams and Enterprise plans.&lt;/p&gt;</description></item></channel></rss>