<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on Tokenise</title><link>https://tokenise.rosvetic.com/tags/evaluation/</link><description>Recent content in Evaluation on Tokenise</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 05 Oct 2026 17:00:00 +0000</lastBuildDate><atom:link href="https://tokenise.rosvetic.com/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>A separate evaluator agent is what makes long agent runs worth the money</title><link>https://tokenise.rosvetic.com/posts/planner-generator-evaluator-harness/</link><pubDate>Mon, 05 Oct 2026 17:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/planner-generator-evaluator-harness/</guid><description>&lt;p&gt;Ask one agent to build a whole application and two things tend to go wrong. It loses the thread as its context fills, and it grades its own work generously. Anthropic&amp;rsquo;s engineering team wrote up a &lt;a href="https://www.anthropic.com/engineering/harness-design-long-running-apps" target="_blank" rel="noopener noreferrer"&gt;harness design for long-running application development&lt;/a&gt;&#10; that attacks both, and the numbers are honest about what it costs.&lt;/p&gt;</description></item></channel></rss>