<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Terminal-Bench on Tokenise</title><link>https://tokenise.rosvetic.com/tags/terminal-bench/</link><description>Recent content in Terminal-Bench on Tokenise</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 03 Oct 2026 12:00:00 +0000</lastBuildDate><atom:link href="https://tokenise.rosvetic.com/tags/terminal-bench/index.xml" rel="self" type="application/rss+xml"/><item><title>That 2-point benchmark lead might just be a bigger server</title><link>https://tokenise.rosvetic.com/posts/benchmark-gaps-infrastructure-noise/</link><pubDate>Sat, 03 Oct 2026 12:00:00 +0000</pubDate><guid>https://tokenise.rosvetic.com/posts/benchmark-gaps-infrastructure-noise/</guid><description>&lt;p&gt;Every model launch comes with a table of coding benchmark scores, and the gaps are often a few points. Anthropic&amp;rsquo;s engineering team published a study that should make you read those gaps more carefully. In its &lt;a href="https://www.anthropic.com/engineering/infrastructure-noise" target="_blank" rel="noopener noreferrer"&gt;infrastructure noise analysis&lt;/a&gt;&#10;, it found that how much compute an eval runs on can move agentic coding scores by more than the leaderboard gap between top models.&lt;/p&gt;</description></item></channel></rss>