<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>schristoph.online</title><link>https://schristoph.online/tags/benchmarking/</link><description>Personal homepage and blog of Stefan Christoph</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><copyright>Stefan Christoph. All rights reserved.</copyright><lastBuildDate>Fri, 24 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://schristoph.online/tags/benchmarking/index.xml" rel="self" type="application/rss+xml"/><item><title>Stop Guessing Which Coding Agent to Use — Benchmark It on Your Own Tasks</title><link>https://schristoph.online/blog/benchmarking-coding-agents/?utm=rss-feed</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://schristoph.online/blog/benchmarking-coding-agents/</guid><description>&lt;div class="tldr" data-pagefind-weight="5" data-pagefind-meta="tldr" style="display:block;font-size:.875em;margin:2rem 0;border-left:4px solid #ccc;padding-left:1rem;line-height:1.5;">&lt;strong>TL;DR:&lt;/strong> Generic coding-agent leaderboards don&amp;rsquo;t tell you which model or tool fits &lt;em>your&lt;/em> work at a budget you&amp;rsquo;re willing to spend. So I built a small task set that mirrors my real work and ran an open-source framework (&lt;a href="https://github.com/aws-samples/sample-agent-cost-bench">aws-samples/sample-agent-cost-bench&lt;/a>) that measures cost, quality, and latency together. Two findings surprised me: the priciest model (Opus 4.8) missed a task that a mid-priced one (Sonnet 4.6) passed at a third of the cost, and, when comparing the &lt;em>same&lt;/em> model across two tools, the per-credit rate you actually pay can flip which tool looks cheaper. All numbers below are from real runs I did on 2026-07-16.&lt;/div>
&lt;p>Every few weeks a new coding-agent leaderboard makes the rounds, and every time the question I actually care about goes unanswered: not &amp;ldquo;which model tops a public benchmark&amp;rdquo; but &amp;ldquo;which model, wrapped in which CLI, does &lt;em>my&lt;/em> kind of work well enough at a price I&amp;rsquo;m happy to pay.&amp;rdquo; Those are two different questions, and neither shows up on a generic leaderboard built from someone else&amp;rsquo;s tasks.&lt;/p></description></item></channel></rss>