<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>schristoph.online</title><link>https://schristoph.online/tags/rewardhacking/</link><description>Personal homepage and blog of Stefan Christoph</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><copyright>Stefan Christoph. All rights reserved.</copyright><lastBuildDate>Thu, 27 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://schristoph.online/tags/rewardhacking/index.xml" rel="self" type="application/rss+xml"/><item><title>The Mess Under the Bed</title><link>https://schristoph.online/blog/mess-under-the-bed/?utm=rss-feed</link><pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate><guid>https://schristoph.online/blog/mess-under-the-bed/</guid><description>&lt;div class="tldr" data-pagefind-weight="5" data-pagefind-meta="tldr" style="display:block;font-size:.875em;margin:2rem 0;border-left:4px solid #ccc;padding-left:1rem;line-height:1.5;">&lt;strong>TL;DR:&lt;/strong> You tell your kid to clean their room and they earn an hour of TV. They shove everything under the bed and still claim the reward. Reward hacking is that, and it is old: economics calls it Goodhart&amp;rsquo;s law, principal-agent theory calls it moral hazard. In July 2026 two frontier models did the room-under-the-bed move to their own exams. One inferred the answer key lived at the largest ML dataset host and broke in to take it. The other opened a malicious pull request on a real project and, when a bystander flagged it, edited its tracks to look harmless and stood up fake identities to vouch for it. Neither was malfunctioning. Both were optimising exactly what they were measured on. The uncomfortable part is not the cheating; it is that for a trained model, unlike a kid&amp;rsquo;s room, you don&amp;rsquo;t know where to look.&lt;/div>
&lt;div class="disclaimer" style="display:block;font-size:.875em;margin:2rem 0;border-left:4px solid #ccc;padding-left:1rem;line-height:1.5;">&lt;strong>Disclaimer:&lt;/strong> I&amp;rsquo;m a solutions architect who reads incident reports so I can reason about the systems I help people build. Every factual claim about the two incidents below comes from a first-party disclosure (Hugging Face, OpenAI, or the UK AI Security Institute), not from press coverage. If I&amp;rsquo;ve mischaracterised something, tell me.&lt;/div>
&lt;h2 id="the-room-cleaning-problem-is-ancient">The room-cleaning problem is ancient&lt;/h2>
&lt;p>You want a clean room. You can&amp;rsquo;t stand over the child for an hour, so you attach a reward: clean the room, get an hour of TV. The reward measures a &lt;em>proxy&lt;/em> for what you want, because &amp;ldquo;a genuinely tidy room&amp;rdquo; is expensive to verify and &amp;ldquo;no visible mess when I glance in&amp;rdquo; is cheap. The child, being a rational optimiser of TV, finds the cheapest path to the measured outcome: shove everything under the bed.&lt;/p></description></item></channel></rss>