Weekly Review — Sep 28–Oct 4, 2026
written by Stefan Christoph
- 8 minutes readEvery week has a few posts that were written separately and still end up arguing with each other. This week it was legibility: the gap between what someone actually does, wants, or is, and the part of it that reaches the person on the other side. Two entries in the asymmetry series (commitment and effort visibility) look at that gap at work and at home. A post on personal branding treats it as a design problem you can solve on purpose. A post on evaluating very long agent runs asks how you would even see whether one worked. And the third Lunch Break Physics explains why the colour above your head comes from light scattered by the air, not from anything the sky reflects.
This Week on the Blog
Being Someone’s Option
When one side is all-in and the other is still exploring, the explorer usually holds more power, and it has nothing to do with being cooler. Rusbult’s investment model predicts commitment from satisfaction, investment and the quality of your alternatives, so all else equal, whoever has better alternatives is less committed, and Waller’s principle of least interest describes the tendency for the less interested side to have more say. The fix for the all-in side is to build a real alternative, cap escalation in advance, and believe what the other side invests over what it says; the exploring side owes honesty about the option structure, a deadline, and no free work it would not pay for.
Your Personal Brand, in the Age of AI
“Build your personal brand” is common advice with no method attached, so this post supplies one: a nine-block Personal Brand Canvas, filled by an AI that interviews you and keeps asking “help whom, to do what?” until the generic answer turns specific. The canvas condenses into a positioning line, a tagline, three or four content pillars and a next action. The harder half is living it: the post argues that the strategy is the easy part and showing up consistently is where most people fall down. The canvas is on GitHub under the MIT licence, with a copy-paste prompt you can use in a general-purpose AI assistant.
Why Is the Sky Blue (And Sunsets Red)?
The sky is not reflecting the ocean; it stays blue over deserts and on days when the sea is grey. Air molecules scatter short wavelengths far more than long ones (Rayleigh scattering goes as one over the wavelength to the fourth power), so blue light near 450 nm is scattered about six times more strongly than red near 700 nm, sending diffuse blue light toward you from across the sky. A sunset is the same mechanism through more air, which strips the blue out of the direct beam and leaves the reds. This is the third Lunch Break Physics.
When Your Benchmark Runs for 88 Hours
On Dwarkesh Patel’s podcast, Noam Brown discussed OpenAI’s reported run of roughly 10,000 agents spending 130 billion tokens over 88 hours on a single problem (as reported, not independently confirmed). The number that should worry a builder is the 88 hours: a benchmark can still score the final answer, but one end-of-run score cannot tell you what the extra agents bought, where coordination failed, or whether the result was worth the time and compute, and Brown himself says OpenAI’s published measurements go up to about 16 agents. The post lays out five things to measure (task quality versus agent count, coordination overhead, wrong-answer concentration, time-to-converge, cost per solved task), how to instrument each, and why evaluation should score task outcomes while tracking token use, latency, coordination overhead and trace-derived process metrics separately.
The Iceberg Problem
Visible effort often counts for more than equally valuable effort nobody sees, and the research behind that is uncomfortable: people value a result more when the work is displayed (the labor illusion), and simply being seen at a desk earns “dependable” and “committed” ratings with no information about output. Maintainers, glue workers and remote workers can be especially exposed, because so much of their best work shows up as nothing going wrong. The fix is honest legibility rather than working louder: narrate outcomes, make prevention countable, name your glue work, and, if you are the one judging, ask what did not happen because of this person.
The Thread This Week
Four of the five posts describe the same machine from different seats. In the commitment post, the signal that matters is investment, and the trap is reading a structural fact (they have alternatives) as a verdict on your worth. In the iceberg post, the evaluator sees the tip and not the mass, so the quiet maintainer looks idle and the visible firefighter looks heroic. In the branding post, generic is invisible, so the work is to make the real thing specific enough to be seen at all. In the evaluation post, a single final score says little about how an 88-hour agent run got there, so you have to decide in advance which traces to keep. Each fix is a mechanism rather than an effort to try harder: build an actual alternative, make the invisible work countable, run the interview until the answer stops sounding like everyone else’s, and instrument the agent run on purpose. The sky post fits more loosely, but it is the cleanest example of the week’s idea that what reaches your eye is not the same as what is out there.
Further Reading
Things worth your time that did not get their own post, all public:
- How We Built LangChain’s Paid Media Agent (LangChain): A detailed build log of a long-running marketing agent in Slack. The two lessons that travel: use models for judgment and code for consistency (LangChain reports that moving calculations into code and removing unnecessary model calls made an early reporting workflow about 40x cheaper and 13x faster), and route proposed changes through human approval before they are applied.
- Why a New Class of AI “Judgement Models” Could Have Big Business Implications (The AI Daily Brief): Small models that estimate a probability for a specific yes-or-no question instead of producing a paragraph of text, with their value as gates depending on how well that probability is calibrated. A useful frame for the cheap gating and routing layer in front of expensive generative calls, which is where many agent decisions actually live.
- The Watchdogs of AGI: Rune Kvist of AI Underwriting Company (Latent Space): The argument that risk, not capability, is now what holds back AI adoption, and that insurance and standards will form a flywheel: standards make risk priceable, and insurance pays you to meet them. Risk has to be made visible before anyone can underwrite it.
- Parallel’s Parag Agrawal: Building a New Web for AI Agents (Sequoia Capital): Agrawal’s claim that human click data is the wrong ranking signal for agents, which need compressed, structured answers rather than result pages, and that an ad-funded web breaks when the visitor is a machine that never sees the ads.
- RL Environments Explained: How AI Agents Learn Real-World Work (Sequoia Capital): Mercor’s Brendan Foody on the shift from crowdsourced preference data to sandboxed environments with verifiable rewards, where agents practice real tasks through tools. His observation that training environments and frontier evals are converging into the same thing is the part to keep.
- Coinbase’s Everything Exchange: Agentic Finance, Stablecoins and Tokenization (No Priors): Brian Armstrong’s case that agent payments do not fit card rails: he says most agent transactions Coinbase sees are under 30 cents, where common per-transaction card fees can make each payment uneconomic, which is why he argues for agent wallets and stablecoins. Treat the numbers as his, but the fee-floor point is a clean way to see why payments for agents are their own design problem.
Until Next Sunday
If there is one habit to take from this week, it is to separate what you can see from what is there, before you judge it. That goes for the colleague whose best work is that nothing broke, for the counterpart who seems cool but simply has alternatives, and for your own profile when it sounds like everyone else’s. Which invisible work in your team deserves to be named out loud this week? Tell me in the comments on LinkedIn.
This is the Weekly Review, and it also goes out as Sunday’s newsletter.
About the Author
Stefan Christoph is a Principal Solutions Architect at AWS, focused on agentic AI, media & entertainment, and helping builders move from demo to production. He writes about AI architecture, developer productivity, and the future of software.
This is a personal blog. Opinions expressed here are my own and do not represent the views or positions of my employer.
🎬 Also available as a blog walkthrough video on YouTube
❤️ Created with the support of AI (Kiro)