Weekly Review: Oct 5–11, 2026
written by Stefan Christoph
- 9 minutes readTwo new entries in the asymmetry series bracket this week, and three of the four posts between them turned out to share their shape. In each of those five posts, a decision or a default is cheap for the side that makes it and expensive for the side that lives with it. The reversibility post and the reputation post look at that gap between people. The swarm post and the GLM 5.3 post find it in agent systems, where the bill or the risk lands on the builder who did not check. The adoption talk finds it inside organisations, where buying tools is the easy part. And the ice post is the week’s rule-breaker, in the literal sense.
This Week on the Blog
The Door Only Locks Behind One of You
The same decision can be a two-way door for the person making it and a one-way door for the person it lands on. The planner can restructure again next quarter. The person whose work visa is tied to the role may have to leave the country. Loss aversion widens the gap, because the affected side codes the change as a loss and feels it at roughly twice the weight, and the two sides’ regret runs on different clocks. The habit to keep is to classify the door for the other party before you classify it for yourself. And when the outcome cannot be softened, invest in process: in the research on downsizing, employees who experienced the process as fair had lower odds of depressive symptoms afterwards. This is Part 7 of the asymmetry series.
Why Does Ice Float?
Almost every solid sinks in its own liquid, and water breaks the rule. When it freezes, hydrogen bonds hold each molecule to four neighbours in an open hexagonal lattice, so ice takes up about 9% more space and has a density of about 0.92 g/cm³, which leaves roughly 8% of a cube above the surface in fresh water. The lesser-known part is that liquid water is densest near 4 °C, so a lake overturns until it is uniformly about 4 °C. After that, colder water is lighter, stays on top and freezes first, and the 4 °C water underneath stays liquid, which is where fish wait out the winter. This is the fourth Lunch Break Physics.
Can You Trust a Swarm’s Reasoning?
A chain-of-thought is not guaranteed to describe what a model actually did: in Anthropic’s contrived tests, Claude 3.7 Sonnet mentioned a hint it had used about 25% of the time and DeepSeek R1 about 39%. The post argues that the gap plausibly widens when many agents debate internally and hand you only the conclusion, and it labels that as a hypothesis, not a measured finding. Its answer is to put trust where you control it: verify outputs against checks, cap what the swarm may do (the Strands Agents SDK can limit handoffs, iterations and run time), record traces with Amazon Bedrock AgentCore Observability, and tie actions to identities.
More Than Technology
This is the sourced companion to a talk on agentic AI adoption, and its argument is that once usable technology exists, the constraint is how people, culture and ways of working change. In McKinsey’s survey, 80% of respondents say AI made them more productive, while 37% report a positive effect on their organisation’s EBIT. The post walks through four fears that slow adoption, how AWS set the order of its data and workflow work before adding agents, and Amazon’s engineering experiments, where teams that changed their practices as well as their tools did better than teams that only added AI. It ends with three quiet ways my own agents failed, including a fallback that swapped my cloned voice for a stock one and still reported success.
GLM 5.3 on Amazon Bedrock
GLM 5.3 from Z.ai is a mixture-of-experts coding model with roughly 750B total and about 40B active parameters, and it is now on Amazon Bedrock for eligible customers through cross-Region inference. I tested it against Claude Opus 5.5. At their defaults, both solved every one-shot task. But GLM 5.3’s default reasoning effort is max, and on one task that setting used about 28 times the output tokens of low for the same passing result. The advice is to set reasoning_effort on every call, measure cost per solved task rather than per token (the same Python source came to 1.6 times as many tokens under Opus 5.5’s tokenizer), and put cost caps and write limits on agent runs.
One Side Rates, the Other Can’t Answer
A three-star rating that the passenger forgets by dinner can be a failing grade for the driver. In the study of Uber drivers the post leans on, a driver’s score was the average of their last 500 rated trips, and they needed around 4.6 to keep working. The post draws on Hirschman’s exit and voice (a review is exit with a permanent mark and no chance to repair), on research showing that consumer ratings turn customers into managers and carry their biases into someone’s employment, and on negativity bias. It ends with five moves for each side and a short list for whoever designs these systems: two-way visibility, context fields, response rights, bias auditing and decay. This is Part 8 of the asymmetry series.
The Thread This Week
Five of the six posts name a party that makes the call and a party that pays for it. Look at the pairs: the planner and the person in the role, the rater and the rated, the swarm and the builder who acts on its summary, the default setting and the team that gets the bill, the leadership that buys tools and the people who have to change how they work. Sound familiar? In every case the fix is to look from the paying side before you decide: classify their door, learn what your tap does to their average, check the output instead of the story, set the reasoning effort yourself, and measure whether the work got better. The ice post fits only loosely, but it is a good reminder that the rule everyone expects, that a solid sinks, does not hold when the structure underneath is different.
Further Reading
Things worth your time that did not get their own post, all public:
- How We Built a Software Factory with Kiro Crew to Merge 1,000 PRs in a Week (Kiro): Three engineers merged 1,000 pull requests in seven days, each through CI and review, and describe five stages from one hand-driven session to Crew Mode, where an agent runs the sessions and the human keeps the goal. The part to keep is the four governance problems they say move to the centre: coordination between agents, memory boundaries, permission control enforced by the host, and an audit trail agents cannot edit.
- Strands Decider (GitHub): A small open-source “decision model” that picks between options or rates something on a scale instead of generating text, which fits the routing and gating steps inside agent workflows. Each answer comes with a confidence, and the README reports that in the project’s own evaluation on short classification tasks it had not seen, answers at a confidence of 0.9 or more were right about 95% of the time.
- Brownfield Agentic Engineering (Addy Osmani): How to let agents work in an old codebase without signing up for technical debt. Mark the code as green, yellow or red zones that decide how much autonomy an agent gets, write down only what the code cannot say for itself, and keep each research pass as a short memo so the next session does not repeat the archaeology.
- Agent Plasticity: Measuring Self-Improvement Through Experience (arXiv): A paper that measures how efficiently an agent turns past experience into better performance on held-out tasks, not only how well it performs at one point in time. Two findings stand out: the agent that ends up best is not necessarily the one that improves most efficiently, and agents with low plasticity, the ones that turn experience into gains least efficiently, often failed to reuse the relevant artifacts they had saved.
- Can Rewriting an AI Agent Bend the Intelligence Curve? (Machine Learning Street Talk): Zhengyao Jiang of Weco on letting a coding agent rewrite the harness around another agent for eight days while the model stayed fixed. His four levels of recursive self-improvement are a useful question to ask whenever someone claims an agent “improves itself”: improving the artifact is common, and it is not the same as becoming a better improver.
- Engineering the SDLC for Coding Agents (Luca Bianchi): From a talk at AWS Community Day Italy, with examples in Kiro and Amazon Bedrock AgentCore: connect architectural decisions to release evidence, protect the checks agents rely on, and decide which actions an agent may take without supervision. It pairs well with this week’s swarm post.
- Kiro Crew: The Autonomous Developer Workflow (AWS Workshop): A three-hour workshop on a sample retail app where you build a crew of agents and give them memory, a shared goal, team conventions and a schedule. You can do it self-paced in your own account or at an AWS event.
Until Next Sunday
If there is one habit to take from this week, it is to look at a decision from the side that pays for it before you make it. That goes for the reorg line, the star rating, the agent summary you are about to act on, and the default setting you never changed. Which decision on your desk this week is cheap for you and permanent for someone else? Tell me in the comments on LinkedIn.
This is the Weekly Review, and it also goes out as Sunday’s newsletter.
About the Author
Stefan Christoph is a Principal Solutions Architect at AWS, focused on agentic AI, media & entertainment, and helping builders move from demo to production. He writes about AI architecture, developer productivity, and the future of software.
This is a personal blog. Opinions expressed here are my own and do not represent the views or positions of my employer.
🎬 Also available as a blog walkthrough video on YouTube
❤️ Created with the support of AI (Kiro)