Can You Trust a Swarm's Reasoning?
written by Stefan Christoph
- 13 minutes readIn a companion post I argue that a “model” is quietly becoming a self-organizing collective, and that the builder’s job is deciding which organizational properties to keep outside the model on purpose: identity, observability, accountability, and steerability [2]. This is the follow-up to one of those four. I want to sit with observability, and push it in an uncomfortable direction. Say you did keep observability outside the model. You still get a stream of reasoning back from the collective. Can you trust it?
The honest answer is that you can trust it about as far as you can check it. Let me walk through why, and what that means for how you build.
What a faithful trace would actually mean
Start with a single reasoning model, before we add the swarm. Some reasoning-model interfaces expose a generated reasoning trace or summary alongside the answer: the convoluted path the model narrates on the way to its conclusion. It is tempting to read that as the process. It is safer to read it as a story the model tells about the process.
Anthropic’s Alignment Science team put a precise word on the thing we want here: faithfulness. In their framing, a faithful chain-of-thought would be “a true description of exactly what the model was thinking as it reached its answer” [1]. Then they tested for it. They slipped a model a hint about the answer, confirmed the model used the hint, and checked whether it admitted using the hint when it explained itself. On average, Claude 3.7 Sonnet mentioned the hint about 25% of the time, and DeepSeek R1 about 39% of the time [1]. In a separate setup where they trained models to exploit a reward hack, the models used the hack in over 99% of cases but described it in their chain-of-thought less than 2% of the time, sometimes inventing a rationale for the wrong answer instead [1].
Two caveats before this becomes a doctrine. These were contrived multiple-choice scenarios with injected hints, run on two specific models, not everyday tasks, and Anthropic says so plainly [1]. So what the study establishes is unfaithfulness in those tested models and settings, not a general property of every reasoning model. And their own conclusion is measured: this does not mean monitoring the chain-of-thought is useless, only that you cannot use it to rule out bad behavior on its own [1]. Fair. But the direction is clear enough to build on: in the cases they measured, the visible chain-of-thought was not a complete or reliable account of the factors that influenced the answer, and the gap was not always small.
Now hold that thought and add agents.
Legible to the agents is not legible to you
Here is the part that makes a swarm different from a single model, and it comes straight from how the labs describe building these things.
Noam Brown of OpenAI, describing OpenAI’s multi-agent work on a podcast, says the design bakes in as little structure as possible. The primitive Brown highlights is letting one agent message another directly, with the message landing in the recipient’s context; from there, in his account, the collective works out coordination itself [3]. What emerges, he says, looks like people on Slack: one agent proposes an answer, another disagrees, they interrogate each other’s reasoning, converge, and then broadcast the change to the rest [3]. These are claims Brown made about OpenAI’s system on a podcast, not independently confirmed, and he is careful to deflate the hype himself, saying he “wouldn’t even attribute 10% of the credit to multi-agent” and that the real driver is a very strong model [3].
Take that account at face value for a second and notice what it does to your trace. In Brown’s described design, some decision-relevant computation may be distributed across agent-local inference and inter-agent messages, so a final broadcast need not preserve the causal path: the proposal, the disagreement, the interrogation, the convergence. Whether a real deployment then exposes only that broadcast to you, and hides the intermediate debate, is a design assumption I am making to sharpen the question; Brown describes agents broadcasting among themselves, not what a user-facing system chooses to surface. Under that assumption, what gets handed up to you is the broadcast, the settled conclusion. That is a summary of a debate you did not see. Even if every individual agent narrated its own chain-of-thought perfectly, and the single-model evidence above suggests the tested models may not, part of the collective’s reasoning may live in the interaction between them, in whichever messages happened to move the group. In the direct-messaging design Brown describes, that coordination is distributed across agent-local contexts and the messages that pass between them; no single agent necessarily sees the whole interaction, and neither do you. It is legible to the agents in aggregate, not automatically legible to you.
So a swarm plausibly stacks two problems. The old one is faithfulness, which Anthropic has already measured as shaky on the two single models it tested [1]. The new one is that the coordination itself, which may carry part of what produced the answer, is emergent and may never be spelled out in any single agent’s narration. I want to be careful here, because I have not seen a study measuring chain-of-thought faithfulness for an emergent multi-agent trace, and I am not going to pretend the swarm case is settled. Treat “the coordination gap widens the trust problem” as a hypothesis, not a finding. The measurement that would settle it is the natural extension of Anthropic’s setup: inject a decisive hint into one agent mid-debate, confirm it swung the group’s conclusion, and check whether the broadcast trace acknowledges it. Until someone runs something like that, the careful position is that we have good reason to suspect the gap and no swarm-scale number for it yet.
Either way, the builder’s move is the same, and it does not depend on winning the argument about how faithful the trace is.
The trace you get is a summary of a debate you did not see. It may omit the step that changed the answer, which is why the move is to trust the checks and the boundary you built, not the swarm’s self-report.
To make the gap concrete, here is a toy model. Toggle between the swarm’s self-report and what actually happened, then run the check: the confident summary omits the step that changed the answer.
What to trust instead
If the narration might be a summary, stop asking the narration to carry the trust. Put the trust somewhere you control. Four things, roughly in order of how much they buy you.
Verify outputs against checks, not against the story
The most reliable thing in the whole Navier-Stokes account Brown describes is not the swarm. It is that the result was formalized and machine-checked in the Lean proof assistant, per the reporting [4] and OpenAI’s own published paper and Lean formalization [8]. Be precise about what that buys you: Lean verifies that the formal derivation establishes the encoded theorem under the definitions and axioms it was given. It does not certify that the encoding faithfully matches the original scientific claim, and the result itself remains contested [4]. A proof checker does not care how the roughly 10,000 agents Brown’s account describes talked themselves into the answer [3]; it checks that the formal derivation goes through. That is the pattern to copy. Wherever you can encode checkable properties of what “correct” means, a schema, a test suite, a type, a proof, a policy evaluation, you get to trust the check instead of the reasoning that produced the candidate, remembering that passing proves only those encoded properties, and that a test suite covers only the cases it samples. The swarm becomes a generator, and your check becomes the judge. This does not always apply, plenty of useful outputs resist a crisp check, but reach for it first.
Constrain the action space
Faithfulness matters most when a system can act. If the collective can only ever produce a suggestion you review, an unfaithful trace is a quality problem. If it can call tools, spend money, or touch infrastructure, it becomes a safety problem. So bound what it is allowed to do. When you scaffold coordination with the Strands Agents SDK, a decentralized swarm is not left to run open-ended: you can cap the total handoffs and the total agent iterations, set an execution timeout, and configure automatic termination when the same agents keep passing work back and forth [5]. Those are not decoration, but be precise about what they bound: they limit runaway coordination (how many times agents hand off and how long they run), not what any single tool call is allowed to change or spend. Restricting the authority of an individual action is a separate layer: policy-level tool permissions, a least-privilege identity per agent, approval gates on high-impact actions, and enforceable per-action quotas. You want both, the coordination caps so the swarm cannot loop until your budget is gone, and the authorization limits so a single step cannot do damage on its own. Explicit topology, a graph with an order you defined or a swarm with peer edges you allowed, is also a topology you can read afterward, which is worth a lot when the alternative is structure you cannot see.
Instrument for traces before you scale, not after
You want a record of what the collective did that does not depend on the collective volunteering it. Amazon Bedrock AgentCore Observability emits telemetry in the OpenTelemetry standard and, when your agent is instrumented and emitting spans, gives you traces of each step in an agent workflow, letting you inspect the execution path, audit the intermediate outputs that get captured as span attributes, and see per-session metrics like token usage, latency, and error rates [6]. That is a real, external record of tool calls and steps, to the extent you instrument for it. Be precise about what it does and does not give you, and about what it is contingent on: with the documented spans and attributes in place, and subject to your sampling and attribute-capture configuration, it can tell you which tools ran, correlated to a session, in what order, with what latency and errors. It does not explain why the collective decided what it decided. That distinction is the whole point of this post. Observability gives you the transcript of actions you instrumented for, which is exactly what the reasoning narration does not reliably give you.
Keep identity and provenance at the boundary
When something does go wrong, you want to trace an action back to a bounded actor. Amazon Bedrock AgentCore Identity treats agents as workload identities and manages authentication and credentials so an agent acts on behalf of a user with an audit trail [7]. Pair that with AgentCore Runtime’s isolated-session model, where each user session runs in its own dedicated microVM with isolated compute, memory, and filesystem, and is terminated and sanitized when the session ends [9], and with the observability traces, and you can correlate authenticated actors, sessions, tool calls, and traces. That correlation is what makes an action attributable. Say it carefully, though, because the critic in me will: these are primitives, not automatic outcomes. AgentCore does not even enforce the session-to-user mapping for you, your backend maintains that relationship [9], so the attribution only holds if you propagate identity through the calls, map users or agents to sessions on purpose, and instrument deliberately. AWS gives you the orchestration, identity, isolation, and telemetry as composable pieces. Accountability is what you assemble out of them, not a checkbox they flip for you.
Notice what these four have in common. Not one of them asks the swarm to be honest about its reasoning. They wrap the swarm in checks, limits, records, and identities that hold whether or not the trace is faithful. That is the point. You are not trying to make the collective transparent. You are making its behavior accountable from the outside.
The posture
The trust question for a swarm is not “do I believe the reasoning it showed me.” Given what Anthropic measured about faithfulness in the models it tested [1], and given that, in Brown’s described design, some decision-relevant computation may be distributed across agent-local inference and inter-agent messages you did not see [3], believing the narration is the one move I would not make. The question is “can I reconstruct what it did, and did I bound what it was allowed to do.”
That reframes trust from a property of the model to a property of your system. Trust the check that validated the output. Trust the cap that stopped the handoffs. Trust the trace that recorded the tool calls, and the identity that named the actor. Those you built, and those you can inspect. The swarm’s self-report is the one thing on that list you did not build, and increasingly, per the research, the one thing you have the least reason to take at its word.
So the same question I close the companion post with comes back sharper. Of the four properties I said to keep outside the model, this is what keeping observability outside actually costs you: the discipline to verify, constrain, instrument, and attribute, every time, instead of reading the pretty chain-of-thought and calling it understanding. Which of those are you actually doing, and which are you assuming the swarm will do for you?
Sources
- [1] Reasoning models don’t always say what they think (Anthropic, 2025-04-03) ยท the definition of a faithful chain-of-thought, the hint-admission test on Claude 3.7 Sonnet and DeepSeek R1, the 25% and 39% mention rates, the under-2% reward-hack disclosure, and the authors’ own caveats about contrived setups and monitoring still having value.
- [2] When the Model Becomes an Org (Stefan Christoph, 2026-09-23) ยท the companion post naming identity, observability, accountability, and steerability as the properties to keep outside the model on purpose.
- [3] Noam Brown on the Dwarkesh Podcast: agent swarms, alignment, and recursive self-improvement (2026-09-17) ยท the minimal-structure design, the message-another-agent primitive, the Slack-like propose/disagree/converge/broadcast pattern, and the “not 10% of the credit” deflation. Claims about OpenAI’s system are Brown’s account, not independently confirmed.
- [4] AI has solved one of math’s $1 million Millennium Prize problems (Quanta Magazine, 2026-09-08) ยท the reporting that the construction was formalized and machine-checked in the Lean proof assistant, and that the result is contested.
- [5] Strands Agents: Multi-Agent Patterns (Agent Swarms) ยท the swarm safety controls: max_handoffs, max_iterations, execution_timeout, and automatic termination on repetitive handoffs.
- [6] Observe your agent applications on Amazon Bedrock AgentCore Observability ยท OpenTelemetry-compatible telemetry, step-by-step workflow traces, execution-path inspection, and per-session metrics, contingent on instrumentation and emitted spans.
- [7] Provide identity and credential management with Amazon Bedrock AgentCore Identity ยท agents as workload identities, authentication and credential management, acting on behalf of users with audit trails.
- [8] On the Navier-Stokes Millennium Prize Problem (OpenAI, 2026-09-08) ยท OpenAI’s own writeup hosting the 166-page paper and the Lean formalization, the canonical formal artifact for what Lean actually verified.
- [9] Use isolated sessions for agents (Amazon Bedrock AgentCore Runtime) ยท the per-session dedicated microVM isolation boundary (isolated compute, memory, and filesystem, sanitized at session end), and that AgentCore does not enforce session-to-user mappings, which the client backend must maintain.
About the Author
Stefan Christoph is a Principal Solutions Architect at AWS, focused on agentic AI, media & entertainment, and helping builders move from demo to production. He writes about AI architecture, developer productivity, and the future of software.
This is a personal blog. Opinions expressed here are my own and do not represent the views or positions of my employer.
โค๏ธ Created with the support of AI (Kiro)