When Your Benchmark Runs for 88 Hours
written by Stefan Christoph
- 21 minutes readA note on where this comes from: I design agentic systems for a living, and my last post was about what happens to org structure when a “model” becomes a self-organizing collective. This one is the sequel I actually needed for my day job. Once the unit you deploy runs for days and talks to itself, how do you tell whether it is any good? The claims about OpenAI’s system are Brown describing it on a podcast, not independently confirmed. Everything else is cited at the bottom.
The number that should worry you is 88, not 10,000
My last post [5] was about structure: once a “model” behaves like a self-organizing collective, old laws of org design start applying inside it, and the builder’s job becomes deciding which properties to keep outside the model on purpose. That was about what the thing is. This post is about a plainer problem that shows up the moment you try to ship one: how do you tell whether it worked?
Here is the framing from the podcast. Dwarkesh sets it up, and Brown does not dispute the shape of it: a system of about 10,000 agents spent 130 billion tokens over 88 hours on one problem [1]. Brown is careful to deflate the hype around it. “I wouldn’t even attribute 10% of the credit to multi-agent,” he says [1]. “The reality is that OpenAI has trained a very powerful model.” [1] Keep that deflation in mind. The interesting part for a builder is not the swarm and not the result. It is the shape of the run.
Think about what “88 hours” does to your test setup. The evaluation habit almost all of us built for single-shot models is: hold out a set of tasks, send each one as a prompt, score the single answer that comes back, average it. That habit assumes the unit of work is one bounded request and one response, judged only at its external output. A run that goes for days, spawns thousands of cooperating agents, and passes messages among them is not one request and one response. It is closer to watching a small organization work than to grading an exam. Sound familiar? It is the same mismatch you hit the first time you tried to unit-test a distributed system with the tools you had for a single function.
And the honest state of the measurement is thinner than the headline suggests. Brown says it directly: “we don’t have very good science on multi-agent scaling up to this kind of scale.” [1] The published plots he refers to go up to about 16 agents, with a default of four in what he calls Ultra Mode [1]. Ten thousand was a single run, not an ablation. He is explicit that they have not run the single-agent baseline on the same problem, and that a proper ablation “at that scale” would be too expensive [1]. So nobody has the curve. That is not a knock on the result. It is the whole reason this is worth writing about.
Why the single-shot benchmark breaks
A single-shot benchmark answers one question: given this input, how good is the output. That is a fine question when the system is one bounded request evaluated at its output. To be precise about what breaks: it is not the single shot as such (you could record a trace from one run) but scoring only the final output. It quietly fails for a long-horizon collective for a few reasons, and it is worth being precise about which.
First, the unit is wrong. You are no longer scoring an answer. You are scoring a process that ran for a long time, and the same final answer can come from a cheap, fast, well-coordinated run or an expensive, slow, thrashing one. A benchmark that only sees the final answer cannot tell those apart, and the difference is most of what you care about in production. Process-supervision research on mathematical reasoning makes the analogous point at the single-model level: grading the intermediate steps rather than only the final answer catches failures that outcome-only scoring misses [7]. Whether that extends to grading arbitrary multi-agent traces is a reasonable hypothesis to build on, not something that work established.
Second, the cost stops being ignorable. When a run is a single bounded request against a fixed model, the cost per request is often small and stable enough to leave out of a quality-only benchmark, even though it still varies with input and output tokens, caching, tool calls, and retries. When a run is thousands of agents over days, cost varies by orders of magnitude between two configurations that score the same on quality. Any evaluation that does not carry cost is measuring half the system.
Third, the failure modes concentrate in a way single-shot scoring hides. If you run one model once, an error is an error. If you run a swarm of copies of the same model that were trained to cooperate, they can be wrong together, agree with each other, and converge confidently on the wrong answer. This echoes a classical result from the classifier-ensemble literature: under the assumptions those studies make, it is independent errors that let majority voting help, and correlated errors erode that gain [8]. Independence is what buys the benefit, not a strict precondition for any ensemble to be useful. That work is about classifier ensembles, not cooperating copies of one language model, so treat wrong-answer concentration in a model collective as a proposed diagnostic rather than an established transfer. Aggregate accuracy can reveal the error rate but not whether failures repeatedly collapse onto the same wrong answer. You have to look at how they got there.
The through-line is that the interesting information moved from the answer into the trace. Which is why the rest of this is about measuring the run, not the reply.
Why the single number misleads. A days-long, thousands-of-agents run is not one request and one response, so measure the run, not the reply: quality vs agent count, coordination overhead, wrong-answer concentration, time-to-converge, and cost per solved task.
A builder’s shortlist: five things to measure, and how
Here is the part you can actually use. These are the metrics that would matter for a long-horizon multi-agent system. Brown does not report them for the Navier-Stokes run, and I want to be clear that standard, off-the-shelf benchmarks for most of them are thin today. That is exactly why you have to instrument for them yourself. For each one I have tried to say not just what it is but how you would capture it.
Task quality versus agent count
This is the curve Brown says is missing: how does outcome quality change as you go from one agent to four to sixteen to more [1]? Getting it honestly is harder than it sounds, because a larger configuration usually also gets more token budget, more parallel compute, and possibly more wall-clock time, so a naive plot confounds agent count with the resources it consumes. Run controlled ablations that fix one resource dimension at a time (token budget, model-compute, billed cost, or wall-clock), holding model, tools, task set, scheduler, and stopping policy constant, and report the dimensions you did not fix; each choice answers a different question, so name which one you are asking. Repeat across seeds, and report confidence intervals rather than single points. Then, separately, plot the unconstrained quality-versus-cost frontier so you can see what more resources buy regardless of how they are spent. The “quality score” in both cases has to be an outcome-level judgment, and the right instrument depends on the task: exact match or executable verification where the valid outcome is unique and mechanically checkable, rubric grading where it is not. Without these curves you are guessing at your operating point. With them, you can see where the returns flatten, which for many domains they appear to do: Brown describes the gains as “slightly sublinear” past a point, and strongly domain-dependent, with math fairly parallelizable and something like novel-writing not [1].
Coordination and communication overhead
In a peer-messaging design, agents send messages to each other, and those messages are tokens you pay for that are not direct work on the task. To make this operational rather than a vibe, define the quantities: let n be the agent count, let M(n) be the communication volume (the total inter-agent message tokens per run at that agent count), pick an outcome proxy for useful work (solved tasks is the bluntest workable one), and report a ratio such as communication tokens per solved task as n grows. If that ratio climbs while quality stays flat, you have found the point where adding agents starts buying you a meeting culture instead of throughput. Be careful not to treat all coordination traffic as waste: some messaging is the mechanism doing its job, which is exactly why you track the ratio against outcomes rather than penalizing messages as such. This is measurable directly from a trace that records agent-to-agent calls as spans.
Wrong-answer concentration
Run the same task many times and look at the distribution of failures, not just the pass rate. If failures cluster (the collective tends to converge on the same wrong answer rather than failing in scattered, independent ways), that is a wrong-answer concentration signal, and it is the one most likely to bite you in production because it looks like confidence. To make it reproducible, define incorrect-output clustering explicitly: canonicalize the final outputs first (normalize formatting, units, and trivially equivalent forms), decide semantic equivalence with the same instrument you use for grading (mechanical comparison where outputs are canonical, an adjudicated equivalence judgment where they are not), then cluster the wrong outputs and report a defined statistic, for example the largest wrong-cluster share or the entropy of the wrong-output distribution conditional on failure, with uncertainty from bootstrapping over the repeated runs. That gives you a number two teams can compute the same way. Keep per-agent pairwise error correlation separate from this: it is a different quantity, it requires agent-level correctness judgments inside each run rather than run-level outputs, and it is the heavier lift, so be clear about which one you are reporting [8]. Either way there is no clean public benchmark for this that I would lean on; you build it from your own repeated runs.
Time-to-converge
Long-horizon is the defining property, so measure it, and define it before you measure it so the number is reproducible. Pick a stability threshold (for example, no agent revises its answer for a fixed observation window, or agreement stays above some level across that window) and call that the convergence point; runs that hit a timeout or are terminated externally without meeting the threshold should be treated as censored, not counted as converged. Brown describes exactly the behavior you would track here: agents propose, disagree, interrogate each other, then one “broadcasts to the other agents” that it has changed its answer [1]. Report wall-clock convergence latency separately from the tokens consumed getting there; they are different quantities and mixing them hides which one is growing. A trace with timestamped agent messages gives you both almost for free.
Cost per solved task
This is a metric worth carrying from day one, provided you compute it honestly and treat it as what it is: one policy-level efficiency estimator for a single operating point, not the whole comparison. Define “solved” with an explicit estimator and threshold (the same outcome-level judgment from above), sum the cost across all attempted runs (successes and failures alike), and divide by the count of solved tasks: dollars per task you actually solved, not per task attempted. The ratio is undefined when nothing was solved, so decide that case in advance: report “no solves at this budget” alongside the total spend rather than inventing a number. Compute the numerator from provider-specific metering (model-specific input and output rates, caching) plus non-model charges like tool calls and infrastructure, rather than assuming token counts alone give you cost. Report it with uncertainty (the estimator is noisy when solve counts are small) and alongside success rate, the quality distribution among successes, and latency percentiles; on its own it hides all four. It complements the other metrics rather than replacing them: a configuration that scores slightly higher on quality but costs ten times as much per solved task is usually the wrong choice, and this estimator makes that visible at a glance, while the quality-versus-cost frontier from the ablations above remains the comparison across operating points.
Notice that some of these are externally observable and some are not. Task quality versus agent count, wrong-answer concentration, and billed cost you can measure from a run’s inputs, outputs, and invoice without a full execution trace. End-to-end completion latency is externally observable too, but convergence time is not the same thing: deciding when the collective actually converged (answer revisions settling, agreement holding) has to be read from the trace, unless the run is terminated the instant the convergence criterion fires. Coordination overhead, convergence time, and any per-agent attribution (who said what, where a cost or an error originated) are process metrics that need the trace. That distinction is where the tooling conversation starts.
To see why one number misleads, here is a toy model you can push around. Slide the agent count and watch the metrics move apart: quality keeps rising while cost per solved task has already turned.
Score outcomes and process, not tokens
The evaluation method has to change along with the metrics. Three things are worth keeping separate: final-output evaluation (did the answer clear the bar), outcome verification (a mechanical check that it is actually correct), and process or trajectory evaluation (was the path itself sound and efficient). The shift is away from token-level string matching toward outcome-level judgment, combined with trace-derived process and efficiency metrics. Token-level scoring (comparing the model’s output to a reference string) was always a proxy, and it gets worse as horizons grow, because there are many long, valid paths to a good outcome and a reference string captures one of them. The consistent rule is the one from the quality curve above: outcome-level evaluation may be exact match, executable verification, or rubric grading, depending on whether the valid outcome is unique and mechanically checkable. For tasks with a canonical, objectively checkable output, exact match or an executable reference test is the right tool and is more reliable than any model grader. Reach for rubric-graded evaluation where the outcome is not unique: define what a solved task looks like, then judge the run against that definition. For a math problem, Brown notes, the line is clean: did you get the integer right, and did you actually solve it rather than read an answer key [1]. For fuzzier work you write a rubric and, increasingly, use a capable model as a grader against that rubric, which introduces its own calibration and reliability errors; studies of LLM-as-judge setups document position bias, verbosity bias, and imperfect agreement with human raters, so validate the grader against human labels and report the agreement or error rate before you trust it [9]. Either way you are scoring the outcome and the path to it, captured from the trace, rather than matching tokens. Grading the path is not a novelty of agent systems: process-supervision work on mathematical reasoning found that judging intermediate steps catches failures that final-answer grading misses [7], and the same logic motivates grading a multi-agent run’s trace, even though that extension has not been directly tested.
One concrete example of tooling adopting this order: the Strands learning path puts its multi-agent patterns lesson immediately before a lesson on evaluating agents, framed as “how to reliably measure the quality and success of your agent systems” [4]. One curriculum does not establish an industry trend, but the sequencing is right: evaluation treated as the step after you build the topology.
Two things make outcome-based evaluation practical rather than aspirational. You need the run recorded in enough detail to grade, and you need to be honest that the grader is itself a component you have to trust and check. On the recording: this is where a tracing primitive earns its keep. Amazon Bedrock AgentCore Observability records instrumented operations as traces and spans in an OpenTelemetry-compatible format, along with built-in metrics like session count, latency, duration, token usage, and error rates, and lets you add custom spans and metrics from your own agent code [2]. A tracing system records what you instrument, not everything that happens: the inter-agent messages show up because you emit them as custom spans, the timestamps that give you time-to-converge come from the spans you create, and the token usage it reports is an input to cost, not the billed cost itself, which you derive separately from provider metering, caching, tool, and infrastructure records. If you host on AgentCore Runtime, each session runs in its own isolated compute, memory, and filesystem [3]; that isolation gives you a clean unit, but it becomes an evaluation unit only if your application assigns one session per run, so distinguish the Runtime isolation from the Observability telemetry and wire the mapping yourself.
I want to be careful about what that does and does not give you. AgentCore Observability is a telemetry primitive, not an evaluation product. It hands you the traces you asked for; deciding what a solved task is, writing the rubric, running the grader, and computing cost per solved task are things you compose on top. And the traces are only as useful as your instrumentation: to attribute a message or a cost to a specific agent and session, you have to propagate identity and map sessions on purpose, the same way you had to keep identity, observability, and accountability outside the model in the last post [5]. The primitive makes the accountable trace possible; your instrumentation is what makes it real.
The horizon problem you cannot outrun
There is one more reason this matters, and it is the sharpest thing Brown said on the topic. As models start operating over longer horizons, the horizon can outrun your ability to test it. The direction of travel here is measured, not hypothetical: METR’s evaluations of the task lengths AI systems can complete find that horizon growing steadily across model generations [6]. Brown’s example: if a model can work effectively over three months but your release cycle is every two months, “you don’t have a way to evaluate the models at the full length of their capabilities before the next model release cycle.” [1] He is careful to add that this is not only a safety concern. It is also a product-quality one: “maybe the product degrades over that time span in ways that we have not had sufficient time to test.” [1]
That is a lab-scale version of a problem you will hit at your own scale much sooner. If your agentic workflow is designed to run for a day, and your test cycle assumes you can observe a full run before you ship a change, the arithmetic stops working the moment the run is longer than the cycle. The mitigations are not exotic, but neither are they free, and each one is itself a hypothesis you have to test before you trust it: evaluate on shorter proxy tasks that exercise the same coordination, but first check that the proxy actually preserves the long-horizon coordination behavior you care about; checkpoint and grade long runs partway through instead of only at the end, but validate that the checkpoint score tracks the eventual outcome on runs you did let finish; and treat cost per solved task as a potential leading indicator, but confirm against completed full-horizon runs that early cost predicts eventual solution probability before you let it drive a release decision. None of that removes the tension. It just keeps you from pretending it is not there.
What I would actually do on Monday
If you are building a long-horizon multi-agent system and want to evaluate it honestly, the checklist is short:
- Record the run, not just the answer. Turn on tracing so each instrumented agent operation and inter-agent call is a span you can query later, and record token usage, agent and session identity, and outcomes as attributes or metrics on those spans. On AWS, AgentCore Observability gives you OpenTelemetry traces and built-in metrics as that substrate [2], but it records what you instrument, so add custom spans for the things specific to your workflow.
- Match the grading instrument to the task. For objectively checkable outputs, keep exact match or executable tests. Where the outcome is not unique, define what “solved” means and judge the trace against a rubric, and treat any model grader as a component you validate against human labels, not an oracle [9].
- Build the quality-versus-agent-count curve yourself, at a fixed budget. Run the same suite at several agent counts holding token or cost budget constant, repeat across seeds, and report confidence intervals. Do not assume more agents is better; Brown’s own account, not independently confirmed, says the gains go sublinear and depend heavily on the domain [1].
- Track cost per solved task from day one, alongside its context. Compute it from real metering over all attempted runs, define the zero-solve case in advance, report it with uncertainty next to success rate, quality distribution, and latency percentiles, and keep the quality-cost frontier for comparing operating points.
- Assume the horizon will outrun your cycle. Grade long runs partway through, use shorter proxy tasks, and read cost early, but validate each of those shortcuts against completed full-horizon runs before you rely on it. Do not wait for a full run you may not be able to afford to observe end to end.
The measurements that would settle whether a 10,000-agent run beats a 1,000-agent one are not ones Brown reports for this system, and he is refreshingly plain that “it is entirely possible that 10,000 humans are better at coordinating than 10,000 agents right now.” [1] You do not have to resolve that debate to build well. You just have to stop evaluating a days-long collective with a tool built for a single reply, and start measuring the run.
So here is the question I would put to anyone shipping one of these: when your benchmark runs for days instead of seconds, what are you actually measuring, and can you put a dollar figure on a solved task? If you cannot, you are grading the answer and ignoring the system that produced it.
Sources
- [1] Noam Brown on the Dwarkesh Podcast: Agent swarms, alignment, and recursive self-improvement (2026-09-17) · the 10,000-agent / 130B-token / 88-hour run, the “not 10% of the credit” deflation, “we don’t have very good science on multi-agent scaling”, the 1/4/16-agent plots and Ultra Mode default of four, sublinear and domain-dependent scaling, the emergent propose-disagree-converge-broadcast behavior, the unmeasured 10k-versus-1k question, and the three-month-horizon versus two-month-release-cycle evaluation gap. Claims about OpenAI’s system are Brown’s account, not independently confirmed.
- [2] Observe your agent applications on Amazon Bedrock AgentCore Observability (AWS documentation) · traces and spans for instrumented operations in OpenTelemetry-compatible format, built-in metrics including session count, latency, duration, token usage, and error rates, and support for custom spans and metrics from your own agent code.
- [3] Amazon Bedrock AgentCore: Use isolated sessions for agents (AWS documentation) · the per-session boundary you map users or agents to, which gives each run a comparable unit to trace and evaluate.
- [4] Strands Agents: Multi-Agent Patterns (Agent Swarms) · swarm handoff and iteration safety controls (max_handoffs, max_iterations, execution timeouts, repetitive-handoff detection) and the following lesson framed as evaluating agents and measuring the quality and success of agent systems.
- [5] When the Model Becomes an Org · my previous post on what changes when a model behaves like a self-organizing collective, and why identity, observability, and accountability are properties to keep outside the model on purpose.
- [6] Kwa, West et al., “Measuring AI Ability to Complete Long Tasks” (METR, arXiv:2503.14499) · measures model capability by the human time-length of tasks completed at a fixed success rate and finds that horizon growing steadily across model generations; the primary reference for long-horizon capability outpacing evaluation cycles.
- [7] Lightman et al., “Let’s Verify Step by Step” (arXiv:2305.20050) · process supervision versus outcome supervision; evidence that grading intermediate steps catches failures that final-answer-only scoring misses, the research grounding for trajectory-level evaluation.
- [8] Kuncheva and Whitaker, “Measures of Diversity in Classifier Ensembles and Their Relationship with the Ensemble Accuracy” (Machine Learning, 2003) · the classical result that ensemble value depends on members failing independently and degrades under correlated errors, the grounding for treating shared-model wrong-answer concentration as a first-class risk.
- [9] Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena” (arXiv:2306.05685) · documents position bias, verbosity bias, and human-agreement rates for model graders, the grounding for calibrating and validating any LLM judge against human labels before trusting it.
About the Author
Stefan Christoph is a Principal Solutions Architect at AWS, focused on agentic AI, media & entertainment, and helping builders move from demo to production. He writes about AI architecture, developer productivity, and the future of software.
This is a personal blog. Opinions expressed here are my own and do not represent the views or positions of my employer.
🎬 Also available as a blog walkthrough video on YouTube
❤️ Created with the support of AI (Kiro)