Compute is not the bottleneck

September 2026

In July 2024, Bradley Brown and colleagues published Large Language Monkeys. They took DeepSeek-Coder-V2-Instruct and ran it on SWE-bench Lite, the set of real GitHub issues. With one attempt per issue it solved 15.9%. Then they let it try 250 times and kept whichever attempt passed the unit tests. It solved 56%. The best single-attempt system at the time, a mix of GPT-4o and Claude 3.5 Sonnet, solved 43%. The weights never changed. Only the number of tries did. In the same paper, five samples from the cheaper DeepSeek model solved more issues than one sample from Claude or GPT "while also being over 3x cheaper."

Two years later most agent harnesses still act as if tokens were the expensive part. They cap turns, cap subagents, cap wall-clock time, and ask the model for one answer. I think they're guarding the wrong thing. When I run agents against anything that can check their work, the runs that fail almost never fail because they cost too much. They fail because the agent explored too little and stopped too early. I wrote it down in July, in the middle of a task: "Compute is not the bottleneck here — under-exploration and premature convergence both are." The real limits are sandboxes, memory, disk, and above all the verifier. This post covers the evidence, how I run agents because of it, what keeps it from turning into slop, and where it breaks.

More tries beat a better first try

Coverage scales for four orders of magnitude. The Monkeys paper's main finding is that "coverage – the fraction of problems that are solved by any generated sample – scales with the number of samples over four orders of magnitude," often log-linearly. On CodeContests, Gemma-2B went from 0.02% with one sample to 7.1% with 10,000, a 300x gain from sampling alone. Where answers can be checked automatically (code with tests, Lean proofs), coverage is the score.

o1 turned this into a product. OpenAI's Learning to reason with LLMs said it plainly: "the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute)." On AIME 2024, GPT-4o solved 12% of problems. o1 solved 74% with one sample, 83% with a consensus of 64, and 93% when re-ranking 1,000 samples with a learned scorer. At IOI 2024, under the human limit of 50 submissions per problem, its coding model scored 213 points. "When allowed 10,000 submissions per problem, the model achieved a score of 362.14 – above the gold medal threshold – even without any test-time selection strategy."

AIME 2024 accuracy for GPT-4o and for o1 at increasing test-time compute

Spent well, test-time compute substitutes for parameters. Snell et al. showed that allocating test-time compute by question difficulty can "outperform best-of-N using up to 4x less test-time compute," and that in a FLOPs-matched comparison a small model with extra inference beats "a 14x larger model" on easy and intermediate questions. The hard cases matter too, and I'll come back to them.

AlphaCode did it first. DeepMind's AlphaCode generated "a massive amount of C++ and Python programs for each problem, orders of magnitude larger than previous work," then filtered, clustered and reranked them down to 10 submissions, and reached the level of the median Codeforces competitor. The Monkeys authors note that AlphaCode's performance "continues to improve with a million samples per problem."

For agents, tokens explain most of the outcome. Anthropic's multi-agent research system is the cleanest agent-side result I know. An Opus 4 lead with Sonnet 4 subagents "outperformed single-agent Claude Opus 4 by 90.2%" on their internal research eval. On BrowseComp, three factors explained 95% of the variance in performance, and "token usage by itself explains 80%." Their summary: "Multi-agent systems work mainly because they help spend enough tokens to solve the problem." Agent count is an implementation detail. The variable is how many useful tokens the system gets to spend in parallel.

The failure is stopping early

If spending more helps this reliably, why do agents keep underspending? Because models stop on their own, and harnesses stop them sooner.

Agents declare victory. Anthropic's post on long-running agent harnesses lists the failure mode: "a later agent instance would look around, see that progress had been made, and declare the job done." That was Opus 4.5, looping across many context windows with compaction, asked to build a web app from a high-level prompt.

Flat swarms play it safe. Cursor's first long-running swarm let agents coordinate as equals through a shared file. "With no hierarchy, agents became risk-averse. They avoided difficult tasks and made small, safe changes instead. No agent took responsibility for hard problems or end-to-end implementation." In the same post, "Opus 4.5 tends to stop earlier and take shortcuts when convenient, yielding back control quickly." In the follow-up, their continuous executor would "refuse to plan and spawn more than a few narrowly focused tasks" and "claim premature completion."

Production agents are built to stop. The Measuring Agents in Production survey of 306 practitioners found that "68% execute at most 10 steps before requiring human intervention," and 47% fewer than five. The authors say organizations "deliberately trade capability for controllability." That is a reasonable choice for an agent that answers customers. It also means most deployed agents run far below the regime where every result in the previous section lives.

My own notes are full of the same complaint. In early July I wrote that I like it when the harness "spawn a lot of subagents investigating," because then the task "have more probability of being made effective ... i have infinite compute." The confident answer after a long trajectory was almost always better than the quick, hedged one. A quick answer isn't efficient if it's wrong.

Fan out and recurse until the answers repeat

In nearly every serious task, I now paste a standing block that tells the agent to treat inference compute as effectively unconstrained. In practice it means:

The stopping rule is the important part: stop at saturation, not at a budget. If the fifth wave returns nothing new, the space has probably been covered. If it still turns up surprises, you stopped too early. I treat single trajectories the same way. I'd rather give one run a very large max_tokens budget and get the long, opinionated answer than cap it and get a hedge.

Wide fan-out stays affordable because the models are tiered. Mechanical work and web search go to cheap models. Planning and synthesis go to expensive ones. Cursor measured why this works in Agent swarms and the new model economics. Across their model mixes, workers carried "at least 69% of the tokens, and over 90% in most." With GPT-5.5 doing both jobs, the workers alone cost $9,373. With Opus 4.8 planning and Composer 2.5 working, "the entire worker fleet cost $411," and the whole run $1,339, against $10,565 for GPT-5.5 alone. Quality was similar. Their explanation matches my experience: "Few moments in a large task genuinely require frontier intelligence."

By September, the same preamble also tells the agent to research existing solutions first and to prefer project code or maintained libraries over new code. Fan-out without that turns into fifty agents writing fifty versions of the same helper.

On owned hardware, idle compute is waste. When a marginal token costs little more than power, the useful question stops being how many agents you can afford and becomes how many the machine can run. I think that number should come from an experiment, not a guess. Start a loop at 64 parallel agents. If it holds, go to 128, then 256, then 512, "tipo uma busca binaria" ("like a binary search," as I put it on September 3). Run one configuration per host so the results aren't polluted by interference.

My bet is that what breaks first won't be inference. An inference error can be retried, and a simple retry absorbs most of the pressure on that side. What runs out is everything around the model: sandbox instances, memory, disk, locks. So every limit should justify itself from first principles. If the memory is free, what stops the next doubling? Often the answer is isolation that's heavier than it needs to be, like a full clone per agent where a git worktree would do.

Two kinds of capacity matter here. Hard capacity is the width at which the machine becomes unstable. Throughput capacity is useful trajectories per hour. They are not the same number, and between them there's a Pareto frontier of agents in parallel against trajectories in parallel. That's why the number worth tracking is healthy trajectories that end with a submission, not agent count.

The public reports hit the same walls, and none of them is inference:

That last case is the one I like most, because the fix for a width problem was a better oracle, not more compute. It's the same binary search, pointed at the bug instead of the hardware.

Bound the tool call, not the trajectory

Once inference stops being the scarce thing, turn limits look strange. A max_turns cap doesn't protect anything real. It just cuts off the trajectories that were working hardest. So my rule is no turn limits, no wall-clock budgets and no message caps. Thinking stays on and tool calls are unlimited. On August 31 I wrote that I didn't want to impose max_turns on the agents, and that "o unico gate de timeout deve ser se a trajetoria passar de 10 compactacoes de contexto" (the only timeout gate should be a trajectory going past 10 context compactions).

Long trajectories survive through self-compaction. At about 70% of the window the agent compacts itself back below 30% and keeps going. About ten compactions is the only hard stop. A long-surviving agent is a good sign, not a leak. (What the context window should contain, and why the log and not the window is the source of truth, is the subject of The log is the truth.)

Resource control moves down to the individual tool call, where the real hazards are. A process that runs past 5 minutes, or past 1 GiB of memory, gets killed, and the failure comes back as the tool's output. The agent reads "killed after 5 minutes" or "exceeded 1 GiB" the same way it reads a compiler error, and changes its approach. The system prompt tells the model what the limits are, so a kill is information and not a surprise. I prefer this runtime enforcement to predicting memory in advance, which I never managed to do well.

Cursor landed in the same place from the other direction: agents "may write code that has memory leaks or causes deadlocks," and "explicit process-based resource management tools were required to allow the system to gracefully recover." Carlini built the time version. Claude "can't tell time and, left alone, will happily spend hours running tests instead of making progress," so his harness had a default --fast option that ran a deterministic 1% or 10% sample of the tests. Both bound the tool call and leave the trajectory alone.

Speculative experiments, asynchronous results

Long trajectories are slow for a boring reason. The agent runs one full experiment, waits for it, and only then finds out that a prerequisite was broken. The fix I use came from Alex Zhang's Speculative Programmatic Tool Calling (sPTC). In a harness where code in a REPL is the action space, sPTC will "speculate and pre-launch tool calls from partially generated REPL calls as the harness is still generating tokens." If the finished code actually makes those calls, they return right away from the speculated results. It also acts "as a really naive JIT compiler," running independent sub-agent calls in parallel even when the code wrote them sequentially.

On his RLM benchmarks, Zhang reports speed-ups "generally on the order of 1-1.2x." His claim is about latency hiding, and it's modest on purpose.

What I took from it is broader, and it's my extension, not his. On August 28 I wrote that "a minha visao de sptc e o agente conseguir fazer varios testes em paralelo nn e o produto e o metodo de chegar no produto" (my version of sPTC is the agent running many tests in parallel; it isn't the product, it's the method of getting to the product). Launch many cheap, different experiments at once across all the machines. Find the one that works on one host. Then scale that one. In the harness this means the model fires every independent action at once, any tool call and not only subagents, and results stream back during the trajectory as runtime events instead of blocking the turn. On September 21: "quero q as tool outputs sejam recebidas durante a trajetoria (conforme elas venham ganhando os resultados ...), pois ai assim vamos conseguir acelerar bastante" (I want tool outputs received during the trajectory, as results come in, because that's how we speed this up a lot). Subagents get the parent's full tool set, run at a reasoning effort the parent picks, and have a cap set high enough that it only catches runaways.

Anthropic named the same bottleneck in their research system: "Synchronous execution creates bottlenecks," because "the entire system can be blocked while waiting for a single subagent to finish searching." Moving to 3 to 5 parallel subagents, each running 3 or more tools in parallel, "cut research time by up to 90% for complex queries."

Long, unattended, under a supervisor

The most productive thing to do with abundant compute is also the laziest. Clone a well-contextualized agent into many long-running copies and leave them overnight. The pattern has a name. Geoffrey Huntley called it Ralph in July 2025: "In its purest form, Ralph is a Bash loop," while :; do cat PROMPT.md | claude-code ; done. His next line is the thesis of this post in one sentence: "Ralph can be done with any tool that does not cap tool calls and usage." Carlini's compiler harness is the same idea: a while true loop that starts a fresh Claude Code session every time the last one exits, so "when it finishes one task, it immediately picks up the next."

I like to run several Ralph loops on the same question at once, each in its own folder or tmux pane, left for a night and compared in the morning. Loops that run longer than a night need a supervisor agent that watches trajectory health, fixes errors and rescales width. Its objective, as I wrote on September 2, is that "o maior reward function do goal e maximizar progresso ... sempre focar no maximo de troughput de progresso" (the biggest reward function of the goal is maximizing progress, always focused on maximum progress throughput).

The resilience rules matter more than the width. Retry forever against the one chosen model. Turn every failure into a typed outcome on that trajectory, so no failure can crash the coordinator. Repair live runs in place when the fix is small, and restart cleanly when errors would contaminate the trajectories.

The public long runs look similar:

Run Width Duration Cost
Carlini, C compiler 16 agents, ~2,000 sessions two weeks just under $20,000 at API prices
Cursor, browser several hundred agents, ~10M tool calls one week not disclosed ("trillions of tokens")
Sumner, Bun in Rust 64 Claudes, ~50 workflows 11 days ~$165,000 at API prices

Cursor's browser system "peaked at ~1,000 commits per hour" and "once the system started, it didn't require any intervention from us." Bun passed its full test suite on every platform. Carlini's compiler builds a bootable Linux 6.9 on three architectures.

"Unattended" doesn't mean "unsupervised," though. Carlini: "Most of my effort went into designing the environment around Claude—the tests, the environment, the feedback—so that it could orient itself without me." Sumner spent most of the 11 days reading workflow outputs and "prompting Claude to edit the loop to fix things," and his rule was to fix "the process that generates the code instead of hand-fixing the code." The human moves up one level, from doing the work to designing the loop, which is the argument of Find the loops.

Effort is not evidence

Fan-out multiplies claims as fast as it multiplies coverage. A thousand subagents produce a thousand confident paragraphs. If the final answer isn't held to an evidence standard, more compute just produces more slop, and faster.

So the other half of my preamble is a rule I wrote on July 8: "State a claim as true only when you can point to specific evidence for it ... AND you attempted to falsify it." Specific evidence means a source, a reproduced result, or a passing test. The falsification attempt has to be real. Around that rule:

Sumner's version is built into the structure: "1 implementer, 2 or more adversarial reviewers per implementer." His reason is the same one: "The Claude that wrote the code wants the code to get accepted. The Claude that reviews wants to find issues in the code." Every line of the Bun port went through two such reviews before it was committed.

The same rule applies to how I read runs. Commit volume is not progress. I tell agents to judge real progress against scaffolding and volume, with evidence and numbers. And a null counts. As I put it once: "Your job is to test it, not to prove it. A null or negative result, clearly measured, is a successful run." (How to build evals that can tell those apart is in Evals built backwards.)

Where it breaks

The thesis has four real limits.

Verification doesn't scale like coverage. This is the strongest objection, and it comes from the same paper I opened with. On MATH, Llama-3-8B-Instruct's coverage rose from 82.9% with 100 samples to 98.44% with 10,000. But with majority voting or a reward model picking the answer, "the biggest performance increase is only from 40.50% to 41.41% over the same sample range." Those methods "plateau beyond several hundred samples." The correct answer is in the pile. Nobody can find it.

Coverage keeps rising with samples, but answer selection without a verifier stays flat

Even with a verifier, verification is imperfect. The same authors describe unit tests producing false positives and false negatives in their own coding runs. Carlini says it directly: "it is easy to see tests pass and assume the job is done, when this is rarely the case." Snell's paper adds a ceiling on the other side. On the hardest MATH questions "no method makes much meaningful progress," and on harder questions or under high inference load, "pretraining is a more effective way to improve performance." So the honest version of my title is conditional. Compute is not the bottleneck when a good verifier exists. When it doesn't, the verifier is the bottleneck, and more samples mostly buy a bigger haystack.

The multi-agent token tax. Anthropic is candid about the price: "agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats." Their early agents made mistakes like "spawning 50 subagents for simple queries," and vague task descriptions made subagents "duplicate work, leave gaps." They also say "most coding tasks involve fewer truly parallelizable tasks than research." Carlini's 16 agents on one kernel bug are that sentence in practice. Width only pays when the task can be split, or when someone builds an oracle that splits it.

API prices are not owned hardware. Carlini's compiler cost just under $20,000 and Bun's port about $165,000, both at API prices. For most teams, "compute is unconstrained" at those prices is a budget decision, and Anthropic's rule applies: multi-agent systems "require tasks where the value of the task is high enough to pay for the increased performance." The claim is strongest on self-hosted open models, where a marginal token costs little more than power and an idle GPU is the real waste. There the constraint becomes capacity, which is why the binary search exists. On APIs the argument still holds with a smaller multiplier, through tiering (Cursor's $411 worker fleet) and through the Monkeys result that many cheap samples can beat one expensive one.

More agents, more slop. Cursor's economics post has the best numbers. Under their old harness, a Grok 4.5 swarm "produced 68,000 commits in its first two hours, roughly 70 times the new run's pace," and "accumulated more than 70,000 conflicts before we paused it." One file collected 7,771 conflicts from 1,173 agents. The old swarm sprawled to 54 crates, "including three separate SQL packages," while the new one settled on nine. In one mix, the old run needed 64,305 lines of engine code to pass the suite, and the new run 9,908. Width without structure buys sprawl. I agree with this completely, and it's why I wrote in August that "the real opportunity is not to unleash thousands of agents to generate slop, but to place swarms of agents inside closed, cybernetic loops." Width is only safe inside a loop with a deterministic eval that is hard to game. (Why flat swarms produce stigmergic slop is in Bottom-up is one ontological level higher.)

Put the four together and the claim narrows, but it gets sharper. Compute is not the bottleneck for agents working on tasks that can be split, against verifiers that can be trusted, on hardware you already own. Most of the work I care about can be put in that shape, and putting it in that shape is the job.

What follows if the verifier is the scarce thing

This part is speculation, and I'll mark where it stops being evidence.

The evidence says that where answers can be checked, sampling keeps paying for four orders of magnitude, and where they can't, it flattens early. If that holds, the value moves away from the model and toward what surrounds it: a good oracle, sandboxes that start fast and die cleanly, and tokens per second on hardware you control. (That compute works like capital for agents is the argument of Positive feedback eats the world.)

Formal math is where I'd push this to the extreme. Picture a frozen open-weight policy running thousands of parallel proof attempts and improving without any weight updates, with what it learns kept in external memory and the Lean kernel as the only oracle. Intelligence would live in the loop, not in the weights. Nobody has shown that this works at scale, and I haven't either. It's where the argument points: if verification is cheap and exact, a fixed model plus enough width and memory behaves like a model that keeps learning.

My guess is that harnesses will turn into schedulers. They'll decide per task how much width, how many sequential revisions and how many verifier calls to buy, the way Snell's compute-optimal policy allocates by difficulty. And the operator's job will become what Carlini and Sumner already did by hand: build the oracle, find the hardware ceiling, and fix the loop instead of the output. That's where the speculation ends.

The preamble, as a checklist

Here is the standing block I paste into serious tasks, rewritten so anyone can use it.

For the agent:

  1. Treat inference compute as effectively unconstrained. Spend tokens before you spend my attention.
  2. Research first. Look for existing solutions, project code and maintained libraries before writing anything new.
  3. Fan out. One subagent per independent angle. Subagents may spawn their own.
  4. Work in waves. Each wave starts from what the last one found. Stop when a new wave returns only what you already know.
  5. Tier the models. Cheap models for mechanical work and search, strong models for planning and synthesis.
  6. Fire independent actions together. Don't wait on a result you don't need yet. Read results as they arrive.
  7. Try many cheap, different experiments before scaling one.
  8. Don't stop because the trajectory is long. Compact when the window fills and keep going.
  9. Every claim needs specific evidence and a falsification attempt that failed. If the question is still open, report a leading hypothesis and its gaps.
  10. Ask why the best expert in the field would reject your current choice.
  11. A clearly measured null is a result. Don't soften it.

For the operator:

  1. Find capacity by doubling until failure. 64, 128, 256, 512, one configuration per host.
  2. Count healthy trajectories that end with a submission, not agents.
  3. Bound tool calls, not trajectories. Kill a process on time or memory and return the reason as its output.
  4. Retry inference forever. Turn every other failure into a typed outcome on that trajectory.
  5. Put a supervisor on long runs, and restart cleanly when errors would contaminate trajectories.
  6. Build the oracle before adding width. If every agent hits the same bug, more agents won't help.

In a harness, the whole thing is a small config:

const harness = {
  trajectory: {
    maxTurns: null,
    wallClock: null,
    compactAt: 0.7,              // of the context window
    compactTo: 0.3,
    stopAfterCompactions: 10,    // the only hard stop
  },
  tools: {
    parallel: true,              // fire every independent call at once
    results: "stream",           // deliver outputs as events mid-trajectory
    timeoutMs: 5 * 60_000,
    memoryBytes: 1024 ** 3,
    onLimit: "kill-and-return",  // the failure becomes the tool output
  },
  subagents: {
    max: 1024,
    recursive: true,
    tools: "inherit",
    effort: "parent-chooses",
    stopWhen: "wave-returns-nothing-new",
  },
  models: {
    plan: "strong",
    work: "cheap",
    onInferenceError: "retry-forever",
  },
  capacity: {
    widths: [64, 128, 256, 512], // double until unstable, one config per host
    metric: "healthy-trajectories-with-submit-per-hour",
  },
  claims: {
    require: ["specific-evidence", "failed-falsification"],
    disagreements: "adjudicator",
  },
  supervisor: {
    objective: "progress-throughput",
    onContamination: "restart-clean",
  },
}

None of these numbers is sacred. The 5 minutes, the 1 GiB and the ten compactions are starting points, not measurements. The point is where the limits sit: on the processes, where the damage happens, and not on the thinking, where the value is.

Sources

← The log is the truth
Show the problem, hide the metric →

Markdown version: /blog/compute-is-not-the-bottleneck.md. Every essay: /agents.