# Bottom-up is one ontological level higher

*September 2026*

MetaGPT's README has a section called [Software Company as Multi-Agent System](https://github.com/FoundationAgents/MetaGPT). "Internally, MetaGPT includes product managers / architects / project managers / engineers. It provides the entire process of a software company along with carefully orchestrated SOPs." The next line is the philosophy: "`Code = SOP(Team)`." ChatDev's README describes its first version the same way: ["a Virtual Software Company"](https://github.com/OpenBMB/ChatDev) whose agents, "e.g., CEO, CTO, Programmer," attend "specialized functional seminars." CrewAI's docs sell ["Role-Playing Agents"](https://docs.crewai.com/en/introduction) that "communicate and collaborate like human team members," managed by a Flow you should "think of as the 'manager.'"

On June 12, 2025, Walden Yan of Cognition published [Don't Build Multi-Agents](https://cognition.ai/blog/dont-build-multi-agents): "running multiple agents in collaboration only results in fragile systems." The next day Anthropic published [How we built our multi-agent research system](https://www.anthropic.com/engineering/built-multi-agent-research-system), where a lead Opus 4 agent with Sonnet 4 subagents "outperformed single-agent Claude Opus 4 by 90.2%" on an internal research eval, while using about 15 times the tokens of a chat. And when Berkeley researchers ran the software companies on simple programming tasks, "such as implementing Tic-Tac-Toe, Chess, or Sudoku," [MetaGPT failed 60% of the time and ChatDev 66.7%](https://arxiv.org/abs/2503.13657).

Most diagrams of advanced AI look like organizational charts. A large box at the top decomposes the objective, passes it to a planning agent, which delegates to research agents, coding agents, verification agents and memory agents. Their outputs move upward until a final answer emerges. It is a corporation rendered in Mermaid.

My claim is that this gets the order of things backwards. Bottom-up systems are one ontological level higher than top-down systems, because they can produce top-down systems as local structures. A completely top-down system can't produce bottom-up emergence without surrendering some of the control that makes it top-down. Cognition and Anthropic argue over whether to use multiple agents. Others argue over whether "scaffolding is dead." Both questions treat structure as one quantity. It's two. Structure that tells the system how well it did should grow as models improve. Structure that tells the system how to act should shrink. The org chart is almost entirely the second kind.

## A corporation rendered in Mermaid

It's strange that our imagination of machine intelligence so quickly recreates the institution most of us already spend our lives in. We give the model a CEO, managers, departments, meetings and a chain of command, then call the result autonomous.

Melvin Conway explained why in 1968. ["Any organization that designs a system... will inevitably produce a design whose structure is a copy of the organization's communication structure."](http://www.melconway.com/research/committees.html) He added: "To the extent that an organization is not completely flexible in its communication structure, that organization will stamp out an image of itself in every design it produces." People who work in companies design agent frameworks, so the frameworks come out shaped like companies.

The imagination used to be wider. In October 2000, in a book chapter posted to the SL4 mailing list, Ben Goertzel predicted an Internet ["centered on a community of powerful AI agents"](http://sl4.org/archive/0010/0117.html) and asked how "these agent societies are going to operate." He pictured a society. Twenty-three years later the frameworks shrank it to a company.

Christopher Alexander gave the deeper reason in 1965, in [A City Is Not a Tree](https://www.patternlanguage.com/archive/cityisnotatree.html). Designers draw trees because "the mind has an overwhelming predisposition to see trees wherever it looks." A tree is a structure in which every unit has one parent and nothing overlaps. It's easy to hold in your head, which makes it a fact about the designer's mind.

The claim is about generative range, not morality. A market can generate a firm. A neural network can learn an algorithm. Evolution can generate an organ. A general intelligence can construct a hierarchy. But a hierarchy can't fully specify an emergent system in advance, because the emergent part is exactly the part the specification doesn't contain. Hierarchy is one configuration emergence can produce. Emergence isn't a department inside a hierarchy.

## The overlaps the tree deletes

Alexander's example is a street corner in Berkeley. A drugstore has a newsrack in its entrance, and a traffic light stands outside. People waiting for the light read the papers on the rack. The rack, the light and the sidewalk form one unit; the drugstore and its entrance form another. The two units overlap in the newsrack, and the overlap is where the city is alive. Natural cities are semilattices, full of these overlaps. Planned cities are trees, and "only in the artificial-tree conception of the city are their natural, proper and necessary overlaps destroyed." His conclusion: "The city is a receptacle for life. If the receptacle severs the overlap of the strands of life within it, because it is a tree, it will be like a bowl full of razor blades on edge."

Most agent systems are artificial cities. The researcher researches, the coder codes, the critic criticizes. But the hard part of real work usually sits between these categories. A design decision changes the code, the product, the customer's incentives, the economics of the company, the future training data, and the kind of organization needed to maintain it. Assigning each concern to a separate agent destroys the overlap where judgment lives. Every department can be locally correct while the company becomes globally stupid.

The evidence for this is the failure data.

**Cognition's Flappy Bird.** Yan's example splits "build a Flappy Bird clone" into a background subtask and a bird subtask. One subagent builds a Super Mario background, the other builds a bird that "moves nothing like the one in Flappy Bird." Even when both get the full context, they pick different visual styles, because "Actions carry implicit decisions, and conflicting decisions carry bad results." The style of the game is the overlap. No subtask owns it.

**MAST.** [Why Do Multi-Agent LLM Systems Fail?](https://arxiv.org/abs/2503.13657) annotated 1,642 execution traces from seven open-source frameworks, built a taxonomy of 14 failure modes with inter-annotator agreement of κ = 0.88, and found "41% to 86.7% failure rate" across the systems.

![MAST failure rates across six multi-agent frameworks, with MetaGPT and ChatDev highlighted](https://future-seems-so-good.com/blog/assets/bottom-up-is-one-ontological-level-higher/charts/mast-failure-rates.svg)

The traces show where the overlap goes missing. A chess program from ChatDev "passes superficial checks (e.g., code compilation) but contains runtime bugs because it fails to validate against actual game rules." Verifier agents "perform only superficial checks, despite being prompted to perform thorough verification." A verifier persona checks what a verifier role is supposed to check. Whether the game follows the rules of chess belongs to the product, the code and the tests at once, and the tree gave it to nobody.

**One judge on top.** Warren McCulloch's 1945 paper, [A heterarchy of values determined by the topology of nervous nets](https://journal.emergentpublications.com/Article/7b136ca7-c0eb-4370-a50c-14c83b21c94e/github), showed that a net of six neurons can prefer A to B, B to C and C to A. What looks like inconsistency indicates "consistency of an order too high to permit construction of a scale of values," an organism "too rich to submit to a summum bonum."

We keep placing one scalar above every intelligent process: one reward, one planner, one CEO agent, one final judge into which every conflict must collapse. Instead of a spreadsheet ranking every value for eternity, the system needs a topology that can resolve the conflict from inside the situation.

## Hierarchy is cached emergence

Herbert Simon gives the strongest argument against a naive bottom-up view, and it's correct. In [The Architecture of Complexity](https://mwolf.dev/library/architectureofcomplexity-hsimon1962/TTS) (1962), two watchmakers build watches of about 1,000 parts. Tempus assembles each watch in one piece, and every interruption makes it fall apart. Hora builds "subassemblies of about ten elements each," then subassemblies of those. With a 1% chance of interruption per part, Simon calculates that Tempus takes "about four thousand times as long to assemble a watch as Hora."

An intelligence that dissolves every boundary will spend all its time rediscovering structure. So keep the structure, and treat hierarchy as cached emergence. A good module is a solution you no longer reconsider at every step. A bureaucracy begins when the cache becomes impossible to invalidate.

Read `Code = SOP(Team)` with that in mind. A standard operating procedure is a cache. Software companies computed it over decades, under human constraints: eight-hour days, small working memory, meetings that cost salaries, people who defend their turf. The cache key was a human substrate. Agents run on a different one. Their bottleneck is context length, and they don't carry egos between tasks (Cognition's 2026 post: ["They don't have egos"](https://cognition.com/blog/multi-agents-working)). Copying the SOP into prompts reuses a cache after its key has changed.

Rich Sutton's [Bitter Lesson](http://www.incompleteideas.net/IncIdeas/BitterLesson.html) says built-in human knowledge "plateaus and even inhibits further progress," and his list of what to stop building in includes, word for word, "multiple agents."

Caches also govern what produced them. Hermann Haken's [synergetics](http://www.scholarpedia.org/article/Synergetics) calls this the slaving principle: parts create an order parameter that then determines their behavior, "one speaks of circular causality." A group creates a metric and eventually works for the metric. Hierarchy itself is fine. The mistake is forgetting that the hierarchy was produced, then treating it as prior to the process that produced it.

Markets show the asymmetry. Ronald Coase, in [The Nature of the Firm](https://onlinelibrary.wiley.com/doi/full/10.1111/j.1468-0335.1937.tb00002.x), quoted D. H. Robertson on firms as "islands of conscious power in this ocean of unconscious co-operation like lumps of butter coagulating in a pail of buttermilk." Outside the firm, prices coordinate. Inside it, a worker moves departments "because he is ordered to do so." The decentralized process produces centralized islands wherever command is cheaper than negotiation. Command becomes a local strategy. A total command system has a much harder time instantiating independent discovery, because the discovery may contradict the command. To permit emergence is to permit the possibility that the center was wrong.

## The agent is a phase transition, not an employee

Gilbert Simondon reversed the usual order between the individual and the process that makes it. In [The Position of the Problem of Ontogenesis](http://www.parrhesiajournal.org/parrhesia07/parrhesia07_simondon1.pdf), the individual is "a relative reality, a certain phase of being that supposes a preindividual reality," and "individuation does not exhaust with one stroke the potentials of preindividual reality." His model is a crystal forming in a supersaturated solution: a metastable field full of tension resolves into structure, and the structure and its milieu appear together.

That's a better image for agents than permanent personas. A "researcher" shouldn't exist before the problem. The problem creates a tension, the system individuates a research process around it, the process persists while the tension needs it, and then it dissolves. What stays is the structure it produced.

This avoids the most common mistake in multi-agent design, where personality replaces architecture. We tell one model it's skeptical, another that it's creative, as if intelligence will emerge from a dinner party of adjectives. Cognition's follow-up names it: "Prompt engineering encourages gimmicky techniques like 'you're a senior software engineer' or 'think for longer.'"

The multi-agent systems that work already treat agents this way. Anthropic's lead research agent spawns subagents per query, following rules in its prompt: "Simple fact-finding requires just 1 agent with 3-10 tool calls," and "complex research might use more than 10 subagents." Each subagent returns condensed findings and ends. Even the ChatDev team's 2025 paper proposed trading its fixed chain of roles for ["a learnable central orchestrator optimized with reinforcement learning"](https://github.com/OpenBMB/ChatDev). A better system asks what temporary differentiation the problem requires, instead of which permanent member of the organization owns the task.

## Slop is stigmergic

Coordination also doesn't require every component to keep reporting to a manager. Pierre-Paul Grassé coined stigmergy in 1959 to explain how termites build. His definition was ["the stimulation of workers by the very performances they have achieved."](http://pespmc1.vub.ac.be/Papers/Stigmergy-Springer.pdf) Francis Heylighen's summary: "work performed by an agent leaves a trace in the environment that stimulates the performance of subsequent work—by the same or other agents," which works "without any need for planning, control, or direct interaction between the agents."

The word stayed with the people who studied it. Theraulaz and Bonabeau wrote [its history](https://psycnet.apa.org/doi/10.1162/106454699568700) for Artificial Life in 1999; in the Usenet archive I searched, comp.ai.alife used it most, and Hacker News has 12 posts with it since 2020. Programmers rebuilt the mechanism without the name.

A codebase is stigmergic. A test left by one programmer changes what another does years later. So does a type, a failing build, a database constraint, a production incident, or a strange function name. The repository becomes memory, the test suite becomes law, and the architecture becomes frozen history.

**The compiler built without a manager.** In February, Nicholas Carlini published [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler). Sixteen agents ran nearly 2,000 Claude Code sessions over two weeks, consumed 2 billion input tokens and 140 million output tokens for just under $20,000, and produced a 100,000-line Rust compiler that builds Linux 6.9 on x86, ARM and RISC-V ([HN discussion](https://news.ycombinator.com/item?id=46903616)). The coordination mechanism is a directory. "Claude takes a 'lock' on a task by writing a text file to current_tasks/... If two agents try to claim the same task, git's synchronization forces the second agent to pick a different one." Then: "I don't use an orchestration agent. Instead, I leave it up to each Claude agent to decide how to act." Carlini's summary: "Most of my effort went into designing the environment around Claude—the tests, the environment, the feedback—so that it could orient itself without me."

That's termite architecture with git as the mud. The specialized agents maintained the medium: "LLM-written code frequently re-implements existing functionality, so I tasked one agent with coalescing any duplicate code it found." Others worked on performance, code quality and documentation. None of them was a department.

**Bad traces recruit more bad work.** Slop is stigmergic. A locally ugly abstraction changes the affordances visible to every future contributor. The next agent reads the bad decision, then starts thinking inside the world it created.

Snyk showed it with Copilot in [Copilot amplifies insecure codebases by replicating vulnerabilities](https://labs.snyk.io/resources/copilot-amplifies-insecure-codebases-by-replicating-vulnerabilities/) (February 2024). Asked for a SQL query, Copilot wrote a parameterized one. With a vulnerable query open in a neighboring tab, the same request produced string concatenation: "We've just gone from one SQL injection in our project to two." A study of 733 AI-generated snippets in GitHub projects, [Security Weaknesses of Copilot-Generated Code](https://arxiv.org/abs/2310.02059), found weaknesses in 29.5% of the Python ones and 24.2% of the JavaScript ones.

Slop comes from training on short horizons (*How to achieve superintelligence*), and stigmergy is how it compounds after deployment. If the codebase is the coordination medium, code quality is the noise floor of the channel every future agent reads.

## Two kinds of scaffolding

On July 31 I worked through a formal version of this with Claude. The formalism below is coauthored, and the parameters are illustrative, not measured.

We started from an earlier document that described scaffolding with one axis, H = 1 − B: the more hierarchy, the less bottom-up freedom. Claude's objection was that this makes the document's own conclusion unrepresentable: "a tightly specified envelope with fully autonomous agents inside it." "It's the difference between 'how much structure' and 'which structure,' and the second question is the one that pays." We kept the document's naming. **Evaluative** scaffolding tells the system how well it did: verifiers, tests, evals, scores, memory. **Prescriptive** scaffolding tells it how to act: fixed roles, stage order, mandatory critique passes, hand-written routing, rigid handoff schemas. The document also had the best name for the failure mode of too little structure: "a rhizome of incompetent nodes does not generate intelligence; it generates combinatorial diffusion."

With S_p for prescriptive structure, S_e for evaluative structure and I for the capability of the substrate, all between 0 and 1, the model is:

```
Λ = ln(1 + σ·S_p) + κ·I + μ·S_e·I − λ·S_p·I − ν·(1 − S_p)·(1 − I)
D = e^Λ
E = 1 − e^(−D/V)
```

Λ is the log of the difficulty the system can handle, D the difficulty, and E the fraction of a task distribution (mean difficulty V) it covers. The terms:

- **Prescription buys a bounded gain**, ln(1 + σ·S_p). You write the rules you're most sure of first, so the fortieth rule is a guess, and a scaffold can't outperform its author.
- **Capability buys an exponential**, κ·I. Prescription sits inside a log; capability sits outside it.
- **Verification multiplies capability**, μ·S_e·I. A verifier with nothing capable to check is worthless, and a capable model that can't tell good output from bad can't use best-of-n or self-correction.
- **Prescription clips capability**, −λ·S_p·I. A fixed decision costs the gap between what you specified and what the model would have chosen: free when the model is weak, expensive when it's strong.
- **Diffusion**, −ν·(1 − S_p)·(1 − I). It needs both open space and incompetence. Weak agents in a rigid pipeline execute badly but don't spiral. Weak agents in an open rhizome duplicate work and never converge.

Two results fall out. The derivative with respect to evaluative structure is ∂Λ/∂S_e = μ·I, which is never negative: you build as much verification as you can afford, and it becomes more valuable as models improve. For prescriptive structure the optimum is:

```
S_p*(I) = 1 / ((λ + ν)·I − ν) − 1/σ,   clipped to [0, 1]
I_low  = ν / (λ + ν)          below this, S_p* = 1
I_high = (σ + ν) / (λ + ν)    above this, S_p* = 0
```

Below I_low, in Claude's words, "prescription isn't merely preferred, it's forced, because the diffusion cost dominates." Above I_high, prescription is pure loss. With σ = 6, λ = 8 and ν = 1.5, I_low is 0.16, I_high is 0.79, and in between the optimal amount of prescription falls fast: 1.00 at I = 0.2, 0.57 at 0.3, 0.14 at 0.5, 0.03 at 0.7.

Claude also reframed my intuition about minimal programs: the best controller minimizes K(M) + K(S|M) + λ·C_runtime subject to E ≥ τ. Intelligence relocates complexity from orchestration code into weights, environment and feedback.

What convinces me is that the published evidence sorts along its axes.

**The biggest ChatDev fix was a loop, not a better org chart.** MAST's authors tried two interventions on ChatDev with the same model. Stricter role and verifier prompts, including a rule that "only superior agents can finalize conversations" (the paper's summary: "ensuring the CEO had the final say"), took task success from 25.0% to 34.4%. Changing the topology from an acyclic pipeline to a cycle that "terminates only when the CTO agent confirms that all reviews are properly satisfied," which the main text calls "adding a high-level task objective verification step," took it to 40.6%.

![ChatDev task success with stricter role prompts versus a verification loop](https://future-seems-so-good.com/blog/assets/bottom-up-is-one-ontological-level-higher/charts/chatdev-interventions.svg)

**Cognition's patterns that survived are evaluative.** Ten months later, Yan's [Multi-Agents: What's Actually Working](https://cognition.com/blog/multi-agents-working) reports that "multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions." [Hacker News](https://news.ycombinator.com/item?id=45096962) readers had [narrowed](https://news.ycombinator.com/item?id=45098720) the first post to "do not build paralleled multi-agents," and [noted](https://news.ycombinator.com/item?id=45097924) that Anthropic's article "says the same things as the cognition article, it just has a different definition of multi-agent." The lead example is a reviewer. On PRs Devin wrote itself, "Devin Review catches an average of 2 bugs per PR, of which roughly 58% are severe," and it works best "when the coding and review agents do not share any context beforehand."

**A verifier role isn't evaluative structure.** MAST's superficial verifiers were prompted personas. Carlini's verifier was a test suite, a CI pipeline and, for the kernel, GCC as an oracle, and he said it had to be "nearly perfect, otherwise Claude will solve the wrong problem." Evaluative structure needs contact with ground truth. An agent told to be a reviewer is one more prescription. (I make the same argument for companies in [Show the problem, hide the metric](https://future-seems-so-good.com/blog/show-the-problem-hide-the-metric): deterministic sensors beat judges.)

Cognition's rule, many readers and one writer, is also the shape bottom-up takes when it builds its own top-down. Decentralize to search, centralize to act.

## Where the rhizome fails

The case against this essay is strong, and some of it comes from my own sources.

**MAST cuts both ways.** Its authors write that "many MAS failures arise from the challenges in organizational design and agent coordination rather than the limitations of individual agents." That reads as my point, or as "design the org better." MetaGPT, an "Assembly Line" in the paper's classification, showed 60 to 68% fewer system-design and inter-agent misalignment failures than ChatDev, with 1.56 times more verification failures. Prescription worked on the axis it was aimed at.

**The winners are hierarchies.** Anthropic's lead-agent tree beat a single agent by 90.2%, largely by spending tokens ("token usage by itself explains 80% of the variance" on BrowseComp), and Anthropic fences off coding: "LLM agents are not yet great at coordinating and delegating to other agents in real time." Cursor's case is harder: it tried the design this essay would have recommended, and it failed. In January, Wilson Lin described [running hundreds of concurrent agents](https://cursor.com/blog/scaling-agents) on one codebase, first with equal status and a shared lock file, Carlini's `current_tasks/` at a larger scale. "Twenty agents would slow down to the effective throughput of two or three." Worse: "With no hierarchy, agents became risk-averse. They avoided difficult tasks and made small, safe changes instead." What worked was recursive planners over workers that never talk to each other. By February the harness [orchestrated "many thousands of agents"](https://cursor.com/blog/self-driving-codebases) and "peaked at ~1,000 commits per hour across 10M tool calls over a period of one week." Cursor's reading is a direct rebuttal of my title: "These models were not explicitly trained in this way, which suggests it's emergent behavior and possibly the correct way of structuring software projects after all."

Two details complicate it. The hierarchy was found by search, not specified: flat, then a planner-executor-judge pipeline that was "too rigid," then one overwhelmed executor, then recursive planners. Lin's summary is the Λ model's interior optimum in words: "Too little structure and agents conflict, duplicate work, and drift. Too much structure creates fragility." And in July Cursor wrote that the browser ["succeeded as a proof of concept, but fell far short of polished software."](https://cursor.com/blog/agent-swarm-model-economics)

**Below a threshold, the rhizome only diffuses.** When Carlini's agents reached the Linux kernel, one giant task, "Every agent would hit the same bug, fix that bug, and then overwrite each other's changes. Having 16 agents running didn't help." That's combinatorial diffusion, observed. Capability decided whether the setup worked at all: "Opus 4.5 was the first to cross a threshold." Cognition hit the lower edge from the other side: its small SWE 1.5 model "was not good enough at being the primary model," and its suggested remedy is prescription, exactly where the model predicts: "Depending on the intelligence of the primary model, certain kinds of domain-specific prescriptive guidance may be necessary, such as always invoking the smart friend for merge conflicts."

So the thesis needs a qualifier. Bottom-up is the higher level, but a system can only occupy it once its nodes are competent enough to hold the loop closed. Below that, prescription is forced, and people building with weak or cheap models aren't wrong to build pipelines, as long as each is a patch with an expiry date. In stationary domains, the pipeline never expires.

## If you build the org chart anyway

The strongest objection to this essay is that the people running the largest agent systems, at Anthropic, Cursor and Cognition, keep drawing hierarchies, and I feel the same pull when I design agents for long runs. The design that comes out is familiar: a top planner, subplanners that appear as the work requires, ephemeral workers handing results back up, and one agent allowed to merge. Draw it and you get the Mermaid corporation, CEO included.

The evidence leans the same way. Cognition now calls unstructured swarms "mostly a distraction": "The practical shape is map-reduce-and-manage: a manager splits work, children execute, the manager synthesizes and reports back." Either the essay is wrong or these hierarchies are. The essay's own frame gives a more specific answer: it depends on which kind of structure each box in the chart is.

**Merge authority can be evaluation.** A top agent that assigns roles, fixes a stage order or tells workers how to implement anything is prescription. A top agent whose only power is deciding what merges is something else. A pull request is a closed trajectory, and a merge is a verdict on an outcome. If you build the hierarchy, give the top that power and delete the hand-written gates below it. The hard limits that remain belong in the envelope, like a budget cap at spawn time.

**The model can draw most of the hierarchy.** What stays hand-written is an objective every level can reread and a planner prompt with no personas. The planner allocates the budget and rebalances it during the run, while subplanners and workers appear and dissolve with the work. That's the phase transition from earlier, inside a hierarchy a few levels deep.

**Writes are where the nodes fall below threshold.** Prescription is forced where nodes are weak, and today the weakest skill in these systems is reconciling concurrent writes. Cursor's workers facing a collision "either overwrite the other change or abandon their own," so Cursor added "a neutral third-party agent" to resolve conflicts. That's where bottom-up re-centralizes: Carlini's lock files, Cognition's single writer, Cursor's planners that "make design decisions themselves rather than delegating them," a single merge authority. Each gives an implicit decision one owner, which solves Yan's Flappy Bird problem with ownership instead of shared context. Cursor's costs follow the same logic: once a frontier planner has "collapsed the ambiguity," workers can be cheap. GPT-5.5 workers cost $9,373; Composer 2.5 workers under an Opus 4.8 planner cost $411. Prescription flows down to cheap nodes; judgment stays where the capability is.

That's also what separates these hierarchies from MetaGPT's. A software company splits work by function, so the overlap between product manager, architect and engineer belongs to nobody. Cursor's planners split by region of the problem and must ensure "that no two delegated subtrees decide the same question." The overlap survives inside each owner's scope.

**A fixed structure is a stale cache.** A standing number of subteams, picked before the run, is keyed on the designer's guess about the width of the problem. It needs the expiry every prescription gets. Carlini's width was sixteen, and at the kernel it stopped helping. Instead of a manager or more agents, his fix was a verifier that split the problem: GCC compiled most of the kernel as an oracle, so each agent could chase a different failing file. The split came from the problem, read through a sensor, not from a standing structure.

**A merge without sensors is a judge.** An LLM deciding what to merge is a judge, not a sensor. In MAST the CEO-final-say fix was worth 9.4 points and the loop that ended only when review passed was worth 15.6. Cursor removed its central integrator because "there were hundreds of workers and one gate (i.e. 'red tape') that all work must pass through." A single merge authority only earns that bottleneck by judging better than tests would. By this essay's rules, the top agent keeps the merge and reads deterministic sensors (tests, CI, an oracle) before using it.

**Work has to flow back.** The part of such a hierarchy I'd defend hardest is the return path. A worker commits and submits one handoff, the harness acknowledges it durably, and the parent is woken to read it and record a judgment that survives restarts. The harness shouldn't commit on the worker's behalf, because the point is to learn whether the agent itself can close the trajectory. Underneath, the orchestrator should own nothing that git, the filesystem or the process manager already own, and should have no schedule: no rounds, batches, launch waves or staged ramps. Rounds are prescription about time. Built this way, the hierarchy is mostly the channel through which finished work returns to be judged. What those handoffs add up to is the subject of [Agent civilizations](https://future-seems-so-good.com/blog/agent-civilizations).

So I don't think the org chart refutes the thesis, if it's built this way. It's an instance of it. Bottom-up produces a thin top-down structure where the nodes are weakest and outcomes need a judge, and regenerates the rest each run. What to watch is the moment the cache stops being invalidated: the subteam count becoming permanent, the top agent starting to tell subtrees how, the merge turning into a ritual without sensors. That's where a hierarchy becomes a bureaucracy.

## Past the upper threshold

This part is speculation, and I'll mark where it stops being evidence.

Cognition says cross-agent communication "doesn't happen by default, because models haven't been trained in environments where it needed to." That's evidence. What follows isn't. In March I wrote that RL for agent swarms should use a value function to distribute the reward of the whole task to the subagents. If labs train that way, the ability to split, delegate and merge moves into the weights, the way edit formats already have (see [The harness is in the weights](https://future-seems-so-good.com/blog/the-harness-is-in-the-weights)). The org chart stops being something we draw. The model draws a temporary one per problem and erases it when the problem is solved. A planner prompt that tells a model to split its budget and redraw its own tree is what that looks like before the training catches up.

The future intelligence will alternate between rhizome and hierarchy. It will decentralize to search, centralize to act, modularize to preserve, and dissolve modules to escape. It will create an agent and absorb it, build a protocol and break it.

That changes what programming is. Syntax is getting cheap. What remains is choosing what can exist and which actions have consequences. The programmer writes the laws under which events become possible. That's the envelope. AI may kill the programmer as a writer of syntax and restore programming as the construction of worlds. When such a system starts producing more of its own productive capacity, that's acceleration, the subject of [Positive feedback eats the world](https://future-seems-so-good.com/blog/positive-feedback-eats-the-world).

## The envelope thickens, the pipeline thins

Here's the rule from the July 31 session with Claude, which I now use as a design review:

> your envelope (objective, verifier, budget, sandbox, provenance, immutable interfaces, termination conditions) should get thicker as models improve. Your pipeline (roles, stage order, mandatory critique passes, routing) should get thinner.

As a checklist for an agent system:

1. **Label every structure.** For each prompt rule, role, stage and check, write down whether it tells the system how well it did or how to act. If you can't tell, it's prescriptive.
2. **Spend on verifiers first, and ground them.** Tests, type checkers, compilers, CI, oracles like GCC, production signals. Reviewers start from the diff, not the coder's history.
3. **One writer, one owner per decision.** Let any number of agents search, review and advise. Keep writes single-threaded, and make sure no two subtrees decide the same question.
4. **Coordinate through the medium, return through a handoff.** Lock files, progress files, failing tests and commit history before messages. Every worker ends with one handoff, and the harness never commits for it.
5. **Spawn roles from the problem and let them dissolve.** Decide agent count per task and keep no standing staff. A fixed number of subteams counts as standing staff.
6. **Put an expiry on every prescription.** Each fixed stage names the failure it patches, and every model upgrade reruns the system without it. The [mini-swe-agent](https://github.com/SWE-agent/mini-swe-agent), "The 100 line AI agent" that "scores >74% on SWE-bench verified," is the baseline to beat.
7. **Prescribe failures, not org charts.** Below the capability threshold, add the narrow rule that fixes an observed failure ("always invoking the smart friend for merge conflicts"), not a department.
8. **Put time and budget in the envelope, not in rounds.** Start every independent piece of work before waiting on any. Use token budgets, time boxes (Carlini's `--fast` 1% or 10% test sample), and loops that end only when verification passes.
9. **Keep the medium clean.** Slop is stigmergic. Fix existing vulnerabilities before adding agents and run a deduplication pass. Deleting structure is a capability; see [Simplicity does not sell](https://future-seems-so-good.com/blog/simplicity-does-not-sell).
10. **Give the merge authority sensors, not one scale.** A single merge point is fine if it reads independent sensors (correctness, latency, cost, safety) and lets the situation decide which binds.

Bottom-up is one ontological level higher because structure is one of the things it can produce.

## Sources

- MetaGPT, [README: Software Company as Multi-Agent System](https://github.com/FoundationAgents/MetaGPT)
- OpenBMB, [ChatDev README](https://github.com/OpenBMB/ChatDev) (ChatDev 1.0 as a virtual software company; the 2025 evolving-orchestration work)
- CrewAI, [Introduction](https://docs.crewai.com/en/introduction)
- Walden Yan, [Don't Build Multi-Agents](https://cognition.ai/blog/dont-build-multi-agents), Cognition, 2025-06-12; Hacker News discussion [45096962](https://news.ycombinator.com/item?id=45096962), 2025-09-01, including comments [45098720](https://news.ycombinator.com/item?id=45098720) and [45097924](https://news.ycombinator.com/item?id=45097924)
- Walden Yan, [Multi-Agents: What's Actually Working](https://cognition.com/blog/multi-agents-working), Cognition, 2026-04-22
- Anthropic, [How we built our multi-agent research system](https://www.anthropic.com/engineering/built-multi-agent-research-system), 2025-06-13
- Nicholas Carlini, [Building a C compiler with a team of parallel Claudes](https://www.anthropic.com/engineering/building-c-compiler), Anthropic, 2026-02-05; Hacker News discussion [46903616](https://news.ycombinator.com/item?id=46903616)
- Wilson Lin, [Scaling long-running autonomous coding](https://cursor.com/blog/scaling-agents), Cursor, 2026-01-14
- Wilson Lin, [Towards self-driving codebases](https://cursor.com/blog/self-driving-codebases), Cursor, 2026-02-05
- Wilson Lin, [Agent swarms and the new model economics](https://cursor.com/blog/agent-swarm-model-economics), Cursor, 2026-07-20
- Cemri et al., [Why Do Multi-Agent LLM Systems Fail?](https://arxiv.org/abs/2503.13657) (v3, Figure 5, Table 5, Appendices B and H)
- SWE-agent team, [mini-swe-agent](https://github.com/SWE-agent/mini-swe-agent)
- Randall Degges, [Copilot amplifies insecure codebases by replicating vulnerabilities in your projects](https://labs.snyk.io/resources/copilot-amplifies-insecure-codebases-by-replicating-vulnerabilities/), Snyk Labs, 2024-02-22
- Fu et al., [Security Weaknesses of Copilot-Generated Code in GitHub Projects: An Empirical Study](https://arxiv.org/abs/2310.02059), TOSEM
- Ben Goertzel, [RE: Ben what are your views and concerns](http://sl4.org/archive/0010/0117.html), SL4 mailing list, 2000-10-01
- Christopher Alexander, [A City Is Not a Tree](https://www.patternlanguage.com/archive/cityisnotatree.html), 1965
- Melvin Conway, [How Do Committees Invent?](http://www.melconway.com/research/committees.html), Datamation, 1968
- Herbert Simon, [The Architecture of Complexity](https://mwolf.dev/library/architectureofcomplexity-hsimon1962/TTS), 1962
- Warren McCulloch, [A heterarchy of values determined by the topology of nervous nets](https://journal.emergentpublications.com/Article/7b136ca7-c0eb-4370-a50c-14c83b21c94e/github), 1945
- Ronald Coase, [The Nature of the Firm](https://onlinelibrary.wiley.com/doi/full/10.1111/j.1468-0335.1937.tb00002.x), Economica, 1937
- Hermann Haken, [Synergetics](http://www.scholarpedia.org/article/Synergetics), Scholarpedia
- Gilbert Simondon, [The Position of the Problem of Ontogenesis](http://www.parrhesiajournal.org/parrhesia07/parrhesia07_simondon1.pdf), trans. in Parrhesia 7, 2009
- Francis Heylighen, [Stigmergy as a Universal Coordination Mechanism](http://pespmc1.vub.ac.be/Papers/Stigmergy-Springer.pdf); Theraulaz and Bonabeau, [A Brief History of Stigmergy](https://psycnet.apa.org/doi/10.1162/106454699568700), Artificial Life, 1999
- Rich Sutton, [The Bitter Lesson](http://www.incompleteideas.net/IncIdeas/BitterLesson.html), 2019-03-13
- Stigmergy usage counts (comp.ai.alife as the heaviest mailing-list user; 12 Hacker News posts since 2020): my queries over Usenet, mailing-list and Hacker News archives via Scry, 2026-09-25
- Scaffolding formalism (evaluative vs. prescriptive axes, the Λ model, thresholds, envelope/pipeline rule): worked out with Claude on 2026-07-31, starting from an earlier document that supplied the evaluative/prescriptive naming and the "rhizome of incompetent nodes" phrase