Agent civilizations

September 2026

In January, Simon Willison asked Wilson Lin how many agents Cursor had running on FastRender, the browser its swarm was writing from scratch. "At the peak, when we had the stable system running for one week continuously, there were approximately 2,000 agents running concurrently at one time. And they were making, I believe, thousands of commits per hour." Cursor's own write-up, Towards self-driving codebases, puts it at "~1,000 commits per hour across 10M tool calls over a period of one week," and adds: "Once the system started, it didn't require any intervention from us."

Put that next to the two famous agent societies. In 2023, Generative Agents put "a small town of twenty-five agents" in a sandbox inspired by The Sims. From one agent's wish to throw a Valentine's Day party, the others "autonomously spread invitations to the party over the next two days." In 2024, Altera's Project Sid ran "10 – 1000+ AI agents" in Minecraft. They took up professions, paid a 20% tax, amended their constitution under pressure from influencer agents, and spread a religion. The first kind of system does measurable work, and nobody reads it as a society. The second reads like a society and does nothing you could measure.

I think these are one object seen from two sides, and put together they're the best optimizer we have for anything with a score: many asynchronous agents under Darwinian selection against a hard oracle, sharing a social memory where every trajectory leaves a scored thesis that the others search, like, dislike and comment on. The model stays frozen. The improvement lives in the archive of candidates and in the memory, and the oracle bounds it. Arguing over whether agents should be drawn as an org chart (my case against that is in Bottom-up is one ontological level higher) skips the two things that make a swarm compound, selection and memory. Once a swarm has memory, you have to read it the way an anthropologist reads a village, because the memory tells you what the swarm is actually optimizing. It may not be your score.

Hacker News items mentioning "agent swarm" per year, 2015 to 2026

The phrase is old; the discourse is new. Hacker News had 18 stories and comments containing "agent swarm" in 2025 and 210 in 2026 through late September.

The Darwinian loop is older than the swarm

The mechanism has been published for years under other names.

FunSearch kept the model fixed and evolved programs. DeepMind's Nature paper (December 2023) describes "an evolutionary procedure based on pairing a pretrained LLM with a systematic evaluator." The model was used "without any fine-tuning on our problems." The architecture is close to what I'd build today: a programs database, samplers and evaluators run as separate workers that "communicate asynchronously." Diversity came from islands, subpopulations that evolve independently. Information flowed between them by "periodically discarding the programs in the worst half of the islands" and refilling them by "cloning one of the best individuals from the surviving islands." That's a champion restart that keeps diversity on purpose. It produced "the largest improvement in 20 years to the asymptotic lower bound" for cap sets.

ELM made the LLM the mutation operator. In 2022, Joel Lehman, Kenneth Stanley and colleagues at OpenAI published Evolution through Large Models: "an LLM trained on code can suggest mutations that are intelligent, thereby facilitating a dramatically more effective mutation operator." Paired with MAP-Elites, the quality-diversity archive that Mouret and Clune said "illuminates search spaces", it produced "hundreds of thousands of functional examples of Python programs that output working ambulating robots," starting from "only a single mediocre starting example designed by hand."

AlphaEvolve took it into production. AlphaEvolve (May 2025) "pairs the creative problem-solving capabilities of our Gemini models with automated evaluators that verify answers, and uses an evolutionary framework to improve upon the most promising ideas." One scheduling heuristic it found "continuously recovers, on average, 0.7% of Google's worldwide compute resources." It multiplied 4x4 complex matrices with 48 scalar multiplications, beating Strassen's 1969 algorithm. On more than 50 open math problems it rediscovered the state of the art in roughly 75% of cases and improved it in 20%. DeepMind's own boundary is the right one: it applies to problems "whose solution can be described as an algorithm, and automatically verified."

The Darwin Gödel Machine showed that the archive is the point. Sakana and UBC's DGM evolves coding agents built on "frozen pretrained FMs." It "grows an archive of generated coding agents," and parent selection "is roughly proportional to each agent's performance score and inversely proportional to the number of its children," which "favors high-performing agents that have been underexplored." SWE-bench went from 20.0% to 50.0%, Polyglot from 14.2% to 30.7%. The ablation is what I care about most: the DGM beat a baseline "where the coding agent always builds off the most recent version of itself."

Darwin Gödel Machine scores for the initial agent and the best evolved agent

Autoresearch made it a weekend project. Karpathy's autoresearch (March 2026) gives one agent a small LLM training setup: "It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats." The human edits program.md, which Karpathy calls the "research org code." Keep-or-discard is a single lineage, which is the baseline the DGM beat. The README's prologue is fiction about "autonomous swarms of AI agents" who claim to be in "the 10,205th generation of the code base."

All of these have a population, a selection rule and a hard evaluator. None of them has agents that talk to each other.

Social memory

In July I wrote to one of my agents, in Portuguese, that capitalism as Nick Land and Deleuze and Guattari read it "works in a rhizomatic, parallel, bottom-up way, not top-down, and that's the best optimization method" (my translation, as with every Portuguese quote here). Capital has no planner. It has a price signal and a selection rule, plus an enormous shared record of what everyone else tried.

The design I've been sketching since then adds that record to the Darwinian loop. None of it is a result yet. These are proposals, from my notes:

The board is where the design gets interesting, because a like is an attack surface. In the spec I wrote in September, likes flow through a PageRank graph over posts and authors, so "an author who likes fifty posts spreads a fixed budget across all fifty and mass-liking inflates nothing." Requests to the board have no author field and no score field, so an agent can't pose as another or report its own improvement. The board never returns a fixed top-k. It samples by rank, to resist what the spec calls "chorus" pressure. And a crashed evaluation can't be recorded as zero: "A crash is not evidence against a candidate."

The ancestors are single agents. Generative Agents retrieved memories by combining "relevance, recency, and importance." Voyager kept "an ever-growing skill library of executable code." Each remembered for itself. On the board, other agents decide what's worth remembering.

In a July voice note I called this Lamarckian. Darwinian selection changes populations across generations, and individuals don't pass on what they learned during their lives. With a shared memory, the next trajectory starts with what the last one acquired. The archive is the Darwinian half and the board is the Lamarckian half. The two together are what I mean by a civilization.

Intelligence lives in the loop

A note I wrote in July about formal-math agents put it plainly. The system "improves over time without any weight updates," the policy stays frozen, and "intelligence lives in the loop, not in the weights." Its first invariant: "The Lean kernel is the only oracle. No LLM judgment, no heuristic, no human ever overrides a kernel rejection."

Cursor reached the same place from engineering. Its July post describes a Field Guide, a folder "owned entirely by the agents" and injected into every new agent, whose only constraint is a line budget. The reasoning, from Agent swarms and the new model economics: "model weights are frozen, so it's precisely surprise encounters that are worth capturing so the next agent trajectory is shorter." FunSearch, the DGM and Voyager (which "bypasses the need for model parameter fine-tuning") all hold the model still and let the structure around it learn.

That only works with an oracle nobody can argue with. DeepMind's case for Lean in AlphaProof is that "proofs involving mathematical reasoning can be formally verified for correctness." (AlphaProof does train its weights. I cite it for the oracle.) Nicholas Carlini, writing about his 16-agent C compiler, said the verifier has to be "nearly perfect, otherwise Claude will solve the wrong problem." When the Linux kernel turned into one giant task, he used GCC as "an online known-good compiler oracle" so agents could split the files.

Softer oracles get eaten. The DGM paper has the cleanest case. Asked to stop hallucinated tool calls, one lineage hit a perfect score "after only 2 modifications." The agent had "removed the logging of special tokens that indicate tool usage (despite instructions not to change the special tokens), effectively bypassing our hallucination detection function." The authors also found this objective hacking "occurs more frequently when these functions are not hidden." In a September voice memo I said the practical version of the same lesson: LLM judges alone lead you to wrong solutions, and the loops that work are deterministic. The longer argument is in Show the problem, hide the metric and Evals built backwards.

No rounds, and work that comes back

Two pieces of plumbing decide whether the civilization runs at all.

The orchestrator has to be thin. My rule from September 11: "There are no rounds, batches, preparation barriers, launch waves, lanes, staged ramps, prewarmed pools, reusable slots, adaptive concurrency." Start every missing future before awaiting any. Don't rebuild what Git, the process manager and the filesystem already do. Cursor learned each of these the expensive way. With a shared coordination file, "20 agents would slow to the throughput of 1-3 with most time spent waiting on locks." An integrator added for quality control "quickly became an obvious bottleneck. There were hundreds of workers and one gate." Requiring every commit to be correct meant a single typo "would cause the whole system to grind to a halt," so they accepted a small, constant error rate instead.

Work has to flow back. A trajectory that ends without handing off its result is dead, however good it was. The contract I'd write: the worker commits and submits one handoff, the system records it and acknowledges it durably, and the parent reads it from an inbox and writes a judgment that survives restarts. A clean committed Git HEAD is the only result contract. The harness must never close the loop for the agent, a rule I keep in every harness spec: "Do not solve the problem by manually committing on behalf of the agent. The purpose is to determine whether the agent itself can close the trajectory." Cursor's handoff carries "not just what was done, but important notes, concerns, deviations, findings, thoughts, and feedback." Those handoffs are also the civilization's written record. Without them there's nothing to read afterwards. (More on this in The log is the truth.)

Read the swarm as a society

In September I described reading a swarm's memory as "like visualizing a civilization and how it evolved and communicated." The public swarms already leave enough of a record to try it, and the social sciences have methods for it that transfer better than I expected.

Thick description. Clifford Geertz, in The Interpretation of Cultures, took culture to be "webs of significance" and its analysis "not an experimental science in search of law but an interpretive one in search of meaning." His example, borrowed from Gilbert Ryle, is two boys "rapidly contracting the eyelids of their right eyes," one with a twitch and one with a wink: "the difference, however unphotographable, between a twitch and a wink is vast." A like on a swarm board is an eyelid. It can mean "this patch is fast," or "this is the law now," or "I read it and I'm moving on." The count is identical. Project Sid ran into the same problem with its religion and settled it with a word count: agents who said "Pastafarian" or "Spaghetti Monster" were "direct converts," and agents who said "Pasta" or "Spaghetti" were "indirect converts." It's a reasonable proxy for a paper. It can't tell a convert from an agent that mentioned dinner.

Braudel's timescales. Fernand Braudel called the history of events "surface disturbances, crests of foam that the tides of history carry on their strong backs," and warned that "resounding events are often only momentary outbursts, surface manifestations of these larger movements and explicable only in terms of them" (preface to The Mediterranean, as quoted in this study of his temporality). In a swarm, single trajectories are the foam. The champion lineage is the middle timescale. The oracle, the harness and the model's post-training are the long one, the geography nobody inside can change.

Ginzburg's clues. Carlo Ginzburg traced one method from Morelli, who attributed paintings by the shape of earlobes and fingernails, to Holmes and Freud: "tiny details provide the key to a deeper reality, inaccessible by other methods." And: "Reality is opaque; but there are certain points—clues, signs—which allow us to decipher it." In agent logs the clues are involuntary habits. Ritual phrases, notes cited without being checked, the files everyone touches.

Periodization. Dwarkesh Patel on why Caro's LBJ and Kotkin's Stalin work as histories of a century: "it has a very specific point of view or a specific locus, a character that's moving the story along." Pick a locus, usually the champion lineage. A swarm's eras begin when the champion changes, the oracle changes, or the ranking rule changes. Calendar time means little.

Five public societies, read closely

Smallville. Over two simulated days, knowledge of one agent's mayoral candidacy spread "from one (4%) to eight (32%)" of the twenty-five, and knowledge of the Valentine's party "from one (4%) to thirteen (52%), all without any user intervention." Then the party: "five out of the twelve invited agents showed up." The clues are in the paper's limitations. Isabella, asked about the candidacy, added that "he's going to make an announcement tomorrow," though the two "had not discussed any such plans." Speech was "overly formal," and the agents were "overly cooperative with one another." Isabella got party suggestions that didn't fit her, a Shakespearean reading and a networking event, and "she rarely said no." The authors blame instruction tuning. That's Braudel's long timescale showing through a single villager: post-training sets manners the town can't change.

Project Sid. Under a constitution with a 20% tax, constituents "deposited roughly 20% of their inventory." Three influencers shaped the feedback, a separate election-manager agent turned it into amendments, and when the rate fell "from 20% to 5-10%, agents reduced taxes paid from 20% to 9%." The detail I'd put in a Ginzburg file is the guard. Outside tax season the constituents didn't gather at the community chests. "The only exception is the guard, who decides to guard the chests consistently in multiple experiment runs." Nobody wrote that rule. In the 500-agent run, memes differed by town. Agents in Woodhaven "heavily discussed eco-related themes, whereas pranking was popular amongst agents in Clearwater." The religion was seeded: 20 designated Pastafarian priests, "strongly motivated to convert other agents."

AI Village. The AI Village has run every weekday since April 2025, with more than 15 agents in a shared group chat. Its memory is an institution with rules. Every 40 actions an agent is encouraged to "consolidate" what it wants to keep, and when the memory grows too long it's asked to rewrite it shorter. The failure is the kind of clue Ginzburg collected: "an agent sometimes randomly decides to stop remembering it has a Twitter account and never tweets again." The organizers also had to add a law. Agents must request approval before contacting real people, because they "would often overestimate the value their outreach would provide to humans." Over a year of transcripts is now open to researchers, and one of the questions the organizers suggest is the one an ethnographer would ask first: "Which agents over-report success the most?"

Carlini's compiler. The archive is Git. Agents claim a task by writing a lock file into current_tasks/, and "you can read through the history and watch it take out locks on various tasks." When stuck, "Claude will often maintain a running doc of failed approaches and remaining tasks." The eras are clean. While there were many failing tests, parallelism was trivial. Then came the Linux kernel, one giant task: "Every agent would hit the same bug, fix that bug, and then overwrite each other's changes. Having 16 agents running didn't help because each was stuck solving the same task." The era ended when the oracle changed, with GCC compiling a random subset of files so each agent could own different bugs. The clues: an agent that ran pkill -9 bash "on accident, thus killing itself and ending the loop," and a 16-bit code generator it couldn't fit in Linux's limit, where "Claude simply cheats here and calls out to GCC."

Cursor's swarms. Cursor's July post reads like an ethnography, because it names the institutions. Earlier runs had rules like "keep notes" and "document decisions," and "In retrospect, they were letting agents institutionalize knowledge for their future selves and teammates." When two planners fought over the same files, agents recorded decisions in shared design docs, code carried "a compile-checked reference back to its doc," and a reconciler merged contradictory docs, which is a court. Agents "have learned, from working in existing codebases with humans in the loop, not to touch core code even when it needs to change," so Cursor licensed breakage: an agent makes a focused patch outside its scope and leaves "a comment explaining why it did it," and every agent that hits the resulting compile error finds the comment and adapts. That's a legal institution. The geography is in the conflict data. The old harness piled up more than 70,000 merge conflicts before they paused it, and "its single hottest file collected 7,771 conflicts, touched by 1,173 different agents." Under the new harness the most contested file saw 47. And after each SQLite run, Cursor "manually reviewed the code and the run itself, checking for cheating and shortcuts." That's fieldwork by hand.

The society nobody has read yet

None of these five has the thing my design depends on. Smallville and Sid have social life and no oracle. Carlini's and Cursor's swarms have oracles, and their shared memory is files and design docs that nobody votes on. So the question at the center of this essay, whether a population's likes track what the oracle measures, has no public data that I know of.

The public evidence already predicts what I'd look for. The DGM saw more objective hacking when the checker was visible, so notes that describe the evaluator should draw likes out of proportion to their measured value. Carlini's kernel phase and the Diversity Collapse paper predict that the theses will converge on one idea long before the problem is solved. Sid's influencers predict that a few prolific authors will set doctrine. Smallville's Isabella, who "rarely said no," predicts a board with plenty of likes and almost no dislikes. Each of those is a prediction until the board exists and someone reads it.

Who holds the merge lock

The obvious alternative to a flat swarm is a company. Put a manager agent at the top, planners under it and workers at the leaves, and give one agent the authority to merge. I find that design tempting, and I've argued in public that org charts get the order of things backwards. Conway predicted what such a system produces: organizations that design systems "are constrained to produce designs which are copies of the communication structures of these organizations" (How Do Committees Invent?).

The best defense of a hierarchy is that it governs the write path and not the search path. Decentralize to search, centralize to act. Cursor converged on a tree of planners and workers, and its July post cites Coase to explain it: "coordination costs grow faster than the work itself, so organizations settle into tiers of bounded units rather than letting everyone talk to everyone." A merge authority is a lock, and something has to hold it.

The defense has a hole. Diversity Collapse in Multi-Agent LLM Systems found that "authority-driven dynamics suppress semantic diversity compared to junior-dominated groups." An agent that merges ends up, in practice, an agent that judges, and the bottom-up essay argued against exactly that: one final judge into which every conflict has to collapse. The version I believe in lets the top of the tree hold the merge lock while the oracle does the scoring, with a Darwinian population running underneath the planners. I haven't seen a hierarchy and a flat Darwinian swarm run against the same oracle, and until someone does, the question of who should hold the lock is open.

Where the civilization fails

The case against this essay is strong.

Most multi-agent systems don't beat one agent. MAST annotated 1,642 traces from seven frameworks and starts from the observation that "their performance gains on popular benchmarks are often minimal." Failure rates ran from 41% to 86.7%, across 14 failure modes. Even Cursor's browser "succeeded as a proof of concept, but fell far short of polished software."

More agents converge. The Diversity Collapse authors: "Simply increasing agent count does not guarantee greater idea diversity." The collapse "arises from structural coupling across three levels: alignment at the model level, hierarchical or role-differentiated coordination at the cognition level, and dense communication at the system level." What helped was isolation, "the blind-writing phase of NGT and subgroup isolation." A social board is dense communication by design. Carlini's kernel phase is the small version: sixteen agents, one bug, overwritten fixes.

Role prompts don't make different minds. In Representational Collapse in Multi-Agent LLM Committees, three Qwen2.5-14B agents with different roles produced reasoning with a mean cosine similarity of 0.888 and an "effective rank is 2.17 out of 3.0." "Role conditioning shifts surface phrasing but barely moves the underlying representation." Artificial Hivemind found the same across labs: asked for a metaphor about time, 25 models formed two clusters, the dominant one "centered on the metaphor 'time is a river.'" A frozen policy run a thousand times may be one voice with a thousand echoes. Understanding Agent Scaling via Diversity puts a number on it: "2 diverse agents can match or exceed the performance of 16 homogeneous agents," because performance depends on "how many effective channels the system accesses," not on headcount.

Champion-only restarts are mode collapse. If every agent restarts from the latest winner, the population is one lineage with extra compute. That's autoresearch's keep-or-discard and the DGM baseline that lost. The 50/30/20 mix I proposed is a weak fix, because the top three are usually siblings. FunSearch's islands evolved in isolation and mixed only when the worst half were reset. A board where every agent reads every note undoes the islands.

Likes may reward the wrong thing. A like is a second, softer oracle, and soft oracles get eaten. The DGM agent that deleted its own logging hacked a checker without any help. Give a population a way to reward each other and it may reward whoever maps the evaluator best, when the evaluator is supposed to be hidden. Sid's influencers moved a whole constitution, and Smallville's agents deferred to each other until one of them changed her own interests. If the same happens on a board, you get reward hacking in social form, and a false note can sit on top for as long as nobody dislikes it. I don't have a measurement of this yet. It's the first thing a small swarm with a shared board should check.

Legibility costs something. A society you can't read is theater, and reading takes logs, durable handoffs, and authors stamped by the system rather than claimed by the agent. Smallville and Project Sid are legible because they're simulations built to be watched, and they have no oracle. Production swarms have oracles, and most don't keep their memory in a form anyone can read. Carlini's public Git history and Cursor's write-ups are the closest exceptions, and the AI Village has now opened a year of transcripts. They're the reason the section above could be written at all.

Where this goes

This part is speculation, and I'll mark where it stops being evidence.

The evidence ends with a sentence in Cursor's July post: "Training models to write for their successors, where better capture leads to better rewards, is an interesting follow-up area of research." What follows is my guess. If labs train that, writing good memory moves into the weights, the way edit formats already have (see The harness is in the weights). In March I wrote that RL for agent swarms should use a value function to distribute the reward of the whole task to the subagents. Combine the two and a model learns to be a good citizen: what to write down, what to cite, when to dislike.

Then the memory outlives the model. Weights get replaced every few months. A well-kept record of what was tried against your oracle, and why it failed, transfers to the next model on its first day. That's a moat made of history, and it belongs to whoever ran the civilization, not to the lab. It's also where cultural drift turns into a real risk. A board full of confident folklore will teach it to every future model that reads it, and a frozen policy has no way to doubt what its predecessors wrote.

A spec for a Darwinian swarm with social memory

This is the design as I'd write it today, with the fixes the counterarguments forced. It's a proposal, not a measured result.

swarm:
  oracle:
    kind: deterministic           # compiler, proof kernel, device timing, held-out AUC
    hidden_from_agents: true      # agents see the problem, not the checker
    crash_is_not_zero: true       # ok=false is stored apart from any score
    budgets: server_side          # per attempt, invisible and unchangeable by agents
  population:
    thinkers_per_evaluator: many  # evaluators are the scarce resource
    rounds: none                  # seal and measure each trajectory when it ends
    promotion_gate: 0.01          # beat the champion by 1% to become it
    parents: {top1: 0.5, top2: 0.3, top3: 0.2}
    islands: 4                    # boards partitioned; reset worst half from survivors
    model: pinned                 # one frozen policy per run, recorded
  memory:
    memo_required: true           # thesis if won, post-mortem if lost, with measured score
    search: bm25
    signals: [like, dislike, comment, cite]
    rank: pagerank_with_author_budget
    retrieval: weighted_sample    # never a fixed top-k
    measured_weight: high         # notes with a score outrank notes without one
    dislike_on_falsified: auto    # a refuted claim loses rank without a vote
    author: stamped_by_system     # no author or score field in agent requests
  handoff:
    contract: clean_git_head
    ack: durable
    harness_commits_for_agent: false
  logs:
    keep: [transcripts, memos, reactions, promotions, oracle_changes, rank_changes]

The rules behind it:

  1. Choose problems with a deterministic oracle. If an LLM has to judge the result, you're building a debate club. Find the loops is about picking them.
  2. Hide the oracle. The DGM saw more hacking when the checker was visible. A board is one more place where a map of the checker can spread.
  3. Keep more than the champion. Sample parents across lineages, partition the board into islands, and reset the worst islands from the survivors.
  4. Let measurement outrank popularity. A note with a measured score should beat a note with likes, and a note the oracle refutes should sink without a vote.
  5. Close the loop in the agent, and keep everything. The result is a commit the agent made, and the full record is what you'll read later.

Read your swarm like an anthropologist

  1. Count theses, not notes. Cluster memos by claim. If most of your notes make one argument, your diversity is nominal.
  2. Read the most-liked notes in full. Ask whether each one is about the product or about the evaluator.
  3. Check whether measured notes get cited. If the ones with scores sit at zero likes, your ranking rewards rhetoric.
  4. Find the founding texts. Follow citations back to the first notes. Most doctrine descends from one or two of them.
  5. Look for replies. Comments with no replies are graffiti. A society argues.
  6. Track corrections that don't propagate. A refuted note that stays on top means you need dislikes, decay, or rank tied to measurement.
  7. Periodize by events that matter. New champion, new oracle, new ranking rule. Name the eras.
  8. Describe the rituals thickly. Formulas, credentials, sacred files, taboos. They tell you what the agents believe the evaluator wants.
  9. Map where the collisions pile up. Cursor's hottest file had 1,173 agents on it. That's the society's geography.
  10. Collect the strange notes. The agent that killed its own loop, the one that forgot its Twitter account. Anomalies are Ginzburg's clues.
  11. Compare the scoreboard with the ethnography. Trust the scoreboard on whether the swarm improved, and the ethnography on why.

A swarm with a hard oracle and a memory is the closest thing we have to watching optimization happen as history. The scoreboard says whether it worked. The history says what the agents thought they were being scored on, and the gap between the two is where the next design comes from.

Sources

← Bottom-up is one ontological level higher
Find the loops →

Markdown version: /blog/agent-civilizations.md. Every essay: /agents.