# Find the loops

*September 2026*

In the second game of the March 2016 match in Seoul, AlphaGo played a move that nobody in the room understood. One commentator called it "a very strange move." The other said "I thought it was a mistake." Lee Sedol left the match room and needed nearly fifteen minutes to answer. Fan Hui, the European champion AlphaGo had beaten five months earlier, said "It's not a human move. I've never seen a human play this move," and then, "So beautiful. So beautiful."

After the game David Silver went back to the control room to see what the machine had computed. As Cade Metz reported in [WIRED](https://www.wired.com/2016/03/googles-ai-viewed-move-no-human-understand/), AlphaGo's network trained on human games estimates how likely a human is to play each move. "For Move 37, the probability was one in ten thousand." AlphaGo knew no professional would play it. It played it anyway, because the part of it trained on millions of moves from games against itself judged that the move would likely succeed. "It discovered this for itself," Silver said, "through its own process of introspection and analysis." DeepMind's [AlphaGo page](https://deepmind.google/research/alphago/) keeps Lee's verdict: "I thought AlphaGo was based on probability calculation and that it was merely a machine. But when I saw this move, I changed my mind. Surely, AlphaGo is creative."

Nine years later the same kind of search went to work on Google's own infrastructure. [AlphaEvolve](https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/) pairs Gemini with "automated evaluators that verify answers." It found a scheduling heuristic for Borg that has been in production for over a year and "continuously recovers, on average, 0.7% of Google's worldwide compute resources." It sped up a matrix multiplication kernel in Gemini by 23%, which cut Gemini's training time by 1%. It got "up to a 32.5% speedup" on FlashAttention, in low-level GPU code that "is usually already heavily optimized by compilers, so human engineers typically don't modify it directly." DeepMind calls the Borg heuristic "remarkably simple" and human-readable. It evolved out of the one Google's engineers already had in production.

In March I tweeted "alphago 37 move on everything." I still mean it literally. What I've learned since is that the search is the easy part.

Most companies read the agent moment as volume: give thousands of agents access to the repositories and let them write. That produces slop, because agents on short horizons with no feedback write plausible code nobody needed (*How to achieve superintelligence*). Move 37 and the Borg heuristic came from the opposite setup: a search inside a closed loop, judged by an evaluation it couldn't argue with. My thesis is that the real opportunity for a company is to put swarms of agents inside closed cybernetic loops with deterministic evals that are hard to hack and tied to long-term free cash flow. Any company past a certain size is already hundreds or thousands of loops converging on free cash flow, so it can be treated as a discrete optimization problem. How to run one loop well is the subject of [Show the problem, hide the metric](https://future-seems-so-good.com/blog/show-the-problem-hide-the-metric). This essay is about the step before it: choosing which loops to close. I think that choice needs a loop of its own, and that building it is the central problem.

## Loops, not slop

In August I wrote a one-paragraph manifesto about this. Its first sentence is still the one I'd keep: "the real opportunity is not to unleash thousands of agents to generate slop, but to place swarms of agents inside closed, cybernetic loops with the right tools, harnesses, guardrails, and exceptionally well-crafted deterministic evaluations—metrics that are difficult to reward-hack, easy to measure, and closely aligned with long-term free cash flow."

What that looks like depends on the loop. In a GPU kernel, the sensor is runtime on fixed inputs and the guard is a correctness suite the agent can't edit, because a swarm rewarded only for speed will find ways to skip work. In a database's query planner, the sensor is latency on a fixed benchmark workload and the guard is that every query returns the same rows. In a customer-facing agent, the sensor has to be behavioral: the agent judged on realistic sessions end to end, not on unit tests of its parts. In each case the first version of the reward has a hole, and finding the hole before the swarm does is most of the work.

The human work moves. From the manifesto: "Developers increasingly become designers of environments, incentives, and 'taste encoded as evals,' giving models a precise smell for what good means and the freedom to discover solutions humans would never conceive." The engineer writes the sensor, the red lines and the promotion rule. The swarm writes the code.

Stafford Beer saw both the opportunity and the usual mistake sixty years ago. In *Decision and Control* (1966) he wrote, [per Wikiquote](https://en.wikiquote.org/wiki/Stafford_Beer), "If cybernetics is the science of control, management is the profession of control." And later in the same book: "We have, over the centuries, devised a management structure for running things, whether firms or whole countries. This structure depends absolutely on the limitations of the human hand, eye, and brain... Yet we insist on retaining the original structures and automating them. In so doing, we enshrine in steel, glass, and semiconductors those very limitations of hand, eye, and brain that the computer was invented precisely to transcend."

Most AI plans I see automate the org chart. They give each department a copilot and each team an agent. That's Beer's mistake with a new substrate. (The longer argument against multi-agent org charts is in [Bottom-up is one ontological level higher](https://future-seems-so-good.com/blog/bottom-up-is-one-ontological-level-higher).) Starting from the loops means starting from the feedback structure of the business: where a signal comes back, how fast, and how much money rides on it. The departments come second.

## Closed loops in production

The public record of loops that found something people hadn't is short, and it has a consistent shape.

**Cooling a data centre.** In 2016 DeepMind and Google's data centre team trained neural networks on "historical data that had already been collected by thousands of sensors within the data centre," with average future PUE (total building energy over IT energy) as the target, according to [Evans and Gao](https://deepmind.google/discover/blog/deepmind-ai-reduces-google-data-centre-cooling-bill-by-40/). The detail that matters is the next sentence. They trained two more ensembles to predict temperature and pressure over the next hour, "to ensure that we do not go beyond any operating constraints." On a live site the system "was able to consistently achieve a 40 percent reduction in the amount of energy used for cooling, which equates to a 15 percent reduction in overall PUE overhead," and "produced the lowest PUE the site had ever seen." These were data centres Google already called "sophisticated," run by people who had spent a decade on efficiency. One model optimized and two models held the red lines.

**Scheduling Borg.** The [AlphaEvolve paper](https://arxiv.org/abs/2506.13131) describes the verifier behind the 0.7%. "We use a simulator of our data centers to provide feedback to AlphaEvolve based on historical snapshots of workloads and capacity across Google's fleet. We measure the performance of AlphaEvolve's heuristic function on an unseen test dataset of recent workloads and capacity to ensure generalization." Only after it beat the production heuristic on held-out data did they roll it out, and "post-deployment measurements across Google's fleet confirmed the simulator results." The sensor was a simulator, the verifier was held-out workloads, and production had the final say. For the Gemini kernels, the paper says the loop cut optimization time "from several months of dedicated engineering effort to just days of automated experimentation."

**Ranking ads.** Ad click prediction is one of the oldest large commercial loops on the internet. In 2014 Facebook's team described a system serving "over 750 million daily active users and over 1 million active advertisers" ([He et al.](https://research.facebook.com/publications/practical-lessons-from-predicting-clicks-on-ads-at-facebook/)). Their headline gain was a model that outperformed its components "by over 3%, an improvement with significant impact to the overall system performance." Two lines from the abstract hold for every loop I know of. "The most important thing is to have the right features." And: "even small improvements are important at scale." A 3% gain was worth a paper because the loop ran over hundreds of millions of people every day.

**Autoresearch.** Karpathy's [autoresearch](https://github.com/karpathy/autoresearch) is the smallest published version I know of. An agent edits one file, `train.py`, trains for "a fixed 5-minute time budget," and is scored on validation bits per byte, "approx 12 experiments/hour and approx 100 experiments while you sleep." The evaluation code in `prepare.py` is "Not modified." What interests me most is where the human went. "You are programming the `program.md` Markdown files that provide context to the AI agents and set up your autonomous research org." The human now writes the environment and the agent writes the code.

![Reported improvements from closed-loop search in production at Google](https://future-seems-so-good.com/blog/assets/find-the-loops/charts/google-closed-loops.svg)

All four share three properties. The sensor reads an outcome (energy, stranded compute, clicks, loss) and not an activity. The constraints are enforced outside the optimizer, by separate models, held-out data or a frozen file. And each reading comes back in minutes or hours, not quarters.

They share a fourth property that is easy to miss. In every case a person chose the problem. Somebody at Google knew cooling was a large line item with thousands of sensors already logging. Somebody knew stranded resources on Borg were worth fighting for. The loop searched. People picked where it searched. At a company with thousands of candidate loops, picking is the bottleneck.

## A company is thousands of loops

Take any company of a few hundred people. An online retailer is a catalog, a pricing engine, a warehouse, an app, a support operation and an acquisition machine at once. Each of those is a loop with a sensor, a controller and an actuator. Conversion on a product page, delivery time, a support agent's resolution rate and the return rate all feed back into what the company does next. And all of them converge on one number: free cash flow.

Beer built a theory of organizations on this. His [viable system model](https://en.m.wikipedia.org/wiki/Viable_system_model), from *Brain of the Firm* (1972), is recursive. As Wikipedia summarizes it, "viable systems contain viable systems that can be modeled using an identical cybernetic description as the higher (and lower) level systems." Each operating unit is itself a viable system, with its own regulation, and signals that the local loop can't handle escalate as "algedonic alerts," alarms and rewards that climb the levels "when actual performance fails or exceeds capability." Under all of it sits W. Ross Ashby's law of requisite variety, which Ashby compressed to ["only variety can destroy variety"](https://en.m.wikipedia.org/wiki/Variety_(cybernetics)) and Beer restated as "Variety absorbs variety." A regulator has to have at least as many distinguishable responses as the disturbances it is trying to absorb.

That law is why I think agents change the economics of a company and copilots don't. A company of a few hundred people serves millions of customers with different phones, networks, budgets and moods. It can't match that variety, so it attenuates. It builds dashboards, segments and averages, and it acts on the average. Beer's example of what averages miss, in *Designing Freedom*, is a hospital patient with a fever. In Wikipedia's paraphrase, "no amount of variety recording the patients' average temperature would detect this small signal." A swarm on a loop is variety on the regulator's side. It can try a thousand variants of a checkout flow, each against a replay of real sessions, where a team can try three.

At the highest level, then, the company is a discrete optimization problem. The manifesto put it this way: "among all possible actions, which ones maximize long-term free cash flow, and what 'AlphaGo Move 37' strategies might emerge from the system that no human would independently imagine?" I don't expect the answer to look like a strategy deck. I expect it to look like the Borg heuristic: simple, legible after the fact, and something nobody would have written. The same paragraph has a sentence I'd still defend: "the loops we choose today may therefore matter more than many decisions once reserved for kings."

## The meta-loop

I haven't seen anyone publish how they choose loops, so here is my proposal. It is a design, not a report on a system, and I've written it so it doesn't depend on whose data it reads.

The first version of the idea came in April: map a company's ontology and look for the parts of the system that are naturally suited to an agent improving them in a loop. By August it had become a loop in its own right. In my notes from August 14, translated from Portuguese: "the main goal is to find a meta-loop, that is, to find problems that look a lot like they can be solved by loops."

A loop worth closing has five properties. It may cross repositories and teams, because the best problems rarely respect the org chart. It has headroom, meaning nobody has squeezed it yet. It has large latent economic value. Its scope is narrow enough for an agent to hold. And it admits a fast, deterministic feedback cycle. Most candidates fail on one of these. A problem with a huge prize and a quarterly signal is a strategy question. A problem with a fast deterministic signal and a small prize is a hobby.

The meta-loop I'd build has three roles.

**Generators.** Explorer agents run long trajectories over everything a company knows about itself: code across repositories, product analytics, logs, customer journeys, the data warehouse, support tickets, and notes other agents have left. Their job is to propose theses shaped like loops, each with a problem statement, a candidate sensor, an actuator and red lines. They should understand code, products and customer behavior.

**An independent verifier.** A separate agent, with no stake in any thesis, tests each proposal by writing and running its own queries, and returns structured evidence. The explorer never computes its own numbers. Every number in the evidence carries the query that produced it, so a person can rerun it.

**A deterministic reward.** Code, not a model, turns the evidence into a score. The best theses become the starting point for the next round of explorers, the same Darwinian promotion the inner loops use.

The scoring should keep apart three quantities that usually get blended. The *prize* is how much money a year the loop could be worth if it were solved. It does not subtract implementation cost, experimentation cost or expected risk. *Headroom* is how much of the prize is still on the table: zero if the loop has already been squeezed, one if nobody has worked on it. *Realized value* is what a swarm has actually demonstrated after it ran. The reward is prize times headroom times *closability*, a term for how cleanly the loop's evaluation can be closed and hammered with parallel compute.

I leave cost out of the prize because cost is exactly what compute is making cheap. Subtract implementation cost and the ranking favors small, easy prizes, which are the ones a human team would have done anyway. Implementation cost also changes every month as models improve, while the size of the prize doesn't. The feasibility question belongs in the closability term, where the sensor's properties measure it instead of a guess.

I multiply by headroom because a loop that has already been squeezed has little left in it. The best candidates tend to be unglamorous: a slow path every customer touches and no team owns. Headroom is also the defense against the most common failure of opportunity-sizing, which is finding the biggest line item in the company and calling it an opportunity.

I keep realized value out of the reward because blending confidence into magnitude lets a weak causal story inflate a number. Whoever can measure a term should measure it, and the terms stay apart so a person can read which one is doing the work.

The explorers also never see the reward function. That's the same rule as in the inner loops: show the problem, hide the metric. An explorer that knows headroom is scored will write theses that sound fresh.

## Soul-driven search

The meta-loop can't be only data-driven, because a company's data is mostly a record of the customers who stayed.

A table of current customers is a list of survivors. Survey the people using a product and you learn what the people who tolerated it like. Rank problems by how many current customers each one touches and you'll underrate exactly the problems that drove people away, since those people are no longer in the table. A purely data-driven explorer reading that table will conclude the problem is small. The conclusion is backwards when the problem is what made the population small.

The textbook version of this is Abraham Wald and the bombers. The popular story has him telling the military to armor where the returning planes weren't hit, because the planes hit there never came back. Bill Casselman's [column for the AMS](https://www.ams.org/publicoutreach/feature-column/fc-2016-06) is a good corrective. Most of that story is "plausible reconstruction." The memoranda Wald actually wrote "are severely technical," and in them "Wald says nothing about what the military should do to improve things." What Wald did was harder than the legend. He estimated vulnerability from damage to the survivors alone, by modeling the planes that didn't return. That is what an explorer has to do with a company's data: model who is missing from the table.

There's a nice symmetry with Move 37. AlphaGo's network trained on human games said one in ten thousand. That network is the data-driven explorer. It tells you what the company, or the professional, would do. What found the move was the part that evaluated outcomes. An explorer that ranks opportunities by what the data shows most often will propose what the company would already have done. I want it to find the moves the data calls unlikely.

So the manifesto asks for search that is "soul-driven, preserving human taste about what feels fast, intelligent, precise, elegant, and genuinely valuable in the reality of the user." In practice that means two things. The explorer prompts carry taste, written by people who use the product: what a fast app feels like, what an insulting offer looks like, what a customer would never want to see. And the verifier's evidence has a required field for the population it covers and the population it doesn't. Users remain what I called "a kind of black hole whose desires we can never model perfectly." The best we can do is admit where the data stops.

## Where the loops break

The objections are serious, and one of them applies to the meta-loop itself.

**Goodhart applies to every sensor.** Goodhart's original 1975 form, [quoted on Wikipedia](https://en.m.wikipedia.org/wiki/Goodhart%27s_law), is sharper than the popular one: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." A swarm is a lot of pressure. Latency is a good proxy for customer happiness until a thousand agents find ways to be fast that users don't value. The inner-loop defenses (hidden verifiers, red lines, rejected hacks read as sensor bugs) are in [Show the problem, hide the metric](https://future-seems-so-good.com/blog/show-the-problem-hide-the-metric). They reduce the collapse. They don't abolish it.

**The meta-loop has its own Goodhart.** Scoring prize times headroom invites explorers to write big, fresh-sounding theses. Leaving cost out of the prize makes that worse, because a grand thesis is never penalized for being expensive. Hiding the reward and routing every number through the verifier helps. But the verifier's queries are choices too. It picks the table, the window and the population, and a verifier that always uses the same cohort will systematically overrate some kinds of problems. Headroom is the hardest term to measure honestly, since "nobody has worked on this" is often a fact about who wrote the logs. I don't have a clean answer. I'd make the verifier's queries a first-class artifact, audit a sample of them by hand, and treat the meta-loop's own track record as its sensor: did the theses it ranked high produce realized value once a swarm ran on them?

**A loop tied to cash flow can harm the other side of the market.** The best recent evidence is an audit of Uber's pricing in the UK. Reuben Binns and colleagues at Oxford analyzed [1.5 million trips from 258 drivers](https://arxiv.org/abs/2506.15278), obtained through data subject access requests, across the introduction of dynamic pricing in 2023. On the Employment Tribunal's definition of working time, average pay per hour fell from £22.20 to £19.06 in real terms, the year before against the year after. On Uber's own narrower definition it fell from £37.01 to £35.91. Uber had advertised a 25% cut. After dynamic pricing the median driver kept 71% of the fare, and on some trips Uber's cut was as high as 50%. The authors' summary: "Post-dynamic pricing, Uber's passengers now pay higher prices, but drivers are not better off."

![Uber driver pay per hour before and after dynamic pricing, from the Oxford audit](https://future-seems-so-good.com/blog/assets/find-the-loops/charts/uber-dynamic-pricing-pay.svg)

That's a closed loop doing what it was built for. It is also the failure mode I worry about most. Every business with two sides has someone in the driver's position: the seller on a marketplace, the merchant on a payments network, the borrower at a lender, the worker on a platform. A loop that raises cash flow by quietly worsening their deal will look good on the company's sensor for a while. Beer's line covers it: ["The purpose of a system is what it does."](https://en.wikiquote.org/wiki/Stafford_Beer) The defense is structural. The other side's outcomes (their earnings, their fees, their time) go in the red lines of any loop that touches pricing, and loops that could move them stay in human hands.

**Cybersyn ran into politics.** Beer's most ambitious application of the viable system model was [Project Cybersyn](https://en.m.wikipedia.org/wiki/Project_Cybersyn) in Allende's Chile, from 1971 to 1973: a network of telex machines in nationalized factories, statistical software to flag indicators outside acceptable ranges, an economic simulator and an operations room. It had one real success. During the October 1972 truck strike, according to a CORFO official quoted on Wikipedia, the telex network helped the government move essential goods with only about 200 trucks. The telex network was also the only component the government used regularly. The project ended with the coup of 11 September 1973, and the military destroyed the operations room. Whatever the sensors could do, what decided the outcome was a level above them, what Beer called System 5, the part that sets policy and balances the whole. Inside a company the same thing is true at a smaller scale. Who chooses which loops get compute is a political decision, and a meta-loop doesn't remove the politics. It makes the choice legible, which is not the same as settling it.

**Autoresearch is not production.** Karpathy's loop optimizes one metric on one GPU with one editable file, and the README is honest that results "become not comparable to other people running on other compute platforms." Production sensors are noisier, users change under you, and a regulator can end the game. What lets a loop survive the move is whatever sits between the swarm and production: for AlphaEvolve, a simulator, held-out workloads and post-deployment measurement. Even AlphaEvolve's 0.7% is a Google-scale number. At a smaller company the same heuristic might not repay the engineers who set up the simulator.

**Some loops should stay open.** Anything that touches a regulatory license, or anything else where one bad action has an unbounded downside and a small upside, stays with people. That's the convexity argument from [Show the problem, hide the metric](https://future-seems-so-good.com/blog/show-the-problem-hide-the-metric), and the meta-loop should score those theses at zero no matter how large the prize looks.

My answer to all of this is the same design principle applied one level up. The meta-loop is a loop, so it gets a sensor (did its top theses produce realized value?), red lines (the other side's outcomes, regulatory exposure) and a person who reads the champions, not just their scores.

## Where this goes

This part is speculation.

The progression I expect is the one the manifesto sketched: from AI as a coding copilot, to self-directed autoresearch on loops people chose, to models that create and improve their own loops. The last step is the meta-loop running without us: explorers finding the problem, the verifier building the evidence, a generated sensor passing its own mutation tests, and a swarm assigned before anyone reads the thesis. If that works, the company becomes what I called "a rhizomatic hive mind in which physical and digital workers build knowledge bottom-up." The memory side of that is in [The log is the truth](https://future-seems-so-good.com/blog/the-log-is-the-truth). What kind of people it takes is in [Hire people who close loops](https://future-seems-so-good.com/blog/hire-people-who-close-loops), and why compute is not the constraint is in [Compute is not the bottleneck](https://future-seems-so-good.com/blog/compute-is-not-the-bottleneck).

I don't know how many years that takes. I do think the gap it opens compounds. A company whose loops improve while its people sleep pulls away from one that is still assigning copilots. The ambition is a system that finds what the manifesto called "the nonhuman strategies" for the company it runs in. Move 37 was one stone on one board. The bet is that every company has many of them, and that the machine that finds them can be built.

## A loop-scoring spec

This is how I'd write the contract between the three roles, for any company. Explorers write a `Thesis`. Only the verifier writes `Evidence`. The reward is a pure function of both.

```ts
type Thesis = {
  id: string
  problem: string                    // what swarms will see; never the reward
  scope: { repos: string[]; teams: string[] }
  sensor: {
    measure: string                  // "p95 latency on a fixed TPC-H workload; rows must match"
    direction: "lower" | "higher"
    deterministic: boolean           // same artifact, same reading
    secondsPerReading: number
    parallelReadings: number         // how many can run at once
  }
  actuator: string[]                 // what agents may edit
  redLines: string[]                 // what they may never touch
  regulatoryExposure: boolean        // licenses, filings, anything a regulator reads
  touchesOtherSide: boolean          // prices, fees, pay or limits for customers, sellers, workers
}

type Measured = { value: number; queries: string[] }  // every number carries its queries

type Evidence = {                    // written only by the verifier
  thesisId: string
  prize: Measured                    // money per year; not net of any cost
  headroom: Measured                 // 0 = already squeezed, 1 = never worked on
  realizedValue: Measured            // money per year shown so far; signed, never clipped
  population: { covered: string; missing: string }
}

const readingsPerHour = (s: Thesis["sensor"]) =>
  (s.parallelReadings * 3600) / s.secondsPerReading

const closability = (s: Thesis["sensor"]) =>
  s.deterministic ? Math.min(1, Math.log10(Math.max(1, readingsPerHour(s))) / 3) : 0

export const reward = (t: Thesis, e: Evidence) => {
  if (t.regulatoryExposure || t.touchesOtherSide) return 0
  if (t.redLines.length === 0) return 0
  if (e.prize.queries.length === 0 || e.headroom.queries.length === 0) return 0
  return e.prize.value * e.headroom.value * closability(t.sensor)
}
```

`realizedValue` is reported next to the reward and never enters it. It is the meta-loop's own sensor: once a swarm has run on a thesis, the gap between its reward rank and its realized value tells you how well the meta-loop is scoring.

`closability` is my choice of shape and the thresholds are arbitrary. A thousand deterministic readings an hour scores 1. Autoresearch's twelve an hour on one GPU scores about 0.36. A non-deterministic sensor, such as an LLM judge, scores 0. Closability comes from measurable properties of the sensor, not from anyone's opinion of how hard the problem feels.

## A loop-candidate checklist

Before a thesis enters the meta-loop, a person should be able to answer yes to each of these.

1. **The prize is sized in money.** Money per year, from a query someone can rerun, with no cost subtracted.
2. **Someone checked the missing population.** The evidence says who is absent from the data, and why. Churned customers, rejected applicants and failed sessions count.
3. **It has headroom for a stated reason.** "Nobody has run a swarm on this" should be a fact about the repository history and the roadmap, not a feeling.
4. **The sensor reads an outcome.** Latency, memory, losses, conversion. Not PRs merged, tickets closed or lines changed.
5. **The sensor is deterministic.** Same artifact, same reading. If a model's opinion decides it, it's a monitor, not a sensor.
6. **Feedback comes back in minutes.** If a reading takes a week, the loop runs about fifty times a year and a swarm has nothing to do.
7. **The scope fits in one agent's context.** Narrow actuator, explicit editable paths, everything else read-only.
8. **The red lines are written down.** What the customer sees, what the other side pays, what the regulator sees. Each one is enforced outside the agent's reach.
9. **The other side of the market is protected.** If the loop can raise cash flow by worsening a customer's, seller's or worker's deal, that path is a red line.
10. **The downside is bounded.** Anything with regulatory exposure stays with people, whatever the prize.
11. **There is a realized-value plan.** How you'll measure, after the swarm runs, whether the prize was real, and who reads the result.
12. **It's soul-checked.** A person who uses the product has read the thesis and agrees the win would feel like a win.

## Sources

- Cade Metz, [How Google's AI Viewed the Move No Human Could Understand](https://www.wired.com/2016/03/googles-ai-viewed-move-no-human-understand/), WIRED, 2016-03-14
- Google DeepMind, [AlphaGo](https://deepmind.google/research/alphago/)
- Google DeepMind, [AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms](https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/), 2025-05-14
- Novikov et al., [AlphaEvolve: A coding agent for scientific and algorithmic discovery](https://arxiv.org/abs/2506.13131), 2025
- Richard Evans and Jim Gao, [DeepMind AI Reduces Google Data Centre Cooling Bill by 40%](https://deepmind.google/discover/blog/deepmind-ai-reduces-google-data-centre-cooling-bill-by-40/), 2016-07-20
- Xinran He et al., [Practical Lessons from Predicting Clicks on Ads at Facebook](https://research.facebook.com/publications/practical-lessons-from-predicting-clicks-on-ads-at-facebook/), ADKDD, 2014-08-24
- Andrej Karpathy, [autoresearch](https://github.com/karpathy/autoresearch), 2026
- Reuben Binns, Jake Stein, Siddhartha Datta, Max Van Kleek and Nigel Shadbolt, [Not Even Nice Work If You Can Get It; A Longitudinal Study of Uber's Algorithmic Pay and Pricing](https://arxiv.org/abs/2506.15278), FAccT 2025
- Bill Casselman, [The Legend of Abraham Wald](https://www.ams.org/publicoutreach/feature-column/fc-2016-06), AMS Feature Column, 2016-06
- Wikipedia, [Viable system model](https://en.m.wikipedia.org/wiki/Viable_system_model); [Variety (cybernetics)](https://en.m.wikipedia.org/wiki/Variety_(cybernetics)); [Goodhart's law](https://en.m.wikipedia.org/wiki/Goodhart%27s_law); [Project Cybersyn](https://en.m.wikipedia.org/wiki/Project_Cybersyn); [Survivorship bias](https://en.wikipedia.org/wiki/Survivorship_bias)
- Wikiquote, [Stafford Beer](https://en.wikiquote.org/wiki/Stafford_Beer) (*Decision and Control*, 1966, pp. 177 and 239; *Diagnosing the System for Organizations*, 1985, p. 99)
- My own notes: "loops, not slop" manifesto (August 2026), notes on loops (April 17 and August 14, 2026), tweet "alphago 37 move on everything" (March 12, 2026)
