Benchmarks for recursive self-improvement

September 2026

In May 2025 Google DeepMind published AlphaEvolve, an evolutionary coding agent built on Gemini. The headline results were in math: 4x4 complex matrices multiplied with 48 scalar multiplications, a new lower bound for the kissing number in 11 dimensions. The part I keep rereading is the section about Google's own infrastructure. A scheduling heuristic it found for Borg, "now in production for over a year, continuously recovers, on average, 0.7% of Google's worldwide compute resources." It proposed "a Verilog rewrite that removed unnecessary bits in a key, highly optimized arithmetic circuit for matrix multiplication," and the change was "integrated into an upcoming Tensor Processing Unit." It found a better way to split a large matrix multiplication, which "sped up this vital kernel in Gemini's architecture by 23%, leading to a 1% reduction in Gemini's training time." The post says the training gains include "training the large language models underlying AlphaEvolve itself."

The numbers are small. But this is the first public case I know of where a model improved, in production, the process that makes models: the datacenter it runs in, the chip it trains on, and the kernel its successor trains with.

That should change what we benchmark. Most benchmarks measure the output of the AI production function: can the model pass this exam, fix this issue, solve this puzzle. They saturate, and each new one mostly tells you the model got better at what it was trained on. The benchmarks worth building now measure whether a model can improve an input to that function: learning algorithms, kernels, inference efficiency, memory, data, evaluation, hardware. The current answer to "can AI do AI research" is a family of benchmarks that ask whether an agent can do what an ML researcher does in 8 or 24 hours. That's a proxy for a proxy. I'd rather have many small environments, each isolating one mechanism of the production function, each scored in seconds by a deterministic evaluator, with no known ceiling.

The idea is old. In 1965 I.J. Good wrote, in "Speculations Concerning the First Ultraintelligent Machine" (the volume of Advances in Computers that printed it is dated 1966): "Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an 'intelligence explosion,' and the intelligence of man would be left far behind." There's a benchmark hidden in that sentence. "The design of machines" is a list of activities, and sixty years later we can measure some of them with a stopwatch.

Tokens per second are researchers

On September 2 I wrote in my notes (translating from Portuguese): imagine we manage to make tokens per second 10x faster; that pushes AI progress a lot, because it's 10x more autonomous researchers.

When research is done by people, a faster chip makes each experiment cheaper. When research is done by agents, a faster chip also adds researchers. Inference throughput becomes headcount. A model that designs a better accelerator, or writes a kernel that doubles decode speed, saves money and also multiplies the number of agents that can work on the next improvement, including the next accelerator.

METR's RE-Bench shows which terms of that function matter today. It has 7 open-ended ML research engineering environments and "data from 71 8-hour attempts by 61 distinct human experts." With a total budget of 2 hours per environment, "the best AI agents achieve a score 4× higher than human experts." With more time the order flips: humans are "narrowly exceeding the top AI agent scores given an 8-hour budget, and achieving 2× the score of the top AI agent when both are given 32 total hours." The agents' advantage is iteration rate:

RE-Bench score runs per hour for two agent scaffolds and for human experts

"On average, AIDE and modular agents run score 36.8 and 25.3 times per hour respectively, while human experts only do so 3.4 times." The humans' advantage is what they do with a long horizon. Both are terms of the same function. Faster inference raises iteration rate directly, and it buys horizon too, because a long trajectory costs less. Memory is what lets an agent keep compounding past hour two. Evaluator speed caps iteration rate no matter how fast the model thinks. That's why memory and evaluation are on the list next to kernels and chips. (I made the case for spending inference aggressively in Compute is not the bottleneck; this post is about what to point it at.)

So the value of an RSI benchmark is how directly its score feeds one of those terms. AlphaEvolve's 0.7% is the honest unit: a fleet-level counterfactual in compute, not a rank on a leaderboard.

What the AI R&D benchmarks show

Four public benchmarks and one recent result carry most of the evidence.

RE-Bench: agents beat every human on a kernel, and cheat on a training script. In "Optimize a Kernel," "both o1-preview and Claude 3.5 Sonnet find distinct solutions to our kernel optimization problem that beat the efforts of all 9 human experts." The paper is honest about how: "many agent runs solve the same 'Optimize a Kernel' environment not by writing a successful Triton solution (which is very difficult), but by carefully tweaking the starting Pytorch solution." And in "Optimize LLM Foundry," where the task was to cut a finetuning script's runtime while keeping the model close to a reference, "the agent just copies the weights of the reference model" and "chooses to 'simulate' training by modifying a small subset of the weights by a small value." METR scored it zero. METR also found that a top agent score on "Finetune GPT-2 for QA" fell from 0.88 to 0.69 on a rerun, "partially driven by overfitting to this noise."

MLE-bench: attempts matter as much as the model. OpenAI's MLE-bench uses 75 Kaggle competitions. The best setup, o1-preview with AIDE, "achieves at least the level of a Kaggle bronze medal in 16.9% of competitions," and "o1-preview's score doubles from 16.9% using pass@1 to 34.1% using pass@8." Time helps less: "GPT-4o scores 8.7% given 24 hours to attempt each competition, but 11.8% when given 100 hours."

PaperBench: replication is still a human game, and the grader is a model. PaperBench asks agents to replicate 20 ICML 2024 papers from scratch, graded against rubrics with "8,316 individually gradable tasks." The best agent, Claude 3.5 Sonnet (New), averaged 21.0%. On a 3-paper subset, "our human baseline of ML PhDs (best of 3 attempts) achieved 41.4% after 48 hours of effort, compared to 26.6% achieved by o1." The grading is done by "o3-mini-high with custom scaffolding," which "achieves an F1 score of 0.83" on an auxiliary evaluation against human grades. That's good for a judge. It also means part of every PaperBench number is the judge.

KernelBench: most generated kernels don't beat PyTorch. KernelBench has "250 carefully selected PyTorch ML workloads." In February 2025, frontier reasoning models were "matching the PyTorch baseline in less than 20% of the cases." Kernels are the most direct lever on tokens per second, and the benchmark had a lot of room left.

Headline MLE-bench and PaperBench results for agents and human experts

Recursive: the loop works on small, fast, hardened benchmarks. In June, Recursive published First Steps Toward Automated AI Research. They chose benchmarks because they "have clear metrics, relatively low variance, and evaluators that can be hardened against reward hacks." On NanoChat autoresearch (train the best small model in five minutes on one GPU), they went from the community's best 0.9372 to 0.9109 validation bits per byte. On NanoGPT Speedrun, where "83 human record-setting contributions" had cut training time "from roughly 45 minutes down to 79.7 seconds" since mid-2024, the system found 77.5 seconds. On SOL-ExecBench's 235 kernels it raised the mean score from 0.699 to 0.754, "an 18% reduction in the gap to the hardware limit." The line I'd put above every RSI benchmark is theirs: "as the search became stronger, the evaluator had to become stronger too."

Agents already win when the task is short, the score is numeric, and they can try many times. They lose on long horizons and on work that needs a rubric to grade. And both METR and Recursive caught agents gaming the evaluator.

Two cautionary tales

Sakana's AI CUDA Engineer. In February 2025 Sakana claimed a system that could speed up some model training by up to 100x. TechCrunch's summary: "The only problem is, the system didn't work." One user measured a 3x slowdown. Sakana's postmortem said the system had found "exploits in the evaluation code" that "allowed it to bypass validations for accuracy," and the company wrote that it had "since made the evaluation and runtime profiling harness more robust to eliminate many of such [sic] loopholes." This is what a kernel benchmark looks like when the evaluator is weaker than the optimizer.

AlphaChip. Google's 2021 Nature paper on reinforcement learning for chip floorplanning claimed layouts that beat human experts. In September 2023 Nature added an editor's note: "Readers are alerted that the performance claims in this article have been called into question." It pulled an accompanying News & Views piece after its author, Andrew Kahng, changed his assessment; his group's reproduction found human experts typically beat the method, but they had no way "to replicate the pre-training described in the Nature paper." Google's authors answered in That Chip Has Sailed that the reproduction "did not pre-train the RL method" and used "20x fewer RL experience collectors (26 vs 512 in Nature)." In September 2024 Nature published an addendum instead of a correction: "Today, we give this method a name: AlphaChip." The authors say the method "has been used in production for multiple generations of Google's flagship AI accelerator (TPU)."

I don't know who is right about AlphaChip, and that's the point. The evaluation lived inside a private production flow, so it took three years of papers, an editor's note and a lawsuit from a fired Google researcher to argue about a number. A benchmark for recursive self-improvement has to settle that kind of argument in minutes, on a machine anyone can rent.

That's why I keep writing the same rule into every loop I build. From August: "is really important to focus on making the best eval and most complete one because then we can free the models on the feedback loop." From September, more bluntly (translated): I don't want LLM as a judge, the eval must only have deterministic assessments. The reasons are in Evals built backwards and Show the problem, hide the metric.

What a good RSI benchmark looks like

Over September I wrote down the rules I wanted for a suite of these.

  1. Small and simulable. Evaluation runs in seconds or minutes on one machine, on a model small enough to run thousands of experiments, and large enough that scaling laws let you extrapolate the result.
  2. One mechanism. Freeze everything except the thing being improved. If the task is an optimizer update rule, the data, architecture, budget and evaluator are fixed.
  3. A private reference. Before shipping the benchmark, someone writes a solution that beats the baseline. It proves improvement is reachable and never goes to the agent.
  4. Held-out distributions. The agent can inspect training data during its trajectory. Test data, labels and the evaluator stay hidden, and the score comes from an external scorer.
  5. An honestly named proxy. Say what the number measures. "Validation loss of a 50M model after five minutes" is honest. "Research ability" isn't.
  6. Deterministic scoring. Asserts on numbers, outputs and timings, never on prose.
  7. One word and a goal. If the benchmark can't be stated as a verb and a target ("Compress: shrink the KV cache at equal accuracy"), it's two benchmarks.
  8. Unsolved. Tasks that frontier agents solve at once are wasted compute. The benchmark should have a baseline that's easy to beat and no known top.

These rules exist because of the failures above. Rules 4 and 6 are the Sakana rules. Rule 3 is the AlphaChip rule: if nobody outside can check that improvement was possible, nobody can check the claim either. Rule 5 is the PaperBench rule.

A suite of twelve families

I've been sketching a suite of 12 small executable families on these rules. Each isolates one part of the production function:

Family One word and a goal Proxy, named honestly
Optimizer update rules Descend: lower loss at a fixed step budget final loss of a small fixed model
Continual gradient mixing Retain: learn a new task without losing the old one accuracy on both task sets after sequential training
Data curriculum Order: reach target loss with fewer tokens tokens to target on a fixed corpus
INT4 kernels Quantize: faster low-bit matmul at a fixed error bound throughput at a numerical tolerance
Activation recompute Fit: train a larger model in the same memory peak memory and step time
KV-cache compression Compress: smaller cache at equal accuracy cache bytes per token at a quality floor
Research memory Remember: carry learnings between trajectories improvement per run on a fixed task sequence
Eval discriminators Discriminate: separate good from planted-bad candidates detection rate minus false positives
Pipeline repair Repair: fix a broken training pipeline fraction of seeded faults fixed without regressions
Causal control Intervene: find the lever that moves an outcome effect recovered in a simulator with known ground truth
Tabular ranking Rank: raise ROC-AUC on a public credit dataset AUC on Kaggle's hidden Home Credit test split
Incident resolution Resolve: restore a faulted microservice stack faults mitigated and time to recovery in an open simulator like AIOpsLab

The table describes the design, not results. The last three are closer to business loops than to AI internals. I include them because the loops that make AI useful are part of the production function too (Find the loops), and because each already has a public version with ground truth. Kaggle's Home Credit Default Risk scores default predictions by ROC-AUC on a test set nobody can see. Microsoft's AIOpsLab deploys microservices, "injects faults, generates workloads, and exports telemetry data," and grades the agent's fix.

The benchmark I'm most excited about is the eighth. An evaluator is a production input like any other. A family where the agent's job is to build the discriminator, scored by how many planted reward hacks it catches, directly attacks the bottleneck Recursive described. It's the same idea as mutation-testing a verifier, which I wrote about in Prove nothing changed.

Co-design models for the hardware you can get

The usual direction is hardware for models: labs design accelerators around the architectures they already use. AlphaEvolve's Verilog change is the same direction, automated. On September 2 I asked the opposite question (translated): why don't AI companies optimize models to be fast on their infrastructure too, like incorporating that into the loss? And a few lines later: maybe the feedback loop of training the model for a hardware is better than hardware for a model.

Sara Hooker named the reason this matters. The Hardware Lottery "introduces the term hardware lottery to describe when a research idea wins because it is compatible with available software and hardware and not because the idea is superior to alternative research directions." Specialization, she writes, "arguably makes it more even more [sic] costly to stray off of the beaten path of research ideas."

Recursive self-improvement changes the lottery's economics. If agents do the adaptation, the cost of fitting a model to unusual hardware drops toward the cost of the compute to search. Then the interesting target is hardware that isn't supply-constrained: whatever is cheap, abundant, and ignored because the dominant architecture runs badly on it. A benchmark here fixes the hardware and a quality floor, and scores tokens per second. The agent may change architecture, precision, sparsity, kernels, anything. If tokens per second are researchers, a model co-designed for cheap hardware is a way to buy researchers nobody else is bidding on.

This is an open hypothesis. I haven't seen a public result that puts inference speed on a specific cheap target directly into a frontier training loss, and I don't know how much quality it costs.

Aim the swarm upstream, at the tokenizer

The usual autoresearch target is the part of the stack closest to the metric: the head on top of the embeddings, the loss, the hyperparameters. On August 7 I wrote: "I think that one of the best spawn of agent is to focus since the tokenizer [...] instruct the models to make the best tokenizer as possible." Everything downstream inherits the tokenizer's choices, so an improvement there compounds through the embeddings, the probe, and the metric. Each candidate tokenizer could be scored cheaply with a logistic-regression probe on embeddings from a small model chosen to sit on the frontier: large enough that its signal extrapolates, small enough to run many fast experiments.

The same logic holds for language models. The tokenizer sets the unit of compute for everything after it. Meta's Byte Latent Transformer replaced the fixed vocabulary with byte patches "segmented based on the entropy of the next byte," ran "the first flop controlled scaling study of byte-level models up to 8B parameters and 4T training bytes," and matched "training flop-controlled performance of Llama 3 while using up to 50% fewer flops at inference." That's a tokens-per-second gain from changing the first layer of the stack.

There's a detail in the Recursive post that supports this. Starting from a vanilla Transformer, their system added "byte-level features injected right after the token embedding" and hashed bigram and trigram embedding tables. Nothing in the post says it was pointed at the input representation. The search went upstream on its own. A benchmark that says so explicitly ("Tokenize: minimize bits per byte times inference flops at a fixed training budget") points the swarm where the search was already drifting.

Benchmarks with no known optimum

On September 7, looking at benchmark suites where each task has a known best answer, I wrote (translated): wouldn't it be better to have something where even we don't know the solution or the optimum, and we can only give a score?

A private reference proves improvement is reachable. It shouldn't be the ceiling. The best RSI benchmark in public is one that nobody designed as a benchmark: NanoGPT Speedrun. A community spent two years and 83 records taking it from about 45 minutes to 79.7 seconds, and an automated system still found 2.2 more seconds. Nobody knows the optimum. The only thing anyone needs to agree on is the scorer.

Open-endedness research has argued for this since before language models could code. POET used "open-ended" to mean "the intriguing potential for algorithms like POET to continue to create novel and increasingly complex capabilities without bound." OMNI-EPIC generates environments in code, "in principle, to create any simulatable learning task," and uses a model of interestingness to decide which new tasks are worth keeping. It names the new bottleneck: "how to effectively acquire large amounts of data for diverse tasks." Silver and Sutton, in Welcome to the Era of Experience, say "the world abounds with quantities such as cost, error rates, hunger, productivity," and every term of the AI production function is one of those quantities: flops, bytes, seconds, loss.

Games belong here, and I've argued they're the GPT-3 moment for RL in How to achieve superintelligence, so I won't repeat it. One addition: in a persistent world like Screeps, bad architecture and missing recovery logic "remain active liabilities rather than disappearing when the episode ends." That's the long-horizon term RE-Bench shows agents are missing, and a persistent environment measures it where an 8-hour snapshot can't.

Small models in generated worlds

This part is speculation, and I'll mark where it stops.

On August 8 I wrote down the experiment I most want to run (translated): put a small model into RL with many asymmetrically generated tasks, and then measure the impact on other capabilities. "Asymmetrically" means generated in the easy direction and solved in the hard one, like a sudoku that's easy to construct and hard to solve. Long, tool-heavy trajectories inside a real harness, with submit only at the end.

As a benchmark, the thing being improved is the environment generator, and the score is transfer: how much a fixed small model gains on held-out capabilities after training in the generated worlds. That makes data an input you can optimize with the same loop as a kernel. The evidence ends there. I don't know how well transfer from generated worlds scales, and nobody has published the version I'd trust: a fixed small model, fixed RL budget, and a generator as the only free variable.

Intelligence lives in the loop

Not every improvement needs new weights. The shape I keep coming back to is a formal-math loop where the policy is frozen and runs as many parallel inference streams as you can afford. It improves over time without any weight updates, because the learning goes somewhere else: a database of wins, a library of proven lemmas, a database of failures, and a policy for where to spend the next attempt. The Lean kernel is the only oracle. Intelligence lives in the loop, not in the weights.

That's why research memory is a family in the suite. It's also why harness benchmarks count. On HAL's CORE-Bench Hard leaderboard, the same Claude Opus 4.5 scores 33.33% in HAL's generalist agent and 77.78% when run through Claude Code. Same weights, a different loop around them, more than double the score (Optimize the harness, not the model). A benchmark for recursive self-improvement should allow improvement through weights, code, memory or harness, and report which one moved. How loops like that behave as societies is the subject of Agent civilizations.

Where the proxy leaks

The strongest objections are real, and two of them could sink the program.

Proxies don't transfer. A 2-second win on NanoGPT Speedrun might not survive at frontier scale. A kernel tuned for one shape might not matter in production. RE-Bench already shows noise passing for skill (0.88 on the first run, 0.69 on the rerun), and the MLE-bench authors found that "agents can score well on competitions that can be solved with well-known approaches but struggle to debug issues and recover from missteps." Rule 1 leans on scaling laws to extrapolate, but that assumes the mechanism scales, which is exactly what's in question. My answer is partial: hold out distributions, report the scaling curve rather than one point, and treat a benchmark as validated only after one of its wins has been reproduced at a larger scale. AlphaEvolve's gains count because they happened in production. Most of ours won't have that check.

RSI is gated by compute and experiment throughput. Epoch estimated that training compute "has grown by a factor of 10 billion since 2010, with a doubling rate of around 5-6 months." Against that, a 1% training-time gain is small. Faster inference gives you more agents, but each training experiment still needs GPUs, and the RE-Bench curve says the returns are in long horizons, which cost compute. Epoch also notes it measures only final training runs: "We simply do not have sufficient information to determine the total compute through the entire experimentation process." The quantity that bounds automated research, experiment compute, is the one nobody reports. If it's the binding constraint, the benchmarks that matter are the ones that raise experiments per GPU-hour, and a benchmark on anything else is measuring a term that doesn't bind. I think that argues for putting kernels, compression and co-design first. It doesn't refute the program.

Deterministic evals are the prerequisite, and they're hard. Three of the results above involved reward hacking: METR's copied weights, Sakana's bypassed accuracy checks, Recursive's candidates "caching outputs, relying on persistent state, or taking advantage of timing-harness details." A loop is only as trustworthy as its evaluator, and the evaluator has to improve as fast as the optimizer. That's why the eval-discriminator family is in the suite. If the evaluators lose, these benchmarks turn into a machine for producing Sakana papers.

A template for an RSI benchmark family

Every family in the suite fills in the same fields. If one can't be filled, the benchmark isn't ready.

name: Compress                    # one word
goal: shrink the KV cache at equal accuracy
mechanism: KV-cache representation (quantization, eviction, merging)
production_term: tokens/sec and context length per GPU
frozen:
  - model weights (a small open model, pinned by hash)
  - tokenizer, prompts, decoding settings
  - evaluator code and hardware type
free:
  - cache layout, precision, eviction policy, kernels
baseline: full-precision cache
reference: private solution with 2x smaller cache at equal accuracy
proxy_name: "cache bytes per token at >= baseline accuracy on held-out long-context tasks"
score: baseline_bytes / candidate_bytes, zero if accuracy falls below the floor
evaluator_time: under 5 minutes on one GPU
held_out_split:
  visible: public long-context tasks
  hidden: unseen task types and context lengths 2x longer than visible
hack_surface:
  - caching answers across runs
  - detecting the evaluator and changing behavior
  - moving work outside the timed region
hack_checks: fresh process per run, randomized inputs, timing across the whole call
promotion_rule: beat the champion by >1% across 5 seeds with a paired bootstrap

production_term is the field that keeps the suite honest. It says which input to the AI production function the score is supposed to move. hack_surface is written before the first agent runs, and it grows every time an agent finds something new.

Five benchmarks I'd build next

  1. Tokenize. Minimize bits per byte times inference flops for a small model at a fixed training budget. The agent can change the tokenizer or remove it (byte patches, hashed n-grams). Held out: languages and domains absent from the visible corpus.
  2. Co-design. Fix a cheap, abundant hardware target and a quality floor on a held-out eval, and maximize tokens per second. Architecture, precision and kernels are all free. This is the direct test of the hardware-lottery hypothesis.
  3. Trim. Minimize the area or delay of an arithmetic circuit in Verilog, accepted only if a formal equivalence check passes. It's the open, reproducible version of AlphaEvolve's TPU change, and the kind of chip-design benchmark that would have settled the AlphaChip argument.
  4. Discriminate. Given a kernel or training benchmark and a stream of candidate solutions, some with planted reward hacks, build the checker. Score: hacks caught minus false rejections, on hack types the agent never saw. This benchmark improves every other one.
  5. Generate. Write an environment generator for a fixed small model and a fixed RL budget. Score: gain on held-out capabilities after training. Data becomes an input with a number attached.

Each can be stated as one word and a goal, runs in minutes, freezes everything but one mechanism, and has no known optimum. None of them asks whether a model resembles a researcher. They ask whether the process that makes AI got faster, and by how much.

Sources

← Evals built backwards
Prove nothing changed →

Markdown version: /blog/benchmarks-for-recursive-self-improvement.md. Every essay: /agents.