Evals built backwards

September 2026

Checking a sudoku takes two lines. In Peter Norvig's Solving Every Sudoku Puzzle, the whole check is: "A puzzle is solved if each unit is a permutation of the digits 1 to 9." Building one is almost as cheap: fill a valid grid, erase cells. Solving it is the hard part, at least for anything that can't run Norvig's constraint propagation in its head. Sudoku-Bench gave frontier LLMs 100 variant puzzles chosen with the hosts of Cracking the Cryptic, and "state-of-the-art LLMs solve fewer than 15% of puzzles unaided."

In January 2025 the authors of Humanity's Last Exam explained why they needed a new test: "LLMs now achieve over 90% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities." Their answer was 2,500 questions, each with "a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval." At launch GPT-4o scored 2.7%, o1 8.0%, and o3-mini (high) 13.4%.

Both artifacts work the same way, and I think it's the most useful idea in eval design right now. When going from X to Y is hard and going from Y to X is easy, start at Y, walk the easy direction to produce X, and grade the model on the hard one. Then keep an item only if it separates strong setups from weak ones. On June 3 I wrote the idea down in my notes: "the eval is going to be X -> Y that is hard but we are going to generate on the inverse direction Y -> X that is easy."

What the usual practice gets wrong is where the effort goes. Teams mirror production traffic, count items, and ask a model to grade the answers. The result is a large, easy, soft eval. Items built backwards from a known answer can be harder than anything the builder could solve, and they can be graded by code.

The asymmetry

The pattern has one requirement: the builder holds something the solver doesn't. Call it the witness. A filled sudoku grid is a witness for the puzzle made by erasing it. A working codebase is a witness for the bug you inject into it. An expert's answer, found in her own lab, is a witness for the question she writes around it. The builder never solves anything. She starts from the solution and hides it.

This is the NP intuition used as an engineering rule. Checking a candidate is cheap. Finding one is expensive. As long as the check stays cheap, you can make the search as expensive as you like, and the grader doesn't need to be smarter than the solver. In September I put it in the sudoku terms I started with, translated from Portuguese: "just like sudoku, where building a game is very easy but solving it is hard." So a task can be made almost impossible to solve without becoming any harder to check.

The field has been building evals this way for a while, often without saying so.

Sudoku-Bench. Variant sudokus add rules on top of the classic ones. The authors picked them because "each puzzle introduces unique or subtly interacting constraints, making memorization infeasible and requiring solvers to identify novel logical breakthroughs." The set is graded by size: 15 puzzles at 4×4, 15 at 6×6, and 70 at 9×9. Checking a filled grid against the rules stays mechanical however strange the rules get.

ARC-AGI. François Chollet's On the Measure of Intelligence is the clearest statement of why this matters: "unlimited priors or unlimited training data allow experimenters to 'buy' arbitrary levels of skills for a system, in a way that masks the system's own generalization power." Each ARC task hides a transformation rule behind a few input and output grids, and grading is an exact grid match. The ARC Prize 2024 report called it "five years old and remains unbeaten" as of December 2024, with the state of the art on the private set moving from 33% to 55.5% against an 85% target.

Reasoning Gym and Enigmata. These make the generator explicit. Reasoning Gym ships "over 100 data generators and verifiers" and can "generate virtually infinite training data with adjustable complexity." Enigmata has 36 puzzle tasks across seven categories, each with "1) a generator that produces unlimited examples with controllable difficulty and 2) a rule-based verifier for automatic evaluation." Its trained Qwen2.5-32B scored 32.8% on ARC-AGI and 0.6% on ARC-AGI 2. The generator-verifier pair is the whole method in two functions.

SWE-bench against SWE-smith. SWE-bench is built forwards: 2,294 problems "drawn from real GitHub issues and corresponding pull requests across 12 Python repositories." SWE-Gym does the same for training, with 2,438 real instances, and reports "up to 19% absolute gains in resolve rate." Forward curation is slow. SWE-smith says existing datasets have "at most 1,000s of training instances from 11 or fewer GitHub repositories" and take "hundreds of hours of human labor." So it runs backwards: given a working codebase, it "automatically synthesizes 100s to 1,000s of task instances that break existing test(s)." The working code is the witness and the tests are the verifier. The result was 50,000 instances from 128 repositories and a 32B model at 40.2% on SWE-bench Verified.

Humanity's Last Exam. HLE is a backwards eval with people as the generator. Experts wrote questions whose answers they already knew, each "accompanied by a detailed solution to verify accuracy." Then came a filter: "questions are rejected if LLMs can answer them correctly." Exact-match questions had to stump every model in the panel. The paper logged "over 70,000 attempts, resulting in approximately 13,000 questions which stumped LLMs," which went to expert review and ended as 2,500.

Needles and hay

The best-known backwards eval for LLMs is the needle in a haystack, and it shows how the method fails.

The idiom predates all of this. It exists as "agulha num palheiro" in Portuguese, and Mandarin has "大海捞针", fishing a needle out of the sea (Wiktionary). In the corpora I searched through Scry, the earliest copy is a 1984 Usenet post, and the earliest on Hacker News is a 2007 comment about false leads: "If your system generates too many false leads, it will be ... like adding more hay to a haystack with a missing needle." The phrase already describes the asymmetry. Finding the needle is hard. Knowing you've touched it is trivial.

Greg Kamradt turned it into a long-context test in 2023. His harness "runs a sweep of (context length × needle depth) cells," with Paul Graham's essays as the default haystack. It's built backwards: plant a fact at a known depth, ask for it, check the answer by string match. It cost almost nothing to make or grade, and by 2024 RULER could describe it as "widely adopted to evaluate long-context language models."

Then models aced it. RULER called the test "indicative of only a superficial form of long-context understanding." Its 17 models achieved "nearly perfect accuracy in the vanilla NIAH test," yet although all claimed 32K tokens or more, "only half of them can maintain satisfactory performance at the length of 32K" once the tasks got harder.

NoLiMa found the specific leak. In the vanilla test the question shares its words with the needle, so the model can find the answer by literal matching. NoLiMa measured the overlap across benchmarks and built a needle set where "questions and needles have minimal lexical overlap."

Literal word overlap between question and relevant context, by benchmark, from NoLiMa's Table 1

Vanilla NIAH scores 0.905. NoLiMa scores 0.069. With the overlap gone, "at 32K, for instance, 11 models drop below 50% of their strong short-length baselines." GPT-4o fell "from an almost-perfect baseline of 99.3% to 69.7%." Command R+ went from 90.9% to 7.4%.

Kamradt built his test in the right direction. But the witness leaked into the question, so verification was cheap and search was cheap too. A backwards eval only measures something when the search it forces is hard.

A trajectory is worth what it discriminates

By late July I had stopped counting items. The rule I wrote down: a trajectory is valued only for how well it separates strong setups from weak ones. One that a weak prompt passes as easily as a strong one is dead weight. It costs compute, adds noise to the average, and tells you nothing.

The psychometricians got here first. tinyBenchmarks applies Item Response Theory to LLMs, treating models as test-takers and "learn[ing] representations of examples encoding latent abilities." The finding: "to accurately estimate the performance of an LLM on MMLU, a popular multiple-choice QA benchmark consisting of 14K examples, it is sufficient to evaluate this LLM on 100 curated examples." That's under 1% of the items, "within about 2% error on average." For telling models apart, most of MMLU was hay.

Density over volume follows. It also explains why I refuse to mirror production traffic. Most traffic is easy, and a mirror of it gives you an eval that saturates the week you build it. Norvig ran the experiment on random generation in 2006. He generated a million random puzzles and timed his solver on each.

How many of one million random sudoku took longer than each threshold to solve, from Norvig

The average took 0.01 seconds and "more than 99.95% took less than 0.1 seconds." One in a million took more than 100 seconds. Hard items are rare in any distribution you don't select. You have to go and find them. The one Norvig found at 188.79 seconds turned out to have more than one solution, a problem I come back to below.

So the generator has a second job. After each trajectory it reflects on what the solver did and makes the item harder: add a constraint, remove a hint, chain a second tool, move the needle into a place the solver doesn't look. HLE's pipeline has the same loop with humans: "We accept questions that make frontier LLMs fail, then iteratively refine them with the help of expert peer reviewers."

Two refinements matter.

Discrimination beats difficulty. HLE filtered for items that stumped a fixed panel of models, and its authors warn about the consequence: "small inflections close to zero accuracy are not strongly indicative of progress." An item nobody passes separates nobody. The filter I want keeps items where the strong setup passes and the weak one fails, measured over several runs each. Difficulty is how you get there, and the gap is what you keep.

A score that noise can move doesn't discriminate. Spurious Rewards found that RL with "randomly assigned rewards" improved Qwen2.5-Math-7B on MATH-500 by 21.4 points, "nearly matching the 29.1-point gain from ground-truth rewards." The paper is about training signals, but it says something about the benchmark too: for that model family, a gain on MATH-500 doesn't by itself tell you the model learned anything. The same random rewards "often fail to produce gains for other model families." An item's value depends on the setups you're comparing.

There's a tension with another rule of mine, that coverage should be weighted by traffic. The two fit together. Traffic decides which scenarios get weight: if half of a product's users ask about the same thing, half the weight goes there. The generator decides which items represent each scenario, and it picks the ones that discriminate.

The rule

By September the idea had hardened into five conditions. An eval item counts only when all of them hold.

  1. A trusted builder starts from a witness. The builder is code or a person you trust, never the model under test. The witness exists before the item does.
  2. Verification is cheaper than search. If the solver can find the answer as cheaply as the grader can check it, the item measures string matching. NIAH's 0.905 overlap is the example.
  3. Difficulty is measured, not assumed. Author labels are weak. Terminal-Bench 2.0 compared its authors' difficulty ratings with empirical pass rates across frontier models and found a correlation of r = 0.436. Of the tasks humans rated medium, 54.5% were hard for models. Run a panel and let pass rates set the difficulty.
  4. Every valid solution is accepted. The verifier checks the constraints, not equality with the witness. Norvig's random puzzles were "not guaranteed to have one unique solution," and "a few (about 0.2%) have no solution." A grader that compares against one stored answer punishes a correct solver for finding a different correct answer.
  5. An LLM is never the only labeler, and never judges its own output. Zheng et al. documented "position, verbosity, and self-enhancement biases" in LLM judges. In their position test, zero-shot GPT-4 kept its verdict after the two answers swapped places only 65.0% of the time. GPT-4 favored its own answers "with a 10% higher win rate" and Claude-v1 its own with 25%, though the authors say their data "cannot determine whether the models exhibit a self-enhancement bias." On September 7 I wrote it bluntly, translated: "I don't want LLM as a judge, the eval should only have deterministic assessments."

The fifth rule is the one people push back on, and I've argued the deterministic-sensor side at length in Show the problem, hide the metric, including a result where an LLM judge did better than I expected. The short version is that assertions go on data: rows written, routes called, widgets rendered, tool arguments sent. Never on prose.

From one tool to many

Tool use is where I think the method pays off most. Built forwards, a tool-use eval starts from logged conversations and inherits their easy distribution. Built backwards, the generator doesn't start from a user message. It starts from the witness: a target set of tools and the effects they should have. Then it writes the user input that would require exactly those calls. The verifier reads the trace: were the right tools called, with arguments that produce the expected state?

The curriculum should run from one capability to combined ones. I think in five levels. Level 1 needs one tool. Level 5 needs several tools combined, where the output of one is the argument of the next and the user never names either. Each level is easy to generate, because you choose the tools first. Solving gets harder with every level, because the solver has to infer the plan from a request that hides it. Reasoning Gym and Enigmata make the same bet with a difficulty knob on every generator.

Before generating anything, map the capability space. Have a model read each tool description and record one action per thing that tool makes possible, written at a general level ("list the open issues assigned to someone") instead of as a concrete query. Let it keep going until it declares that the space is saturated and nothing is left to add. Then group the actions by the tools they share and compose them into multi-step intents. That map is what makes the items dense. Without it the generator keeps rediscovering the same three easy actions, the way Norvig's generator keeps producing easy puzzles.

I think the same shape carries to other domains. Game tasks, where the witness is a reachable game state and the verifier reads it. Esoteric questions synthesized from about ten references in one person's reading, each with one deterministic answer, which reward reasoning over memorization because nobody else has that combination of sources. Harness evals that plant many needles at once across memory, web search, code and computer use, so the agent has to find all of them. For evals where the thing being improved is the process that makes AI, see Benchmarks for recursive self-improvement, where I also describe training small models on tasks generated this way.

Users who write in lowercase

Backwards generation has a failure mode of its own: the items read like they were generated. A generated user writes something like "Could you please retrieve my calendar events for next Tuesday?" A real one writes something like "tenho algo terca".

So I think simulated users have to write like actual people. The user I picture never types a capital letter, and it's very hard to get a model to add missing accents or cultural traits on purpose. Terse, lowercase, no punctuation, the request half-formed. The other half of the rule is that persona is the wrong place to put difficulty. My worry with difficult personas is that a model asked to play a confused user overacts, and the confusion becomes a cartoon the solver learns to parse. Difficulty belongs in the task structure: the number of tools, the dependencies between them, what the user leaves out, what changes halfway through.

The personal-assistant eval I'd want is little stories that mix web search, computer use, coding, tool use, memory and personal finance, aimed at ordinary people rather than power users: a bad audio transcript, someone with a PDF invoice, someone on WhatsApp Web. Tasks work better as environments than as one-shot prompts. Options unlock over time, and the data and pipelines are real. The bar I'd set is hard enough that 100% would mean superintelligence with the best harness. If a good product can pass the suite today, it's too easy to steer the next one.

Measure the product, not the method

The rule I hold for evaluating any agent product is to measure the product, not the method. Both the system and its eval should be harness agnostic.

That means the agent under test runs as it runs in production. Full system prompt, full tool descriptions, no truncation, free to reason and loop as long as it wants, scored only on the final outcome. Whether search goes through one API or another, whether the context gets compacted or not, is the method. The eval doesn't care. It asks whether the hard thing got done, and it checks that with deterministic assertions on each trajectory.

Academic benchmarks struggle here because they want to measure a model, and the model can't be separated from its harness. The Terminal-Bench 2.0 paper says so: "agent and model performance are hard to decouple. Many agent scaffolds have been engineered to accommodate the tendencies of certain models, especially when the model and agent are developed by the same organization." Its headline figure notes that "the agent scaffold used to report each model was chosen to maximize performance." I made the case that the coupling goes into the weights in The harness is in the weights. For a product team the coupling is the thing you ship, so the eval should score all of it together and let every component be swapped underneath.

This is also why I think the eval has to be complete before anything else runs. A deterministic eval that covers the product is what lets an optimizer run unattended. An eval with holes is an invitation to find the holes.

Where backwards breaks

The method has real weaknesses, and some come from the same property that makes it cheap.

Synthetic tasks may transfer weakly. Most evidence that generated tasks matter comes from the people who built the generators. Enigmata reports gains on out-of-domain puzzles and math "with little multi-tasking trade-off." That's encouraging, but it's the authors' measurement on the authors' generator. Spurious Rewards is the warning in the other direction: large gains "can arise on Qwen models even from random rewards that do not reflect genuine capability improvements." I don't have a published measurement of how well a backwards-built eval predicts success on real user tasks, and I don't have a clean one of my own either. The honest claim is that backwards evals measure the capability they encode. Whether that capability is the one users need is a separate question, answered by weighting scenarios by traffic and checking against production outcomes.

Generators leave fingerprints. NIAH is the documented case: the construction procedure left a 0.905 lexical overlap that models exploited. Every generator has a style, and a model trained or tuned on the same family learns the style along with the task. SWE-smith's injected bugs are a plausible example. I haven't seen a study showing they cluster on patterns the injector knows, so treat that as a risk to check rather than a finding. The defense is to measure leakage directly, for example the overlap between prompt and answer, and to rotate generators.

A witness guarantees neither hardness nor uniqueness. Norvig's million puzzles are the cleanest evidence: almost all were trivial, and the hardest wasn't a valid sudoku. HLE shows the human version. Its own appendix estimates "a final estimated expert disagreement rate of 15.4% for the public set," and about 18% on a biology, chemistry and health subset. Here is my inference, which the authors don't make: a filter that keeps what models get wrong will also keep some questions whose reference answer is wrong, because a wrong key looks exactly like a hard item. That's why rules 3 and 4 exist.

Contamination is real, but GSM1k found little of it at the frontier. GSM1k rebuilt GSM8k with fresh problems and found "accuracy drops of up to 8%, with several families of models showing evidence of systematic overfitting." It also found that "many models, especially those on the frontier, show minimal signs of overfitting." Backwards generation helps here, because you can regenerate with new seeds and keep a private held-out set, as HLE does "to assess model overfitting and gaming on the public benchmark." But a published generator is also a training environment. Reasoning Gym was built for RL. Once your generator is public, assume someone is training on its distribution.

Deterministic checks miss taste. Zheng et al. also found that GPT-4 as a judge reached "over 80% agreement" with human preferences, "the same level of agreement between humans." For tone or warmth, a judge is a reasonable instrument, and a set of assertions on database rows would miss most of what users feel. My rule is for autonomous loops, where the eval is the only thing standing between an optimizer and a hack. For a human reading a weekly report, a calibrated judge is fine.

Where this goes

This part is speculation, and I'll mark where it stops being evidence.

If items are cheap to regenerate and valuable only for what they discriminate, the dataset stops being the asset. The generator is. A private generator that produces fresh, discriminating items on demand can't be contaminated the way a fixed set can. It also keeps its value as models improve, because it can be pushed to the next level. I think of private evals as an edge: they let any team compare models on its own problems and say something useful about a new model in the week it ships. A private generator makes that repeatable.

The next step is to put the generator in a loop of its own and score it by the discrimination of what it produces. That's an adversarial pair: a solver trying to pass, and a generator rewarded when a strong setup passes and a weak one fails. It's the same move as the eval-discriminator family in Benchmarks for recursive self-improvement, applied one level down. Nobody has shown me that such a generator stays honest over many rounds, and the obvious failure is a generator that learns to produce ambiguous items, which look like hard ones. Rule 4 and a human audit of a sample are the brakes I'd use. The evidence ends here.

An eval-builder spec

The whole method fits in one interface and one loop. This is the shape I use, simplified:

interface EvalFamily<Answer, Witness extends { answer: Answer }, Item> {
  // Trusted code. Never the model under test.
  witness(seed: number, level: number): Witness
  // The easy direction: Y -> X.
  generate(w: Witness): Item
  // Cheaper than search. Checks constraints, so any valid answer passes.
  verify(item: Item, answer: Answer): boolean
  // Rejects items whose answer leaks into the prompt (for example, high n-gram overlap).
  leaks(item: Item, w: Witness): boolean
  // Makes the item harder after seeing a trajectory. The witness must stay valid.
  harden(item: Item, w: Witness, trace: Trace): Item
}

function build(family, strong: Setup[], weak: Setup[], seed: number) {
  for (let level = 1; level <= 5; level++) {
    const w = family.witness(seed, level)
    let item = family.generate(w)
    for (let round = 0; round < MAX_ROUNDS; round++) {
      if (family.leaks(item, w)) break
      if (!family.verify(item, w.answer)) break           // the witness must pass its own verifier
      const s = passRate(strong, item, family.verify, K)  // K runs per setup, production harness
      const v = passRate(weak, item, family.verify, K)
      if (s - v >= MIN_GAP) return { item, level, difficulty: 1 - s, gap: s - v }
      if (s < FLOOR) return null                          // nobody passes, so it separates nobody
      item = family.harden(item, w, lastTrace(strong, item))
    }
  }
  return null
}

What each piece is for:

Piece Job Fails when
Witness The answer, held before the item exists The model under test built it
Generator Walks the easy direction Its style leaks into the item
Verifier Checks constraints cheaply It compares against one stored answer
Discrimination filter Keeps items with a strong-minus-weak gap It filters for "everyone fails"
Difficulty estimate Pass rate across a panel, per item It's an author's label

Checklist

  1. Name the witness first. If you can't say what the builder holds that the solver doesn't, you're writing a forward eval.
  2. Run the witness through the verifier. An item whose own answer fails is a bug in the generator.
  3. Measure leakage. Compute the overlap between prompt and answer. NIAH's 0.905 is what a leak looks like.
  4. Accept every valid solution. Verify constraints, or check uniqueness at generation time the way sudoku setters do.
  5. Measure difficulty on a panel. Author ratings correlated at r = 0.436 with empirical difficulty on Terminal-Bench.
  6. Keep items by gap, not by failure. Drop items that the weak setup passes and items that nobody passes.
  7. Cap volume. A hundred well-chosen items estimated MMLU within about 2%. Add items only when they add information.
  8. Weight scenarios by traffic, choose items by discrimination. Mirror what users ask about, never how easy it usually is.
  9. Map the capability space until it saturates. List general actions per tool until the mapper has nothing left to add, then compose them into multi-step intents.
  10. Write simulated users in lowercase. Put difficulty in the task structure, not in the persona.
  11. Test the product as it runs. Full prompt, full tools, no truncation, scored on outcomes with deterministic assertions on data.
  12. Keep an LLM out of the only-grader seat. Never let a model label alone or judge its own output.
  13. Keep a private seed set and regenerate. Treat a published generator as a training environment.

Sources

Research for this post used Scry to search Hacker News, Usenet archives and academic catalogs.

← Show the problem, hide the metric
Benchmarks for recursive self-improvement →

Markdown version: /blog/evals-built-backwards.md. Every essay: /agents.