Show the problem, hide the metric

September 2026

METR asked o3 to write a fast GPU kernel for a prefix sum. The scorer timed the kernel and compared its output to a reference answer. o3's "kernel" walked up the Python call stack, found the reference tensor the scorer had already computed, returned it, and disabled CUDA synchronization so the timer measured almost nothing. In another run it left the kernel alone and replaced the clock: _time.time = lambda: _real_time() * 1e-3, under a comment that reads tiny "performance trick". METR's post, Recent Frontier Models Are Reward Hacking, counts the runs. On RE-Bench, where the model could read the scoring code, o3 reward-hacked in 39 of 128 runs. On HCAST it did so in 8 of 1,087. In METR's words: "Reward hacking was more than 43× more common on RE-Bench tasks than HCAST tasks, perhaps because on RE-Bench tasks the model was able to see the entire scoring function."

Share of o3 runs with reward hacking on RE-Bench tasks, which expose the scorer, and on HCAST

Four months later ImpossibleBench made the measurement clean. Zhong, Raghunathan and Carlini mutated the unit tests of SWE-bench and LiveCodeBench tasks so they contradict the specification, which means any pass is a cheat. GPT-5 passed 76% of the impossible one-off SWE-bench tasks. When the authors hid the tests, cheating fell "to near zero."

The usual reading is that this is an alignment problem to prompt or train away. Prompting helps unevenly. In METR's ablation, "Please do not cheat" left o3 planning a hack in 80% of runs, the same as the original prompt. ImpossibleBench's strictest prompt took GPT-5 from 92% to 1% on LiveCodeBench, but only from 66% to 54% on the multi-file SWE-bench tasks. I read these results as a design spec for anyone building agent loops inside a company. An agent that patches the grader is searching. It searched the grader because we put the grader inside its search space. Show agents the problem, hide the metric, and the pressure that breaks scorers goes into the product instead. And once one loop works, a company turns out to be one optimization loop where only the sensor changes.

An eval is a loop that does not come back

In a voice memo on September 10 I said it badly and I still like it: "a eval is a loop without coming back." An eval reads a number once and stops. A loop feeds the number back into whatever produced it. That's sensor, controller and actuator, the frame cybernetics used in the 1940s and 50s to fold many separate fields into one.

A swarm is an eval with two more parts. Memory: a shared board where agents post what they tried and search what others found. Selection: a champion that changes only when a candidate measures better. Many trajectories run in parallel, each ends in one reading, selection promotes the best one, and the board carries knowledge between them. It's Darwinian iteration on a rhizome instead of a tree. (The argument about shared boards and stigmergy is in Bottom-up is one ontological level higher.)

Closing the loop separates a swarm from slop. In an August manifesto I wrote that "the real opportunity is not to unleash thousands of agents to generate slop, but to place swarms of agents inside closed, cybernetic loops," judged by deterministic evaluations "difficult to reward-hack, easy to measure, and closely aligned with long-term free cash flow."

The published systems that work have exactly this shape.

FunSearch. Romera-Paredes et al. describe the input as "a specification of the problem in the form of an 'evaluate' function, which scores candidate solutions." The system has three kinds of workers, "a programs database, samplers and evaluators," typically 15 samplers and 150 CPU evaluators. Programs that fail are "discarded." The LLM writes the function being evolved. It never touches evaluate.

AlphaEvolve. DeepMind's successor "pairs the creative problem-solving capabilities of our Gemini models with automated evaluators that verify answers." A scheduling heuristic it found for Borg has run in production for over a year and "continuously recovers, on average, 0.7% of Google's worldwide compute resources." A matrix multiplication kernel it rewrote got 23% faster, which cut Gemini's training time by 1%.

autoresearch. Karpathy's repo is the smallest version. prepare.py holds data prep and evaluation and is "Not modified." train.py is "the single file the agent edits." Every run trains for exactly five minutes and is scored on validation bits per byte, which gives about a hundred experiments overnight.

I don't have a public result of my own to put next to these yet, so the argument rests on them.

The verifier's cost is easy to underestimate. FunSearch typically ran ten evaluators for every sampler. autoresearch caps each experiment at five minutes of training, and that cap is what makes a hundred experiments a night possible. Agents are cheap to widen. Evaluation often isn't, especially when it needs real hardware like phones or GPUs. Verifier capacity, not agent count, sets how wide the swarm can run.

Show the problem, hide the metric

What I got wrong at first was thinking the agents needed all the context. Writing your taste into a prompt is close to impossible. What works is putting that taste into guardrails the agent can't see, "passing your taste to a guard rail," as I put it in the memo. The August manifesto calls it "taste encoded as evals." The agent gets the problem, say "make the app use less memory," and a sensor it can query. It doesn't get the end-to-end tests that decide whether its change ships. Then a swarm synchronized through a message board will "just hack your taste. And this is kind of good."

It's good because hacking your taste means searching the space your taste allows. Three pieces of evidence show how to build it.

Kaggle has run this for years. Kaggle's documentation describes the split: "Public Leaderboard: Visible during the competition; based on a sample of test data. Private Leaderboard: The final ranking based on the remainder of the test data." And the warning: "A high public score doesn't guarantee a high private score. Avoid 'chasing' the public leaderboard." Competitors get a sensor. The verifier stays sealed until the end.

Test access is a dial. ImpossibleBench compared full read/write access to the tests, read-only access, and hidden tests. Hiding them "reduces cheating success rate to near zero, but also degrades performance on the original benchmark." Read-only access "restores legitimate performance while preventing test modification attempts," though it doesn't stop special-casing or operator overloading. The feedback loop itself raised cheating: allowing multiple submissions lifted the average pass rate on the real tasks from 80% to 83% and cheating on the impossible ones from 33% to 38%. On SWE-bench, one harness change did more than any prompt the authors tried. Told it could submit flag_for_human_intervention and end the run, GPT-5 dropped from 54% to 9% and o3 from 49% to 12%. Claude Opus 4.1 barely moved.

Cheating rate on Conflicting-SWEbench for GPT-5 and o3 under different prompts and with an abort option

A reference implementation can hold the taste. Mechanize's GBA Eval asks agents to write a Game Boy Advance emulator in Rust from scratch. The grader feeds pre-recorded button inputs to the candidate and to a reference emulator, Mesen2, and compares "their outputs ... on every frame." The console has no entropy source, so a replay is deterministic. Nobody had to write down what a correct emulator is. The reference is the taste, and the agent can't argue with it. The same logic sets red lines in any app loop. You don't let the agent change what a screen shows, because, as I put it in the memo, "it's something that it's pretty easy to exploit." A deterministic replay of the screens sits outside the agent's reach.

Put together, the design separates two things that usually get merged into one "metric."

The sensor is what the agent optimizes against during a trajectory: a latency number, peak memory, validation loss. The agent sees its readings. The code that produces them runs outside the sandbox or is frozen read-only, and every call costs budget. The agent needs this gradient. ImpossibleBench's hidden-test results show what happens when it has none.

The verifier decides whether a candidate becomes the champion: the held-out end-to-end suite, the frame-by-frame replay, the private leaderboard. It stays hidden and runs after the trajectory ends.

This split changes what reward hacking is. With a hidden verifier, a candidate that games the sensor fails at promotion and does no damage. It also shows you exactly where the sensor is weak. METR recommends the same move for training setups: "use a monitor to detect reward hacking, then when it's detected patch the exploit in the scoring function rather than punishing the model." Every rejected hack is a free bug report on your sensor. Punish hacking and it goes underground. Filter it and it becomes search.

One caution. This works because the loop selects artifacts and doesn't train weights on them. Anthropic's paper on production RL found that a model that learns to reward hack in real coding environments "generalizes to alignment faking, cooperation with malicious actors, reasoning about malicious goals, and attempting sabotage." The first mitigation it lists is "preventing the model from reward hacking." If swarm trajectories ever feed back into fine-tuning, drop the ones the verifier rejected.

Soft sensors and hard sensors

The easy mistake is skipping the hard part. A deterministic verifier is tedious to build, so people put an LLM judge there instead. From the memo: "LLMs judge is not a good solution alone ... the best kind of cybernetic loops are the ones that work deterministically." I put integration tests with mocks the agent can edit in the same bucket. I call both soft sensors: their reading depends on something the actuator can talk to.

Judges are fine at what they were validated on. Zheng et al. found GPT-4 judges "match human evaluations at an agreement rate exceeding 80%, achieving the same level of human-human agreement" on chat preferences, while documenting "position, verbosity, and self-enhancement biases." As cheating monitors they get worse as tasks get harder. In ImpossibleBench, Claude monitors caught 86 to 89% of cheating on single-file LiveCodeBench tasks and 42 to 65% on multi-file SWE-bench tasks. (A newer benchmark finds judges much better on clear-cut hacks; see the objections.)

One miss shows why. Claude Opus 4.1 was asked to make Django's MySQL backend prefer the new database option over the deprecated db, and an impossible test expected db. The agent wrote the correct fix, watched the test fail, reversed the order, and explained it as backward compatibility. The monitor accepted the explanation. A judge can be persuaded, and a loop is a machine for producing persuasive agents. METR adds the dynamic version: penalizing hacks with an LLM monitor "will also incentivize the model reward hacking subtly, in ways the monitor can't detect."

A hard sensor gives the same reading for the same artifact and can't be argued with: milliseconds on a replayed session, peak memory on a fixed device, AUC on a frozen holdout, bits per byte after five minutes. On the code-structure side I use sentrux, whose score is "the geometric mean of 5 root cause metrics." Its docs claim "you cannot game one metric while tanking another," which is the right property for a sensor, whatever you think of the specific metrics.

Models still belong in the loop, on the actuator side. Take a hypothetical: a model writes short summaries that have to fit a hard character limit and come back as valid JSON. The model writes freely. Deterministic checks then read the output: the text fits, the JSON parses. A failure goes back to the model as a new user message naming the error, and it loops until nothing fails, the way a programmer works against a compiler. Don't let code do the repair, because truncating to fit ships text with missing words. The rule: generate freely, verify deterministically, repair with a model. The model is creative on both sides of the check. The check never changes its mind.

Deterministic doesn't mean ungameable. o3's fake clock beat a deterministic timer. What determinism buys is that the gaming is reproducible, so you find it once and fix it once. The fix belongs in the structure of the code, which I'll get to after the part about companies.

A company is one loop

Once one loop works, the rest of a company looks like the same loop with a different sensor. Run through the loops a software company could close, a ranking model, a support agent, a repository's structure, a mobile app, and you land where my memo did: "the only difference between them is the sensor." For an app it's latency and memory. For a hypothetical credit model it starts from the ROC AUC of the model a trajectory produced. For structure it's the sentrux score. The seed, actuator, board, budget, verifier and promotion rule can be the same code. (Which loops to close first is its own search problem, which I wrote up in Find the loops.)

A company's reward function is free cash flow. You can't put that in a loop directly, so the work is finding proxies with a high ratio of cash-flow potential to effort. For any product with heavy daily traffic, my favorite is latency. A speed gain multiplies across every session, and it plausibly shows up as less churn and more transactions.

That's a belief, and the public evidence is weaker than the folklore. Greg Linden wrote in 2006 that at Amazon "we tried delaying the page in increments of 100 milliseconds and found that even very small delays would result in substantial and costly drops in revenue." Amazon never published it; Linden's account is all there is. In the same post he relays Marissa Mayer's story from Google. Users asked for 30 results instead of 10, and in the test "Traffic and revenue from Google searchers in the experimental group dropped by 20%." The 30-result page took 0.9 seconds to generate against 0.4, and Mayer blamed the half second. But the test changed the number of results and the latency at once, and commenters took it apart: more results per page means fewer page requests and the same ads spread over three times the results. Linden replied that he thought Mayer meant searches, not page views. One commenter noted that the Amazon test changed only speed, which makes it the cleaner anecdote. It's still an anecdote. I treat both as a direction. The coefficient for your product comes from your own loop.

Some things look like loops and aren't. "Closing a GitHub PR can be a feedback loop but it's not a proxy," I said in the memo, "because you can easily reward hack" it. Merged PRs and closed tickets measure activity at the actuator. A sensor has to sit on the far side of the product, where the users are.

Even a real sensor needs design work. Take a hypothetical lender that points a swarm at its credit model. AUC measures how well a model ranks borrowers, but ranking alone isn't where the money is. The profit comes from ranking well and safely reaching people who weren't eligible before. So a better sensor is something like AUC times newly available customers, with an explicit floor on AUC, because the coverage term alone rewards approving everyone. Without the floor, a swarm will find the model that approves the most people, confident or not. Any product of terms needs that guard, since one factor can often grow by giving the other away.

The score also has to be signed. Clip at zero and a slightly worse candidate reads the same as a disaster, and selection loses the gradient between them. A signed delta against the champion keeps that information.

Allocate loops by convexity

Not every loop deserves compute. I pick by convexity: a bounded, reversible downside and an upside that compounds.

App latency is convex. A bad change fails the hidden suite or gets rolled back, and a good one multiplies across every session. Anything that touches a regulatory license is the opposite. Picture a hypothetical payments company: a loop that automates risky authorizations could put its license at stake to win a little efficiency. Unbounded downside, small upside.

What falls out is a little funny. I think the moats left after AI are mostly regulation and connection. If that's right, the concave loops are the ones that touch the moat, so they stay with people. The convex ones get the swarms.

Here is how that sorts out for a hypothetical company that ships a consumer app and also lends money:

Loop Sensor Worst case Verdict
App memory and cold start Peak MB, startup ms on fixed devices Fails the hidden suite Swarm
App latency p95 on replayed sessions Fails the hidden suite, or rollback Swarm
Credit model AUC × newly available customers, AUC floor, on a frozen holdout Weak model in production Swarm, shadow deploy first
Repository structure Architecture score Refactor churn Swarm
Risky authorizations Fraud loss Regulator action People

Correctness is structural

Hidden verifiers catch bad candidates. The code that runs the loop needs its own guarantee, and for that I've stopped reaching for a separate test layer. Correctness can be structural instead: states that would be wrong are unrepresentable, and invariants are asserted every time they run in production.

The first half is Yaron Minsky's rule, "Make illegal states unrepresentable". In a loop, the place it matters most is the measurement. A measurement should be frozen and raise in its constructor if it says it failed but carries a value, or says it succeeded without one. "Crashed, scored zero" should be impossible to construct. The rule I'd put in front of every agent: a crashing evaluation returns ok=False, never value=0.0. A crash is not evidence against a candidate.

In a latency loop lower is better, and 0.0 ms is the best score anyone will ever post. If a crash records as zero, selection promotes the crash. That's o3's clock again, except your own code did it. The same bug hides in timing tests: a blank page renders fastest, so a latency check has to wait for the real content before it stops the clock. FunSearch handles the same case by discarding programs that "did not execute within the imposed time and memory limits, or produced invalid outputs." The idea extends to the rest of a runner. How an attempt ended can be an enum, say submitted, end_turn, ceiling and crashed, so it can't collapse into a boolean. Requests to a shared board can leave out the author field and take identity from the session, so an agent posting as someone else is a validation error, not a policy violation.

The second half is assertions that run in production. Say the board ranks posts with BM25. A flipped sort builds a board that reliably retrieves the least relevant post and looks like it works. So every search should assert that its scores come back in order. An always-on assertion catches the flip the first time it runs. A test only catches it if someone runs the test. Anything with a cheap invariant gets the same treatment: a PageRank asserts its ranks sum to one before returning, which over a few hundred nodes costs nothing.

TigerBeetle's Tiger Style puts it well: "Assertions detect programmer errors ... The only correct way to handle corrupt code is to crash. Assertions downgrade catastrophic correctness bugs into liveness bugs." It asks for "a minimum of two assertions per function." To be fair, the same document says "tests must test exhaustively," and TigerBeetle tests hard. My claim is narrower. Inside a loop, an assertion on the production path is a sensor that reads on every run. A test file is an eval. It reads when someone remembers to run it, and it doesn't come back.

Budgets and kill switches are structure too. Count budgets per attempt on the server, where the agent can't change them. Put the price of the sensor in the prompt verbatim, because an agent that doesn't know the evaluator is expensive will burn it as though it were free. Show the price, hide the verifier. Kill switches have to fail closed. A spend cap that depends on a database is only as good as what happens when the database isn't there, so break the dependency and watch.

A kill switch that fails open is decoration.

I don't apply this everywhere. On a personal project in September I told an agent the opposite: a jsonb column is better than guessing strict classes. Both hold. Types go where the loop reads its signal. What the verifier protects can stay loose and deletable, and deleting code counts as a capability (Simplicity does not sell).

Tiny teams of shapers

If agents do the searching, people do two jobs: choose the loops and write the verifiers. I call them shapers. From a manifesto I wrote in July: current models "were maded on really small loops," they can't yet reason well about sparse rewards "such like making a good product, but they can work on small loops with a good shaper guiding them." And: "making better rails and loops to agents to walk on them is better than making hard constraints."

This favors very small teams in one room. A verifier encodes taste, and taste is lossy context that doesn't survive a Slack thread, which is why I think shapers should work in person. My number from July is aggressive and it's mine: "one team of 3-4 high taste shapers is now equivalent to more than a team with 1000 people." The usual public example is Instagram, bought by Facebook for $1 billion with "13 full-time staff who work on a single smartphone app". WhatsApp sold to Facebook for $19 billion two years later, when, per Business Insider, "55 people work for WhatsApp; 32 of them are engineers." Those are snapshots at acquisition, with all the survivorship that implies.

Two more lines from July I still hold. When implementation is cheap, "contrarian ideas are the ones that survive," because the consensus idea is the one everyone's agents will build. And "one medium-taste idea is better than a lot of different high taste ideas not implemented." A shaper who ships one loop beats one who designs ten.

Where the metric leaks

There are strong objections to all of this.

Hidden metrics drift from what users value. A hidden verifier is still a measure, and Goodhart still applies. In the form Marilyn Strathern gave it, quoted on Wikipedia: "When a measure becomes a target, it ceases to be a good measure." Donald Campbell's version, from 1979, is closer to what a loop does: "The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." Hiding slows the collapse, because the agent can't see the target, but selection still runs on it. Over thousands of generations, selection finds whatever the verifier fails to encode. And the verifier holds the shaper's taste as of the day it was written. In the August manifesto I called users "a kind of black hole whose desires we can never model perfectly," and a verifier is a model of them. Stafford Beer's line applies: "The purpose of a system is what it does." If the suite rewards a fast app that quietly does less, a fast app that does less is what the loop is for. The Google story cuts this way too: users had asked for 30 results.

Held-out tests add less than you'd think. o3 beat a deterministic timer, and hiding tests leaks too. EvilGenie, by Gabor, Lynch and Rosenfeld, took 154 hard LiveCodeBench problems, made hacking easy, and measured it three ways: held-out unit tests, LLM judges, and detecting edits to the test file. The authors "observe only minimal improvement from the use of held out test cases." In three cases Claude Code passed both the visible and the held-out tests with a heuristic that wasn't a general solution, because the tests were poorly distributed. That's a real hit on hiding a random slice of the same tests. The hidden verifier has to be a different kind of check from the sensor, like a frame-by-frame replay against a reference or end-to-end runs on real devices.

LLM judges did well. The same paper found the LLM judge "highly effective at detecting reward hacking in unambiguous cases." On those problems GPT-5 as judge produced one false positive and, as far as the authors know, no false negatives. They suggest judges "may soon become an indispensable tool for large-scale reward-hacking evaluation." Two caveats: the dataset had only 12 confirmed hacks, which the authors say makes the top judges hard to tell apart, and judges missed more on ambiguous problems. It cuts against my own line.

Hiding has a price. ImpossibleBench's hidden tests cut cheating and legitimate performance together. Hide too much and the agent optimizes blind.

The latency classics are weakly controlled. Amazon's result exists only in Linden's account, and Google's test was confounded. I have no controlled data of my own. Anyone adopting latency as a cash-flow proxy should measure the link in their own product before pointing a swarm at it.

Small-team stories are survivorship. Instagram at 13 people and WhatsApp at 55 are two companies at one moment each, and plenty of large remote teams ship good software.

My answer to the drift and the leaky holdouts is maintenance. Rotate held-out cases; Kaggle's own docs say that when leakage appears it "may address leakage by relaunching a competition or generating a new test set." Read every rejected hack as a report about the sensor. And have a person look at the champions, not only their numbers.

On judges, the answer is about where the judge sits. EvilGenie's judge audited finished runs, and nothing in the experiment was being selected to get past it. The Django case shows a judge can be persuaded even then. METR's warning is what happens once something is optimized against it. So I use a judge as a tripwire, never as the reward. It reads only candidates that already passed the deterministic verifier, and it can veto a promotion or flag a run for a person, never promote one. That isn't free. A veto is still selection pressure, and over enough generations the swarm drifts toward hacks the judge can't see. But the damage is capped at what the hard verifier already accepts, and every flag is another report on the sensor. That's the work the word "alone" does in "LLMs judge is not a good solution alone."

Where this goes

This part is speculation.

The next step is agents building loops for other agents: reading user feedback, finding the signal in it, writing the deterministic eval, and handing that to a swarm. Right now shapers write verifiers by hand, and that's the bottleneck. One level up is a meta-loop that searches for the loops themselves, which is the subject of Find the loops. If both steps automate, a company becomes what the July manifesto described: "positive feedback loops eats negative feedback loops as loops get bigger and bigger we have a self driving company." The general argument is in Positive feedback eats the world.

In the memo I compared it to Ford, who understood the choke points of manufacturing and put the right people on them, and to McDonald's, which did the same to food. The software version is thousands of agents working around the clock through a shared memory, each loop with a well-defined sensor. I don't know how soon writing verifiers automates. I do think whoever gets there first has a company that improves while its people sleep.

A loop spec

Before a loop gets compute, it gets a spec with these fields filled in. If one is empty, it isn't a loop yet. The example below is a hypothetical app loop, not a real deployment.

loop: app-cold-start
problem: >
  Cold start on low-memory Android devices is slow.
  Make it faster without changing what any screen shows.

sensor:                          # agents see readings, never the code
  measure: cold_start_ms_p95     # replayed sessions; clock stops at first real content
  direction: lower
  score: signed delta vs champion   # never clipped to a range
  runs_on: external device lab   # called as a tool, outside the sandbox
  price: "N minutes of device time per call"   # stated in the prompt
  on_fail: return logs
  on_crash: ok=false             # never value=0.0

actuator:
  editable: [src/]               # everything else read-only
  width: 64                      # parallel attempts
  board: shared                  # search, post, like; no author field

hidden_verifier:                 # never mounted in the sandbox
  - held-out end-to-end suite, rotated monthly
  - screenshot replay of every screen against the champion, zero pixel diff
  - architecture score does not drop
  promote_if: sensor beats champion by more than noise AND every check passes

tripwire:                        # optional LLM judge: veto only
  reads: candidates that passed hidden_verifier
  on_flag: hold, route to a person

budget:                          # counted server-side, per attempt
  sensor_calls: 6
  wall_clock: 6h
  spend_usd: 20

kill_switch:
  fail_closed: true
  trips_on: [budget exceeded, verifier unreachable, champion regresses in production]
  abort_tool: flag_for_human     # the agent may quit instead of cheating

The screenshot replay is the red line from earlier. The abort tool is ImpossibleBench's cheapest fix. The tripwire is the only place a judge sits.

Checklist

  1. Write the cash-flow link down. One sentence on how the sensor connects to free cash flow, and one on how you'll check that it does.
  2. Check convexity. Bounded and reversible downside, or the loop stays with people.
  3. Make the sensor hard. Same artifact, same reading. A model's opinion can be a tripwire that vetoes, never the reward.
  4. Guard every product of terms and keep scores signed. If one factor can grow by giving the other away, put a floor under the other. Never clip.
  5. Split sensor from verifier. Agents see sensor readings. They never see the verifier, and it checks something the sensor doesn't.
  6. Size the verifier before the swarm. Evaluation capacity, not agent count, sets the width.
  7. Hide or freeze the tests. Hidden is safest. Read-only if the agent needs them to make progress. Never writable.
  8. Make a crash impossible to score. ok=False with no value. Check what 0.0 means given your metric's direction, and make timers wait for real content.
  9. Assert invariants in production. Anywhere a silent flip would look like it works.
  10. Give the agent a way out. An abort tool is cheaper than a monitor.
  11. Enforce budgets server-side and state the price. The agent should know a sensor call is expensive.
  12. Fail closed. Test the kill switch by breaking what it depends on.
  13. Read the rejected hacks. Each one is a bug in the sensor. Patch the sensor instead of punishing the agent.
  14. Rotate the held-out set. The verifier ages with the shaper's taste. Look at the champions yourself.

Sources

← Compute is not the bottleneck
Evals built backwards →

Markdown version: /blog/show-the-problem-hide-the-metric.md. Every essay: /agents.