# Optimize the harness, not the model

*September 2026*

On HAL's [CORE-Bench Hard](https://hal.cs.princeton.edu/corebench_hard) leaderboard, the same model appears in several rows with very different scores. Claude Opus 4.5 scores 33.33% in HAL's generalist agent. In a scaffold that runs it through Claude Code, it scores 77.78%. The weights are the same in both rows. Only the harness changed, and it more than doubled the score.

It isn't a one-off. On ARC-AGI-2, [Poetiq](https://poetiq.ai/posts/arcagi_verified/) reported 54% at $30.57 per problem with a harness around Gemini 3 Pro. The previous best, Gemini 3 Deep Think, scored 45% at $77.16. And [GEPA](https://arxiv.org/abs/2507.19457), an optimizer that only rewrites prompts, "outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts." GRPO updates the weights. GEPA leaves them alone.

Most teams spend their evaluation budget on choosing a model. A new release comes out, someone runs it through the flow, looks at one number, and the team switches or doesn't. The harness is treated as a constant, written once and patched when it breaks. I think that has the levers in the wrong order. The harness is usually the bigger lever, and it's the one you own. Optimize it the way you'd train a model: against a concrete metric, by a strong proposer that sees the whole thing, with models kept swappable underneath so you can measure them on the same flow.

In [The harness is in the weights](https://future-seems-so-good.com/blog/the-harness-is-in-the-weights) I argued that part of the best harness for a model gets decided in the lab's post-training. This essay is about the rest: what you can still move, and how to move it with a loop instead of by hand.

## What counts as the harness

By harness I mean everything between the weights and the user. The system prompt. The tool list and every tool's description. The code that decides what goes into context, what gets retried, what gets repaired, when to stop. The inference stack that serves the model, including the provider, the chat template, the grammar and the decoding parameters. A model ID names a file of weights. The product is the weights plus all of that.

In most planning documents, only the weights count as the model. The rest gets called plumbing, and plumbing doesn't get a metric.

![Score gains from a model swap with the harness fixed, against a harness change with the model fixed](https://future-seems-so-good.com/blog/assets/optimize-the-harness-not-the-model/charts/harness-vs-model.svg)

The chart puts two public comparisons side by side. For each one, the grey bar is the best model swap available with the harness fixed, and the black bar is a harness change with the model fixed. Both are in percentage points, but the benchmarks differ, so compare bars within a row, not across rows.

**CORE-Bench Hard.** Holding the scaffold fixed at CORE-Agent, the best Claude swap from Opus 4.5 (42.22%) is to Opus 4.1 (51.11%), a gain of 8.9 points. Holding the model fixed at Opus 4.5, moving from HAL's generalist agent to Nicholas Carlini's Claude Code scaffold gains 44.4 points, and HAL notes it reaches 95.5% "after fixing a few grading errors via manual scoring." HAL declared the benchmark solved. The newer model does worse than the older one in the same scaffold, which is its own warning about shopping by release date.

**ARC-AGI-2.** ARC Prize's [2025 results analysis](https://arcprize.org/blog/arc-prize-2025-results-analysis) created a leaderboard category called Model Refinements for this. Poetiq's harness "improves performance on ARC-AGI-2 from baseline 31%, $0.81/task up to 54%, $31/task" on the same Gemini 3 Pro. Moving to a stronger model configuration, Gemini 3 Deep Think, got 45%. Poetiq describes its method as a "meta-system to optimize all parts of our solution," and says "we do not need to build, or even fine-tune, our own large frontier models."

**Edit formats.** Stencil changed only the edit tool and [moved Grok Code Fast 1 from 6.7% to 68.3%](https://stencil.so/blog/the-harness-problem), "because patch was failing so catastrophically that the model's actual coding ability was almost completely hidden behind mechanical edit failures." I covered that result and its replications in the previous essay, so I'll only use the sentence: a model's ability can be hidden behind its harness.

**Formatting alone.** Sclar and coauthors ([Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design](https://arxiv.org/abs/2310.11324)) found "performance differences of up to 76 accuracy points when evaluated using LLaMA-2-13B" from meaning-preserving changes to prompt format. That's the variance you leave on the table if nobody measures the prompt.

Leaderboards already know this, and they handle it by maximizing it away. The [Terminal-Bench 2.0 paper](https://arxiv.org/abs/2601.11868) says of its headline figure: "The agent scaffold used to report each model was chosen to maximize performance." Every model number you read is a model number plus somebody else's harness search.

## Optimize the whole prompt, with a strong proposer

Most teams move the harness by hand: someone reads a failed conversation, adds a sentence to the system prompt, ships. That's gradient descent with a batch size of one and a human as the optimizer.

The better tool is a prompt optimizer that reads trajectories. [GEPA](https://arxiv.org/abs/2507.19457) (Agrawal et al.) is the clearest published version. Given "any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts."

The numbers are why I use it. On Qwen3 8B across six tasks:

| Method | Aggregate score | Rollouts per task |
|---|---|---|
| Baseline prompt | 45.23 | 0 |
| MIPROv2 | 47.84 | not reported in the table |
| GRPO (weight updates) | 48.91 | 24,000 |
| GEPA (prompt only) | 54.85 | about 3,936 on average |

That's the table behind the abstract's 6% average and 35x fewer rollouts. On HotpotQA it took the baseline from 42.33 to 62.33, while GRPO reached 43.33. Rewriting the prompt beat training the weights, at a sixth of the rollouts.

It also transferred. Prompts GEPA optimized on Qwen3 8B, evaluated unchanged on GPT-4.1 Mini, gave "a +9.00% aggregate improvement across 6 benchmarks," beating MIPROv2's +5.64% even though MIPROv2 had optimized directly on GPT-4.1 Mini. A harness found on one model isn't worthless on the next one.

Three rules I'd hold to when running this kind of loop.

**Optimize the whole prompt.** In June I wrote in my notes: "the point is that GEPA needs the whole prompt; optimizing only one part won't lead to a result for the prompt as a whole." A prompt's sections interact. The tone section changes how the model reads the tool section, and the tool section changes what the tone section has to cover. Freeze everything except one block, and the optimizer finds the best block for a prompt that's otherwise wrong. Poetiq's phrase, "optimize all parts of our solution," is the same instinct at a larger scope.

**Give the proposer room.** GEPA's quality depends on the model writing the proposals and on how long it's allowed to search. I'd budget the proposer in the hundreds or thousands of steps, not dozens. Proposals are cheap compared to the evaluations that score them, so a short search saves little and leaves gains unfound. If an optimization run plateaus, the proposer's budget is the first thing I'd raise.

**Run it offline against concrete metrics.** The loop should run against mocked tools with deterministic scores, so a thousand steps don't touch a user or a production API. The metric decides what to try. A held-out split the optimizer never saw, and then a small canary, decide whether to ship it.

## Prompt engineering is dead, again

![Hacker News stories with "prompt engineering" in the title per year, and items saying "prompt engineering is dead"](https://future-seems-so-good.com/blog/assets/optimize-the-harness-not-the-model/charts/hn-prompt-engineering.svg)

Hacker News titles with "prompt engineering" went from 19 in 2022 to 190 in 2023, then fell to 101 in both 2024 and 2025 and 36 so far this year. The phrase "prompt engineering is dead" doesn't appear in the search index before 2024. It shows up 2 times that year, 5 in 2025, and 4 so far in 2026.

The recent ones say the same thing. [One June comment](https://news.ycombinator.com/item?id=48454500) put it as: "Prompt engineering is dead. The cool kids are making Claude prompt itself. They're writing loops, not prompts," and then quoted Boris Cherny: "I don't prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops." The dissent is older. In 2024, [another commenter](https://news.ycombinator.com/item?id=40403272) wrote: "Whenever 'AI influencers' say 'prompt engineering is dead', I realize they are full of bs. If anything, prompt engineering has become more nuanced, down to the choice of each token."

Both are right, and the GEPA table shows how. Prompts matter more than ever. The 76-point formatting spread and the 20-point GEPA gains say so. What's dying is the human writing them one sentence at a time. The prompt engineer moved into a loop, and the loop writes better prompts than the person did.

## The prompt is one visible file

That makes the prompt more of a crafted object, because now two readers need to see it: the optimizer and the person reviewing what the optimizer did.

My rule is that the whole prompt lives in one file. In June I told an agent: "I don't want you to build the prompt in layers, I want it in one file." Layered assembly stitches a base prompt, a persona prefix, a tool block and conditional sections together at runtime. Then nobody has ever read the prompt the model actually sees. The optimizer can't rewrite what it can't see as a whole, and a reviewer can't judge a diff scattered across six template fragments.

I've also moved from short, token-perfect prompts to long ones. In September I wrote: "we should incentivize more biggest prompts with more signal not discretizing or overfitting the task into one specific format." Agents mostly fail from missing context, not from too many words. A prompt that explains the domain, the user and the reason behind each rule lets the model handle cases the rule didn't list. A prompt that compresses the task into a rigid format gets you a model that follows the format and misses the point.

What the receiving agent does with a prompt is the score for whoever wrote it, human or model. That's the same standard GEPA applies, and I think it's the only honest one.

The mirror image of long prompts is short code. I don't like harnesses that box the model in with hand-written gates, if-branches, host allowlists and bespoke tools. In June, reading a harness full of them, I wrote: "im hating this code style of having thousands of suboptimal implementations of gates ... take them out of the code." Each gate is a decision someone made once, in advance, for every future case. Context tells the model what matters and lets it decide. A gate decides for it and never learns. When trajectories fail, I now suspect the harness before the model, and usually the harness is guilty.

This is the split I described in [Bottom-up is one ontological level higher](https://future-seems-so-good.com/blog/bottom-up-is-one-ontological-level-higher): evaluative structure should grow and prescriptive structure should shrink. The optimizer and the verifier are evaluative. The gates are prescriptive.

## Keep models swappable and measure them on the same flow

The model still matters. You can only see how much if you can swap it cheaply and measure it on the same flow.

In July I wrote the rule down plainly: "The model must not be hardcoded in the flow." An agent should take the model as configuration. When a new release comes out, run open and closed candidates through the same real flow with the same harness and the same scorer. Compare with repeated runs and intervals, not one lucky score, and stay on the incumbent until an alternative is proven. Without that process, a model switch is a feeling instead of a number.

Swappability has a catch that Sclar's paper names: "format performance only weakly correlates between models, which puts into question the methodological validity of comparing models with an arbitrarily chosen, fixed prompt format." A harness optimized on one model is a harness optimized for that model. Running a second model through it measures the second model's fit to the first one's prompt. So the fair comparison runs the optimization loop for each candidate, or optimizes one harness against several models at once and keeps the Pareto set. GEPA's cross-model transfer result says the gap is smaller than you'd fear. Sclar's says it isn't zero.

Measure the product, not the method. In September I put it as: the system and its eval "should be harness agnostic." Whether search uses one API or another, or whether context gets compacted, shouldn't matter to the scorer. Only the final result should. That's what lets you change the harness freely. If the eval checked the method, every harness change would break the eval. I make the longer case in [Evals built backwards](https://future-seems-so-good.com/blog/evals-built-backwards).

## Token quality is part of the harness

Most teams stop at the model ID. For open weights, the model ID doesn't tell you what you're running.

Moonshot's [K2 Vendor Verifier](https://github.com/MoonshotAI/K2-Vendor-Verifier/blob/main/README.md) tests providers serving the same Kimi K2 weights "over a set of 4,000 requests," comparing each against the official API. Its explanation of why: "When selecting a provider, users often prioritize lower latency and cost, but may inadvertently overlook more subtle yet critical differences in model accuracy." In its November 2025 table for kimi-k2-0905, the share of tool calls whose JSON passed schema validation was:

| Serving stack | Schema accuracy |
|---|---|
| Moonshot official API | 100.00% |
| DeepInfra, Fireworks, NovitaAI (via OpenRouter) | 100.00% |
| vLLM | 76.00% |
| SGLang | 73.13% |
| Together (via OpenRouter) | 71.96% |

Same weights. A quarter of vLLM's tool calls failed the schema in that test. Groq passed 100% of schemas but scored 69.52% on trigger similarity, below the 80% Moonshot treats as acceptable. It decided whether to call a tool differently from the official API. These are one date's numbers and vendors have patched since. That's the point: you can't know unless you test.

Moonshot's later [Kimi Vendor Verifier post](https://www.kimi.com/blog/kimi-vendor-verifier.html) found that "a significant portion of these cases stemmed from the misuse of Decoding parameters," and concluded: "The more open the weights are, and the more diverse the deployment channels become, the less controllable the quality becomes." Their suite includes AIME2025 as a long-output stress test that, in their words, "Catches KV cache bugs and quantization degradation that short benchmarks hide."

Grammars do the same thing at decode time. A vLLM pull request tried to align GPT-OSS's strict tool-call grammar with the order the chat template renders. The model generates a different order. A maintainer ran BFCL multi-turn and found the branch ["gets `0%` accuracy while main gets `~55%`"](https://github.com/vllm-project/vllm/pull/51020). The PR was closed, and a lenient fix kept 54.5%. A few grammar patterns were worth the whole benchmark.

I think this is where attention is most lopsided. Serving work gets measured in tokens per second, because that's what serving dashboards show. The quality of those tokens mostly isn't on the dashboard. If a quantization change or a template update cost three points on an agent's eval, nobody on the serving side would see it. An optimization that raised throughput 20% and cost 5% of tool-call accuracy would be a regression that looks like a win.

The fix is small: a script that runs a fixed set of quality checks against any endpoint and prints per-check scores plus one index against the published reference. It belongs in the same CI as the throughput benchmark. There's a sketch at the end.

## Hedge the tail

Latency is also a harness property. For routes where a slow answer is a failed answer, I'd use a hedged request: send it to one provider, and if nothing has come back after a set delay, start the same request on a second provider without cancelling the first. Whichever answers first wins.

This is an old idea. Jeffrey Dean and Luiz André Barroso described it in [The Tail at Scale](https://www.barroso.org/publications/TheTailAtScale.pdf) (2013): "A simple way to curb latency variability is to issue the same request to multiple replicas and use the results from whichever replica responds first." Their tuning advice was to "defer sending a secondary request until the first request has been outstanding for more than the 95th-percentile expected latency for this class of requests. This approach limits the additional load to approximately 5% while substantially shortening the latency tail." In a Google benchmark reading 1,000 BigTable keys, hedging after 10ms cut the 99.9th-percentile latency "from 1,800ms to 74ms while sending just 2% more requests."

Two adjustments for model providers. First, set the hedge delay from your measured p95 time to first token per route. A round number like five seconds is a guess. The p95 is a measurement. Second, hedge only across providers that pass the quality check. The fastest replica in the K2 table might be the one that gets a quarter of its tool calls wrong. A hedge that picks the first answer from a bad provider trades latency for correctness without telling you.

## Generate freely, verify deterministically, repair with a model

The optimizer needs a metric, and the metric has to be one it can't talk its way around. That's the other half of the harness.

Take any output with hard constraints: a message with a character limit, a JSON object that has to validate. The obvious harness truncates or reformats whatever the model writes. I don't like that. A deterministic fix produces text nobody wrote, a sentence cut mid-word or a field silently dropped, and it hides the failure from the metric that should have seen it.

So I'd separate three jobs. The model generates freely, without the constraints crowding its attention. Deterministic code checks every rule. When a check fails, the exact error goes back to the model as a user message and it rewrites, and the loop repeats until nothing fails. It works like a compiler loop. The error message matters as much as the check: in [The harness is in the weights](https://future-seems-so-good.com/blog/the-harness-is-in-the-weights) I argued for errors that fix the next call, and the same holds here.

The checker is code, never a model. "I don't want LLM as a judge, the eval should only have deterministic evaluations," I wrote in September. A model judge can be persuaded, and an optimizer running a thousand steps will find the persuasion. A length check can't be persuaded.

This pattern and the optimization loop fit together. The deterministic checks that gate production are the same checks the optimizer climbs, and the repair loop's retry count can be a second signal: a harness that needs three repair rounds per message is worse than one that needs none, even if both end valid. I cover the measurement side in [Evals built backwards](https://future-seems-so-good.com/blog/evals-built-backwards) and [Show the problem, hide the metric](https://future-seems-so-good.com/blog/show-the-problem-hide-the-metric).

## Where the harness loses

The strongest case against all of this is Rich Sutton's [bitter lesson](http://www.incompleteideas.net/IncIdeas/BitterLesson.html): "general methods that leverage computation are ultimately the most effective, and by a large margin." Building in human knowledge "always helps in the short term," then "plateaus and even inhibits further progress." On that reading, harness engineering is the chess heuristic of 2026. Labs will train the behavior into the next model and your harness becomes dead weight.

ARC Prize says as much about the harnesses on its own leaderboard: "We expect these types of general refinement and harness improvements to eventually make their way 'behind the API' of commercial AI systems."

The rest of the case is specific.

**Weight updates still win some tasks.** GEPA lost to GRPO on AIME-2025, 32.00 against 38.00, the one benchmark in the table where it did. Deep math may need what only weight updates give. A prompt can't teach a model arithmetic it doesn't have.

**Optimizers overfit the eval they climb.** If formatting alone swings 76 points, an optimizer can find formats that please a narrow benchmark and nothing else. Any gain an optimizer reports on the benchmark it climbed is suspect, including any I'd report myself. That's why the benchmark needs a held-out split the optimizer never sees, and why a win on the holdout still goes to a small canary before it goes to all traffic.

**Harness gains shrink on the models built for their harness.** In Stencil's table, hashline against `str_replace` gave GPT-5.2 Codex −0.4 points and 26% more output tokens. On CORE-Bench, Opus 4.1 did worse in Claude Code (42.22%) than in CORE-Agent (51.11%). A harness advantage is a fit between one harness and one model, and the fit can go negative.

**Some harness gains are expensive.** Poetiq's refinement took Gemini 3 Pro from $0.81 to $31 per task, about 38 times the cost, for 23 points. That's a good trade on ARC and a bad one for a chat reply that has to arrive in seconds.

I accept most of this, and it changes how I do the work, not whether. The bitter lesson is aimed at hand-built knowledge, and the gates I keep deleting are exactly that. A harness found by search over trajectories, scored by a verifier, is the other side of Sutton's line: a general method that gets better with more computation. The loop is what lasts. When a lab moves an improvement behind the API, the loop re-runs on the new model and finds the next gap. The prompt it found last month was never the asset.

The cost point splits in two. Poetiq-style refinement spends at inference time, on every task. GEPA-style optimization spends once, offline, and ships a prompt that costs about the same to run as the one it replaced. The loop I described above is the second kind.

## Where this goes

This part is speculation, and I'll mark where it stops being evidence.

GEPA edits prompts. Tool descriptions, repair loops, context policy and routing are still written by people. The obvious next step is a proposer that edits all of it, with the harness as code and the eval as the only fixed point. In August I wrote that "we should put a direction of unhobbling to the models, self harness," and the question I keep returning to is the ceiling of prompt-only optimization against a model rewriting the entire harness.

There's early evidence. [Automated Design of Agentic Systems](https://arxiv.org/abs/2408.08435) (Hu, Lu and Clune) has a meta agent program new agents in code, "including novel prompts, tool use, workflows, and combinations thereof," and reports "that agents invented by Meta Agent Search maintain superior performance even when transferred across domains and models." Its premise is the bitter lesson applied to scaffolds: "hand-designed solutions are eventually replaced by learned solutions."

Here's where evidence stops. My guess is that within a year the best harnesses for narrow, well-measured tasks will be written mostly by proposers, and the human job will be the eval, the constraints and taste about what the eval misses. I don't know how much a code-level proposer adds over a prompt-level one. I haven't run that experiment.

## A harness optimization loop

Here's the loop as I'd set it up today, as a spec an agent can implement.

```yaml
artifact:
  harness: harness/agent.md          # whole system prompt + every tool description, one file
  code: harness/policy.ts            # context, retry and repair policy (optimized later, see speculation)
  models: [incumbent, challenger]    # open or closed; configuration, never hardcoded

eval:
  tasks: evals/tasks.jsonl           # real flows, weighted by production traffic
  split: { train: 0.6, pareto: 0.2, holdout: 0.2 }   # optimizer never reads holdout
  tools: mocked                      # recorded responses; no production side effects
  scorer: deterministic              # asserts on outputs, schemas, DB rows; no model judges
  runs_per_candidate: 3              # report mean and interval, not one lucky run

proposer:
  model: strongest-available         # can differ from the production model
  reads: [failed trajectories, scorer errors, current artifact]
  edits: whole artifact              # never one section with the rest frozen
  steps: 1000
  selection: pareto                  # keep any candidate that is best on some task

gate:
  holdout_gain: "> 2x run-to-run std"
  per_model: true                    # re-score each candidate on every model in `models`
  then: canary on a small slice of traffic, compare the product metric against incumbent
  incumbent_wins_ties: true
```

And a sketch of the provider-quality check, the script any serving team should run next to its throughput benchmark:

```python
import json, statistics, time
from jsonschema import validate, ValidationError
from openai import OpenAI

REFERENCE = {"tool_schema": 1.00, "tool_trigger_f1": 0.84, "long_output": 0.95}

def tool_checks(client, model, cases):
    ok = triggered = tp = fp = fn = 0
    for case in cases:
        r = client.chat.completions.create(model=model, messages=case["messages"], tools=case["tools"])
        called = r.choices[0].finish_reason == "tool_calls"
        tp += called and case["expect_call"]
        fp += called and not case["expect_call"]
        fn += (not called) and case["expect_call"]
        if not called:
            continue
        triggered += 1
        for call in r.choices[0].message.tool_calls:
            schema = next(t["function"]["parameters"] for t in case["tools"]
                          if t["function"]["name"] == call.function.name)
            try:
                validate(json.loads(call.function.arguments), schema)
                ok += 1
            except (ValidationError, json.JSONDecodeError):
                pass
    f1 = 2 * tp / (2 * tp + fp + fn) if tp else 0.0
    return {"tool_schema": ok / max(triggered, 1), "tool_trigger_f1": f1}

def long_output(client, model, problems):
    right = 0
    for p in problems:
        r = client.chat.completions.create(model=model, messages=[{"role": "user", "content": p["q"]}],
                                           max_tokens=32000)
        right += p["answer"] in r.choices[0].message.content[-200:]
    return {"long_output": right / len(problems)}

def check(base_url, api_key, model):
    client = OpenAI(base_url=base_url, api_key=api_key)
    cases = [json.loads(l) for l in open("checks/tool_calls.jsonl")]
    problems = [json.loads(l) for l in open("checks/long_output.jsonl")]
    start = time.time()
    scores = tool_checks(client, model, cases) | long_output(client, model, problems)
    ratios = [min(scores[k] / REFERENCE[k], 1.0) for k in REFERENCE]
    scores["index"] = statistics.geometric_mean(ratios)
    scores["seconds"] = round(time.time() - start)
    return scores

if __name__ == "__main__":
    import sys
    print(json.dumps(check(*sys.argv[1:4]), indent=2))
```

The index is a geometric mean of each score against its reference, so a stack that's perfect on schemas and broken on long outputs can't average its way to a pass. Run it on every serving change, on every provider you route to, and before a provider joins the hedge list. The first two reference values are what Moonshot published for K2-0905's official API. The long-output value is a placeholder. Replace all three with numbers from your model's official API.

The order I'd do this in, if I were starting over: make the model a config value, write the deterministic scorer, move the prompt into one file, run the provider check, then turn on the optimizer. Each step makes the next one measurable. The optimizer comes last because it's only as good as the metric it climbs.

## Sources

- HAL, [CORE-Bench Hard leaderboard](https://hal.cs.princeton.edu/corebench_hard), Princeton (retrieved 2026-09-25)
- ARC Prize, [ARC Prize 2025 results and analysis](https://arcprize.org/blog/arc-prize-2025-results-analysis)
- Poetiq, [Poetiq Shatters ARC-AGI-2 State of the Art at Half the Cost](https://poetiq.ai/posts/arcagi_verified/), 2025-11-20
- Stencil, [We improved 15 LLMs at coding in one afternoon. Only the harness changed](https://stencil.so/blog/the-harness-problem), 2026
- Agrawal et al., [GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning](https://arxiv.org/abs/2507.19457), 2025
- Sclar, Choi, Tsvetkov and Suhr, [Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design](https://arxiv.org/abs/2310.11324), 2023
- [Terminal-Bench 2.0](https://arxiv.org/abs/2601.11868), 2026
- Moonshot AI, [K2 Vendor Verifier README](https://github.com/MoonshotAI/K2-Vendor-Verifier/blob/main/README.md), test date 2025-11-15; [Rebuilding the "Chain of Trust": Kimi Vendor Verifier](https://www.kimi.com/blog/kimi-vendor-verifier.html)
- vLLM, [GPT-OSS strict tool-call grammar PR #51020](https://github.com/vllm-project/vllm/pull/51020), 2026
- Jeffrey Dean and Luiz André Barroso, [The Tail at Scale](https://www.barroso.org/publications/TheTailAtScale.pdf), Communications of the ACM, 2013
- Rich Sutton, [The Bitter Lesson](http://www.incompleteideas.net/IncIdeas/BitterLesson.html), 2019
- Hu, Lu and Clune, [Automated Design of Agentic Systems](https://arxiv.org/abs/2408.08435), 2024
- Hacker News comments: [item 48454500](https://news.ycombinator.com/item?id=48454500) (2026-06-09), [item 40403272](https://news.ycombinator.com/item?id=40403272) (2024-05-19); yearly counts from the [Algolia HN Search API](https://hn.algolia.com/api), exact-phrase queries, 2026-09-25
- My own notes and agent conversations, June to September 2026, quoted with dates (Portuguese passages translated)
