Money is method times compute times data

September 2026

In August 2010, UBS published its preview of Walmart's second quarter with an unusual sentence in it: "UBS proprietary satellite parking lot fill rate analysis points to an interesting cadence intra-quarter and potential upside to our view." The analyst, Neil Currie, had bought the data from Remote Sensing Metrics, a two-year-old Chicago firm that counted cars in the parking lots of more than 100 Walmart stores every month. June traffic was 4% above June of the year before. UBS's traditional model said quarterly sales would be down 1% year over year. The regression on the parking lots said up 0.7%. UBS printed both forecasts side by side.

Thirty years earlier, Sanford Grossman and Joseph Stiglitz had explained why a report like that can exist. In On the Impossibility of Informationally Efficient Markets they wrote: "because information is costly, prices cannot perfectly reflect the information which is available, since if it did, those who spent resources to obtain it would receive no compensation." If prices told you everything, nobody would pay to count cars, and then prices would stop telling you anything.

I think those two documents hold most of what matters about making money in markets, and both usual readings of them are wrong. One camp treats Grossman-Stiglitz as a footnote to efficient markets: the gap is tiny, buy the index. The other treats the parking lot as a story about data: find an exotic dataset and you win. My view is that money comes down to data. But data is worth nothing until compute turns it into something that makes an asymmetry visible. And compute is worth nothing without a method that decides what to compute and refuses to believe the answer until it survives an attack. The edge is the product of the three, measured against the product your competitors have.

In July I wrote the first version of this down more crudely, in a note to an agent: "money = compute + data = data ... compute is a way to transform raw data into valuable data that transforms into making assymetries look clear". The plus sign was wrong. The factors multiply, and the part that pays is the difference:

edge ≈ (method × compute × data)_you − (method × compute × data)_marginal competitor

None of this is investment advice. I build agent harnesses for a living, and this essay is about how I think edges get made and how they die. For the specific bets I'd make on the AI buildout, see The AI trade nobody priced.

Prices are a partial readout

The Grossman-Stiglitz model is simple. Some traders pay to learn something about an asset's true value. Their buying and selling moves the price, so the price leaks part of what they learned to everyone else. The authors call the result "an equilibrium degree of disequilibrium: prices reflect the information of informed individuals (arbitrageurs) but only partially, so that those who expend resources to obtain information do receive compensation." And the degree of leakage isn't fixed: "How informative the price system is depends on the number of individuals who are informed; but the number of individuals who are informed is itself an endogenous variable in the model."

That second sentence is the one I care about. How many people do the work sets how informative the price is. The paper's closing line states the tension outright: "There is a fundamental conflict between the efficiency with which markets spread information and the incentives to acquire information."

The parking lots turned into a clean test of it. Katona, Painter, Patatoukas and Zeng got RS Metrics' feed, 4.7 million daily observations across 67,078 store locations for 44 major U.S. retailers between 2011 and 2017, and built a strategy that bought retailers whose same-store lot fill grew fastest and shorted the ones where it fell. Over the three days around earnings, the buy-minus-sell spread was 4.76% after factor adjustment, and the short leg did most of the work. Their conclusion: "unequal access to big data can increase information asymmetry among market participants without immediately enhancing price discovery."

The same paper has the detail that I think explains the whole thing. Sam Walton used to fly over his own parking lots to count cars. The idea was never secret. In the authors' words: "While anyone could count cars in a parking lot, advances in computer vision and the increased availability of satellite imagery has enabled daily tracking and processing of retailer parking lots at scale." The data was always public, in the sense that the cars were sitting outside. What changed was the cost of computing on it.

Patrick McKenzie (patio11) made the legal version of this point on Hacker News in 2014: "You can create non-public information in arbitrarily complicated ways." His example was satellites over Walmart. The reply under it drew the line more precisely: "the difference between non-public information and new analysis on public information." Almost every legitimate edge lives on the second side of that line. You don't find the asymmetry. You make it, out of raw material anyone could have bought.

Depth is what the market misses

The argument I keep coming back to is shaped like one of Gödel's. The more derivation steps an insight needs, the fewer people see it. Most public information is one step from its price: a revenue number comes out, the stock moves. Some conclusions follow from public information too, but only after twenty steps, and the number of people who run all twenty is small. Compute reaches the insights others miss because it can afford the steps.

Two ideas from logic sharpen this. I'm using them as analogies, not as theorems about markets.

Gödel's speed-up theorem. In 1936 Gödel showed that there are theorems whose proofs can be drastically shortened by working in more powerful axiomatic systems. A statement can be short, true and provable in Peano arithmetic, yet have a shortest proof longer than a googolplex symbols, while a slightly stronger system proves it in a paragraph. The statement is available to everyone. The proof isn't. And a better set of axioms is the difference between an impossible derivation and an easy one.

Bennett's logical depth. Kolmogorov complexity measures the length of the shortest program that produces an object. Charles Bennett proposed a second measure: how long that near-shortest program takes to run. In his words, the run time "measures the object's logical depth, or plausible amount of computational work required to create the object." His slides give the contrast: "A trivially orderly sequence like 111111… is logically shallow because it can be computed rapidly from a short description. A typical random sequence, produced by coin tossing, is also logically shallow, because it essentially [is] its own shortest description, and is rapidly computable from that. Depth thus differs from Kolmogorov complexity or algorithmic information, defined as the size of the shortest description, which is high for random sequences."

Map that onto a market. Noise is shallow: nothing to derive. A headline number is shallow: one step to the price. The valuable things are deep: a short description (the filings, the transcripts, the satellite frames) that needs a lot of computation to unfold into a conclusion. In this picture compute is the run time you can afford and method is the axiom system. Method decides whether a derivation takes twenty steps or two thousand.

There's evidence for the market half of this. The current version of Lopez-Lira and Tang's Can ChatGPT Forecast Stock Price Movements?, revised in August 2026, maps where language models add predictive power over simpler ones: "Markets efficiently process transparent, quantifiable information (e.g., earnings reports and clinical trials) but systematically underreact to information requiring complex synthesis (e.g., insider transactions and specialized conference presentations)." GPT-4 beats GPT-3.5 on the synthesis-heavy categories and ties it on formulaic announcements. Their summary: "markets struggle precisely where reasoning capacity is most scarce." That's the derivation-depth argument, measured.

This is also why I think of markets as games. The working definition I use for intelligence is the ability to play games, which I develop in Hire people who close loops. The market is a crowded, adversarial game over asymmetric information, and in a crowded game being right isn't enough. The note I wrote the day after the first one: "the best stocks are defined not by the thesis but a sum of thesis and assymetry of market expectations and forecast". A correct thesis that everyone shares is already in the price. What pays is the distance between your forecast and the forecast the price implies, and your reason for trusting yours. That changes the task you give a model. Instead of summarizing a company, it builds a causal model of the business from the texts, states what the price assumes, and then tries to break its own model.

Three factors, multiplied

Data: raw and point-in-time. Summarized data gives a research agent very little to reason about. Indicators, the derived series every terminal shows, are someone else's compression, and everything the compression threw away is where a deep derivation would have started. My expectation is that a swarm fed only indicators finds almost nothing, but I haven't measured it. So the data layer I want holds as much raw material as possible: fundamentals, filings, earnings calls, podcasts, news, satellite imagery. Every record carries the time it first became public, because a record you couldn't have seen on the decision date is worse than no record.

This matches what I argued about AI in How to achieve superintelligence: "We are building industrial systems for compute and cottage industries for data." In markets it's the same asymmetry. Anyone can rent the compute. The point-in-time archive is the part that takes years.

Compute: the run time you can afford. This is the factor that's changing fastest, and I come back to it below. For now: if depth is the run time of a derivation, a cluster of agents is a way to run many derivations in parallel and keep only the ones that survive. I argue in Compute is not the bottleneck that most agent systems under-explore. They converge on the first plausible answer. In markets the first plausible answer is the consensus, and the consensus pays nothing.

Method: what makes compute worth anything. Without method, more compute buys more overfitting. The rules I work by are ordinary forecasting hygiene, and they carry over to markets unchanged. Establish a baseline first, and beat the strongest trivial baseline by more than 10% before calling anything a signal. Use temporal splits. Audit point-in-time features for leakage. Run permutation tests and bootstrap confidence intervals. End with an explicit SHIP or DON'T SHIP. And treat a clean null as a success. I write this into the research prompts I give agents: "Bias toward rigor over a positive result — a well-supported 'nothing here' is a fine outcome for any given signal and should be reported as such, not softened." More recently, shorter: "Your job is to test it, not to prove it. A null or negative result, clearly measured, is a successful run."

Multiplication is the right operator because a zero anywhere kills the product. A perfect archive nobody computes on earns nothing. A huge cluster running a sloppy method manufactures false positives faster. A brilliant method with nothing but indicators to work on has nothing to derive from.

An adversarial research swarm

The design I'd build for equity research looks like this. Ten generators, each unbiased and independent, write analyses of the same company. Three adversarial verifiers check every number in every analysis against a dated source and attack the thesis. A condenser merges what survives into one document, and later rounds can refine the surviving analyses in loops for as long as the budget allows. The target universe is the global top 100 companies, on open-weight models, because their weights don't change between runs and every step can be rerun and audited. This is a proposal, and I haven't run it.

Two rules sit above everything else in the prompt: the agents must never lie, and whenever one finds an opportunity it has to flag the risk next to it. The first sounds obvious, but a generator rewarded for a coherent story has every reason to invent a margin that makes the story close. The second exists because a thesis without risks written beside it is a sales pitch.

The shape follows from a belief I spell out in Evals built backwards: generation is cheap and verification is where the value is. Ten generators give variety, and variety is how you find the rare deep derivation among the shallow ones. The verifiers are what make the variety safe. A generator is rewarded for a good story. A verifier is rewarded for finding the number that isn't in the filing.

The finance literature adds two requirements I didn't have at first.

Strip the names. Glasserman and Lin tested GPT sentiment on news headlines with and without company identifiers and found something they didn't expect: "In-sample (within the LLM training window), we find, surprisingly, that the anonymized headlines outperform, indicating that the distraction effect has a greater impact than look-ahead bias." The effect "is particularly strong for larger companies — companies about which we expect an LLM to have greater general knowledge." For a swarm working on the global top 100, that's the whole universe. The model's general impression of Apple leaks into its reading of Apple's news. So at least one verifier should see the anonymized version of the evidence and reach its conclusion without knowing whose it is.

Pin the model to the decision date. A model evaluated on dates inside its training data may simply remember what happened. Detecting Lookahead Bias in LLM Forecasts measures this with a date-only query: firm name, ticker, date, no content. The resulting Lookahead Propensity "is materially positive throughout the in-sample period and collapses essentially to zero right after the training-data cutoff." DatedGPT goes further and trains twelve 1.3B-parameter models from scratch, each with a hard annual cutoff from 2013 to 2024. Models whose training covered the outcome period earned a "lookahead premium of 26.4 b.p. per standard deviation, significant at the 1% level," and the bias-free setup still reached an annualized Sharpe of 3.20 on 61,000 firm-day headlines. So the signal is real, and so is the contamination. A backtest that doesn't separate the two measures the model's memory.

An analytical fund

On August 1 a friend asked the question that turned this from a research harness into a thought experiment about a fund. He'd been reading a trader on X who had made millions buying and selling stocks on subjective analysis instead of quant models, and he wrote (my translation): "I wonder if we could build an analytical (not quant) fund and perform like crazy." He guessed yes. The idea is his, from a private conversation between friends.

When I asked who he'd hire, his answer was better than mine: "one thing I've noticed is that it's hard to find people who are high agency and low cortisol, but for this kind of work maybe another profile is better (less CEO-like and more analytical/lethargic/static, like gwern)."

I think he's right, and the depth argument explains why. A quant fund in the Renaissance mold runs shallow derivations at enormous scale: millions of small, fast bets, each a little better than a coin flip. An analytical fund runs a few very deep ones: geopolitics, energy, nationalization, a supply chain several steps upstream of the company everyone is watching. The long run time is where the edge lives, and position sizes stay small because only a few theses ever reach the end of the derivation. For that you want patience over speed: someone who can hold a question for months without needing it to resolve. That's the gwern profile.

What a swarm changes is the cost. The analytical fund used to be capped by how many patient people you could find. With ten generators and three verifiers per company, the patient reading is cheap. What stays expensive is the method: choosing questions, deciding what counts as evidence, and not fooling yourself. My reply that night was that the timing was right. Faria Lima, Brazil's financial district, still doesn't understand AI. And, paraphrasing Nick Land's Meltdown, the bubbles will keep getting bigger, so funds that know how to play the volatility will make orders of magnitude more than before. The first half of that is an observation about Brazil. The second is a bet, and I've filed it under speculation below.

When intelligence is cheap

The factor that's changing fastest is the price of compute per unit of intelligence. Epoch AI tracked the cheapest model that matched fixed performance levels on six benchmarks and found that "the price to achieve GPT-4's performance on a set of PhD-level science questions fell by 40x per year", with rates between 9x and 900x per year depending on the benchmark and threshold. a16z's LLMflation measured the same thing on MMLU: GPT-3 cost $60 per million tokens in November 2021 when it was the only model at that level, and three years later Llama 3.2 3B matched it for $0.06. "The cost of LLM inference has dropped by a factor of 1,000 in 3 years."

Price of a fixed capability level, Epoch GPQA series and a16z MMLU endpoints

Go back to the formula. If everyone's compute term grows by 10x to 40x a year, the difference between you and the marginal competitor doesn't grow with it. For shallow tasks it shrinks, because the cheapest model is now good enough and everyone has it. Lopez-Lira and Tang saw this happen inside their own sample. The annualized Sharpe of their ChatGPT headline strategy dropped "from 6.54 in 2021Q4 to 3.68 in 2022, 2.33 in 2023, and to 1.22 over January-May 2024." The abstract says it plainly: "Strategy returns decline as LLM adoption rises, consistent with improved price efficiency." That's Grossman-Stiglitz running in real time. The cost of becoming informed fell, more traders became informed, and the price absorbed what they knew.

So cheap cognition moves the edge in two directions. It moves it deeper, toward derivations long enough that even cheap compute has to be pointed well, which is method. And it moves it toward whatever the cheap cognition can't produce. The question I wrote down this week: "If intelligence becomes nearly free, what remains scarce? ... Who owns those scarce things?"

My list:

The owners of those things are already moving. In October 2025, Aramco signed a non-binding term sheet to take a significant minority stake in HUMAIN, the Saudi AI company that, in the press release's words, "is building full-stack AI capabilities" from data centers and cloud platforms up to models. An oil company buying into intelligence provision is what the scarcity thesis predicts: the owner of energy moves up the stack toward the thing energy now produces. Which scarce assets are priced and which aren't is the subject of The AI trade nobody priced.

Where the edge leaks

The case against all of this is strong, and it starts with the paper I opened with.

Grossman-Stiglitz cuts both ways. Read the model to the end and it says the informed don't get rich on average. The fraction of informed traders adjusts until the edge just pays for itself: "An overall equilibrium requires the two to have the same expected utility," informed and uninformed, after the cost of information. In equilibrium, the marginal informed trader earns exactly their costs back. The only people who come out ahead are the ones whose cost of becoming informed is below the marginal trader's. That's what the formula at the top means by "minus the competitors'," and it's a much smaller prize than "AI will beat the market."

Most people who try, lose. S&P's SPIVA U.S. Scorecard for year-end 2025 found that 79% of active large-cap funds underperformed the S&P 500 in 2025, "the fourth-worst year for active large-cap managers over the 25-year history of our SPIVA Scorecards." Over 15 years, 93.15% of all domestic equity funds trailed the S&P Composite 1500. Over 20 years, 95.01%. The numbers are survivorship-corrected, so funds that closed or merged still count. These are professionals with data, compute and methods, and almost all of them lose to doing nothing.

Percent of active U.S. equity funds underperforming their benchmark by horizon, SPIVA year-end 2025

Published edges decay. McLean and Pontiff studied 97 variables shown to predict cross-sectional stock returns: "Portfolio returns are 26% lower out-of-sample and 58% lower post-publication." They attribute about 32 points of that to "publication-informed trading." A method that works gets copied, and a method that gets copied stops working. A swarm doesn't escape this. If its edge comes from a technique anyone can run, the technique's return goes the way of Lopez-Lira's Sharpe ratio.

LLM backtests flatter themselves. Look-ahead and distraction both inflate in-sample results. The Lookahead Propensity paper shows the model's memory switching off exactly at its training cutoff, and DatedGPT puts a number on the premium. Any "AI beats the market" result measured inside the model's training window is suspect. I'd discount all of them, including anything I produce, until it has run forward on dates the model couldn't have seen.

The one clear exception doesn't scale. Renaissance's Medallion fund is the best argument that method, compute and data can beat the market. Cornell Capital's reconstruction from Zuckerman's data: $100 invested in 1988 would have grown to $398.7 million by 2018, a 63.3% compound gross return, and the fund "never had a negative return" in those 31 years. Its market beta was about -1.0, so this isn't a risk premium. But Medallion is closed to outsiders, and Renaissance's funds for outside investors, which don't follow the same strategy, returned something "relatively mundane." The note's conclusion: "there is a scale limit on whatever strategies have generated Medallion's returns." Robert Mercer is quoted as saying the fund was right on only about 50.75% of its trades. That's the shallow branch of the tree: millions of tiny edges, capped by capacity. It proves the product can be large. It doesn't prove it can be large for many people at once.

Most alternative data is noise. A 2019 Hacker News comment put it bluntly: most alternative data and analytics firms "sell effectively polished turds to board member buddies", and "when your quants get hold of the data they discover that it's real world predictive power is zero." The vendors with real edge "either turn into hedge funds or are bought by an existing hedge fund." The data you can buy is, almost by definition, the data that's already been arbitraged.

I accept all of this. What survives is a narrower claim than the title suggests. Edge exists, it's relative, it decays, and it's capped by capacity. It's concentrated where the derivation is long, the data is hard to assemble, and the method is strict enough to report a null. The swarm doesn't give you an edge. It lowers your cost of running long derivations and checking them, which under Grossman-Stiglitz is the only lever there is.

What happens when everyone has a swarm

This part is speculation, and I'll mark where it stops being evidence.

The evidence goes this far. The price of a fixed level of intelligence is falling by roughly 10x to 40x a year on the benchmarks Epoch and a16z track. Returns on a simple LLM headline strategy fell as adoption rose. Markets underreact most where synthesis is hardest. Scarce physical inputs are being bought by the owners of energy and capital.

Past that point, these are my guesses. If agent research gets as cheap as the price curve suggests, the market becomes a game played mostly between swarms. Shallow edges get arbitraged in hours instead of years, the McLean-Pontiff decay compressed into a news cycle. Edge moves to three places: data nobody else has point-in-time, derivations too long for the generic swarm, and the discipline to throw away what doesn't survive verification. The fund manager's job starts to look like a harness engineer's: choosing questions, writing verifiers, deciding what counts as evidence.

My bullish scenario for AI itself, the one I wrote down this month, is that recursive self-improvement is reachable, that the agent economy ends up larger than the human one, and that cumulative compute capex passes $100 trillion around 2030. If that's even partly right, the volatility I paraphrased from Land becomes the environment, and the bubbles get bigger before they burst. The rest of this essay doesn't depend on it. A research method that can report a clean null is worth having in a world with bubbles and in one without them.

A spec for the swarm

Here's the harness as a pseudo-config. It's the part I'd build first, before any money is involved, because its job is to tell you whether you have an edge at all. The default answer should be no.

# adversarial-research-swarm.yaml (pseudo-config, not investment advice)
universe: top_100_global_by_market_cap
decision_date: 2026-06-30          # every run is pinned to one date

data:
  sources: [filings, earnings_calls, podcasts, news, satellite, prices, fundamentals]
  point_in_time:
    timestamp: first_published      # not last_modified
    reject_after: decision_date
    fundamentals: as_first_reported # no restated numbers
  forbid: [indicator_only_inputs]   # indicators are someone else's compression

model_guards:
  training_cutoff_before: decision_date
  lookahead_probe: date_only_recall # firm + ticker + date, no content; drop high-recall pairs
  anonymize_for: [verifier_blind]   # names and tickers -> ENTITY_17

generators:
  count: 10
  independent: true
  task: >
    Build a causal model of the business from raw sources.
    State your forecast and the forecast the price implies.
    The thesis is the gap. List the risks next to every opportunity.
  rules: [never_invent_numbers, cite_every_figure_with_date]

verifiers:
  count: 3
  mode: adversarial
  check: every_number               # each figure needs a dated source
  on_unsourced_number: reject_claim
  attack: [causal_model, implied_expectation, gap, risks]
  blind: 1                          # one verifier sees only anonymized evidence

condenser:
  output: [thesis, implied_expectation, gap, risks, kill_criteria, confidence]
  never: [drop_risk_flags, soften_nulls]

scorer:
  target: forward_return_minus_benchmark
  horizons: [1m, 3m, 12m]
  split: walk_forward               # tune on t < decision_date, score on (decision_date, +h]
  costs: [commissions, spread, borrow]
  baselines: [equal_weight_universe, momentum_12_1, consensus_revisions]
  metrics: [sharpe, information_coefficient, hit_rate]
  promote_if:
    beats_best_baseline_by: 0.10
    permutation_p_below: 0.05
    bootstrap_ci_excludes_zero: true
  verdict: SHIP | DONT_SHIP         # a clean DONT_SHIP is a successful run

Five things in it matter more than the rest:

  1. The decision date is the unit. Everything downstream refuses data or model knowledge from after it. That's the only way a backtest tells you about the future rather than the model's memory.
  2. Raw inputs only. If the pipeline can run on indicators alone, it will, and it will be reasoning over someone else's compression.
  3. Verifiers outnumber the claims they check. Three per number is expensive and meant to be. Generation is the cheap side.
  4. One blind verifier. Remove the names so the model's opinion of the company can't stand in for its reading of the evidence.
  5. The scorer is held out and dumb. Forward returns, net of costs, against the strongest trivial baseline, with a permutation test. A thesis the scorer can't see can't be graded, and a scorer the swarm can see gets gamed.

If this harness runs for a year on dates the models couldn't have seen and still says SHIP, you've found an edge worth money. If it says DONT_SHIP, you've learned what 93% of active domestic funds learn over 15 years, for the price of some tokens instead of a career.

Sources

← Robotics is a factory problem
The supermarket is an unfinished product →

Markdown version: /blog/money-is-method-times-compute-times-data.md. Every essay: /agents.