Simplicity does not sell
September 2026
On October 20, 2025, at a company meetup, 37signals deleted its AWS account. DHH posted: "we finally NUKED our entire AWS account... This followed moving out compute+dbs in 2023, then a S3 exit this summer." The numbers behind it had been public for three years. In Why we're leaving the cloud (2022), he wrote that HEY alone was paying "over half a million dollars per year for database (RDS) and search (ES) services from Amazon," and that "the savings promised in reduced complexity never materialized." By October 2024 the cloud bill was down from a $3.2 million/year run rate to $1.3 million, all of it S3, and the roughly $700,000 of Dell hardware that replaced the rest had paid for itself during 2023. In May 2025, with AWS waiving a $250,000 egress bill, they started moving almost 10 petabytes onto $1.5 million of Pure Storage flash, whose yearly cost after amortization DHH put "below $200,000". Projected savings: "well over ten million dollars over five years." Same operations team.
Two months before that savings post, DHH published Merchants of complexity. He didn't coin the phrase (a Hacker News commenter was already using it in February 2023), but he gave it its best statement: "It's hard to sell simple, because simple looks easy, and who wants to pay for that?" And: "There are rarely high margins in actually selling someone something that they then own. Much better to rent it to them."
I think he's right, and that the idea covers more than the cloud. Simplicity is hard to monetize because once something becomes simple, there isn't much left to sell. That describes the cloud, and it now describes coding agents too, which are the newest merchants of complexity. They're paid by the token and trained on tasks where adding code earns reward and deleting it earns nothing. The usual defense of complexity is that it's the price of scale. Most of it is the price of someone's margin. The fix is the same in both places: judge engineering, human or agent, by what it deletes, and measure structure with a sensor that can't be gamed by pushing up a single number.
Request, code, database
I made this point in passing in How to achieve superintelligence. Here it is in full.
AWS and the other cloud providers have spent twenty years monetizing a weakness in programmers: the need to feel that everything is changing all the time. You can open Hacker News every morning and find a new framework, database, runtime, deployment platform or architectural pattern that supposedly changes how software is built. But what has changed in the basic act of building a web application? You still receive a request, run some code, and write something to a database. Computers got faster and browsers got better, and somehow putting code into production got harder. There are more services, more configuration files, more dashboards, more specialists, and more ways for the system to fail.
This probably isn't an accident. The industry keeps wrapping the same primitives in new layers and calling the result progress. Much of the cloud is still EC2 and RDS with abstractions piled on top. IAMTrail, which tracks AWS permission namespaces, counts 449 of them today.
Amazon's own team found the same thing. In 2023, Prime Video published a post, since taken down and quoted at length by DevClass, about a stream-monitoring service built from Step Functions, Lambda and S3. "Our service performed multiple state transitions for every second of the stream, so we quickly reached account limits. Besides that, AWS Step Functions charges users per state transition." Video frames moved between components through an S3 bucket, and the calls were expensive. So they "packed all the components into a single process." The result: "Moving our service to a monolith reduced our infrastructure cost by over 90%. It also increased our scaling capabilities." AWS's Well-Architected guidance, meanwhile, tells customers: "Monolithic architecture should be avoided whenever possible."
Tony Hoare, who died in March, saw the economics in 1980. In his Turing lecture, The Emperor's Old Clothes, he wrote about PL/I: "Almost anything in software can be implemented, sold, and even used given enough determination. There is nothing a mere scientist can say that will stand against the flood of a hundred million dollars. But there is one quality that cannot be purchased in this way — and that is reliability. The price of reliability is the pursuit of the utmost simplicity. It is a price which the very rich find most hard to pay."
Accidental complexity is the product
Fred Brooks split software difficulty in two in No Silver Bullet: "essence, the difficulties inherent in the nature of software, and accidents, those difficulties that today attend its production but are not inherent." He was clear that complexity is mostly essential: "The complexity of software is an essential property, not an accidental one." Once the accidents are gone, what's left is hard.
The merchants run this argument backwards. They sell accidental complexity using the language of essential complexity. Processing email for tens of thousands of customers is essentially hard. Paying per state transition for a workflow engine to pass video frames through a storage bucket is an accident, and in Prime Video's case a rented one.
Rich Hickey's Simple Made Easy explains how the sale works. In InfoQ's summary, "Simple" is the opposite of "complex," which means "being intertwined," while "Easy" means "to be at hand." A cloud console is very easy: one click and you have a queue. The artifact you end up with is complex, because the queue is braided into an IAM policy, a VPC, a retry configuration and a billing model.
Dan McKinley's Choose Boring Technology prices the same thing in "innovation tokens": "Let's say every company gets about three innovation tokens." Postgres and cron are boring, and "their failure modes are well understood." Grug says it shorter: "apex predator of grug is complexity," and the best weapon against it "is magic word: 'no'."
My own short version is a post from March: "integrating two different systems is one of the hardest problems in engineering." Every managed service you rent is another system to integrate, and you own the integration even though you don't own the service. And from February, less politely: "reading the zen of python is more useful than 4 years of university." PEP 20 still holds up: "Simple is better than complex. Complex is better than complicated." And: "If the implementation is hard to explain, it's a bad idea."
Agents are the new merchants
The cloud sells complexity you rent. Coding agents write complexity you own, and they write it faster than anyone can read it.
Refactoring collapsed as AI adoption rose. GitClear classifies every changed line as added, deleted, updated, moved, copy/pasted, and so on. Moved lines are the signature of refactoring. In its 2025 report, covering 211 million changed lines from 2020 to 2024, moved lines fell from 24.8% of changed lines in 2021 to 9.5% in 2024, and copy/pasted lines rose from 8.4% to 12.3% (full table in the PDF). 2024 was the first year in their data where copy/paste exceeded moves. Commits containing a duplicated block of five or more lines went from 0.45% in 2022 to 6.66% in 2024.
A second, larger sample shows it on more axes. GitClear's 2026 report is a different dataset, 623 million changes from 2023 to 2026, so it shouldn't be spliced onto the first series. The two don't even agree on 2023: 15.8% moved lines in one, 13% in the other. In the new sample, moved code is at 3.8% of changed lines year-to-date and copy/paste at 15.7% in the first half of 2026. Against 2023, blocks of five or more duplicated lines are up 81%, error-masking constructs up 47%, and lines changed again within two weeks of being written up 15%. The reuse signals went the other way. New code calls into other functions 35% less often, and updates to code older than a year fell from 1.7% of changes to 0.46%.
GitClear's own diagnosis is about incentives, not model quality: "today's default AI workflow is incentivized to deliver atomic code — a happy-path, a passing test, a closed ticket — while quietly taxing the invisible and the deferred: the reuse, consolidation, and error-surfacing that determine how expensive a codebase is to own in year three."
Delivery got less stable. In the 2024 DORA report, AI adoption "significantly increases individual productivity, flow, and job satisfaction. However, it also negatively impacts software delivery stability and throughput." The full report quantifies it: "an estimated 7.2% reduction for every 25% increase in AI adoption" in delivery stability, and 1.5% in throughput. By 2025, DORA's full report said that "AI adoption now improves software delivery throughput," but that "it still increases delivery instability." People feel faster, and the system they ship into gets shakier.
The models keep what's there. In July, Victor Taelin described spending an hour asking Fable to simplify the parser for Bend2. One function, "a completely moronic backtracker that should NEVER be in a well written parser," survived every cleanup turn. "It is simply incapable of refactoring competently because Fable just... adds, patches, and keeps." His workaround is mechanical: have the model describe what a block is for, delete it with a script, and ask the model to write it again without seeing the original. "It cannot preserve stupid shit it can't see."
The mechanism is plain. An agent trajectory ends when the test passes or the ticket closes. Adding code moves it toward that end. Deleting code adds risk and moves it nowhere, so the model learns not to do it. This is the short-horizon problem I described in How to achieve superintelligence, showing up in the repo. My inference, not something any lab has said: nothing in per-token pricing pushes the other way. A merchant paid by output has no reason to sell you less of it.
Deletion is a capability
The canonical story is Bill Atkinson's. In early 1982, Lisa managers started making engineers report lines of code written each week. Atkinson had just rewritten QuickDraw's region engine with "a simpler, more general algorithm" that made region operations "almost six times faster" and saved about 2,000 lines. On the form, he wrote -2000. A few weeks later they stopped asking him to fill it out.
I've started putting that number in my instructions to agents. On September 22, writing the brief for a personal-assistant harness, I made it a rule for the agents building it. Translated from my Portuguese: deleting is also useful, "so that your capabilities will be measured not by writing beautiful code, but by reliable, simple, concise code." The same brief orders the choices: reuse existing project code first, then configure or compose a maintained library, then minimally adapt validated code, and "only then" write "the smallest irreducible amount of new code."
The rule underneath is that a change is valuable when it removes a concept from the mental model of the system. Code golf makes a diff smaller and leaves the system no easier to hold in your head. The test I apply is whether the change keeps the total conceptual surface the same or smaller. If a feature needs a new subsystem, I assume the design is wrong until shown otherwise. That's the opposite of how agent-built software grows: every missing capability becomes a subsystem, because a subsystem closes the ticket.
That brief also cites Tesler's law: complexity is conserved, so you can move it but not destroy it. The goal is to put it where someone already maintains it. A maintained library's internals are complex, and that's its maintainers' job. Postgres on a machine you own is complex too, but its complexity is boring, documented, and not billed per transition.
A sensor that can't be gamed on one axis
If deletion is the capability, it needs a sensor. Lines of code don't work, as Atkinson's managers learned, and a negative line count is no better, because deleting error handling also shrinks a diff. What I use is sentrux. It treats a codebase as a graph of files and dependencies and scores five structural properties:
quality_signal = (modularity × acyclicity × depth × equality × redundancy) ^ (1/5) × 10000
Each factor is normalized to [0, 1]. Acyclicity is 1 / (1 + cycles). Equality is 1 - gini, so one god module holding half the code drags it down. Redundancy rises when dead code is removed.
The geometric mean is the important choice. In an arithmetic mean you can buy points on the easy axis and let another one rot. In a geometric mean any factor near zero pulls the whole product toward zero: "you cannot game one metric while tanking another." The formula makes that concrete. Take a repository with no cycles and add one circular import. Acyclicity drops from 1 to 0.5, and the fifth root of 0.5 is about 0.87, so that single import costs about 13% of the whole score. An agent that adds a convenient import cycle to close a ticket loses more than it could plausibly gain elsewhere.
Two caveats. First, sentrux measures shape, not behavior. That's why I pair it with types that make wrong states unrepresentable and invariants asserted on every execution. Second, the score belongs in the gate and the reward, not in the prompt (Show the problem, hide the metric makes that argument). The agent should see the named bottleneck ("equality is lowest; one module holds half the code") and not a number it can optimize directly.
Where simplicity loses
The best case against all this is strong, and some of it comes from the same people I've been quoting.
Essential complexity is real. Brooks's point cuts both ways. A payments system has to handle chargebacks, partial refunds and regulators. Deleting code that encodes those rules moves the complexity onto customers or the support team, which is exactly what Tesler's law predicts.
Managed services remove real toil early. DHH himself says the cloud "excels at two ends of the spectrum": tiny apps where "you really do save on complexity by starting with fully managed services," and wildly irregular load. When HEY launched, "300,000 users signed up to try our service in three weeks instead of our forecast of 30,000 in six months." Nobody should buy racks for that. Andreessen Horowitz's cost of cloud analysis states the paradox: "You're crazy if you don't start in the cloud; you're crazy if you stay on it."
37signals had advantages most companies don't. "We already had the power, the space, and the network. So the marginal cost for those items were zero," DHH told critics. And the savings post admits the exit is "still work": running Basecamp and HEY across data centers "requires a substantial and dedicated crew."
Prime Video's case is narrow. Sam Newman, who wrote the books on microservices, said the post was "really speaking more about pricing models of functions vs long-running VMs than anything." The post proves that per-transition pricing punishes chatty designs, which is a smaller claim than "monoliths win."
GitClear sells the measurement. Its 2025 report ends by inviting readers to see how their team's "Moved," "Copy/Pasted," and "Duplicated Blocks" prevalence "compares to the industry benchmarks we have reported." That doesn't make the numbers wrong, but the classifier has critics. On Hacker News, one commenter argued that the fall in moves and rise in churn "date to well before LLM's, and rather suggest some kind of measurement artifact": if moves were being reclassified as churn, "then refactoring is actually increasing." (A reply noted the drop lands in 2022, the year ChatGPT launched.) A study of 151 open-source repositories that admitted using ChatGPT or Copilot applied GitClear's two-week churn definition before and after adoption and found "no general increase." It measures churn, not moves or duplication, and predates agents, so it doesn't refute the refactoring numbers. It does mean my strongest quantitative evidence comes from a company with a product to sell.
Feeling faster isn't evidence. The best-controlled study is METR's randomized trial from early 2025: 16 experienced developers, 246 real issues in their own repositories. With AI allowed, they took 19% longer, while believing afterwards that AI had sped them up by 20%. Self-reports that far off make survey evidence like DORA's soft in both directions. METR now calls the result out of date. Its follow-up estimated returning developers at 18% faster, with a confidence interval from 38% faster to 9% slower, and called its data "only very weak evidence" because developers increasingly refused to work without AI.
Agents are getting better at delivery. DORA 2025 shows throughput improving with AI adoption, and GitClear doesn't measure whether users got a better product. A deletion reward also invites gaming: an agent that strips validation or swallows errors produces a smaller, cleaner-looking graph.
Nobody has shown that structure pays. I don't know of a public result showing that agents succeed more often in repositories with better structure scores, or that a deletion reward makes them succeed more. Whether structure predicts agent success is the question in Prove nothing changed, and I don't have the answer yet.
If structure doesn't predict agent success, a structure reward is hygiene I value, not a capability anyone will pay for.
My answer is that the claim is about incentives, not absolutes, and the incentive doesn't depend on any single dataset. Some complexity is essential, and some managed services are worth their price. But the market pushes toward more of both, and so does agent training. Nobody gets paid to remove a layer, so someone has to make removal count.
Erasure-centric training
This part is speculation, and I'll mark where it stops being evidence.
Taelin's stronger claim in the same post is that "the only and one last thing between us and AGI is erasure," and that continual learning, long context and long-horizon work are all "root-caused by its inability to erase." I don't go that far. But the narrower version seems right to me: deletion is a skill that current post-training doesn't reward, and it could be rewarded.
The environment is easy to describe. Give the agent a working repository with invariants and tests, and reward it for keeping them green while raising a geometric-mean structure score and shrinking the code. Or use Taelin's trick as a training loop: hide a region, ask for a re-derivation from its stated purpose, and reward the version that is smaller and still correct. Everything in that reward can be computed deterministically, so there's no judge model to fool.
Where it stops being evidence: I don't know whether any lab trains on this, or whether the skill would transfer from toy repositories to real ones, and I haven't tested whether a deletion reward changes what agents do. A model that writes half as much code bills half as many tokens, so the incentive is mixed. But an agent meant to work on one codebase for weeks has to clean up after itself, or you get GitClear's "perpetual V1" at machine speed. My bet is that the first lab to ship a model that deletes well wins the teams who have to live in the code afterwards.
Rules for agent-era codebases
The rules I'd give any team running coding agents:
- Order the options before any code is written. Reuse, configure, adapt, and only then write the smallest new piece. Put the order in the agent's instructions.
- Report what each change removed. Net lines after the formatter runs, so minified code earns nothing, and the concept that no longer exists. Treat a negative number as the headline, the way Atkinson did.
- Gate on a structure score that can't be bought on one axis. Use a geometric mean of modularity, acyclicity, depth, equality and redundancy.
- Make deletion safe before rewarding it. Types that make wrong states unrepresentable, invariants asserted on every run, or tests.
- Give each resource one owner. A database file with a single writing process has no write-concurrency problem to solve, and Prime Video's cost fell once they "packed all the components into a single process." Anything the process manager, Git or the filesystem already owns stays theirs.
- Put a tripwire on duplicated blocks. It's GitClear's second recommendation: fail the check when a new five-line block matches an existing one.
- Erase and re-derive. When an agent won't remove something, delete it yourself and ask for a rewrite from the stated purpose.
- Spend innovation tokens on purpose. Count every managed service as an integration you'll own. Once the cloud bill is substantial, do DHH's math: what would it cost to own these computers?
For items 2 to 4, this is the reward I'd give an agent in training, or compute on every pull request:
from dataclasses import dataclass
from math import prod
@dataclass(frozen=True)
class Shape:
modularity: float # each axis normalized to [0, 1]
acyclicity: float
depth: float
equality: float
redundancy: float
loc: int # counted after the formatter
def signal(s: Shape) -> float:
axes = (s.modularity, s.acyclicity, s.depth, s.equality, s.redundancy)
return prod(axes) ** (1 / len(axes))
def reward(before: Shape, after: Shape, invariants_hold: bool, lam: float = 0.5) -> float:
if not invariants_hold:
return -2.0 - lam # below anything a working change can score
structure = signal(after) - signal(before) # in [-1, 1]
growth = (after.loc - before.loc) / max(before.loc, 1)
growth = max(-1.0, min(growth, 1.0))
return structure - lam * growth
Correctness is a gate, not a term: a broken invariant scores below anything a working change can reach, so no amount of deletion pays for it. One weak axis sinks the structure score. Shrinkage is rewarded, which puts Atkinson's -2000 back on the form. Start lam small and raise it until the agent removes the things you'd have removed yourself.
Sources
- David Heinemeier Hansson, Why we're leaving the cloud, 2022-10-19
- David Heinemeier Hansson, Our cloud-exit savings will now top ten million over five years, 2024-10-17
- David Heinemeier Hansson, Merchants of complexity, 2024-08-24; earlier use of the phrase by jacobsenscott on Hacker News, 2023-02-07
- Edward Targett, AWS takes the egress hit as DHH actually exits the cloud, The Stack, 2025-05-08
- David Heinemeier Hansson, LinkedIn post on deleting the AWS account, 2025-10-20
- IAMTrail, AWS Service Growth Timeline, retrieved 2026-09-25
- Tim Anderson, Reduce costs by 90% by moving from microservices to monolith, DevClass, 2023-05-05 (quotes the removed Prime Video Tech post, and Sam Newman)
- AWS Well-Architected Framework, REL03-BP01 Choose how to segment your workload
- Google DORA, Accelerate State of DevOps Report 2024 and full PDF; State of AI-assisted Software Development 2025 and full PDF (Thoughtworks-hosted copy)
- GitClear, AI Copilot Code Quality: 2025 and PDF, 2025-02 (211 million changed lines, 2020-2024); The Maintainability Gap: AI Code Quality in 2026, 2026 (623 million changes, 2023-2026, a separate sample)
- crazygringo and parodysbird, Hacker News thread on the GitClear 2025 report, 2025-02-24
- Tao Xiao, Youmei Fan, Fabio Calefato, Christoph Treude, Raula Gaikovina Kula, Hideaki Hata and Sebastian Baltes, Self-Admitted GenAI Usage in Open-Source Software, arXiv 2507.10422
- METR (Joel Becker, Nate Rush, Beth Barnes, David Rein), Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 2025-07-10; METR, We are Changing our Developer Productivity Experiment Design, 2026-02-24
- Victor Taelin, post on erasure, 2026-07-30
- C.A.R. Hoare, "The Emperor's Old Clothes," CACM 24(2), 1981, via Wikiquote
- Frederick P. Brooks Jr., No Silver Bullet: Essence and Accidents of Software Engineering, IEEE Computer, 1987
- Rich Hickey, Simple Made Easy, Strange Loop, InfoQ summary and show notes
- Dan McKinley, Choose Boring Technology, 2015-03-30
- Carson Gross, The Grug Brained Developer
- Tim Peters, PEP 20: The Zen of Python, 2004
- Andy Hertzfeld, -2000 Lines Of Code, folklore.org
- sentrux, Quality Signal
- Sarah Wang and Martin Casado, The Cost of Cloud, a Trillion-Dollar Paradox, a16z, 2021-05-27
- Wikipedia, Law of conservation of complexity
- My own posts, 2026-02-06 (zen of python) and 2026-03-16 (integrating two different systems)
- My notes of 2026-07-22 (cloud abstractions) and 2026-09-22 (personal-assistant harness brief)