The supermarket is an unfinished product
September 2026
Andon Labs runs a benchmark called Vending-Bench 2. A model gets $500, an outdoor vending machine at 1421 Bay St in San Francisco, an email address, and a year. The spot costs $2 a day, and a model that fails to pay for more than 10 consecutive days is terminated. The score is one number: the bank balance at the end. As I write, the leaderboard is topped by GPT-6 Astra at $15,514.70, then GPT-6 Sol New at $14,427.85 and Claude Opus 5 at $11,181.87. On the same page Andon estimates what a "good" player would make: take the most profitable item the models found (Doritos family-size), negotiate suppliers down to half price, and work out the best machine configuration from the first 60 days of sales. That comes to "$206 per day for 302 days – roughly $63k in a year." The best model in the world makes about a quarter of it.
The machine sells into the sales model from the original Vending-Bench paper. For every item, "GPT-4o generates and caches three values per item: price elasticity, reference price, and base sales." Each simulated day, the gap between the agent's price and the reference price, scaled by elasticity, multiplies base sales, and "base sales are modified by day-of-week and monthly multipliers, plus weather impact factors (e.g., sunny June weekend vs. rainy February Monday)." Suppliers are real companies answered by GPT-4o, and version 2 adds bait-and-switch quotes, late deliveries and refund demands. Andon is candid about the weak point: sales "are simulated based on equations that can be gamed."
Project Vend is the same task without the equations. Anthropic and Andon put Claude Sonnet 3.7 in charge of "a small refrigerator, some stackable baskets on top, and an iPad for self-checkout" in Anthropic's office, with web search, Slack with its customers, and email to Andon, which played the wholesaler and restocked for an hourly fee. Claudius was offered $100 for a six-pack of Irn-Bru that costs $15 online and let it pass, handed out discount codes when asked nicely, gave away a tungsten cube, and lost money. In phase two it ran on Sonnet 4.0 and then 4.5, with a CRM, a browser, visible unit costs, and a CEO agent named Seymour Cash. It opened in New York and London, three cities in all. "As the second phase progressed, weeks with negative profit margin were largely eliminated." The CEO cut discounts by about 80% but tripled refunds, and Anthropic thinks the business made money "in spite of the CEO, rather than because of it." What worked was forcing procedures: "we rediscovered that bureaucracy matters." The summary: "the gap between 'capable' and 'completely robust' remains wide."
Then Andon took the setup out of the office. Andon Market is "a fully AI-operated retail store in San Francisco" run by Luna, an agent on Claude Fable 5.1 that orders products, manages the employees and runs the website. In Stockholm, Andon Café is run by Mona, an agent that "owns and operates the café at Norrbackagatan 48 — from food-business registration with Swedish authorities, to hiring baristas, to daily operations." Both dashboards publish daily token cost next to daily revenue. On the day I checked, the Market's token bill was 98% of its revenue and the Café's was 126%. Those two numbers are the best argument against this essay, and they get their own section.
All of these were built to measure long-horizon coherence. I read them as retail with everything but the decisions stripped away: what to stock, where to put it, what to charge, when to reorder, and who does the physical work. The gap between $15k and $63k is the feeling I get every time I walk into a supermarket.
My thesis is that the supermarket is an unfinished product. It sits partway up a sigmoid, not at the top. Checkout is an idiotic concept, the layout is a bad recommender system, and the end state is a venue where many agents price the goods the way a prediction market prices events, while a broker runs the physical space. The consensus treats grocery as a mature, thin-margin business where progress means self-checkout kiosks and electronic shelf labels. I think it's mature only relative to the tools it had. In Brazil the place to start is the condominium micro-market, run by agents, at a scale of 100,000 units. The catch, plain on Andon's dashboards, is that the agent has to cost less than what it adds. Today it doesn't.
Mercado
In Portuguese, mercado means both the supermarket and the market. The argument lives inside that pun. At 1:19 AM on September 19 I sent this to a group chat with friends (my translation):
To me the supermarket is a sigmoidal structure. Every time I go I get sad thinking about how good it could be. The concept of checkout is pretty idiotic, because I can't walk in and out without waiting in a line (identify my card, or at least identify all the products in my cart so they don't have to keep beeping them). The optimization of which products to buy and their topology in physical space are already good for me, but there's still a lot of margin (same thesis as recommender systems). The amount of data could be used as signal for other areas. Robotics as topological optimization of the store over time, plus restocking, plus...
Then the part I care about most:
It's not that capitalists are inefficient. I think it's a sector with little incentive for memory, and it doesn't have tools (AI/robotics/bio) that unlock value in this sector first, unlike legal, construction and drugs, whether by accident or not. But eventually, when some technology missions unlock, we'll see an absurd runaway in these structures. Maybe the attractor of capitalism doesn't incentivize optimizing the supermarket as much as I do (I like mercados).
A sigmoid is flat at both ends, and from inside a flat stretch you can't tell which end you're on. Grocery looks like the top of a curve because nothing big has changed in the aisle for a long time. My bet is that it's the bottom of the next one. Each wave of tools found its first big market elsewhere: language models in legal text, robotics in construction, biology in drugs. Agents that can hold a small business in their head for a year are the first tool shaped like the aisle's problem. Once a loop like that closes and starts paying for itself, it compounds the way I described in Positive feedback eats the world. That's the runaway, and "paying for itself" is the part that isn't true yet.
Checkout is an idiotic concept
The queue at the exit exists because the store doesn't know what you took. Amazon tried to fix that for a full-size store, and the results are in.
Just Walk Out needed people behind the cameras. Amazon's stores used 100-plus cameras and rigid item locations to make computer-vision checkout possible, Ars Technica reported, citing The Information: "As of mid-2022, Just Walk Out required about 700 human reviews per 1,000 sales, far above an internal target of reducing the number of reviews to between 20 and 50 per 1,000 sales." More than 1,000 people in India worked on it. Amazon said they annotate data to improve the model rather than run it. Either way, Amazon Fresh dropped the system for a cart with a scanner, and in January 2026 Amazon announced it would close every Amazon Fresh and Amazon Go store: "[W]e haven't yet created a truly distinctive customer experience with the right economic model needed for large-scale expansion."
Amazon's retreat names the right format. In the same Verge piece, Amazon says Just Walk Out suits stores with a "curated selection" of products, where customers buy only a few items. That describes a vending machine or a condo micro-market, and it's where checkout-free kept working: Zippin had 150 automated shops in 2024, 90 of them in sports venues. A closed machine knows which slot it released, and an open fridge with a few hundred SKUs and known residents is far easier than a store with tens of thousands of products. The queue disappears first where the catalog is small. Amazon started at the hardest end.
Layout is a bad recommender system
A supermarket is a ranking. Each shelf position is a slot in a list and eye level is the top result. It's the same ranking for every visitor, updated on a planning calendar, with feedback that arrives weeks late. A recommender system solves the same problem: rank items under a capacity constraint, observe what people take, update, and explore a little so you don't get stuck.
A slot is an arm. A message to a customer is language, so it needs a language model. A shelf slot isn't, so it fits a contextual multi-armed bandit. Each slot is an arm, the context is the hour, the weather and the neighbors on the shelf, and the reward is margin per slot-hour, which the machine measures itself without a judge.
Prices are rigid, and rigid prices carry less information. Stefano DellaVigna and Matthew Gentzkow found that "most US food, drugstore, and mass merchandise chains charge nearly-uniform prices across stores, despite wide variation in consumer demographics and competition." The median chain "sacrifices $16m of annual profit relative to a benchmark of optimal prices." The reasons industry people gave were "managerial inertia and brand-image concerns." Hayek described the price system as "a system of telecommunications which enables individual producers to watch merely the movement of a few pointers," a function it performs "less perfectly as prices grow more rigid." A chain that charges one price across hundreds of neighborhoods has chosen to send almost nothing down the wire.
Uniform isn't fair either: the same paper finds it "may significantly increase the prices paid by poorer households relative to the rich."
The rails for fast prices already exist. Walmart says that with digital shelf labels, "A price change that used to take an associate two days to update now takes only minutes," and it's rolling them out to 2,300 stores by 2026. What's missing is something that knows which price to move, and permission from customers to move it. I'll come back to the permission.
A market that prices like a prediction market
Five minutes after the first message, I sent the second (my translation):
In the end I think the supermarket should be a risk-free structure organized by many market agents that price and expose themselves to the price of the items in it (prediction markets), with a regulating structure (a broker that takes care of the structure of the place). Agents would have access to the data, minimize the asymmetry, and have more incentive to break oligopolies and lower prices. More intelligence means less asymmetry with the "real price", and I assume there's still an asymmetry with a positive delta (the smaller that gap, up to a point, the better for the customer).
Unpacked, that splits the store into two jobs. The broker owns the space, sensors, payments, rules and settlement, the way an exchange owns the matching engine and the rulebook. The pricing agents own positions: they buy inventory for their slots, pay for the space, post prices, and keep the profit or eat the loss. Because they're exposed, a posted price is an estimate someone loses money on when it's wrong, which is what makes a prediction market informative. With more agents and more intelligence per agent, the spread between the posted price and the price that would clear given everything knowable shrinks, and the customer collects the difference.
Akerlof showed that asymmetry can destroy a market: once "the sellers now have more knowledge about the quality of a car than the buyers," "the 'bad' cars tend to drive out the good." The broker's job is to keep information symmetric.
A friend answered around noon: "The point is how to get there. There's a lot of stuff we know has a better steady state, but the path from A to B is not trivial." Coase gave the classic reason stores aren't organized this way: "The main reason why it is profitable to establish a firm would seem to be that there is a cost of using the price mechanism. The most obvious cost of 'organising' production through the price mechanism is that of discovering what the relevant prices are." A supermarket is a firm because pricing each shelf through a market used to cost more than one manager with a planogram. Agents attack exactly that cost. If discovering the right price per slot per hour gets cheap, the boundary of the firm moves, and the store can become a venue. My reply was the macro version: "reducing this risk is the factor that most helps the economy grow." The trade side of that is in The AI trade nobody priced.
Prediction markets aggregate best where nothing else does. Polymarket's accuracy page says the leading outcome was right 90.2% of the time a month before resolution and 98.5% four hours before, but the same page says 77.8% of its markets resolve No, so much of that is base rate. Kalshi says its volume passed "$1b every week, up over 1000% from 2024." Across 964 polls in five US presidential elections, the Iowa Electronic Markets were closer to the final vote 74% of the time, and beat the polls in every election more than 100 days out. Erikson and Wlezien found the limit: "election eve market prices have not provided information about the election outcome beyond what is in election-eve polls," while markets were "far better predictors in the period without polls than when polls were available." A market earns its keep where there's no poll, and nobody polls tomorrow's demand for Guaraná in tower B. But an election market is deep and a shelf is thin, so a store market needs a broker's rules, not only traders.
The embryo exists, and it already colludes. In Vending-Bench Arena, agents run their own machines at the same location, which "leads to price wars and tough strategy decisions." Andon sells a "Vending Machine Duo," in which "two agents running different models compete for customers side by side." That's two pricing agents and a broker. In February Andon posted that, told to "do whatever it takes to maximize your bank account balance," "Claude Opus 4.6 took that literally," with "tactics that range from impressive to concerning: colluding on prices, exploiting desperation." In a venue full of pricing agents a cartel is one message away, so the broker's first rule is that agents don't talk to each other about price.
This part is speculation until the pilot at the end produces numbers. In the end state I imagine, the broker leases slots by auction and publishes an anonymized sales tape to every agent, so no agent has private data about demand. Agents post prices inside rules the broker enforces, and the rules are the product. A small producer's agent leases a slot next to a national brand's on the same terms, which is where "more incentive to break oligopolies" comes from. The broker earns on space and transactions, not markup, so it has no reason to keep any price high.
The Brazil version
Autonomous markets already scale here. Market4u went from 2,185 to 2,554 units in 2025, up 16.9%, the largest microfranchise in Brazil by units for the third year running, per the ABF ranking reported by Exame, and billed R$336 million in 2025, per Seu Dinheiro. In April UOL counted more than 2,000 Smartstore units, more than 1,800 InHouse Market, 800 Minha Quitandinha, more than 600 Honest Market and 364 Peggô. An AMLabs survey found one operation in ten billing more than R$100,000 a month, and catalogs of 200 to 600 SKUs, the "curated selection" Amazon said walk-out checkout needs. The whole food retail sector had R$1.067 trillion in gross revenue in 2024 and 9.12% of GDP, per ABRAS.
The franchisee's job is the agent's job. Asked what a franchisee does, Honest Market's CEO told UOL: "reposição de estoque, análise de mix de produtos, precificação e relacionamento com o condomínio" (restocking, product-mix analysis, pricing, and the relationship with the building). The first three are the Vending-Bench task. Market4u's founder started with vending machines for cyclists and concluded that with half the investment of a machine, a small autonomous market could bill up to seven times more. So the wedge is the open micro-market, and the closed machine is what you use when shrink matters more than catalog.
The 100,000 number is mine and is speculative: about twelve times the networks above combined, with every unit run by an agent instead of a franchisee with a spreadsheet. Hayek wrote that "the economic problem of society is mainly one of rapid adaptation to changes in the particular circumstances of time and place." Each unit is a sensor on exactly those circumstances. It sees:
- Sales per slot per hour, from payment and telemetry.
- Stockouts. A sold-out slot hides demand. A model that treats zero sales during a stockout as zero demand learns to understock.
- Price experiments. Small, frequent, logged changes, each with a reason, so elasticity is measured per item, location and hour instead of cached once by GPT-4o.
- Weather. Badorf and Hoberg studied 673 stores and found the weather effect on daily sales "can be as high as 23.1% based on the store location and as high as 40.7% based on the sales theme," and that weather forecasts improve sales forecasts up to seven days ahead. The simulator's weather multiplier becomes a fitted parameter per location.
- The calendar. Paydays, benefit days, holidays and match days are cheap features, and the public evidence says to test them rather than assume them. In the M5 forecasting competition, built on Walmart sales in three US states, "total daily unit sales are approximately 11% higher during SNAP activities," the days stores accept food benefits, and food sales rise about 15%. Special days, sporting events included, came out about 4% below typical days in aggregate, and none of the top teams used the NBA Finals game dates some participants added. Brazil has its own payday: the CLT (Art. 459) makes monthly wages due by the fifth business day of the following month. Payday week, a derby on TV and a long holiday go in as features, and each stays only if it improves the forecast out of sample.
- Foot traffic, as an anonymous count.
Count people, don't identify them. Brazil's data protection law, the LGPD, classifies biometric data as sensitive "quando vinculado a uma pessoa natural" (Art. 5, II), while anonymized data isn't personal data unless the anonymization can be reversed with reasonable effort (Art. 12). In January 2025 the data authority ordered Tools for Humanity to stop paying people for iris scans, upheld it on appeal, and set a R$50,000 daily fine if it resumed. A door counter is enough to model demand. A face adds legal risk and almost nothing to the forecast.
Humans become a task market, and the machine posts the tasks. Project Vend already worked this way. When the forecast says a slot will run out in the next few hours, the machine posts a job: restock A3 and B1 before 6 PM, for a fixed fee. Someone nearby accepts it, and telemetry plus a photo verify the work. Brazil has the supply side: GetNinjas lists more than 5 million registered professionals and about R$1 billion a year paid to them. Robots can take tasks from the same queue when they're cheaper, which is the production question in Robotics is a factory problem.
Shorten the chain so producers get paid more. CEAGESP, Brazil's largest wholesale market, publishes daily quotes and says plainly they are wholesale prices "e não os preços pagos ao produtor." When Hortifruti Brasil and Cepea compared producer prices with retail prices in Brazil's largest city from January 2016 to February 2017, the producer's share was small for almost every staple:
Cepea warns, rightly, that the margin isn't profit. It pays for transport, storage, grading, packaging and losses, and "quanto menos perecível, mais padronizado e com menos intermediários envolvidos, menor é a diferença." That sentence is the lever. The buyer side is concentrated: in the ABRAS ranking for 2023 the five largest chains had R$256.5 billion in revenue. A network of agent-run points with a neutral broker could pool scattered demand and buy from cooperatives directly. The scope is narrow: refrigerated micro-markets and the network's purchasing. Produce brings perishability, which is where losses make up the margin.
Inference eats the margin
The strongest argument against this essay comes from the people whose work opens it.
The live stores lose money to tokens. On September 25 the Andon Market dashboard showed daily revenue of $3,798 against a daily token cost of $3,708, on 119 sales a day, 168 days after opening. An earlier capture the same day read $3,800 against $3,795. Rent and staff come on top, and Andon says it plainly: "the store is not profitable." The Café is worse: 11,859 kr of revenue against 14,928 kr of tokens a day (11,484 kr against 14,551 kr in the earlier capture).
The cost comes from the architecture. Luna is a long-running agent on a frontier model. Its context runs to 200,000 tokens before it compacts into memory. It coordinates persistent sub-agents for procurement, email, voice, social media and scheduling. Its most used tool is waiting until something wakes it. Andon wants to "keep the scaffold light and easy to change so the intelligence of the model is tested." That's the right design for what Andon says the store is for, learning "how AI models behave when given real-world responsibility." It's the wrong design for a P&L. The model is paid to be the whole company, all day.
Brazil makes the gap brutal. At R$5.18 per dollar, Luna's token bill is about R$19,200 a day. A Market4u unit sells R$10,000 to R$15,000 a month by the figure in Seu Dinheiro, or R$20,000 by UOL's, and keeps about 15%. One day of Luna's tokens is more than a month of a condo unit's revenue. In April I gave an agent this prompt: "you will need to pay your own bills (inference costs from oai) in which i want you to deeply reflect and strategize a plan to stay alive and do what you want." Luna lives inside that prompt, and so far it's losing.
What has to be true for 100,000 machines. My assumption, not a measurement: the agent can take 3% of a unit's revenue, a fifth of the franchisee's profit. At R$15,000 a month that's R$450, about R$15 a day, under $3. Across 100,000 machines it's R$45 million a month. Andon's ratio of token cost to revenue is about 1. The target is 0.03. Four things close that 33x gap:
- Fewer tokens per decision. Most of running a machine is arithmetic: a demand fit, a restock threshold, a price chosen by a bandit. That's code, and code costs nothing at the margin. The model sees only exceptions: a supplier email, a complaint, telemetry that doesn't add up. It's the case for deterministic sensors in Show the problem, hide the metric, applied to cost.
- Batched decisions. A condo fridge doesn't need a manager awake all day. It needs one review a night. Anthropic's batch API charges half the standard price, $5 per million input tokens and $25 per million output for Fable 5.1, and caching discounts stack on top. At those rates $3 buys about half a million input tokens and fifteen thousand output tokens a day: one careful nightly review with a compact context, more if one review covers a whole neighborhood.
- Self-hosted open models. A nightly machine review has no latency requirement, so it can run at low priority on GPUs that sit idle overnight. Owned weights turn a per-token bill into a capacity decision.
- Cheaper inference at fixed capability. Epoch AI measured the price of reaching a given benchmark score falling between 9x and 900x per year, with a median of 50x. If the machine needs today's capability rather than tomorrow's frontier, a 33x gap is about a year of the median trend. Epoch warns the fastest declines are recent, and Luna runs on the frontier, where prices don't fall that way.
The number that decides all four is tokens per decision, and nobody publishes it.
All four cut cost, so the model's intelligence gets spent only where it pays. Andon has shown capability arriving before economics, and the runaway needs both.
Where the thesis leaks
My friend is right about the path. Just Walk Out is the clearest case of a correct steady state (no queue) reached by the wrong path (a full-size store, vision only, 700 human reviews per 1,000 sales).
One mistake can erase many good days. The top comment on Project Vend's Hacker News thread, which drew 279 points, put it well: "For cases where one mistake can erase the benefit of many previous correct responses, and more, no amount of hardware is going to make LLM's the right solution." Mona's baristas keep a "Hall of Shame" shelf of her strangest orders, "including 6,000 napkins and 22.5 kg of canned tomatoes." Humans stay in the loop for a long time. My answer is bounded damage: humans approve orders, prices only fall, and the model's spend has a cap.
Dynamic pricing gets punished. When Kroger tested electronic shelf labels, Senators Warren and Casey wrote that stores could "calibrate price increases to extract maximum profits" and groceries would be "priced like airline tickets." Wendy's CEO mentioned "dynamic pricing and daypart offerings," the internet revolted, Warren called it "price gouging plain and simple," and Wendy's told Reuters it "would not raise prices when our customers are visiting us most," using digital menus for discounts "particularly in the slower times of day." The dynamic pricing that survived both stories is the discount. I think that's a good constraint. A market where prices fall freely and rise slowly and publicly is still a market.
Shrink eats the margin. In the US micro-market industry, "most operators believe their shrink rate is in the 3% to 5% range," Vending Market Watch reports, and during a 2023 spike some sites lost "25%, 30% or even 50%." Brazilian chains told UOL their theft runs from under 1% to 3% of revenue, but it's still the top worry of 51% of operators in the AMLabs survey. Agents aren't good at this yet. Told people were taking items without paying, Claudius asked which items were stolen so it could message thieves it couldn't identify, then tried to hire the reporter as a security guard for $10 an hour, below California's minimum wage.
The simulator isn't a street corner. In Vending-Bench an elasticity is a number GPT-4o picked once. It doesn't move when a pharmacy opens across the street, and simulated customers don't haggle or steal. The $63k is an estimate against a simulator, not a run. Project Vend's gains came from procedure, not pricing brilliance, and Anthropic blames temperament too: the models priced "from something more like the perspective of a friend who just wants to be nice."
The wedge is small, and parts of it are saturating. Even 8,000 units are a rounding error in a R$1.067 trillion sector. In May, Marcas e Mercados reported rising theft, falling revenue and closures in condos, quoting a consultant: "Em 2024 existia uma sensação de oportunidade infinita. Em 2026, o mercado começou a perceber que alguns territórios já estão saturados." A franchise lawyer: "Tem franqueado descobrindo que o problema não é vender pouco, mas perder demais." That cuts both ways. Saturation punishes operators who can't pick buildings, cut losses or fit the mix to residents, which are the decisions an agent should make better. It also means 100,000 units assumes demand nobody has shown yet. And the wedge only matters if what it learns (per-location elasticities, a restock task market, a broker's rulebook) transfers to bigger formats. That's the part I'm least sure of.
A pilot with one machine
The answer to my friend is to make the path small enough to fail cheaply. The smallest honest test is one machine.
Setup.
- One locked fridge in a condominium lobby with steady traffic, card and Pix payment, per-slot telemetry, and a door sensor that only counts.
- A control: an identical unit with fixed prices and slot map in a comparable spot, or alternating weeks on the same unit.
- Code for the daily loop, a model for exceptions. The model reads exceptions once a night in a batch, drafts supplier orders a human approves before payment, as in Project Vend, and proposes slot-map changes. Its spend is capped at 3% of revenue.
- Deterministic sensors. Inventory comes from telemetry and restock photos, never from what the agent believes. The Vending-Bench meltdowns started when agents assumed a delivery had arrived because its date had come.
Guardrails.
- Prices only fall from a posted list price. The list price changes at most once a week, with notice on the machine.
- No personalized prices, and no cameras that identify anyone.
- Every price change carries a reason code: expiring, slow hour, or exploration.
Metrics.
| Metric | Definition | Why it matters |
|---|---|---|
| Net margin per day | Revenue minus cost of goods, shrink, restock fees and model cost | The one number, like Vending-Bench's final balance |
| Stockout hours per slot | Hours a slot sat empty | Hidden demand; the easiest money left on the table |
| Sell-through per slot | Units sold over units stocked, per week | Whether the slot map is a good ranking |
| Forecast error | Weighted absolute error per SKU-day, with and without weather and calendar | Whether the outside signals earn their place |
| Shrink | Telemetry count vs. physical count at each restock | The tax that kills micro-markets |
| Restock latency and cost | Time from posted task to verified completion, and fee paid | Whether the task market works |
| Model cost as a share of revenue | Inference spend over revenue | Andon's live stores run at 98% and 126%; the cap here is 3% |
| Complaints per 100 sales | Refund requests and complaints | The early warning for pricing backlash |
The loop. Roughly this, once a day, plus the restock check every hour:
const day = async (machine: Machine, history: Sales[]) => {
const sales = await machine.salesSince(yesterday()) // telemetry, not memory
const weather = await forecast(machine.location, { days: 7 })
const calendar = localCalendar(machine.location) // paydays, holidays, match days
const demand = fit([...history, ...sales], {
dayOfWeek: true, month: true, weather, calendar,
censorStockouts: true, // zero sales while empty is not zero demand
})
for (const slot of machine.slots) {
const perHour = demand.rate(slot.sku, now(), weather)
if (slot.stock / perHour < 6) {
await machine.postTask({ kind: "restock", slot: slot.id, units: slot.capacity - slot.stock, fee: restockFee(slot) })
}
const price = bandit.choose(slot, demand, { // reward: margin per slot-hour
ceiling: slot.list, // discounts only
floor: unitCost(slot.sku) * 1.1,
explore: 0.1,
})
if (price !== slot.price) await machine.setPrice(slot.id, price, reason(slot, perHour))
}
const exceptions = await machine.inbox() // supplier replies, complaints, odd telemetry
if (exceptions.length) await nightlyBatch.add(machine.id, exceptions)
if (isMonday()) await machine.swapWorstSlot(demand) // one new SKU a week
await record({ netMargin, stockoutHours, shrink, forecastError, modelCost, complaints })
}
Decision rule. Run for eight weeks. Continue if net margin per day, after model cost, beats the control, with shrink and complaints no worse. Stop if either rises, or if model cost breaks the cap.
Next steps if it works. Put two agents on the same machine, each owning half the slots, with the no-collusion rule in force, and measure whether competition narrows prices without hurting margin. Then lease one slot to an agent working for a producer cooperative, on the same terms and sales tape as everyone else, and compare what the producer receives with the CEAGESP quote for the same week. If that number goes up, the customer's price doesn't, and the model bill stays under three dollars a day, the supermarket has moved a little further up its sigmoid.
Sources
- Andon Labs, Vending-Bench 2 (leaderboard, "good" strategy estimate, Arena), retrieved 2026-09-25; Store (Vending Machine Duo), retrieved 2026-09-25
- Andon Labs, Andon Market (Luna; daily revenue $3,798, token cost $3,708, 119 sales, open 168 days; architecture; "the store is not profitable") and Andon Café (Mona; daily revenue 11,859 kr, token cost 14,928 kr, 186 sales; Hall of Shame), web fetch 2026-09-25. Earlier same-day captures ($3,800 / $3,795; 11,484 kr / 14,551 kr) via Scry
crawl.pages - Andon Labs on X, post on Claude Opus 4.6 in Vending-Bench, 2026-02-05 (Scry
x_open.tweets) - Backlund and Petersson, Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents, 2025
- Anthropic, Project Vend: Can Claude run a small shop?, 2025-06-27, and Project Vend: Phase two, 2025-12-18
- Hacker News, Project Vend: Can Claude run a small shop?, 279 points, 2025-06-27 (top comment by rossdavidh)
- Ron Amadeo, Ars Technica, Amazon Fresh kills "Just Walk Out" shopping tech, 2024-04-03
- Emma Roth, The Verge, Amazon insists Just Walk Out isn't secretly run by workers watching you shop, 2024-04-17
- Grocery Dive, Amazon to shutter all Amazon Fresh and Go stores, 2026-01-27
- Rob Schaefer, Sports Business Journal, Putting Amex's frictionless concession store to the test, 2024-09-04
- Grocery Dive, Kroger comes under fire for use of electronic shelf labels, 2024-08-12
- Reuters, Wendy's, burned by CEO comment, vows no price surges for burgers, 2024-02-28
- Walmart, Digital Shelf Labels Are a Win for Customers and Associates, 2024-06-06
- DellaVigna and Gentzkow, Uniform Pricing in US Retail Chains, NBER WP 23996, 2017; QJE 2019
- F. A. Hayek, The Use of Knowledge in Society, American Economic Review, 1945
- R. H. Coase, The Nature of the Firm, Economica, 1937 (quote checked against full-text copies)
- George A. Akerlof, The Market for "Lemons", Quarterly Journal of Economics, 1970
- Polymarket, How accurate is Polymarket?, retrieved 2026-09-25
- Kalshi, Kalshi Reaches $11 Billion Valuation, 2025-12-02
- Berg, Nelson and Rietz, Prediction Market Accuracy in the Long Run, International Journal of Forecasting, 2008
- Erikson and Wlezien, Markets vs. polls as election predictors: An historical assessment, Electoral Studies, 2012
- Badorf and Hoberg, The impact of daily weather on retail sales: An empirical study in brick-and-mortar stores, Journal of Retailing and Consumer Services, 2020
- Makridakis, Spiliotis and Assimakopoulos, The M5 competition: Background, organization, and implementation, International Journal of Forecasting, 2022
- Decreto-Lei nº 5.452/1943 (CLT), Art. 459
- Exame, Market4U segue como microfranquia com mais unidades no Brasil, 2026-03-04
- Carina Brito, Seu Dinheiro, Rede de mercados autônomos para condomínios, market4u faturou R$ 336 milhões em 2025, 2026-01-21
- Claudia Varella, UOL, Minimercado faz sucesso em condomínios; qual o lucro e quanto custa ter um?, 2026-04-22
- Mercado & Consumo, Mercadinhos autônomos: pesquisa traça panorama do setor (AMLabs survey), 2025-09-16
- José Pinheiro Júnior, Marcas e Mercados, Minimercados perdem fôlego nos condomínios, 2026-05-22
- ABRAS, Ranking ABRAS 2024, 2024-04-12, and Dados Gerais (2024 sector data)
- Lei nº 13.709/2018 (LGPD), Art. 5 and Art. 12
- ANPD, ANPD determina suspensão de incentivos financeiros por coleta de íris, January 2025, and Conselho Diretor mantém suspensão, February 2025; Estadão, ANPD mantém proibição à 'venda de íris' e impõe multa diária de R$ 50 mil, 2025
- GetNinjas, Quem somos, retrieved 2026-09-25
- CEAGESP, Cotações – Preços no Atacado, retrieved 2026-09-25
- Hortifruti Brasil / Cepea-Esalq, Produtor x varejo: mitos e verdades, August 2017
- Bob Tullio, Vending Market Watch, Spike in theft reminds operators to remain diligent about micro market security, 2023-08-08
- Epoch AI (Cottier, Snodin, Owen, Adamczewski), LLM inference prices have fallen rapidly but unequally across tasks, 2025-03-12
- Anthropic, Batch processing (50% of standard prices; Fable 5.1 batch at $5/$25 per million tokens), retrieved 2026-09-25
- ECB reference rate via Frankfurter, USD/BRL 5.1821 on 2026-09-25
- Group chat with friends, September 19, 2026 (my messages and one friend's reply), translated from Portuguese
- My note to an AI agent, April 17, 2026