# A mirror, not a thermometer

*September 2026*

On 15 October 2008, 23 new listings appeared on a Brazilian review forum for paid sex. A model that knows the weekday, the month and the holidays expected 8. In twenty-two years of data it is the most anomalous day there is, 5.2 standard deviations above the calendar, and it was the day the Ibovespa lost 11.4%, two days after gaining 14.7%. The VIX closed at 69.

I wanted that day to mean something. Men watch the stock market lose 11% in an afternoon and go out that night: the story writes itself. Then I looked at the days before it. From 11 to 14 October the forum recorded no listings at all. Four days of zero, in a market that normally sees several a day, is what a website looks like when it's down, and the series has 56 days like that, in blocks. The 15th was the backlog. On the 172 days since 2004 when the VIX closed at 40 or above, the forum's activity was ordinary (p = 0.64).

![The forum's strangest day came right after four empty ones](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/crash-day.svg)

I started this project with a narrower question: does the price of paid sex in Brazil follow inflation? Clients on this forum have written down what they paid, week after week, since the early 2000s. If those prices tracked the cost of living, the forum would be a price index nobody else has, a thermometer for a corner of the economy that official statistics don't reach.

It isn't one. The price barely moves, and when it does, it moves for reasons of its own. In real terms it fell by a fifth to a third while nearly every other personal service in Brazil got more expensive. I then tried 30 ways of forecasting the economy with the forum's numbers, and none survived a correction for having tried 30 times. What the forum turned out to be is a mirror. It shows, more sharply than I expected, when Brazilian families gather, how fast a payment system spread, how violent each state is, when people started working from home. About where the economy is going, it says nothing the official numbers don't already say.

The answer to the first question fits in four lines:

```
real price   =  nominal price ÷ price level
nominal:        flat for about a year, then a jump of R$50
price level:    up a little every month
so between jumps, every month of inflation comes out of the real price
```

I did this over four days in September with a coding agent, on a laptop with 8 GB of RAM. The agent wrote the crawler and the analysis code. I asked the questions, and when the first round of answers came back I told it to stop following my list and chase whatever the data suggested. By the end there were 316 statistical tests in one registry, corrected together for multiple comparisons, and 95 of them survived. Most of the survivors describe how the market works inside. Few connect it to the economy outside. That split is the main result, and the essay is organised around it.

I don't name the forum, the sites I compared it with, or anyone on them. Phone numbers were cut to an area code and a salted hash before any analysis, and no working name appears next to a complaint. Every number below describes what clients wrote. None of it measures the women they wrote about.

## What the forum records

The forum has a subforum for each of 13 states, and inside each, a thread per provider: a woman who works on her own, or a venue, which Brazilians call a clínica or a privê. A client opens the thread after a visit, and its header follows a fixed template: what he paid, for how long, the provider's phone, the date, her working name. I call these headers listings. Most replies are reviews by other clients, and they follow a form too: a verdict, a score out of ten (since about 2015), the services offered, the price, and free text. When I started, the forum counted about 100,000 threads and 1.4 million posts.

The verdict is one of four words: positive, neutral, negative, or mancada, slang for a night that went wrong, whether a no-show, a bait-and-switch or a theft. In the reviews I collected, 86% are positive, 6% neutral, 6% negative and 1% mancada.

I had planned to run sentiment analysis on the text. The reviewers had already done it: every review states its verdict. The verdicts are so polarised that a plain bag-of-words classifier trained on them scores an AUC of 0.995 on a random holdout and 0.976 on the two most recent years, which it never saw. Sentiment was the one part of the plan that needed no cleverness.

The data has two layers, and they came out of the crawl very differently. The listings are complete: 96,444 threads from 2002 to 2026, 53,919 of them with a price. The conversation is a sample: 42,889 posts from 3,625 threads, of which 31,514 are structured reviews and 8,929 say on which day the visit happened. An older scrape, from August, had the listings and the text of only 407 reviews, all from 2026. How the sample came to be 3% of the forum, and why that turned out to be enough, is the next section.

![All 96,444 listings, and a sample of the conversation behind them](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/coverage.svg)

## A crawler inside a browser tab

The forum sits behind a bot check, the kind that shows "Just a moment…" while it decides whether you're a person. The agent's first attempt, from a browser tab of its own, sat on that page and never got through. It stopped there, told me the check needed a real person, and said it wouldn't try to automate it. I passed it once, by hand. Everything after that ran inside that tab.

That shaped the design. The tab held the cookies that had passed the check, so the crawler had to live in the tab too: a script running in the page, not a program on my machine. There was no account and no login. Listings and review threads are visible to any visitor. Photos and parts of some posts are shown only to registered members; those stayed hidden, and they turn up in the data as a cluster of reviews full of "hidden content" notices.

![The crawler lived inside an ordinary browser tab](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/crawler.svg)

The queue lived in IndexedDB, the browser's own database: pending jobs, the key of every page already seen, and a priority class for each job. Between four and twelve workers took jobs from it, fetched pages with the tab's cookies, and parsed the HTML in the page into rows of topics, posts, prices, phones and dates. They backed off on HTTP 429 and 5xx errors. When the bot check came back, or the network dropped, every worker paused and a single probe tried again once a minute.

The part I'd keep in any crawler is the acknowledgment rule. Finished pages waited in an outbox in memory and shipped in gzip batches of 300 pages, or every ten seconds, to a receiver on 127.0.0.1: 85 lines of Python that wrote each batch to its own file and replied. Only after that reply was a job deleted from the queue, in the same IndexedDB transaction that saved the new jobs its page had spawned. A crash, a reload or a closed tab could lose pages, but only pages whose jobs were still queued, so they would simply be fetched again. Nothing was marked done before it was on disk. The receiver's files are append-only and nothing downstream edits them. Every table in this essay is rebuilt from them by a build step that deduplicates, parses and deflates, which is the argument of [The log is the truth](https://future-seems-so-good.com/blog/the-log-is-the-truth) applied to a crawl.

The order mattered even more. The crawler fetched every listing page first, because each listing says how many pages its thread has. Then it shuffled the thread pages within three classes (individual providers first, then venues, then general discussion) and worked through each class in random order. A crawl stopped at any moment would leave a random sample of pages within each class. I didn't know yet how much that would matter.

![The browser set the crawler's speed as often as the server did](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/throughput.svg)

Speed was a negotiation, and the browser was the other party as often as the server was. Six workers with a short pause between requests gave 173 pages a minute. Ten gave 422. Twelve gave 108: latency doubled, and we took that as the server asking us to back off. Then the tab went into the background and throughput fell to 35 pages a minute, with no errors at all. Chrome throttles timers in hidden tabs. Once a tab has been hidden for five minutes, chained timers wake at most once a minute, and the pause between requests was a timer. With the pause removed, so that only the number of workers set the pace, the crawl went back to 218.

Saving every page to IndexedDB as it arrived held it to 84 even with the tab in front, because each write was a transaction and the transactions queued behind one another. A second agent rewrote the storage so that pages stayed in memory and only the queue changes were written, which lifted the rate to 330. Its version had also left the outbox without a size limit, so a long block would have grown it until the tab died; the cap of 3,000 pages in the diagram is that fix. At 330 pages a minute, the bot check was back within two minutes.

That was at 2:26 on a Sunday morning, and the block held for the rest of the night and most of the next day. The crawler kept probing once a minute. The watcher the agent had left running sent a "no data" alert every hour, and all of them arrived together the following evening. By then the block had lifted, and the crawler, still probing, had resumed on its own at 23:09. When the agent reloaded its tab a few minutes later, not knowing it was already running, nothing was lost: the pages it hadn't shipped still had their jobs in the queue. From then on it ran at four workers, 144 pages a minute, about two and a half requests a second, and the check never came back.

Two bugs are worth describing because neither made any noise. The first was in pagination. The crawler worked out how many topics a listing page holds by reading an offset from the first matching link it found. On the largest forum that link pointed inside a thread, where a page holds 15 posts, while listing pages hold 30 topics. So it walked the listing in steps of 15 and, having planned the right number of pages, stopped halfway. Nothing failed. Half of the biggest forum was simply never requested. A repair pass compared each forum's page count with the offsets it had already seen and queued the 1,692 listing pages that were missing.

The second was in the receiver. `JSON.stringify` leaves the characters U+2028 and U+2029 unescaped, which is valid JSON. Python's `str.splitlines()` treats them as line breaks, along with `\x85` and a few others. A review containing one of them came out as two broken lines, and the receiver rejected its whole batch with an HTTP 400. The fix was to split on `"\n"` and nothing else.

Twenty minutes later, at 23:32, that run stalled too, and I never fully pinned down why. The receiver was alive but so slow that a status request took 2.2 seconds, and the browser had stopped sending. The next morning the agent had to work out what had happened from the files on disk and the queue in the browser, which is the problem [Intelligence is a trajectory](https://future-seems-so-good.com/blog/intelligence-is-a-trajectory) describes: a system can be sharp at every step and still lose the thread over days, unless the record it rereads is good. It asked Cursor to reopen the browser panel, and the call hung. Hours later there were 335 batches on disk instead of 193: the panel had apparently opened, and the crawler, its queue intact, had picked up where it stopped. Then the laptop started crashing and I shut everything down.

The analysis had the same trouble at a smaller scale. With 8 GB of RAM, swap reached 7.3 GB, and regressions that expanded their fixed effects into dummy columns were killed by the operating system. Absorbing the fixed effects instead of building them, the way Sergio Correia's reghdfe does, solved it.

The crawl was planned to take the whole forum in a night. It ran, on and off, for two days, and got 5,001 thread pages, about 3% of the total, with some 150,000 still in the queue. All of it fits in 27.7 MB of gzip. Because of the shuffle, that 3% is a random sample of pages within each class of thread. If the crawler had walked the threads in order, 3% would have meant a few forums complete and nothing about the rest. The sample is 71% São Paulo because São Paulo is most of the forum, and anything I say about a small state rests on a few hundred reviews.

## The price that forgot to rise

The listings are complete, so the price index is the most solid thing in this essay. Each listing gives a price, a session length, a state and a venue type. A hedonic regression on those features, with a fixed effect for each year, gives a like-for-like price for every year. From 2004 to 2026 it roughly doubled, from 100 to 218. Over the same years the IPCA, Brazil's official consumer price index, went to 327, its personal-services group to 439, and the minimum wage to 640.

![Prices here doubled while everything around them tripled](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/prices-vs-everything.svg)

Deflated by the IPCA of each listing's own metropolitan area, the real price fell 32.8% from 2004. Deflated by the national index and measured from 2005 to 2025, it fell 21%. That's the range I quote, a fifth to a third, depending on the deflator and the base year. The fall wasn't steady: 9% from 2004 to 2014, 25% from 2014 to 2019, when the recession pushed inflation past 10% and nominal prices sat still, and 7% from 2019 to 2025. Venues did worse than independents. Using the prices quoted in 3,757 venue reviews, with a fixed effect per venue, their nominal price rose 91% between 2005 and 2026.

Year to year, the forum's price has nothing to do with inflation. The correlation between its annual changes and São Paulo's IPCA is −0.28 (p = 0.21, 22 years). Full pass-through is rejected (p < 0.001), and the two series aren't cointegrated (Engle–Granger p = 0.77). Across 13 states, annual price growth doesn't follow the state capital's inflation (β = −0.28, p = 0.13). Price levels don't follow local income either: the income elasticity across 14 states is 0.10 (p = 0.48). And adding the forum's index to a simple autoregression doesn't improve a one-quarter-ahead forecast of the IPCA, of services inflation, or of personal services, where it makes the forecast significantly worse.

The mechanism is in the most common price.

![The going rate moved twice in 22 years](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/steps.svg)

The modal price for an hour was R$150 until 2013, when it jumped to R$200, and R$200 until 2026, when it jumped to R$300. Two moves in twenty-two years. The first came after 60% of accumulated inflation, the second after 104%. Carried forward by the IPCA, the R$150 of 2004 would be R$490 today. Between steps the real price only falls, and the steps never catch up. By 2025 the R$200, R$250 and R$300 tiers each held about 14% of listings: the second step was under way before it showed up in the mode.

Some of the adjustment happened in minutes instead of reais. Listings of 45 minutes or less went from 15% in 2015 to 30% in 2026, and sessions longer than 75 minutes from 24% to 5%. Time also got dearer at the margin: half an hour used to cost 33% less than an hour and now costs 40% less, and two hours went from a 34% premium to 63%. The index holds the length of the session fixed, so the fall I report is net of this. The headline price of "a session" held still partly because the session got shorter.

Against food, the fall is steeper.

![An hour in São Paulo used to cost almost a food basket. Now it costs a third](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/baskets.svg)

In 2004 an hour in São Paulo cost 0.87 of DIEESE's basic food basket for the city. In 2022 and 2023 it cost 0.26, and in 2026 it costs 0.33. The basket followed food prices while the price of an hour sat at R$200 to R$250 for more than a decade; the correlation of their annual changes is −0.27. Income tells the same story. An average monthly income in São Paulo bought 7.4 one-hour sessions in 2012 and 14.5 in 2026, and a minimum wage bought 1.7 in 2004 and 5.4 in 2026.

The fall wasn't even across the market.

![The middle of the market fell furthest](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/dispersion.svg)

In reais of 2026, the median hour went from R$477 to R$300, the 90th percentile from R$800 to R$612, and the 10th from R$188 to R$155. The middle lost 37%, the top 24% and the bottom 18%. Price inequality compressed until 2014, when the Gini reached 0.22, and has since opened again to 0.30: after 2020 the top started rising and the bottom didn't.

Internationally the fall is unusual, though not unique. Cunningham and Kendall, working with a large American review site, found the average real hourly price rose 27% between 1998 and 2008, from $273 to $348. The Economist analysed 190,000 profiles in 84 cities across 12 countries and found the opposite between 2006 and 2014, from $340 an hour to $260. Brazil looks like The Economist's sample, not the American one.

![Cheap in dollars, much less cheap in hours of work](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/ppp.svg)

In dollars at purchasing-power parity, an hour here cost $102 in 2008 and $107 in 2025: flat, while incomes grew. The median advertised price is $428 at PPP in London, $546 in Paris, $494 in Barcelona and $209 in Amsterdam. In hours of average income Brazil isn't cheap: 9.2 hours in 2025 and 15.2 in 2008, against 13 to 17 in London, Paris and Barcelona and 4.8 in Amsterdam. Measured in hours of minimum-wage work, an hour cost 92 in 2005 and 40 in 2025. Seven points from different sources can't support a model, so I'll stop at the description: cheap in dollars, much less cheap in hours of work, and getting cheaper in both.

## Sticky, but not the way the models say

A price that sits still for years and then jumps is a textbook object. Most macroeconomic models handle it with Calvo's rule: each month, every price has the same small chance of being reset, whatever happened since the last reset. It's a convenient assumption, and whether real prices behave that way has a literature of its own. Nakamura and Steinsson's five facts about American consumer prices are the usual reference: regular prices last 8 to 11 months, the median change is about 8.5%, and about a third of changes are cuts.

The forum has something those studies can only get from scanner data: the same seller's price, again and again, over years. I used 737 pairs of consecutive listings with the same phone and the same working name. The phone alone isn't enough, because agencies advertise several women on one number. The implied duration of a price is 12.1 months, a little longer than American retail. The median absolute change is 25%, three times theirs. 38% of changes are cuts, close to their third. And increases don't follow the inflation accumulated since the last change (p = 0.30).

![Prices reset when they drift from the market, not when the calendar says](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/calvo.svg)

Calvo's rule fails in a specific way. If the chance of a reset were the same every month, two listings two years apart would be much more likely to show a new price than two listings two months apart. In a complementary log-log model of the chance of a change, the coefficient on elapsed time should be 1. Here it's 0.19 (p < 0.001 against 1). What predicts a reset is distance from the market. Providers whose price has drifted from the going rate reset more often (p = 0.007), those below it raise (p < 0.001), and a reset closes about 44% of the gap. A provider 50% below the market raised her price by 59% on average; one far above it barely moved. That's state-dependent pricing, in steps of R$50, indifferent to accumulated inflation.

65% of changes are multiples of R$50, and I think the steps are the menu cost. My reading, which the data is consistent with but can't prove, is that a price of R$200 is a social fact as much as a number: clients compare it with what they paid last time and with every other thread on the page. Moving to R$217 because the IPCA says so would be illegible. Moving to R$250 is a decision, and people put decisions off.

![Baumol's cost disease hit every personal service except this one](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/baumol.svg)

The comparison that surprised me most is with other personal services. Baumol's cost disease says that services whose productivity can't rise (his example was the string quartet; a haircut works as well) get steadily more expensive relative to everything else, because their workers' wages have to keep up with the rest of the economy. Brazil's IPCA shows it cleanly. From 2005 to 2025, in real terms, manicure rose 39%, waxing 63%, domestic work 48%, the hairdresser 15%, beer at a bar 30%, and the personal-services group as a whole 32%. The forum's price fell 21%. The gap is significant at the usual 5% against every item except cinema, which rose 9% (p = 0.22), and against manicure and waxing it survives the correction for all 316 tests.

Here is a service that is almost pure labour, whose productivity can't have changed, in a country where wages rose, and its real price fell for twenty years. The data doesn't tell me why. The candidates I can see are all on the supply side. New providers enter in good years, not bad ones: the real minimum wage and the entry of new providers move together (r = 0.59, p = 0.006), the opposite of what an opportunity-cost story predicts. One listing in five now comes from someone on tour from another state. Listings that mention one of the big advertising platforms in the title charge 9.7% less, and those mentions went from almost none before 2015 to 10–15% of posts after 2019. The Economist blamed its own decline on the financial crisis, migration and the internet. Each of these would push the price down. None of them is a test.

## Families empty the market

Once the macroeconomic question came back empty, the calendar turned out to be the strongest signal in the data.

![They write about it on Monday. It happened on Friday or Saturday](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/week.svg)

People write during the working week and go at the end of it. I know when they go because 8,929 reviews say so ("yesterday", "last Friday", "on the 24th"). Portuguese makes this harder than it sounds: Monday to Friday are *segunda* to *sexta*, and *segunda* is also the word for "second", as in "the second time", so the parser counts a weekday only with "-feira" or an article in front. Visits peak on Saturday (18.0%) and Friday (17.4%). Sunday is the emptiest day at 6.1%, less than half of a uniform week's 14.3%. The posts describing those visits peak on weekday afternoons, between two and five. The price doesn't notice. No weekday has its own rate, in the listings or in the prices quoted in reviews: the price is fixed and the demand bunches.

![Christmas and Carnival empty the market about equally](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/holidays.svg)

Holidays empty the market, and which holidays is the interesting part. I compared each date with the same weekday in the same month across 22 years. Religious dates cut new listings by 32.5% and secular festivities by 38.1% (both p < 0.001). New Year's Day is the emptiest day of the year (−71%), followed by Christmas Day (−64%) and Good Friday (−53%). By the visit dates written in reviews the drop is deeper: 85% on Christmas, 87% on New Year's Day. Carnival is revealing. Its Saturday dips by an amount indistinguishable from noise (−18%, p = 0.08), but Monday and Tuesday fall 39% and 53%. Even Valentine's Day, which Brazil celebrates on 12 June, cuts listings by a third. The week before each religious date, as a placebo, shows nothing (+6%, p = 0.06).

![More Catholic states don't take the holy days off more](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/religion.svg)

If religious dates emptied the market because of faith, the drop would be deeper where there are more Catholics. It isn't. Across 11 states, ten more points of Catholics change the drop by 2.6 points (p = 0.71). The share of Evangelicals, the share with no religion, and Good Friday on its own give the same null, and the 2022 Census repeats it. Eleven states is little power, so this is weak evidence of absence. But the simplest reading is the one the secular dates already suggest: what empties the market is a family at the table, whatever the occasion.

![Heaven is a good review. Hell is a coin flip](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/heaven-hell.svg)

The sacred does show up in the text, as praise. 4.7% of reviews use heavenly words and 1.1% infernal ones. "Com louvor" is 98% positive, as is "divina", and "deusa", "anjo" and "paraíso" sit at 94–95%, against 86% for all reviews. "Inferno" is a coin flip. The exception is "pecado", sin, which is 95% positive: "um pecado de mulher" is a compliment. None of this vocabulary rises near religious dates (p = 0.78), and it's fading slowly, from 5.0% of reviews in 2006–15 to 4.0% in 2021–26.

"Louvor" also taught me to read before counting. My first lexicon filed it under religion, as in gospel music. A sample of reviews showed it was nearly always "aprovada com louvor", approved with honours, the phrase used when a thesis passes. After that, every lexicon was checked against examples before it became a variable.

Money that arrives on a known date moves the market; one-off money doesn't. In the five days after the fifth working day of the month, the legal deadline for paying wages, there are 4.4% more listings (95% CI 1.9% to 6.9%). Batches of income-tax refunds come with 3.7% more (p = 0.03), which doesn't survive the correction. The Christmas bonus and the one-off withdrawals from the workers' severance fund in 2017 and 2019 show nothing against placebo dates. The full moon, which I added as a placebo, does nothing (+1.2%, p = 0.44), as it should, and neither do rain, heat or cold in 11 capitals. September is the busiest month (+5.6%), May and June the slowest.

Football barely registers. When the biggest club of a state capital plays, local activity falls 4.7% that day (p = 0.01), and a win or a loss changes nothing the day after. Of the national team's results, only "the day after a draw" survives the correction: +15%, with nothing after wins or losses and no mechanism I can think of. I count it as a false positive.

## The clock ran a natural experiment

The forum stamps every post with one clock, its own. Until 2019, Brazil's South, Southeast and Centre-West moved their clocks forward an hour in summer, and the Northeast didn't. If people post by their local time, posts from the daylight-saving states should appear an hour earlier on the forum's clock during summer time, relative to the Northeast in the same month. They appear 0.78 hours earlier (95% CI −1.2 to −0.4), statistically indistinguishable from a full hour (p = 0.31).

![In summer time, São Paulo shows up an hour early on the forum's clock](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/summer-time.svg)

That result matters mostly because it threatened another one. After 2020 the share of posts written in office hours, Monday to Friday between nine and six, rose by 3.6 points, and posts in the small hours fell by 2.4. It looked like remote work. But daylight saving ended in 2019, and that alone shifts the clock. Redone only for March to September, months that never had summer time, the rise holds: 38.6% of posts in office hours in 2015–19, 43.0% in 2021–26 (p < 0.001). Remote work moved part of this market into the working day, or at least moved the writing about it. Mood follows the clock too. Reviews written in the small hours are less positive than afternoon ones (84.5% against 86.7%, p = 0.002).

![Since the pandemic, more of it happens during office hours](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/office-hours.svg)

## A few hundred men write most of it

A review forum measures the people who write, and they are a narrow group.

![One reviewer in ten wrote more than half of all reviews](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/voice.svg)

Among 14,594 reviewers, the Gini coefficient of reviews written is 0.70. The most active 1% wrote 16% of all reviews and the top 10% wrote 55%. The median reviewer has 14 reviews on his profile, and the tail follows a power law with an exponent around 1.8, like most online communities. When I say "clients" in this essay, I mostly mean these men.

They get kinder with time. A review from an account less than a month old is positive 78% of the time; from an account seven or more years old, 92%, holding the year fixed. Accounts created in the previous week wrote 7.1% of reviews, and only 77% of those are positive. I went looking for fake praise from new accounts and found the opposite: people who sign up to complain. The first post in a thread is positive 71% of the time, and posts after the hundredth 92%, because long threads exist when the provider pleases. When the first review is negative, the thread gets 35% fewer replies and 38% fewer views (p < 0.001, net of price, forum and year), which mixes reputation with the quality the review revealed.

![Bad nights get written up faster](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/anger.svg)

Bad nights get written up faster. Among 8,839 reviews that give a visit date, 69% of the negative ones were posted within a day, against 60% of the rest, half a day sooner on average (Mann–Whitney p < 0.001).

Who writes matters as much as who's written about. The reviewer's identity explains 17% of the variance in verdicts, the same as the provider's (16%). A positive review raises the chance that the next one is positive by 11.5 points beyond the thread's average. That fits herding, and it fits quality that changes over time; the data can't tell them apart.

To see how much a reviewer's temperament contaminates the ratings, I fitted a two-way model on the connected core of the network, 22,813 reviews by 4,735 reviewers of 1,375 providers, in which each reviewer has a severity and each provider a quality. Severity is a stable trait: random halves of a reviewer's history agree (r = 0.48). People who pay more are harsher (p = 0.001). That mattered for the result I most wanted to check.

![Satisfaction peaks in the middle of the price range](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/inverted-u.svg)

Satisfaction rises with price and then falls. In the raw data, 79% of reviews are positive in the cheapest decile (about R$134 in today's money), 92% in the fifth (R$289), and 83% in the most expensive (R$731). Harsh reviewers clustering at the top could produce that shape on their own. After removing each reviewer's severity the curve keeps its shape (quadratic term p = 0.016, against 0.001 raw), though with all 316 tests corrected together it sits right on the line (q = 0.05), and the drop at the top is halved. Part of the disappointment at R$700 is the kind of man who pays R$700. A Dawid–Skene model, which treats reviewers as noisy annotators, reclassifies only 7% of providers: with 86% of reviews positive, the information is in who is harsh, not in mislabelled verdicts.

What buys satisfaction is intimacy. Kissing carries a 12% price premium and is the strongest single predictor of a positive review. Oral sex without a condom carries no price premium at all (−0.7%, p = 0.43) and still raises the chance of a positive review. For public health, the risk isn't priced.

Failure leaves a signature. After a thread that almost nobody answers (the bottom quarter by replies), the same phone line comes back 21% sooner and, 9 points more often, under a different working name (p = 0.002). The price doesn't change. The line rebrands. This uses only phones with at most four listings and no agency words in the titles, and a phone line is still not a person.

The forum is also getting quieter. The median number of replies to a review thread fell from 3 in 2005 to 1 in 2025, and threads nobody answers rose from 25% to 35%. The median review shrank from about 1,234 characters in 2011 to 445 in 2026. A thread's median life, from first post to last, went from 51 days in 2004–10 to 2 in 2017–21 and 5 in 2022–26. More expensive providers keep their threads alive longer (a hazard ratio of 0.83 per log point of price).

## The words know the price

![What the market pays for, and what it marks down](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/bodies.svg)

Clients describe bodies, and the descriptions carry prices. On 30,293 reviews with a price, a regression of the real price on the words the client uses, with subforum and year effects, gives the premia in the chart. Implants carry the largest (+10.3%). "Plus-size" (−13.8%), "mature" (−13.0%) and descriptions of Black skin (−9.8%) carry the largest discounts, and all three survive the correction; tattoos are −4.0%. These are the client's words, not measurements, and the premia aren't causal. A client who writes "mature" may be describing a kind of listing as much as a woman. But the pattern has the shape of Brazil's racial pay gap in ordinary jobs, measured here in the market's own currency.

![The words in the reviews sort providers into markets](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/text-map.svg)

The text knows more than the list of attributes. I turned each provider with two or more reviews into a vector (TF-IDF over her reviews, reduced to 64 dimensions with an SVD) and asked a ridge regression to predict her real price from the words alone. Out of sample it explains 49% of the variance (permutation p = 0.01, 1,464 providers). A UMAP projection of the same vectors sorts providers into recognisable markets. The most expensive cluster (median R$602) talks about flats in Moema and the Jardins, reception desks, and showing ID on the way up. Massage and the ad-and-house cluster are the cheapest, at R$245 each. Complaints form a region of their own, at R$326 and 26% positive. Across all the attributes, kissing is the one most correlated with price (ρ = 0.30), followed by the score (0.28), and the first principal component of the whole correlation matrix is an axis of intimacy and satisfaction.

![The working names never age](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/names.svg)

The working names are frozen. I matched the first names of 34,326 listings against the 2010 Census, which counts how many Brazilian women with each name were born in each decade. The typical birth year of the names in the listings is about 1992, and it barely moves: 1991.9 for listings from 2004–09, 1993.2 for 2020–26, about a year of drift while the calendar moved sixteen. That's a property of the names, not of anyone using them. The repertoire (Bruna, Amanda, Camila, Fernanda) is a fixed fashion, the names of women born in the 1980s and 1990s, chosen in 2005 and in 2025 alike. Nicknames that barely exist in the civil registry, like Duda, Gabi, Nanda and Bruninha, went from 6% of listings to about 13%.

!['kkk' overtook 'rs' in 2013](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/laughter.svg)

Some of the mirror is just language. "kkk" overtook "rs" (short for *risos*, laughs) in 2013, as mobile messaging spread, and emojis appear from 2019. "Carinhosa", affectionate, has been falling since 2005 and is half as common now (trend p < 0.001). The explicit search for the *namoradinha*, the girlfriend experience, stays flat at about 5% of reviews.

## Geography is most of the structure

![In São Paulo the premium follows the price of land, loosely](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/sp-map.svg)

Within São Paulo, the premium follows land prices, loosely. I placed 981 listings in a district using the place names in their titles and reviews (ambiguous names like Sé, Saúde and Liberdade had to come with "bairro", "metrô" or "região"), and estimated each district's premium against the same year, session length and format, shrunk toward the city mean. Moema is +45%, Jardim Paulista +27%, Vila Mariana +7%; Tatuapé is −9%, the old centre around República −21%, and Santo Amaro −39%. Across the 17 districts with enough listings, the premium follows the sale price per square metre in a 2019 property dataset (Spearman 0.46, p = 0.06). As a predictor of property prices it's useless: in leave-one-out validation it adds nothing reliable to the distance from the centre (p = 0.17). What does follow the square metre is the density of listings (ρ = 0.40, p < 0.001, 91 districts). The market sets up where there's income and movement, and charges more in Moema and the Jardins.

![One listing in five now comes from someone on tour](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/tours.svg)

Providers travel, and more than they used to. The share of listings whose phone has another state's area code went from about 5% in the 2010s to 21% in 2024–26. Touring providers charge 30% more than local ones in the same city, year and venue type (p < 0.001). São Paulo is the big destination, with a net inflow of 829 listings; Minas Gerais (−377) and Rio (−235) are the largest exporters, and the busiest single route is Rio to São Paulo.

Tours don't bring their clients along. I built the network of reviewers and threads, 13,026 nodes and 18,940 edges in its main component. Among the reviewers of providers working in their home state, 98.4% are local; among those of touring providers, 96.6%, and only 0.6% come from her home state. On the Rio to São Paulo route, none do. The network's 51 communities (modularity 0.77) split first by state, then by venue against independent, and only last by price band. The mutual information between community and state is 0.53, against 0.04 by chance. 76% of reviewers with five or more reviews write at least 90% of them about one format, and only 5% of those with three or more review in more than one state. The market's gatekeepers, by betweenness centrality, are simply its most active reviewers (ρ = 0.98 with degree).

![Where the state is more violent, more nights end badly](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/violence.svg)

Some of what goes wrong belongs to the place. The share of reviews reporting a mancada is 6.7% in Pernambuco and 0.9% in São Paulo, it has been rising over the years (p < 0.001), and across 11 states it follows the homicide rate in the health ministry's mortality records (Spearman ρ = 0.67, p = 0.02). Pernambuco, Espírito Santo, Goiás and Bahia are high on both, São Paulo and Santa Catarina low. After correcting for all 316 tests the q-value is 0.07, and eleven points are few, so I read it as suggestive. If it holds, the market has no scam culture of its own. It takes on the violence of wherever it is.

Health shows a similar pattern and fails the harder test. States where more reviews report unprotected services have more acquired syphilis (ρ = 0.79, p = 0.02, 8 states), but within each state the two series don't move together over time (p = 0.16). Acquired syphilis exploded in Brazil, from 8.6 cases per 100,000 people in 2011 to 122.1 in 2024, and the forum shows no jump to match. That's an association between places, and nothing more.

## The outside world mostly stays outside

![Popes, protests and World Cups barely register](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/events.svg)

The events I expected to see are mostly invisible. I tested each against about 300 placebo windows of the same length, on the same dates in other years. That matters more than it sounds. Give a single day its own dummy and robust standard errors will call it significant almost by construction, because that day's residual is forced to zero. With placebo inference, only the first COVID quarantine clears the bar (−23%, p = 0.005). The 2014 World Cup (host cities against the rest, −16%, p = 0.43), Brazil's 1–7 against Germany (−42%, p = 0.25), the Rio Olympics, World Youth Day, seven editions of Rock in Rio, the June 2013 protests, the impeachment vote, the truckers' strike, the 8 January riots and twenty election Sundays can't be told apart from noise. Rio's events sit on a baseline of less than one listing a day in the state, and their sign flips from one edition to the next. The anomalies say the same: the twelve most unusual days in the series fall within a day of an event as often as random days do (17% against 18%).

![Everyone crashed in the spring of 2020. The forum came back highest](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/covid-world.svg)

COVID is the exception, and at first it looks the same everywhere. In April and May 2020 new listings fell 26.4% against 2019, close to the median fall in Google searches for "escort" across 11 countries (−28.8%). The difference is the recovery. By late 2021 the forum was above its 2019 level, while in most countries searches were back where they started. The fall didn't follow state restrictions: with state and day fixed effects, ten more points on Oxford's stringency index change listings by −1.4% (p = 0.86). The quarantine hit the whole country at once.

![Betting went from nothing to everywhere. The market didn't move](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/bets.svg)

The betting boom is a clean null. Sports betting was legalised in December 2018 and exploded between 2021 and 2024. If it drained money from the same men, the states that search "bet" most would have lost more listings after 2021. They lost 15% per standard deviation of search interest, with a randomization p of 0.27 across 13 states. Betting barely appears in the text, at most 0.25% of posts in any year.

![Clients adopted Pix twice as fast as the country](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/pix.svg)

Pix is where the outside world gets in fastest. The central bank launched its instant-payment system on 16 November 2020. Reviews mentioning Pix reached half of their eventual plateau 12 months after launch; Pix transactions per person in the whole population took 25. To reach 80%, it took 14 months against 36. Cash mentions halved between 2016–19 and 2022–26, and cards fell by about 40%. After 2024 mentions of Pix fall while national use keeps rising, which I read as Pix becoming too ordinary to mention. [Improving Brazil Through Cybernetics](https://future-seems-so-good.com/blog/improving-brazil-through-cybernetics) counts Pix, with Embrapa and the Real, among the feedback loops the country got right, and this corner of the economy took it up at twice the country's speed. Across states the link between mentions and use is weak (ρ = 0.22, p = 0.53), because outside São Paulo there are few reviews a year.

![Brazil stopped adding motels in 2014](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/motels.svg)

Motels are the market's infrastructure, and the federal company register lets me rebuild their history. Companies whose main registered activity is "motel" went from 2,041 in 1995 to a peak of 5,988 in 2014, and there are 5,902 now. Openings fell from 280 a year in 2000–09 to 157 in 2016–25, and closures rose from 56 a year to 156. Brazil stopped adding motels in the years the forum's real price fell fastest. By state, though, the stock of motels doesn't follow listings (p = 0.79), so the two share a decade and nothing else I can show.

The rest is a list of things that don't move the forum, and I report it because each was a question someone could reasonably ask. Sildenafil's patent expired in Brazil in June 2010 and the pill became several times cheaper: mentions of Viagra stay around 0.2% of reviews, repeat sessions don't jump, and the share and the premium of two-hour sessions don't change. PrEP entered the public health system in December 2017, and mentions of unprotected services don't jump either. Formula 1 weekends at Interlagos, the Agrishow farm fair and Oktoberfest in Santa Catarina, all with mostly male crowds: nothing. The construction of the Belo Monte dam and the Comperj refinery: indistinguishable from placebos. A Bartik-style shock from soy and iron-ore prices: nothing. Air passengers and foreign tourists arriving in each state, the peso as an instrument for Argentine tourists (too weak to use), the divorce wave after the 2010 constitutional amendment (239,000 divorces in 2010, 348,000 in 2011), construction employment, the pandemic's emergency aid: none leaves a robust mark. Neither do the VIX, the Ibovespa, the shares of Match Group, Netflix and Ambev, or page views of Wikipedia's article on prostitution. The move from venues to rented flats never happened: mentions of flats and Airbnb peaked at 12% of reviews in 2013 and are 1.6% now. The Venezuelan migration, from about ten people in 2011 to 732,000 in 2025, shows up in at most 0.13% of listings in any year, and I only ever counted it in aggregate.

Two outside forces do register. One is state unemployment. Each extra point cuts the number of new providers being reviewed by 9.0% and listings by 8.5% (p < 0.001 and 0.004), without moving the price (p = 0.28). With men's and women's unemployment in the same model, men's carries the effect (−11% per point, p = 0.05, short of the correction), and women's doesn't push more women into the market (p = 0.11). The forum measures the client side, and clients disappear when their jobs do. The other is GDP, and it's the one that nearly fooled me.

## Can it forecast the economy?

![In sample, the real price moves with GDP and against inflation](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/mirror.svg)

The last question I asked of the data was the one that would make it worth something to someone else: can it see the economy before the official numbers do? Quarterly GDP doesn't move the number of listings (p = 0.67), but it moves the real price, with an elasticity of 0.94 (p < 0.001; 1.20 without the pandemic). Over 78 quarters, the year-over-year change in the real price correlates with GDP growth at 0.54 and with inflation at −0.46. Both links survive in the same regression (GDP +0.83, IPCA −1.17, both p < 0.001), and even the nominal price, before any deflating, moves with GDP (r = 0.55). In sample, the real price Granger-causes GDP (p = 0.001). Had I stopped there, I'd have had an alternative-data product.

Two things should make anyone suspicious of that paragraph. Year-over-year growth rates are smooth series, and two of them will correlate whenever they share a few long swings. Here they share two: the recession of 2015–16 and the inflation of 2021–22, the two deep troughs in the chart. And a correlation in sample is not a forecast. What matters is whether the forum tells a forecaster something that GDP's own past doesn't.

So I ran the test a forecaster would. The targets were quarterly GDP, the central bank's monthly activity index (IBC-Br), consumer confidence, the Ibovespa, and unemployment and real income by state. The signals were the forum's real price, its listings and its new providers, for the current period and the next. For each pair I compared an autoregression with and without the forum's signal, re-estimated from a rolling origin in 2012 (2016 for the states), and scored every forecast with the Clark–West test, which allows for the noise of estimating the extra terms. That's 30 chances.

![30 chances to predict the economy in real time. None survived](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/forecast.svg)

Five of the 30 came out with a nominal p below 0.05, all of them with the real price as the signal: GDP now and next quarter, the IBC-Br now and next month, and state real income next quarter. After correcting for 30 tries, none remain; the smallest q is 0.12. The best of them, the GDP nowcast, cut forecast error by 4% (Clark–West p = 0.025). Without 2020–22 the same model is 10% worse than the benchmark. And when the benchmark also knows inflation, the gain shrinks to 3% and loses its significance (p = 0.06), and for next quarter it turns into a loss of 2.5% (p = 0.26).

![Put inflation in the benchmark and the forum's edge disappears](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/forecast-ipca.svg)

So the forum's price moves with the economy, but whatever it knows about the economy, a forecaster already has from GDP's own past and the IPCA, which the statistics office publishes every month. It reflects the cycle and doesn't see ahead of it. It doesn't help with the thing I started with either: added to an autoregression, the forum's index leaves the forecast error of the IPCA where it was (a relative error of 1.01).

## 316 tests

The analysis came in three rounds. The first tested the questions I brought: the calendar, big events, football, politics, inflation. The second tested questions the data raised on its own: names, payments, clocks, reviewers, bodies. The third went after natural experiments and about forty outside sources, from the national statistics office and the health ministry to the company register, the central bank, Google Trends and a Kaggle dataset of São Paulo apartments. Every test from every round went into one registry with its estimate, p-value and confidence interval, and the whole registry was corrected at once with Benjamini–Hochberg.

![316 tests: a spike of real effects over a floor of noise](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/pvalues.svg)

Of 316 tests, 119 have p < 0.05, 95 survive the correction at 5% and 69 at 1%. The histogram of p-values has the shape you hope for, a spike near zero over a flat floor. The height of that floor (Storey's estimate) says that about half the questions had nothing to find.

![Inside the market most tests found something. Outside it, few did](https://future-seems-so-good.com/blog/assets/a-mirror-not-a-thermometer/charts/registry.svg)

Split the registry by what each test asked and the pattern of this essay appears. Of 88 tests about how the market works (prices, reviews, reviewers, words), 50 survive. Of 73 about the calendar and the clock, 31. Of 155 about shocks from outside (the economy, markets, events, laws, the weather), 14. And of those 14 I believe about half. The first quarantine, the club match days, the two unemployment results, the density of listings against land prices and the in-sample link between GDP and the real price look real. The others don't. The day after the national team draws has no mechanism. An Irish law of 2017 that criminalised buying sex comes with a rise in the share of foreign women in Brazilian listings, with no channel, while three other foreign laws did nothing. The correlations with DraftKings shares and with searches for "tigrinho", an online slot game, are two series rising together after 2022. Beverage production was the pandemic and disappears without 2020–22. The ratio of men to women was urbanisation and disappears with a population control. The real minimum wage and the entry of new providers move together because the good years of 2006–14 pushed both. And capitals run by mayors of the centre see 5 points less annual real price growth than those run by the left (randomization p = 0.006, q = 0.02), with mayors of the right showing almost the same gap just short of the correction (q = 0.07), in a panel of 13 capitals with no close-election design and no channel I can imagine. That one is a lead to investigate, not a finding.

The registry also records the tests that died, and those taught me the most. Each of these changed a conclusion:

- A dummy for a single day, with robust standard errors, gives p ≈ 0 by construction. Every event was redone with placebo windows, and all but one disappeared.
- Trending series correlate for free. Platform mentions and complaints about fake photos correlated at −0.62 in levels, as if the platforms had cleaned up the ads. In first differences the correlation is −0.05 (p = 0.83).
- Posts sampled by page come in clumps of fifteen from the same thread and day. Daily counts use one record per thread per day, or the significance inflates.
- Thirteen capitals are too few clusters for clustered standard errors, so the political tests use randomization.
- Local effects are measured against each state's own baseline, not the national average.
- *Segunda* is also an ordinal, so the visit-day parser needed an article or "-feira".
- The same phone isn't the same person, so price pairs require the same name as well.
- "Louvor" was "aprovada com louvor", not gospel music. Lexicons are audited on samples before they become variables.
- Small counts make spurious elasticities. A first pass found a "significant" link between mentions of music genres and their attention on Wikipedia. There were 380 mentions of genres in 42,889 posts, and a minimum-count filter killed it.
- Treatments that vary by year, like election campaigns, need permutation across years, not standard errors by day.
- Every national correlation after 2020 was redone without the pandemic. Beverages and OnlyFans were only COVID.

Some sources never made it in. Night-time lights from satellites need a NASA login and tens of gigabytes a month. The big American and British review sites have no public microdata, so the international comparison uses published averages and current advertised prices. I excluded a popular annual report on adult-site searches on principle, because its categories include terms that sexualise minors. Police records have no standard series on fraud by state, so I used deaths by assault instead. The electoral court's portal blocks automated downloads, so the mayors come from public tables that reproduce its results. Inside Airbnb publishes only its current snapshot, the national hotel survey exists only for 2011 and 2016, there's no dated calendar of large conventions, and the forum has no subforum in Rondônia or Pará, where the big dams were built.

## Where the mirror leaks

The case against this essay is strong, and I've made most of it in passing. Here it is in one place.

It's a clients' forum. It measures the men who write, a narrow group whose composition changes over time: one reviewer in ten wrote more than half of the reviews. The women's side, what they earn, what they keep, what they risk and whether they chose to be there, is absent. A price paid isn't an income. The Economist made this point about its own decline: platforms cut out brothels, agencies and pimps, so workers' incomes may have fallen less than prices.

The review sample is 3% of the forum's pages and 71% São Paulo. Tests across states rest on 10 to 15 units, and most of them have little power. The violence result has eleven points and a q of 0.07.

Much of it is words, not measurements. Bodies, skin colour, payment methods and services come from what a client chose to write. "Mature" is an adjective, not an age.

The visit date exists only when a reviewer writes it, in 8,929 of 31,514 reviews, and the men who write "yesterday" may not be typical.

A phone is not a person. The rebranding and pricing results use only phones with few listings and a repeated name, which makes them cleaner and smaller.

The correction doesn't make the survivors true. At 5%, about five of the 95 survivors should be false. I named several candidates above, and some of the ones I believe could be among them.

The index holds duration, venue and state fixed. It can't hold fixed what happens in the room, and if the services changed, some of the fall in the real price is a change in the product.

And the forum can't see the part of the market that moved elsewhere. Platform mentions rose from almost nothing to 10–15% of posts, the forum gets quieter every year, and whatever moved to platforms and messaging apps may price very differently.

I accept all of this. What survives is narrower than the title suggests. The forum's prices are sticky in a way the standard model doesn't predict. They fell in real terms while comparable services rose. And they carry no forecasting information about the economy beyond what GDP's past and the IPCA already carry. The forum's activity follows the family calendar, paydays and local clocks closely, and it registers a few outside shocks: the pandemic, unemployment, Pix. That's a mirror. It isn't a thermometer.

## What a mirror is worth

This part is speculation, and I'll mark where it stops being evidence.

The evidence goes this far. A forum written by a narrow group of clients reproduces, with good precision, the rhythm of family life, the legal pay cycle, the clock of each region, the spread of a payment technology and, loosely, the violence of each state. Its prices behave like prices in no textbook. And when asked to forecast the economy it fails in the way alternative data usually fails: a correlation in sample, nothing out of sample once the benchmark knows the obvious.

Past that point, these are my guesses. I think corpora like this one are worth more as social sensors than as economic ones. Official statistics are thinnest exactly where this data is thickest: informal work, behaviour people don't report to surveys, the speed at which something spreads through a population that doesn't answer questionnaires. Pix reaching this market in half the time it took the country is the kind of fact a central bank would want and couldn't get from its own data. So are the hours when people are free and which holidays empty the streets.

The second guess is about cost. This project took four days, a laptop and an agent. The crawl, the 316 tests and the 37 charts cost less than a week of one person's attention. When anyone can run 316 tests on a forum over a weekend, the number of false findings built on alternative data will grow at the same rate, unless the registry and the correction come with the pipeline. [Money is method times compute times data](https://future-seems-so-good.com/blog/money-is-method-times-compute-times-data) puts the rule for research agents this way: "A null or negative result, clearly measured, is a successful run." For agent-run research I'd go further. A run that doesn't report its dead tests hasn't finished.

The third is about privacy. The same pipeline, without the hash, would be a tool for finding people. Hashing the phones before any analysis cost a few lines of code. It should be the default in any crawler an agent writes, not a decision someone has to remember.

The last is about time. The mirror is fading. Replies per thread went from three to one, a review is about a third of the length it was in 2011, and the market's conversation is moving to platforms and private messages that nobody can crawl. If there's anything left to learn from this kind of data, the window is closing, and the next version of this essay may not be possible.

## A spec for the harness

Here's the whole pipeline as a pseudo-config. Its job is to tell you whether a forum is a thermometer, and the default answer should be no.

```yaml
# forum-mirror.yaml (pseudo-config)
question: does a forum's price carry information about the economy?

collection:
  runs_in: ordinary_browser_tab        # bot check passed once, by a person
  account: none                        # guest-visible pages only
  queue: indexeddb                     # pending jobs, seen keys, priority classes
  order: [all_listing_pages, shuffled_thread_pages_by_class]
  workers: {min: 4, max: 12, back_off_on: [429, 5xx, bot_check]}
  timers: none_in_hot_path             # hidden tabs wake once a minute
  outbox: {in_memory: true, max_pages: 3000, batch: 300, every_s: 10}
  delete_job_when: receiver_acknowledges
  receiver: {host: 127.0.0.1, format: gzip_jsonl, files: append_only}
  split_lines_on: "\n"                 # splitlines() also breaks on U+2028

privacy:
  phones: area_code + salted_hash      # before any analysis
  names: working_names_only
  never: [name_next_to_complaint, identify_individuals, sources_that_sexualize_minors]

units:
  daily_counts: one_record_per_thread_per_day
  visit_day: parsed_from_text          # posting date is not visit date
  same_seller: same_phone AND same_name
  lexicons: audited_on_samples

inference:
  dated_events: placebo_windows(n=300)
  few_clusters: randomization
  trending_series: first_differences
  post_2020: rerun_without(2020-2022)
  small_counts: minimum_count_filter
  fixed_effects: absorbed              # 8 GB of RAM

forecasting:
  split: rolling_origin(start=2012)
  benchmarks: [own_lags, own_lags + ipca]
  tests: [clark_west, diebold_mariano]
  correct_for: number_of_tries

registry:
  every_test: [family, estimate, p, ci, n]
  correction: benjamini_hochberg(q=0.05, across=all_tests)
  report: [survivors, survivors_you_dont_believe, tests_that_died]

verdict: MIRROR | THERMOMETER          # MIRROR until a held-out forecast says otherwise
```

Five things in it matter more than the rest:

1. A job leaves the queue only when the receiver says its page is on disk. That one rule turned a crawler that crashed, stalled, got blocked and was reloaded by mistake into one that never lost a page it had counted.
2. The shuffle. Randomising thread pages within classes is what made 3% of the forum a sample instead of a fragment. I didn't know the crawl would stop at 3%, and I didn't need to.
3. Placebo windows for anything with a date. Robust standard errors on a one-day dummy will find an effect every time. The placebo is what removed the World Cup, the Olympics and the pope.
4. One registry, one correction, including the dead. Writing down all 316 tests is what makes it possible to say which 95 survived, and to say out loud that about half of the outside survivors are probably false.
5. A benchmark that already knows the obvious. The forum beats an autoregression on GDP some of the time. It doesn't beat one that also knows inflation, and inflation is published every month.

If this harness runs on another forum and says THERMOMETER, someone has found an indicator worth publishing. On this one it says MIRROR. That's a smaller claim than the one I started with, and the only one I can defend.

## Sources

- Guillermo A. Calvo, [Staggered prices in a utility-maximizing framework](https://doi.org/10.1016/0304-3932(83)90060-0), Journal of Monetary Economics 12(3), 1983
- Emi Nakamura and Jón Steinsson, [Five Facts about Prices: A Reevaluation of Menu Cost Models](https://doi.org/10.1162/qjec.2008.123.4.1415), Quarterly Journal of Economics 123(4), 2008
- William J. Baumol and William G. Bowen, *Performing Arts: The Economic Dilemma*, 1966; William J. Baumol, [Macroeconomics of Unbalanced Growth: The Anatomy of Urban Crisis](https://www.jstor.org/stable/1812111), American Economic Review 57(3), 1967
- Scott Cunningham and Todd D. Kendall, [Prostitution 2.0: The changing face of sex work](https://doi.org/10.1016/j.jue.2010.12.001), Journal of Urban Economics 69(3), 2011; the 1998–2008 price figures are from the working-paper version, Prostitution 2.0: The Internet and the Call Girl, 2009
- The Economist, [More bang for your buck: How new technology is shaking up the oldest business](https://www.economist.com/briefing/2014/08/07/more-bang-for-your-buck), 2014-08-07
- Yoav Benjamini and Yosef Hochberg, [Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing](https://doi.org/10.1111/j.2517-6161.1995.tb02031.x), Journal of the Royal Statistical Society B 57(1), 1995
- John D. Storey, [A direct approach to false discovery rates](https://doi.org/10.1111/1467-9868.00346), Journal of the Royal Statistical Society B 64(3), 2002
- Todd E. Clark and Kenneth D. West, [Approximately normal tests for equal predictive accuracy in nested models](https://doi.org/10.1016/j.jeconom.2006.05.023), Journal of Econometrics 138(1), 2007
- Francis X. Diebold and Roberto S. Mariano, [Comparing Predictive Accuracy](https://doi.org/10.1080/07350015.1995.10524599), Journal of Business & Economic Statistics 13(3), 1995
- Robert F. Engle and C. W. J. Granger, [Co-integration and Error Correction: Representation, Estimation, and Testing](https://doi.org/10.2307/1913236), Econometrica 55(2), 1987
- A. P. Dawid and A. M. Skene, [Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm](https://doi.org/10.2307/2346806), Journal of the Royal Statistical Society C 28(1), 1979
- Vincent D. Blondel, Jean-Loup Guillaume, Renaud Lambiotte and Etienne Lefebvre, [Fast unfolding of communities in large networks](https://doi.org/10.1088/1742-5468/2008/10/P10008), Journal of Statistical Mechanics, 2008
- Leland McInnes, John Healy and James Melville, [UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction](https://arxiv.org/abs/1802.03426), 2018
- Bruce M. Hill, [A Simple General Approach to Inference About the Tail of a Distribution](https://doi.org/10.1214/aos/1176343247), Annals of Statistics 3(5), 1975
- Timothy J. Bartik, [Who Benefits from State and Local Economic Development Policies?](https://doi.org/10.17848/9780585223940), W.E. Upjohn Institute, 1991
- Sergio Correia, [Linear Models with High-Dimensional Fixed Effects: An Efficient and Feasible Estimator](https://scorreia.com/research/hdfe.pdf), working paper, 2016
- Thomas Hale et al., [A global panel database of pandemic policies (Oxford COVID-19 Government Response Tracker)](https://doi.org/10.1038/s41562-021-01079-8), Nature Human Behaviour 5, 2021
- Chrome for Developers, [Heavy throttling of chained JS timers beginning in Chrome 88](https://developer.chrome.com/blog/timer-throttling-in-chrome-88), 2021
- WICG, [Private Network Access](https://wicg.github.io/private-network-access/), draft specification
- IBGE, through [SIDRA](https://sidra.ibge.gov.br) and its [data APIs](https://servicodados.ibge.gov.br/api/docs): IPCA and its sub-items by metropolitan area; quarterly and state GDP; PNAD Contínua (unemployment by state and sex, construction employment, income); Censo 2010 and 2022 (religion, population by sex and age); first names by decade of birth (Censo 2010); Estatísticas do Registro Civil; PIM-PF (beverages); Desigualdades Sociais por Cor ou Raça no Brasil, 2022
- DIEESE, [Pesquisa Nacional da Cesta Básica de Alimentos](https://www.dieese.org.br/cesta/) (São Paulo), via [Ipeadata](http://www.ipeadata.gov.br); Ipeadata, minimum wage
- Banco Central do Brasil: [Pix statistics](https://dadosabertos.bcb.gov.br/dataset/pix) by municipality; IBC-Br and exchange rates from the [SGS](https://www3.bcb.gov.br/sgspub/)
- Ministério da Saúde: SIM (deaths by assault) and SINAN (acquired syphilis) via [DATASUS](https://datasus.saude.gov.br); Vigitel
- Receita Federal, [Dados Abertos do CNPJ](https://dados.gov.br/dados/conjuntos-dados/cadastro-nacional-da-pessoa-juridica---cnpj), September 2026 snapshot
- [B3](https://www.b3.com.br) (Ibovespa); [FRED](https://fred.stlouisfed.org) (VIX, commodity prices); Yahoo Finance (share prices)
- [ANAC](https://www.gov.br/anac/pt-br/assuntos/dados-e-estatisticas) (air passengers); Ministério do Turismo, [international arrivals](https://dados.turismo.gov.br/dataset/chegada-de-turistas-internacionais); [UNHCR](https://www.unhcr.org/refugee-statistics/) (Venezuelan refugees and migrants); [Inside Airbnb](https://insideairbnb.com)
- [Google Trends](https://trends.google.com); [Wikimedia pageviews](https://pageviews.wmcloud.org); World Bank, [World Development Indicators](https://databank.worldbank.org/source/world-development-indicators) (PPP conversion factors, GDP per capita)
- [Open-Meteo](https://open-meteo.com) historical weather for 11 state capitals; public match results for the Brasileirão (2003–2024) and the national team; [Jolpica F1](https://github.com/jolpica/jolpica-f1) calendar
- [São Paulo real estate, April 2019](https://www.kaggle.com/datasets/argonalyst/sao-paulo-real-estate-sale-rent-april-2019) (Kaggle); São Paulo district boundaries from [GeoSampa](https://geosampa.prefeitura.sp.gov.br), the city's geodata portal, via [codigourbano/distritos-sp](https://github.com/codigourbano/distritos-sp)
