# Hire people who close loops

*September 2026*

When Facebook bought Instagram in April 2012 for about $1 billion, Instagram had [13 full-time staff who work on a single smartphone app](https://www.bbc.com/news/business-17666032). When it bought WhatsApp in February 2014 for $16 billion plus [$3 billion in restricted stock](https://about.fb.com/news/2014/02/facebook-to-acquire-whatsapp/), WhatsApp [employed just 55 people](https://www.geekwire.com/2014/facebook-buys-whatsapp-messaging-service/) and had more than 450 million monthly users. Put those numbers next to François Chollet's definition from [On the Measure of Intelligence](https://arxiv.org/abs/1911.01547): "The intelligence of a system is a measure of its skill-acquisition efficiency over a scope of tasks, with respect to priors, experience, and generalization difficulty." By that definition, those were two very intelligent companies.

My thesis is that agents have turned those two companies from outliers into the default shape, and hiring hasn't noticed. When agents do the typing, the scarce person is the one who closes the whole loop alone: research, engineering and product, with a sensor at the end that says whether it worked. That person picks problems for their structure, not their prestige, and leaves a public trace you can check. The usual filters still measure the old bottleneck (years of a framework, a degree), and the usual rebellion against them swaps in brainteasers and vibes. Both miss. Measure how fast someone turns a new game into skill. The selection research says you can get closer to that than most companies try.

## The typing is done

I think the future belongs to very small teams that can hold a lot of context together, because every person you add makes shared context harder to advance. When implementation is cheap, contrarian ideas are the ones that survive, since the obvious ones get built by everyone at once. An organization that stays with the status quo gets eaten by one with a faster positive feedback loop. The unit I believe in is three or four high-taste people running a few loops that automate most of the work. The humans are shapers: they think about sparse reward functions and steer the loops. A team like that can now outproduce an organization many times its size.

"Shaper" is about reward horizons. Current models were trained on short loops, so they're good when the reward is dense (a test passes, a query returns) and weak when it's sparse and far away, like "make a product people keep using." A person who can hold the sparse reward in their head and break it into dense loops an agent can run gets the output of the whole swarm. A person who only works inside one dense loop is now competing with the agent. I tweeted in December: "The agent did the typing; I did the thinking. That's probably the right division of labor." Hire for the thinking. The typing is already on your payroll as tokens.

A corollary I keep coming back to: one medium-taste idea that ships beats many high-taste ideas nobody implemented. A shaper is someone who closes. I've argued that a company is a set of loops you can [find and rank](https://future-seems-so-good.com/blog/find-the-loops), each of which should [show agents the problem and hide the metric](https://future-seems-so-good.com/blog/show-the-problem-hide-the-metric). This essay is about the people on top of those loops.

## Intelligence is the ability to play games

When I think about who to hire, I want to skip the usual path, where a person proves they can program and then proves it again with years of a framework. Intelligence is the ability to play games, and the people I want are smart in that wider sense, not only at programming. I don't want nerds. I want real hackers.

The nerds I mean are the ones for whom programming is a credential: they know the framework, pass the LeetCode round, and have never shipped anything nobody asked for. That profile made sense when implementation was the bottleneck.

"The ability to play games" is my compression of Chollet. A game is an environment with rules you didn't write and feedback you can't argue with. Skill at one game is a stock. The rate at which you pick up the next one is the flow, and Chollet's point is that the flow is the intelligence: "skill is merely the output of the process of intelligence." So evaluate [the trajectory](https://future-seems-so-good.com/blog/intelligence-is-a-trajectory), how someone moves through a new problem, not a snapshot of what they know. (I made the model version of this argument in *How to achieve superintelligence*.) People who have played many games have the meta-skill of reading a new one fast, and a company is a very large game with bad documentation. Follow that far enough and capitalism is the final game.

## Real hackers

The Mentor wrote "The Conscience of a Hacker" on January 8, 1986, shortly after his arrest, and Phrack published it in [issue 7](https://phrack.org/issues/7/hackers-manifesto.html) that September. It's remembered for its ending:

> Yes, I am a criminal. My crime is that of curiosity. My crime is that of judging people by what they say and think, not what they look like. My crime is that of outsmarting you, something that you will never forgive me for. I am a hacker, and this is my manifesto.

The middle is more useful for hiring. The bored kid finds a computer: "It does what I want it to. If it makes a mistake, it's because I screwed it up. Not because it doesn't like me." And the refrain from the adults is: "Damn kid. All he does is play games."

That's a closed loop: honest feedback, an owned error, iteration with nobody grading. The adults were looking at the intelligence and calling it a waste of time.

In [The Word "Hacker"](https://www.paulgraham.com/gba.html) (April 2004), Paul Graham says that to programmers the word "connotes mastery in the most literal sense: someone who can make a computer do what he wants—whether the computer wants to or not," and that the disobedience comes with the package: "They may laugh at the CEO when he talks in generic corporate newspeech, but they also laugh at someone who tells them a certain problem can't be solved. Suppress one, and you suppress the other." In [Great Hackers](https://www.paulgraham.com/gh.html) (July 2004) he adds the property that matters most for small teams: "Great hackers tend to clump together."

In a July draft I wrote that the hacker wanted to understand the system from below, while the modern programmer is rewarded for occupying a narrow position inside a stack nobody understands. "The hacker was domesticated into a software employee." Agents may reverse this. When syntax is automated, the narrow position disappears and the older job comes back: seeing the whole machine and bending it. (More in [Bottom-up is one ontological level higher](https://future-seems-so-good.com/blog/bottom-up-is-one-ontological-level-higher).)

**Taste is part of the definition.** Graham's [Taste for Makers](https://www.paulgraham.com/taste.html) (February 2002) begins with a friend who teaches at MIT and is inundated with applications from would-be graduate students: "A lot of them seem smart," he said. "What I can't tell is whether they have any kind of taste." And to the relativist: "Saying that taste is just personal preference is a good way to prevent disputes. The trouble is, it's not true." My working definition, from the July draft: taste is the perception of a global property before you can enumerate the local causes. My blunter belief is that taste can't be taught: either a person has it or they don't. I'll come back to why that sentence is dangerous.

## What the selection research says

The anti-credential instinct is older than agents, and some of it has been measured.

**Work samples beat pedigree, though by less than people quote.** Schmidt and Hunter's 1998 meta-analysis in Psychological Bulletin summarized 85 years of research on 19 selection methods. In their [Table 1](https://talytica.com/wp-content/uploads/2016/10/Schmidt-and-Hunter-1998-Validity-and-Utility-Psychological-Bulletin.pdf), work sample tests have the highest validity for predicting job performance (.54), followed by general mental ability tests and structured interviews (.51 each) and job knowledge tests (.48). Unstructured interviews are at .38. Years of job experience are at .18. Years of education are at .10. Age is at -.01.

In 2022, Sackett, Zhang, Berry and Lievens [showed](https://doi.org/10.1037/apl0000994) that most of those numbers had been inflated by overcorrecting for range restriction. Their follow-up in [Industrial and Organizational Psychology](https://www.cambridge.org/core/services/aop-cambridge-core/content/view/A20984B138319E3D432E643978BF026D/S175494262300024Xa.pdf/div-class-title-revisiting-the-design-of-selection-systems-in-light-of-new-findings-regarding-the-validity-of-widely-used-predictors-div.pdf) puts the old and new estimates side by side:

![Validity of selection methods for job performance, Schmidt & Hunter 1998 against the Sackett et al. 2022 revision](https://future-seems-so-good.com/blog/assets/hire-people-who-close-loops/charts/selection-validity.svg)

The ranking mostly holds and the magnitudes shrink. Structured interviews move to the top (.42), then job knowledge tests (.40), work samples (.33) and general mental ability (.31). Years of experience fall to .07. The same paper cites a meta-analysis of 113 twenty-first-century studies that puts general mental ability at .23.

So a CV is close to noise, and raw-intelligence testing isn't the answer either. The winners follow what Sackett's group calls the "sample" strategy: "Conducting a careful job analysis and designing measures to sample job behaviors ... emerges as more effective, on average, than a 'sign' strategy of identifying psychological constructs judged as relevant to the job in question." If the work is closing loops with agents, the test should be closing a loop with agents.

**Brainteasers predict nothing.** In 2013 Google's head of people operations, Laszlo Bock, told [The New York Times](https://www.nytimes.com/2013/06/20/business/in-head-hunting-big-data-may-not-be-such-a-big-deal.html) that when Google compared tens of thousands of interviews with later performance, "We found zero relationship. It's a complete random mess, except for one guy who was highly predictive because he only interviewed people for a very specialized area, where he happened to be the world's leading expert." On puzzles, in a passage [quoted on HN](https://news.ycombinator.com/item?id=9484526): "brainteasers are a complete waste of time ... They don't predict anything. They serve primarily to make the interviewer feel smart." The one predictive interviewer is the most useful detail: a domain expert judging people in his own domain. That's a work sample with a human sensor.

**Jane Street's puzzles are a different game.** Jane Street keeps a [puzzles page](https://www.janestreet.com/puzzles/) that explains its puzzle-heavy culture: "The act of solving puzzles, though that might seem abstract, is intrinsic to the work we do at Jane Street." For a trading firm a probability puzzle is close to a work sample, which is why it can work there and fail at Google. The test has to be a slice of the actual game.

**Esoteric signals filter through self-selection.** Graham's [The Python Paradox](https://www.paulgraham.com/pypar.html) (August 2004): "if a company chooses to write its software in a comparatively esoteric language, they'll be able to hire better programmers, because they'll attract only those who cared enough to learn it." Google ran the famous version that year: a billboard on Highway 101 and banners at Harvard Square [said only](https://www.npr.org/templates/story/story.php?storyId=3916173) "{first 10-digit prime found in consecutive digits of e}.com", with no company name. The answer, [7427466391](https://news.ycombinator.com/item?id=6377129), led to a harder puzzle and then to a recruiting page. The same company later found brainteasers useless in interviews, and both findings can be true. The billboard didn't evaluate anyone. It changed who showed up.

**"High agency" is a 2018 word.** Much of this now travels under that label. George Mack's thread ["1/ HIGH AGENCY"](https://x.com/george__mack/status/1068238562443841538), posted November 29, 2018, says he had been thinking about the concept "every week for the last two years since I heard @ericweinstein discuss it on @tferriss' podcast." Hacker News' own search has no match before a passing "high-agency" in a [February 2017 comment](https://news.ycombinator.com/item?id=13705420). The first comment I found that defines the term is from [March 2018](https://news.ycombinator.com/item?id=16657652): "a high agency individual looks for and finds solutions." Mack's thread [reached HN](https://news.ycombinator.com/item?id=18640273) that December. Here is how often the phrase appears on HN per year, counted with HN's Algolia search (stories and comments; a few are about "high agency environments" and not people). Algolia's yearly totals for all items are approximate, so the last column is too:

| Year | HN items matching "high agency" | All HN items (approx., millions) | Per million HN items |
|---|---:|---:|---:|
| 2017 | 1 | 2.51 | 0.4 |
| 2018 | 4 | 2.64 | 1.5 |
| 2019 | 3 | 2.90 | 1.0 |
| 2020 | 9 | 3.01 | 3.0 |
| 2021 | 6 | 3.37 | 1.8 |
| 2022 | 7 | 3.46 | 2.0 |
| 2023 | 13 | 3.81 | 3.4 |
| 2024 | 43 | 3.18 | 13.5 |
| 2025 | 121 | 2.84 | 42.6 |
| 2026 (to Sep 25) | 111 | 2.20 | 50.4 |

HN's total volume was roughly flat over this period, so the raw counts tell the same story as the normalized ones. The phrase is about thirty times more common in 2025 than in 2018, and 2026 has nearly matched 2025 with three months left. Nearly all the growth comes after 2023, which is when I'd date the shift to coding agents. That fits my thesis, and it's a warning. Once a trait becomes a hiring slogan, candidates learn to perform it.

## The 10-minute challenge

The screen I'd use for this kind of person is a challenge with one property: someone with the profile solves it in ten minutes, and someone without it needs a day. It can't require knowledge of cybernetics or any other canon, only intelligence. Sometimes the person shouldn't even need to code. It measures the ability to play a game they haven't seen.

It's the asymmetry I want from evals for agents, where [generation is hard and verification cheap](https://future-seems-so-good.com/blog/evals-built-backwards), with time as the discriminating variable. Ten minutes against an eight-hour day is roughly fifty to one. Scores saturate, because strong and mediocre candidates both get there eventually. Time is Chollet's denominator, experience spent per unit of skill, and it spreads people out.

That ratio needs a specific structure: an obvious path that works but is slow (brute force, grinding cases, hand-tuning), and a structural insight that collapses the search. The insight can't be trivia, or you're testing whether someone read the same book as you. It has to be something a person who plays many games would notice by looking at the system.

A sketch: a tiny simulated shop with one price to set, a demand curve you can't see, and a sales log that reports each result three steps late. Maximize profit over 200 steps. The slow path hill-climbs on the latest number, oscillates because of the delay, and spends hours on smoothing heuristics. The fast path notices the delay in the loop and compensates for it, or fits the curve from the log. Someone who sees the loop is done in minutes. Someone who sees only a number burns the afternoon, with or without an agent.

Agents should be allowed, because the job will have them. That's the real design constraint: if a frontier agent solves the challenge alone, it measures typing and must be thrown away. So every challenge is calibrated before a candidate sees it. Run an agent alone with a fixed budget, run it with someone you know is the profile, and run it with a competent programmer who isn't. Keep it only if the first fails, the second is fast, and the gap to the third is large. Measure the sensor before you trust it.

I haven't run this yet, and the fifty-to-one ratio is a design target, not a measurement.

## The job post is a filter

Before the challenge, someone has to find you. Every time I've asked an agent to draft a job description, it came back too corporate. The post I want has no challenge and no product pitch in it. It's dense with references from deep in the iceberg of esoteric knowledge, so the people you want find it on their own and like it.

That's the Python paradox moved from languages to references. A post dense with Ashby, Pask, the Phrack manifesto, Factorio and Move 37 is noise to most readers and a signal to a few, and the few who recognize half of it are the ones who apply.

The other half is the trace. The four traits I look for rarely appear together: people who close the whole loop alone across research, engineering and product; who choose problems out of structural interest, not prestige; who leave a verifiable public trace; and who started early and unsupervised, with game mods, bots or tools built because they wanted them to exist. Talent that leaves no trace is illegible. And the negative filter: résumé optimization is the opposite of what I'm looking for.

The trace is also the fairest version of the Mentor's standard, "judging people by what they say and think, not what they look like." You can read someone's repository without knowing where they studied.

A third source is the social graph: rank people who follow at least three of a small set of accounts you respect (a few researchers, a forecasting group, a couple of philosophy accounts), corrected against popularity. It's a weak signal, but it finds people who never apply to anything.

## The first ten set the culture

If I could put one sentence at the top of any hiring plan, it would be this: the first ten people define the technical culture and the ceiling of what the team will build.

The mechanism is a feedback loop. In a July draft I borrowed Hermann Haken's slaving principle, where interacting parts produce an order parameter that then governs the parts: "A company culture emerges from behavior and then constrains who gets hired." The first hires are the parts. After them, the culture does the hiring, through referrals and through who feels at home in the interviews. Graham's clumping is the same loop from the inside, and it compounds early mistakes as well as early wins.

The PayPal alumni are the famous case. According to [Wikipedia](https://en.wikipedia.org/wiki/PayPal_Mafia), after eBay bought PayPal, "within four years all but 12 of the first 50 employees had left," and the group went on to found or develop LinkedIn, Palantir, SpaceX, YouTube, Yelp and others. It's the picture most people have of a first ten that compounds, mine included. I'll come back to what else that page says.

Per loop, the number is small: three or four shapers, all in the same room. People weren't built to hold shared context through a screen; we hunted together. Co-location is the cheapest way to share context. The first ten are two or three loops' worth of shapers, and whatever they tolerate becomes the culture.

## Two kinds of loop

On August 1 I asked a friend who he would hire for deep analytical work, and his answer changed how I think about the profile:

> uma coisa que tenho reparado é que é dificil encontrar gente high agency low cortisol
>
> mas pra esse tipo de trabalho, talvez um outro perfil seja melhor (menos tipo CEO e mais analítico/letargico/estático tipo gwern)

In English: it's hard to find people who are high agency and low cortisol, and for this kind of work another profile may be better, less CEO-like and more analytical, lethargic, static, like [gwern](https://gwern.net/about). A few minutes later I added: "a vida do tupiniquim high iq low cortisol high taste no brasil nn é facil." Life isn't easy for the high-IQ, low-cortisol, high-taste Brazilian.

His point is that "loop-closer" hides two archetypes. The operator has high agency, ships fast and stays calm doing it. The calm is the rare part, because agency with high cortisol turns into thrash. The analyst is patient and deep, happy to spend a month inside one question and write the definitive thing about it. A product loop with a daily sensor needs the first. A research or investment loop, where the reward arrives in quarters and the main risk is acting too early, needs the second.

A ten-minute challenge favors the operator. A gwern-like analyst might spend a day on it and produce a better answer than anyone asked for. So the challenge has to match the loop, and one archetype for every loop is overfitting to whoever writes the post, in this case me.

## Where the filter leaks

Some of my own instincts are the weak points.

**Successful founders are middle-aged.** For a long time I preferred very young people, in their early twenties or younger. The best evidence I know points the other way. Azoulay, Jones, Kim and Miranda used U.S. Census administrative data on growth-oriented startups in [Age and High-Growth Entrepreneurship](https://www.nber.org/papers/w24489): "successful entrepreneurs are middle-aged, not young. The mean founder age for the 1 in 1,000 fastest growing new ventures is 45.0." Also: "Prior experience in the specific industry predicts much greater rates of entrepreneurial success. These findings strongly reject common hypotheses that emphasize youth as a key trait of successful entrepreneurs." And in Schmidt and Hunter's table, age predicts job performance at -.01, which is zero.

I concede the preference. Founders aren't early employees, so the paper doesn't settle the question, but the preference was never backed by evidence. It was a proxy for properties: a lot of exploration, little sunk cost in a career, not yet domesticated into a narrow position. The trace and the challenge measure those directly, and a 45-year-old who still has them should get the same shot. A 2018 HN comment says it better: "a founder is a high agency individual, are you really going to let age (whether it be too old or too young) [stop you?](https://news.ycombinator.com/item?id=18214936)" The evidence I'm leaning on has its own tension here. Sackett's group warns that "a number of the predictors at the top of our list in terms of validity are not generally applicable for entry-level hiring," and names work samples and job knowledge tests. If you insist on very young hires, you're choosing the population where your best tools work worst.

**Age filters are also illegal.** In Brazil, [Lei 9.029/1995](https://www.planalto.gov.br/ccivil_03/leis/l9029.htm) prohibits "qualquer prática discriminatória e limitativa para efeito de acesso à relação de trabalho" on grounds that include age. In the U.S., the [ADEA](https://www.eeoc.gov/age-discrimination) protects applicants 40 or older, and a policy that applies to everyone can still be illegal if it has a negative impact on that group and isn't based on a reasonable factor other than age. An age range in a job post or a hiring plan is a liability even if nobody acts on it. I'm not a lawyer, but that's enough to keep age out of both.

**Taste filters produce homogeneity.** This is the strongest objection, because the mechanism that makes an esoteric post work is the one that breaks it. People recognize their own references. A post full of Land and Deleuze selects people who read Land and Deleuze, and the first-ten loop amplifies whatever the first filter let through. The PayPal page has the detail I promised: "Most of the members attended Stanford University or the University of Illinois Urbana-Champaign." My model of an anti-credential network came mostly from two universities. "Taste can't be taught" is exactly the sentence that lets an interviewer call their own bias taste. And Sackett's group reports sizable subgroup gaps even for work samples, so moving from CVs to challenges doesn't make selection neutral.

Ashby has the right principle. His law of requisite variety says ["only variety can destroy variety"](https://en.wikipedia.org/wiki/Variety_(cybernetics)), and I think it applies to teams: a hard program can only be good with a lot of variety, meaning ideas that are profound, even esoteric, though not necessarily orthogonal. Ten people who read the same canon have low variety, however deep the canon. So references must be a door and never a wall: the challenge must be solvable by someone who recognizes none of them, the references should span unrelated fields (games, biology, music, markets, security), and you should measure who the post brings in.

**Esoteric posts exclude good people.** A post that reads like a riddle filters out people who are busy or don't read English comfortably. The trace filter misses people with no trace for reasons unrelated to talent, like an NDA or a job that leaves no free evenings. My friend's gwern-like analyst may never answer a post written in CEO energy. A 2012 [HN comment](https://news.ycombinator.com/item?id=4302874) has the risk and the fix: "unless you're extremely certain that your tests hit the sweet spot, it seems really dangerous to make them a mandatory first filter," and instead "allow someone to 'skip the queue' by sending a test solution along with the application." So build two doors: the esoteric post, and a plain one that says: here's the challenge, solve it, we'll talk.

**Tiny-team stories are survivor stories.** The 13-person teams that failed don't get BBC articles. A small team is a precondition for this kind of output, not a guarantee of it.

## Speculation: challenges decay

Every challenge has a half-life. The calibration step, where an agent alone must fail, gets harder to pass with each model release. Eventually the only part an agent can't do alone will be the part I care about most: choosing which loop to close and which sensor to trust. Challenges will then have to test problem choice directly, say by showing a candidate a messy system and asking which of five loops they'd close first, scored against what happened when someone closed each one. Hiring becomes benchmark maintenance, with private challenges rotated like a held-out eval and retired when models solve them. The last interview question might become "write a challenge that separates people by fifty to one." I have no data for any of this yet.

## A hiring loop spec

Each stage has a sensor, and you measure the whole process as a loop.

**1. Two doors, no proxies.** One dense, esoteric post and one plain post with the challenge link. Neither mentions age, degree or years of experience. Both say candidates may skip the queue by submitting a solution.

**2. Public-trace check.** Score each candidate from links, not a CV. Hide name, school and age from the scorer where possible.

| Trait | 0 | 1 | 2 |
|---|---|---|---|
| Closes the loop | Only one layer (research, code or product) | Two layers | Research, engineering and product in one artifact, with a result |
| Problem choice | Toy problems, tutorial clones, prize-driven | Solid but local | A problem chosen for its structure, often against the obvious path |
| Trace | None, or only certificates | Some public work, thin reasoning | Dense writeups or repos that show original thought |
| Unsupervised start | Only assigned work | Some side projects | Built hard things before anyone asked, with depth growing over time |

**3. The 10-minute challenge.** Design and calibration:

```yaml
challenge:
  shape: slow_path_works + structural_insight_collapses_search
  requires: no trivia, no domain jargon, coding optional
  agents: allowed
  scorer: deterministic (simulator score + wall-clock time)
  calibration:           # run before any candidate sees it
    agent_alone:         { budget: 1h, must: fail_or_score_below_target }
    agent_plus_profile:  { must: reach_target, time: "<= 15 min" }
    agent_plus_typical:  { target_time_ratio_vs_profile: ">= 20x" }
  retire_when: agent_alone reaches target
  variants: operator (fast, daily-sensor loop) | analyst (open-ended, graded on depth)
```

**4. Paid work sample.** Half a day on a real loop with agents, paid. Scored by an expert in that loop, the only kind of interviewer Bock found predictive.

| Dimension | What a 2 looks like |
|---|---|
| Loop closed | Something runs end to end and a sensor says whether it worked |
| Sensor | Deterministic and hard to game, chosen by the candidate |
| Delegation | Agents did the typing; the candidate's own time went to structure and verification |
| Deletion | They removed something and can say why |
| Writeup | A stranger can read it and rerun the result |

**5. Structured interview.** The top predictor in the 2022 revision (.42). Same questions for everyone, answers scored against anchors written in advance. Four questions I'd use: Tell me about a loop you closed that nobody asked you to close. What was the sensor, and how could it have lied? What did you delete? Which problem did you refuse because it was local?

**6. Measure the hiring loop.** Log every stage score. At six and twelve months, check which stages predicted how people actually did, and cut the ones that didn't. Google learned from its own data that its interviews were a "complete random mess." Yours may be too.

## A sample job post

This is the esoteric door, written the way I'd actually post it:

> **loop closer**
>
> the agents do the typing now. we need the people who decide what the typing is for.
>
> you close loops alone: you research the problem, ship the answer, and put a sensor on it that can't be flattered. you pick problems for their structure. you've read the mentor ("my crime is that of curiosity") and felt accused. you know why only variety can destroy variety, why move 37 looked like a mistake, why a factorio base that works is not a factorio base that scales, and why a tool can be convivial or not. you think capitalism is the final game and you want to play it well.
>
> we don't read CVs. we don't care about your degree, your age or your years of anything. we read what you've made in public. we judge people by what they say and think, not what they look like.
>
> there is a challenge at [link]. if you see it, it takes ten minutes. if you don't, it takes a day, and that's fine, it just means this is not your loop. agents allowed. coding optional.
>
> if you recognized fewer than half of the references above, apply anyway. the challenge doesn't care what you've read.
>
> small team, same room, three or four people per loop. high agency, low cortisol.

## Sources

- BBC News, [Facebook's Instagram deal: Can one app be worth $1bn?](https://www.bbc.com/news/business-17666032), 2012-04-10 (13 full-time staff; web fetch)
- Facebook, [Facebook to Acquire WhatsApp](https://about.fb.com/news/2014/02/facebook-to-acquire-whatsapp/), 2014-02-19; GeekWire, [Facebook buys WhatsApp messaging service for $16 billion](https://www.geekwire.com/2014/facebook-buys-whatsapp-messaging-service/), 2014-02-19 (55 employees; web fetch)
- François Chollet, [On the Measure of Intelligence](https://arxiv.org/abs/1911.01547), 2019 (definition on p. 27; PDF text)
- The Mentor, [The Conscience of a Hacker](https://phrack.org/issues/7/hackers-manifesto.html), Phrack 7, 1986-09-25 (written 1986-01-08)
- Paul Graham, [The Word "Hacker"](https://www.paulgraham.com/gba.html), 2004; [Great Hackers](https://www.paulgraham.com/gh.html), 2004; [The Python Paradox](https://www.paulgraham.com/pypar.html), 2004; [Taste for Makers](https://www.paulgraham.com/taste.html), 2002
- Frank L. Schmidt and John E. Hunter, [The Validity and Utility of Selection Methods in Personnel Psychology](https://talytica.com/wp-content/uploads/2016/10/Schmidt-and-Hunter-1998-Validity-and-Utility-Psychological-Bulletin.pdf), Psychological Bulletin 124(2), 1998, Table 1 (PDF text)
- Sackett, Zhang, Berry and Lievens, [Revisiting meta-analytic estimates of validity in personnel selection](https://doi.org/10.1037/apl0000994), Journal of Applied Psychology, 2022; and [Revisiting the design of selection systems in light of new findings](https://www.cambridge.org/core/services/aop-cambridge-core/content/view/A20984B138319E3D432E643978BF026D/S175494262300024Xa.pdf/div-class-title-revisiting-the-design-of-selection-systems-in-light-of-new-findings-regarding-the-validity-of-widely-used-predictors-div.pdf), Industrial and Organizational Psychology 16, 2023, Table 1
- Adam Bryant, [In Head-Hunting, Big Data May Not Be Such a Big Deal](https://www.nytimes.com/2013/06/20/business/in-head-hunting-big-data-may-not-be-such-a-big-deal.html), The New York Times, 2013-06-19; brainteaser passage as [quoted on HN](https://news.ycombinator.com/item?id=9484526)
- Jane Street, [Puzzles](https://www.janestreet.com/puzzles/)
- NPR, [Google Entices Job-Searchers with Math Puzzle](https://www.npr.org/templates/story/story.php?storyId=3916173), 2004-09-14
- George Mack, ["1/ HIGH AGENCY"](https://x.com/george__mack/status/1068238562443841538), 2018-11-29
- Pierre Azoulay, Benjamin Jones, J. Daniel Kim and Javier Miranda, [Age and High-Growth Entrepreneurship](https://www.nber.org/papers/w24489), NBER Working Paper 24489, 2018
- Wikipedia, [PayPal Mafia](https://en.wikipedia.org/wiki/PayPal_Mafia); [Variety (cybernetics)](https://en.wikipedia.org/wiki/Variety_(cybernetics))
- Brazil, [Lei nº 9.029, de 13 de abril de 1995](https://www.planalto.gov.br/ccivil_03/leis/l9029.htm), art. 1; U.S. EEOC, [Age Discrimination](https://www.eeoc.gov/age-discrimination)
- Hacker News comments linked inline; yearly "high agency" counts from the [Hacker News Algolia search API](https://hn.algolia.com/api/v1/search?query=%22high%20agency%22&hitsPerPage=1) (`nbHits` for the quoted phrase with `numericFilters=created_at_i>=START,created_at_i<END` per calendar year, 2026 through September 25), normalized by the approximate `nbHits` for all items in the same window, run 2026-09-25
- Private conversation with a friend, 2026-08-01

*Research for this post used [Scry](https://scry.io) to find Hacker News comments.*
