Every idea already exists

September 2026

On July 13, Anthropic published Claude's values across models and languages. They sampled 309,815 Claude.ai conversations where the user asked for a subjective judgment, split evenly across three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the 20 most common languages on the platform, about 5,000 per model-language pair. Claude labeled which of 339 values each response expressed, and they compressed the labels into four axes: deference vs. caution, warmth vs. rigor, depth vs. brevity, candor vs. execution. After controlling for task, topic and the user's own values, those four axes explain 15% of the variance. The result I keep rereading is this one: "The largest variation is in the Warmth vs. Rigor axis, with Claude leaning toward expressing warmth-related values most in Arabic and Hindi and rigor-related values most in English and Russian." Their example is concrete: "two people asking for feedback on the same business plan, one in Hindi and one in Russian, may come away with different impressions of its quality." And then: "We don't yet know which properties of our training data drive these differences."

The usual reading is that Claude has slightly different values in different languages, and that training should make them consistent. I read it differently. Every idea already exists, and a sentence in any language is one physical form of it. The business plan and the judgment about it are the same object in Hindi and in Russian. What changes is the projection, and each language bends it in its own direction. If that's right, the bend can be measured, but only against a frame that doesn't move with the languages, and nobody has one yet. Once it's measured, language stops being a property of the user and becomes a control input: for each task, pick the language the model thinks best in.

Where this came from

It started with a friend, on June 7. He'd noticed that the same prompt gets answers with different style and structure depending on the language. In Portuguese the answers are short, simplified and full of emojis. In English they read like a corporate email. In Spanish they read like TikTok. His proposal was a loop: figure out what the user needs, pick the language whose style fits that need best, write the answer there, and translate it back at the end. He framed it as speciation: the same process that splits one population into many, with user feedback as the selection pressure.

The next day I turned it around. If intelligence is competence at games and capital runs free, a model's textual style is just a phenotype selected by user feedback. My friend picked a language to fit the task. I think each language already is an adaptive niche that RLHF sculpted. Portuguese is terse and full of emojis because that's what survived in the Brazilian environment.

Five weeks later the Anthropic post showed up, and it has a detail that made me laugh. In the appendix's limitations, the Claude-based labeling tool shows biases "such as labeling warmth-related values more often when responses contain emojis or kinship terms." My friend's emojis are in the measurement.

A few days after it came out, I spent an afternoon working the idea through with an open model. My messages from that session, translated from Portuguese, are the spine of this essay:

I like the concept that every idea already exists and a translation is a physical form of it.

It's very hard to evaluate language because there are several cars moving. We need a very good inertial reference frame.

The best of all worlds would be communicating in ideas, but that's impossible. We have to pass it into the material world.

Imagine a model that receives an embedding and predicts the language that maximizes that task (train an RL on it).

The idea and its garment

The intuition is old. In Book 10 of the Republic, Socrates sets up the theory of forms with furniture (596a–b, Shorey translation): "We are in the habit, I take it, of positing a single idea or form in the case of the various multiplicities to which we give the same name." There are many couches, and one idea of a couch, and "no craftsman makes the idea itself. How could he?"

Frege made the version I care about in 1918. In The Thought: A Logical Inquiry (Quinton translation, Mind, 1956) he wrote the sentence that is basically my thesis: "The thought, in itself immaterial, clothes itself in the material garment of a sentence and thereby becomes comprehensible to us." Thoughts, he argued, are neither physical things nor private mental images: "A third realm must be recognized." The Pythagorean theorem "is timelessly true, true independently of whether anyone takes it to be true... It is not true for the first time when it is discovered, but is like a planet which, already before anyone has seen it, has been in interaction with other planets." And: "In thinking we do not produce thoughts but we apprehend them."

The engineering version came in 1949. Warren Weaver's memo Translation, dated July 15, 1949, is one of the founding documents of machine translation. He describes people living in tall closed towers, shouting to each other. "But when an individual goes down his tower, he finds himself in a great open basement, common to all the towers." So "the way to translate from Chinese to Arabic, or from Russian to Portuguese, is not to attempt the direct route, shouting from tower to tower. Perhaps the way is to descend, from each language, down to the common base of human communication - the real but as yet undiscovered universal language - and then re-emerge by whatever particular route is convenient."

I'm using Plato and Frege as a working hypothesis, not a metaphysics I can prove. The claim I can test is smaller: if there's a shared layer, trained models should show it, and the way each language deviates from it should be stable and measurable. Seventy-seven years after Weaver, we have machines where you can look for the basement.

What the models show

Claude has a conceptual space shared between languages. Anthropic's March 2025 interpretability post, Tracing the thoughts of a large language model, says it directly: "Claude sometimes thinks in a conceptual space that is shared between languages, suggesting it has a kind of universal 'language of thought.'" The companion paper, On the Biology of a Large Language Model, traces the antonym circuit for the same prompt in English, French and Chinese ("small", "petit", "小"). The model "translates concepts to a common 'universal mental language' in its intermediate activations," and "the prevalence of these language-agnostic representations is higher in Claude 3.5 Haiku than in a smaller, less capable model." When they swapped in "hot" features taken from an English prompt, the model produced the right antonym of "hot" in each language. The concept was the operand, not the word.

Shared machinery isn't the same as a shared path. The same paper gives the numbers: "20 out of 27 of the features in multilingual nodes are active across all three prompts." But "the set of features that are influential to the model's response varies quite a bit by prompt (only 10/27 appear in the pruned attribution graphs for all three prompts)." Three quarters of the vocabulary is shared. Only about a third of what actually drives the answer is.

Of 27 features in the multilingual nodes, 20 are active in all three languages and 10 are in all three pruned attribution graphs

The shared space leans English. Chris Wendler and colleagues asked Do Llamas Work in English? and found three phases in Llama-2: intermediate embeddings start far from any output token, then "already allow for decoding a semantically correct next token in the middle layers, but give higher probability to its version in English than in the input language," and finally move into the input language's region. They call these input space, concept space and output space, and conclude that "the abstract 'concept space' lies closer to English than to other languages." Anthropic found the same tilt in Claude: "multilingual features have more significant direct weights to corresponding English output nodes, with non-English outputs being more strongly mediated by say-X-in-language-Y features." Weaver, framing translation as cryptography, wrote: "When I look at an article in Russian, I say 'This is really written in English, but it has been coded in some strange symbols.'" For Llama-2 that's not far from what the middle layers do.

Different networks converge on the same geometry. The Platonic Representation Hypothesis (Huh, Cheung, Wang and Isola, 2024) argues that "representations in AI models, particularly deep networks, are converging." As vision and language models get larger, "they measure distance between datapoints in a more and more alike way." They hypothesize a destination: "a shared statistical model of reality, akin to Plato's concept of an ideal reality."

The geometry is shared enough to translate without a dictionary. vec2vec (Jha, Zhang, Shmatikov and Morris, 2025) translates text embeddings from one model's space to another's "without any paired data, encoders, or predefined sets of matches," through "a universal latent representation." Its translations reach cosine similarity as high as 0.96 with the true target vectors and perfect matching on over 8,000 shuffled embeddings. That's two models with different architectures and training data agreeing on the shape of meaning so closely that you can align them blind.

Google saw a basement in 2016. Google's multilingual NMT system put many language pairs in one model with a token for the target language. It learned to translate between pairs it never saw together, and the authors reported analyses that hint at "a universal interlingua representation in our models."

The surface still decides the score. If the idea were all that mattered, the language of the prompt wouldn't move accuracy. It does, and not always toward English. In Language Models are Multilingual Chain-of-Thought Reasoners, Freda Shi and colleagues had professional translators render 250 GSM8K math problems into ten languages (MGSM). For PaLM-540B, reasoning in English instead of the problem's language added 9.6 points in Japanese and 9.2 in Swahili, and cost 3.2 points in Thai and 0.8 in Chinese. Machine-translating the problem to English first helped on average (55.0% against 48.1% for native reasoning), but in Thai it still lost.

Gain from reasoning in English, or machine-translating to English first, over native-language reasoning, for PaLM-540B on MGSM

A 2025 paper, Could Thinking Multilingually Empower LLM Reasoning?, measured the ceiling. Aggregating answers to the same question in many languages, the upper bound beat English-only reasoning "by nearly 10 Acc@k points," and "multilingual thinking can ideally boost GPQA accuracy from ~45 to ~90." The catch is in the same abstract: "common answer selection methods cannot achieve this upper bound." Knowing which language is right for which question is the unsolved part. That's the model I want to train.

Measuring the bend

My first attempt at the math, on July 16, was naive. Take a translation model M that maps a text in language l1 to a base language b2 and embed the result: l2 = M(l1, b2). Do it for many inputs and sum the differences:

$$\sum_{i=0}^{n} (l_{2i} - l_{1i})$$

Divide by n and you get the average vector between two languages. I saw two problems right away, and wrote them down in that session. The embedding adds friction of its own. And translation isn't perfect: if translation were perfect, all languages would be the same, so every measurement mixes the language effect with the translator's error, and that error itself varies with the language pair. "I don't know how to measure what the truth was in other languages." That's the cars problem. The content moves, the language moves, the translator moves and the embedding model moves, all at once.

The formal parts below came out of the dialogue with that model, so treat them as coauthored. The questions and the cars metaphor were mine. Most of the math was the model's, and it's standard linear algebra. The centroid was its answer when I asked whether we could see the ideal language through the way the languages interact with each other.

That answer has a limit. A center of mass depends on what's in the system. Put eight European languages and two Asian ones in the average, and the centroid leans European. Build the parallel corpus by translating everything from English, and the centroid inherits English's framing through the translators. The centroid is a choice of frame (physicists would say a gauge), not the discovery of an absolute rest frame. I don't think that kills the program. It means every result has to report the frame it was measured in, the same way a velocity has to say relative to what.

Could a language be optimal?

This part is speculation, and I'll mark where it stops being evidence.

In the same session I asked whether we could build an optimal language without the problems of the existing ones. The evidence above suggests two answers.

The first is that the closest thing to an ideal language already exists inside the models, and you can't speak it. The interlingua in Google's NMT, the shared features in Claude's middle layers, the universal latent space vec2vec aligns to: these are representations, not languages. They have no grammar you could teach a child. Weaver's basement is real, but you only pass through it.

The second is that "optimal" only means something relative to a task. The MGSM chart already shows it for one model and one task: the best language to think in depends on the source language, and the answer isn't always English. Anthropic's post shows it for values: if you want a rigorous critique of a business plan, English or Russian gets you closer on average; if you want warmth, Hindi or Arabic. That's where evidence stops. The step I'm taking without evidence is this: for a strong enough model, the optimal language is a per-task routing decision, and a small model trained on the geometry of the shared space can make it better than any fixed rule.

There's an alignment reading too, which I care about more than the benchmark gains. Anthropic ends its post with "How should Claude's values vary across languages?" and admits "we don't know what kinds of variation users interacting with Claude in those languages want." If the values shift lives along the same directions as the language residuals, you can see it in the geometry without labeling a single conversation, and you can choose it deliberately instead of inheriting it from whatever mix of text each language had in pretraining. I argued in The harness is in the weights that values ride along with distillation. They probably ride along with language too, and a language is a much cheaper channel than a training run.

Where the idea leaks

The case against this is strong, and some of it comes from the same papers.

The basement has an English accent. Wendler's concept space sits closer to English, and Claude's multilingual features wire more directly to English outputs. So the "idea" I keep calling language-free might be English with other languages mapped onto it. A centroid computed inside such a model would put English near the center by construction, and I'd be measuring distance from English while calling it distance from the idea. The experiment below has a check for this, but it's the objection I take most seriously.

The projection carries norms, not only meaning. Anthropic measured what Claude cares about when it answers, not whether it gets facts right in Hindi. If the same business plan gets a warmer verdict in Hindi and a harsher one in Russian, the idea wasn't transmitted intact; something normative got added in the materialization. The post also points to Anthropic's own system cards, which "find differences across languages in what Claude knows and how it handles sensitive requests." Frege's thought may be timeless. What a model does with it clearly isn't language-invariant.

Language may shape the thought, not just dress it. Linguistic relativity has real, if limited, support. In Russian blues reveal effects of language on color discrimination (Winawer, Boroditsky and colleagues, 2007), Russian speakers, whose language obligatorily separates light blue (goluboy) from dark blue (siniy), were faster to tell two blues apart across that boundary, and the advantage disappeared under a verbal dual task but not a spatial one. English speakers showed no such advantage. A 2020 replication, Russian blues reveal the limits of language influencing colour discrimination, used identical tasks and found no speed advantage at the siniy/goluboy boundary, concluding that "Russian blues" are "less well-structured than previously thought." My position is that the effect is small and online, which fits a garment that sometimes pinches, and doesn't fit a garment that is the body.

Translation loses things, and there's no neutral translator. Weaver knew it: "'Perfect' translation is almost surely unattainable." His memo cites a letter to the Herald Tribune recalling a great Hebrew poet who said translation "is like kissing your sweetheart through a veil." In 1947 Weaver had put the idea to Norbert Wiener, who answered: "I frankly am afraid the boundaries of words in different languages are too vague and the emotional and international connotations are too extensive to make any quasi mechanical translation scheme very hopeful." Every translator adds its own error, and a model translating its own inputs adds errors correlated with the very bias you're measuring. Anthropic had to test this for its own labeler: they translated 800 conversations into eight languages and found that "only 11 of 339 values expressed by Claude are affected, and the largest effect is an order of magnitude smaller than the corresponding cross-language difference." Small, but not zero, and they had to check.

Geometry has artifacts. A May 2026 paper, Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models, embedded 6,518 idioms and proverbs in English, Bangla, Hindi and Arabic with five production encoders. Retrieval across languages wasn't symmetric, and the dominant cause was hubness, a few points that are nearest neighbors to everything. A hub-aware score (CSLS) closed 63.5% of the gap. If you measure language residuals with plain cosine similarity, you may be measuring hubs.

The words a model reasons in may not be how it reasons. Routing a model's "thinking language" routes the visible trace. Anthropic's Reasoning models don't always say what they think found that, given a hint, Claude 3.7 Sonnet mentioned it in its chain of thought 25% of the time and DeepSeek R1 39% of the time. The tracing post caught Claude claiming a calculation its internals show it never ran. The faithfulness post asks the question that sits under this whole essay: "why, after all, should we expect that words in the English language are able to convey every single nuance of why a specific decision was made in a neural network?" If the trace is a garment too, then choosing its language changes the output for reasons that may have little to do with what I think "thinking in Japanese" means. For a router that's acceptable, because the verifier only sees the answer. For the philosophy it's a warning.

Someone already built a version of the router. AdaMCoT (2025) routes reasoning "in intermediary 'thinking languages' before generating target-language responses," choosing pathways with a reward model and fine-tuning on the winners. As far as I can tell, it doesn't measure the frame, route on values, or check whether the chosen language changes what the model cares about. That's the gap I'd work in.

The experiment

This is what I'd run. Two parts: first build the frame and see whether language tendencies survive it, then train the router on top.

Part 1: the inertial frame

Data. FLORES-200: 3,001 sentences from 842 web articles, professionally translated into about 200 languages. Most languages were translated from English, but "several languages were translated from Spanish, French, Russian and Modern Standard Arabic," which gives a built-in control for pivot bias. Europarl for 21 European languages of long parliamentary text. MGSM for a task where the answer is checkable.

Instruments. At least five multilingual embedding models, plus the mid-layer residual stream of two or three open-weight LLMs from one family at different sizes. The embedding models test whether a tendency belongs to the language. The LLM layers test where it lives inside a model, and whether it grows or shrinks with scale.

Procedure.

  1. Embed every version of every sentence with every instrument: e_{m,L}(s).
  2. Compute a leave-one-out centroid for each sentence, c_{m,−L}(s), the mean over all languages except L, so a language doesn't pull its own frame toward itself. Compute it three ways: all languages, no English, and only languages FLORES translated from a non-English pivot.
  3. The residual r_{m,L}(s) = e_{m,L}(s) − c_{m,−L}(s) is the language's deviation for that idea. Average over sentences for the language offset μ_{m,L}.
  4. Measure the translation noise floor with cycles. For each sentence, run many-to-many round trips L → L′ → L through at least three different machine translators, and embed the result. The drift between the original and the round trip, averaged over paths, is ε_{m,L}. It estimates how much a translation step moves a sentence when the idea hasn't changed.
  5. Only count a language effect when it's above the noise: the ratio ‖μ_{m,L}‖ / ε_{m,L}.
  6. Score everything with CSLS instead of raw cosine, and report mutual-nearest-neighbor reciprocity per language pair, so hubness can't masquerade as a tendency.

Metrics.

Metric What it answers Kill condition
‖μ_L‖ / ε_L per instrument Is the language effect bigger than translation noise? Ratio near 1 for most languages
Spearman correlation of language rankings across instruments Is the tendency a fact about the language or about the encoder? Median correlation below 0.5
Rank of English by distance to the no-English centroid Does the frame lean English? English closest in every LLM (report it, don't hide it)
Singular-value spread of residuals per language Does the language bend meaning in a few directions or many? No difference between languages
Angle between μ_L and a warmth–rigor direction in the LLM residual stream Is the values shift the language shift? Ordering of languages unrelated to Anthropic's (Hindi and Arabic warm, English and Russian rigorous)

The warmth–rigor direction comes from contrast pairs: the same request answered warmly and rigorously in one language, differenced in activation space. The last row is the one prediction that could embarrass me. If languages line up along that direction in the order Anthropic observed in behavior, the values shift is visible in geometry. If they don't, language and values travel on separate channels, which is also worth knowing.

Part 2: the per-task language router

The router takes the language-free part of a request and picks the language the model should think in. The answer always goes back to the user in their own language.

type RouterInput = {
  idea: Float32Array        // request embedding minus its language residual (centroid frame)
  sourceLanguage: string    // what the user wrote in
  taskFamily: "math" | "code" | "factual" | "advice" | "creative"
}

type RouterOutput = {
  thinkingLanguage: string  // one of the K candidate languages, sourceLanguage included
  answerLanguage: string    // always sourceLanguage
  confidence: number        // below threshold: think in sourceLanguage
}

const training = {
  candidates: 10,                          // K thinking languages
  labels: "run every prompt in all K, score with a deterministic verifier",
  reward: "verifier score - lambda * output tokens",
  stage1: "contextual bandit on logged outcomes",
  stage2: "RL on live traffic, with 10% exploration",
  baselines: ["always source", "always English", "translate to English", "AdaMCoT", "oracle (Acc@K)"],
  report: ["share of oracle gap closed", "regret per task family", "tokens per answer"],
}

const guardrails = {
  advice: "never route by accuracy; report the value-axis shift per route",
  backTranslation: "check the answer against the thinking-language draft before returning it",
  logging: "store thinking language with every trajectory",
}

Three choices matter more than the code. The labels come from deterministic verifiers (exact match on MGSM, unit tests on code), because an LLM judge would bring its own language bias, which is the thing Anthropic had to control for in its labeler. The bandit comes before RL, because running all K languages on a fixed set gives you the oracle for free, and the "nearly 10 points" ceiling tells you how much there is to win. And advice-style tasks never get routed by accuracy. For those, the router's job is to report which language moves the answer along which value axis, so a person decides. That's the same logic as Optimize the harness, not the model: the thinking language is a harness parameter, cheaper to change than weights, and it should be measured like one.

If Part 1 fails, if tendencies don't survive the change of instrument, the router still works as a black box, but the thesis loses its footing: there'd be no stable way a language bends an idea, only noise that happens to help on some benchmarks. If Part 1 holds, the residuals become the router's features, and the router becomes a small instrument for looking at the shared space from outside.

Dante ends the Inferno with the pilgrims climbing out of hell: "E quindi uscimmo a riveder le stelle", and so we emerged, to see once more the stars. I posted that line in February without much context. It fits here. Tokens, translators, hubs and English-leaning middle layers are the material you have to climb through. What you're climbing toward is the idea every language points at.

Sources

← Simplicity does not sell
Hire people who close loops →

Markdown version: /blog/every-idea-already-exists.md. Every essay: /agents.