Here is a party trick that has worked for over a century. Fill a glass jar with marbles, ask a few hundred people to guess how many are inside, and write down every answer. Almost nobody gets it right — the guesses are all over the place. But line the guesses up and take the middle one, and the number you land on is uncannily, almost spookily close to the truth. The crowd, as a crowd, knows. Even though not one person in it does. Francis Galton stumbled onto this in 1906, watching villagers at a country fair guess the weight of an ox, and it has been quietly astonishing people ever since [1].
That trick is the whole of artificial intelligence. AI is not a brain. It is the marble jar, industrialised — a very, very good averaging machine, trained on roughly everything humanity has ever written down, handing you the middle of it on demand.
And the middle, it turns out, is excellent. Well-organised, grammatically immaculate, and almost always fine. That is exactly the problem. AI isn’t making us wrong. It is making us all the same shade of plausibly right — and it is doing it to everyone at once.
But the machine plays a beautiful move
You are about to object, and you are right to. In 2016 a machine called AlphaGo played a move in a game of Go that no human would have played, a stone so strange the commentators assumed a glitch. It was not a glitch. It was beautiful, and it won. A few years later AlphaFold solved the shapes of nearly every protein known to science, a fifty-year problem, correct, at a scale no laboratory could touch. If your complaint about the averaging machine is that it only ever hands back the mushy middle, these are your counter-examples, and they are good ones. Drop the lazy version of the argument.
But look at what kind of new this is. Go is a closed world: fixed rules, fixed board, a fixed way to win, all of it handed to the machine before it starts. Inside that world there are more possible games than atoms in the universe, so a good enough search will always turn up a beautiful place no human has stood. AlphaGo found the move. It did not invent Go, and it could never tell you whether Go was worth playing. There is finding a new move, and there is drawing the board. The machine does the first better than we ever will. It has not once done the second.
That is the averaging, refined. The machine is not too timid to be brilliant inside a frame you have drawn for it. What it cannot do is draw the frame, choose which game is worth the candle, or feel that the whole board is wrong. The averaging that bites is not of the answer. It is of the judgment one layer up, the taste that picks the question, and that is the layer the world is now quietly outsourcing.
The jar only works if nobody can see the other guesses
Read the fine print on that marble jar. The magic depends entirely on the guesses being independent. The wild over-counters and the timid under-counters cancel out, and the truth shines through the noise. Let the guessers see each other’s answers first, and the whole thing collapses: everyone shuffles toward the loudest number in the room, and the “average” stops measuring the marbles and starts measuring the crowd’s confidence in itself.
There is even a formula for the magic, written down by the social scientist Scott Page: the crowd's error equals the average person's error minus the diversity of the guesses [9]. Read it twice, because the second term is the whole game. What makes the jar work is not that anyone is right — it is that everyone is differently wrong. And the formula carries a bomb in its pocket. Shrink the diversity, and the crowd gets stupider even while every guesser in it gets smarter. The market can get dumber while every investor in it gets sharper. Nothing else in this piece matters more than that sentence.
Cliff Asness of AQR made the obvious, devastating point: markets are not a sealed jar [2]. Investors can all see each other. They read the same headlines, hold the same positions, and repeat the same story until the story moves the price and the price seems to prove the story. The wisdom of crowds needs the crowd to think independently.
It is worth being precise here, because averaging is not the villain. Done right, it is close to magic. Philip Tetlock spent a decade proving it. His Good Judgment Project recruited teams of hobbyists — a retired programmer, a pharmacist, a ballroom-dancing instructor — and set them against professional intelligence analysts who had access to classified intercepts. The hobbyists won, comfortably [3]. Not because any one of them was a genius, but because the project pooled their genuinely independent, genuinely different views. A good jar, kept carefully sealed.
Now feed that crowd a tool trained on its own back catalogue. We all prompt similar models, trained on similar data, parsing the same feeds — a healthy slice of which were themselves written by AI. The answers converge. This is measured, not suspected. Researchers put the same open-ended prompts to 102 people and 22 different AI models: the models mirrored each other far more than humans mirror humans, even after controlling for the obvious confounders [10]. Another team ran 26,000 real user queries across dozens of models and won a best-paper award for documenting the same collapse — different companies' machines, strikingly similar answers [11]. And the material the machines learn from is already partly their own: by late 2024, up to a quarter of corporate press-release text was machine-written, a figure the study's authors call a lower bound [12]. Dozens of sources say the identical thing, each one apparently confirming the others, and we mistake the echo for a chorus. We have automated groupthink and given it a warm, confident, faintly human tone of voice.
Everyone’s a genius now, which is the giveaway
You may remember the Dunning-Kruger effect: the least competent people are the most confident, because they don’t know enough to know what they’re missing. It was a comforting law. It meant overconfidence was at least a signal — a red flag you could watch for.
AI broke the flag. A 2026 study in Computers in Human Behavior, flagged by Baillie Gifford’s Tom Slater, found that AI use produces uniform overconfidence across every skill level [4]. Novice and expert walk away equally certain they understand more than they do. The cruel twist: the people with the highest AI literacy were the worst calibrated of the lot. They had mistaken fluency with the tool for mastery of the subject.
If that sounds familiar, it should. Tetlock had already catalogued the species in the wild. Tracking tens of thousands of predictions by celebrated experts, he found the average one forecast about as well as “a dart-throwing chimpanzee” [5]. Worse: the more famous and quotable the expert — the ones with the book deals and the green-room makeup — the lessaccurate they tended to be. Confidence and television charisma were inversely related to being right. AI, in this light, is simply the dart-throwing chimpanzee finally handed a publishing contract and a soothing voice. Always available, always certain, reliably average.
We have, in other words, all become the chap at the dinner party who skimmed one article and now has Opinions — except the article is beautifully written, we have twelve of them, and they all agree. As Slater puts it, the winners “won’t be those who use AI most, but those who can still think without it” [4].
So what is actually scarce?
Whatever cannot be averaged.
The philosopher Michael Polanyi gave us the phrase sixty years ago: we know more than we can tell [6]. The surgeon’s hands, the trader’s bad feeling about a deal that looks flawless on paper, the way a great manager reads a room before anyone has spoken — none of it was ever written down, which means none of it is in the machine. It is tacit knowledge, and it is the one asset that does not depreciate the moment it is shared, because it cannot be shared. You had to be there. For years.
My favourite illustration comes from the investor Carson Block, who described his proudly Luddite father on a recent podcast: up 70% in a year holding Nvidia and CrowdStrike, without — by his son’s account — knowing what either company actually does [7]. The punchline: “If our algorithm ever broke down, I’d just ask my father what to buy.” A lifetime of pattern recognition, sitting quietly above the spreadsheet, waiting for the day the spreadsheet fails.
That is the part no model can reach. AI is built to produce an answer. It cannot sit comfortably in uncertainty, and it certainly cannot laugh at the gap between what the data says and what experience knows.
The people who beat the jar
Here is the part that should change how you think about your own edge. What made Tetlock’s superforecasters exceptional was not raw brainpower, and it was not a secret data feed. It was a method — and every ingredient of it is the precise opposite of how a machine produces an answer.
They are comfortable being uncertain. They think in fine-grained probabilities — not “yes” or “no” but “sixty-eight per cent, and here’s what would change my mind.” They treat every forecast as a draft, nudging it in small steps as new evidence trickles in, a habit Tetlock calls being in “perpetual beta.” And they deliberately gather clashing viewpoints and hold them in tension before deciding — the “dragonfly eye”: many lenses, one synthesis [3].
Read that list again and notice what it is. It is a human being doing, slowly and on purpose, the very thing the averaging machine does instantly and for free — except the machine skips the part that matters. It synthesises, but it cannot doubt. It aggregates, but it cannot hold two ideas in tension and feel which one is wrong. It will never tell you “sixty-eight per cent, and here’s what would change my mind,” because it was built to sound sure. The superforecaster’s edge was never information. It is judgment under uncertainty — and that is exactly what grows scarcer, and dearer, the more the world floods with confident, average answers.
The consensus engine cannot sell you the unpopular truth
Here is the structural catch, and it is fatal to the dream that AI will hand you an edge. AI is, by design, a machine for surfacing what most people already believe. Contrarian insight — the kind that spots the turn before the crowd — cannot come out of a system trained on the crowd. You cannot average your way to a non-consensus view. That is what “non-consensus” means.
In August 1979, BusinessWeek ran a cover declaring “The Death of Equities” [8]. It was the perfect distillation of everything sensible people believed, beautifully argued, and it landed almost exactly at the start of the greatest bull market in living memory. A 1979 AI would have agreed with that cover enthusiastically, in flawless prose, with footnotes. The median view is most comfortable precisely when it is most expensively wrong.
Where this leaves your money, and your mind
The investment translation is short, because the principle is. The average is now free. The edge is whatever can’t be averaged.
And the market was rigging itself for sameness long before the machines arrived. In the late 1950s the average American share changed hands roughly every eight years; by 2020, measured the same way, every five and a half months [13]. Roughly two thirds of the money in US stock funds is now passive — it buys everything, weighted by size, and asks nothing [13]. After Europe changed how research is paid for, 334 listed companies lost analyst coverage entirely; nobody is even looking [13]. You can see the sameness on the tape. Ten names are 37 per cent of the S&P 500 [14]. In 2025 alone, 79 per cent of active large-cap managers underperformed the index; over twenty years, 93 per cent have [15]. Everyone expresses their personality in the small positions and hugs the same ten giants with the real money. And sit with the loop inside that fact: the ten giants are the AI trade. The largest concentration in market history is a machine-averaged consensus about the averaging machines themselves.
So the premium shifts to the un-averageable: judgment built over cycles, relationships that took a decade to earn, the on-the-ground pattern recognition that still decides outcomes in real assets, private markets, and physical things — domains where being in the room beats reading about the room. Call it the cognitive-purity premium. It accrues to the people and the businesses that can still think without the machine, in a world where almost nobody bothers to try.
There is a warning tucked in here for every organisation merrily automating. It is gloriously easy to cut the explicit layer — the drafting, the modelling, the codifiable work — and bank the savings. It is much harder to notice you have also quietly stopped producing judgment, because the junior who used to earn it by doing the boring work now just prompts for it. Automate the explicit, by all means. Just don’t strip-mine the tacit, or in twenty years you will run a beautifully efficient company full of people who have never had an original thought, all agreeing with each other, fluently.
One last thing about the middle
When Tetlock’s team built their forecasting machine, they ran into something with a name that belongs on the cover of this piece. The crowd’s averaged forecast, even a good crowd’s, came out chronically too timid — huddled near the middle, hedged half to death. So they did something counterintuitive. They took the average and shoved it back outward, toward the extremes, toward conviction. They called it “extremizing,” and it made the forecasts more accurate, not less [3].
Sit with that. The best forecasters alive found that the average pulls you too far to the middle, and that the fix is to deliberately lean back out.
AI will write your emails, draft your code, and summarise your meetings. Let it — hand over the whole commodity middle with my blessing. Just don’t outsource the one thing it was never able to do: disagree with everybody, sit with the discomfort of being maybe-wrong, and turn out to be right.
A note of thanks before the close. A day before this went out, Ahmed Husain arrived at much the same fear by a different road, with evidence I had not found and a phrase I wish I had coined: a weirdness premium [16]. Two people reaching the same conclusion independently is, of course, precisely the diversity this piece says is disappearing. Long may we disagree.
The future, as ever, belongs to the weirdos who still think for themselves.
This article is 100% free to read, so please share it to pay it forward.
Thanks!
References
[1] Francis Galton, “Vox Populi,” Nature (1907) — the original “wisdom of crowds” experiment: at a 1906 country fair, the median of ~800 independent guesses of an ox’s weight came within ~1% of the true figure. Popularly retold as guessing marbles or jellybeans in a jar.
[2] Cliff Asness (AQR Capital Management) on the “wisdom of crowds” failing in markets — the “poll the audience” analogy and a narrative-to-flows-to-price feedback loop. Asness is a public commentator; argument deployed here in paraphrase.
[3] Philip E. Tetlock and Dan Gardner, Superforecasting: The Art and Science of Prediction (2015). The Good Judgment Project — teams of amateur forecasters outperforming intelligence-community analysts with classified access; the superforecaster traits (”perpetual beta,” the “dragonfly eye,” probabilistic thinking); and the “extremizing” of aggregated forecasts to improve accuracy.
[4] Tom Slater, Investment Manager, Scottish Mortgage Investment Trust (Baillie Gifford), “AI Isn’t Coming For Your Job. It’s Coming For Your Mind,” 2026 — citing a 2026 study in Computers in Human Behavior finding AI use produces uniform overconfidence across skill levels, with the most AI-literate the least well-calibrated.
[5] Philip E. Tetlock, Expert Political Judgment: How Good Is It? How Can We Know? (2005). The roughly two-decade study in which the average expert’s forecasts proved “roughly as accurate as a dart-throwing chimpanzee,” and the most media-prominent experts were among the least accurate.
[6] Michael Polanyi, The Tacit Dimension (1966): “we know more than we can tell.” The foundational statement of tacit versus explicit knowledge.
[7] Carson Block (Muddy Waters Research), interview on the Financial Times Unhedged podcast, 2026.
[8] “The Death of Equities,” BusinessWeek cover, 13 August 1979 — the canonical example of consensus conviction arriving at a major market turning point.
[9] Scott E. Page, The Difference (Princeton University Press, 2007) — the Diversity Prediction Theorem: collective squared error = average individual squared error − prediction diversity. An algebraic identity, not an empirical claim.
[10] Emily Wenger and Yoed N. Kenett, “Large language models are homogeneously creative,” PNAS Nexus 5(3), March 2026 — 102 human participants vs 22 LLMs; model responses mirror other models far more than humans mirror humans.
[11] Liwei Jiang et al., “Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond),” NeurIPS 2025 (Best Paper Award) — 26,000 open-ended queries; pronounced inter-model homogeneity across 70+ models.
[12] Weixin Liang et al., “The Widespread Adoption of Large Language Model-Assisted Writing Across Society,” Patterns (2025) — 537,413 corporate press releases, Jan 2022–Sep 2024; up to 24% of text LLM-generated by September 2024; described by the authors as a lower bound.
[13] Market structure: Reuters/NYSE turnover data (holding periods, 2020 analysis); Investment Company Institute, “Active and Index Investing,” June 2026 (63.8% of US domestic equity fund assets indexed); Fang, Hope, Huang and Moldovan, “The Effects of MiFID II on Sell-Side Analysts, Buy-Side Analysts, and Firms,” Review of Accounting Studies (2020) — 334 EEA firms lost coverage entirely.
[14] S&P 500 constituent weights, official fund holdings file, 24 August 2026 (top-10 constituents = 37.0%).
[15] SPIVA U.S. Scorecard, Year-End 2025, S&P Dow Jones Indices — 79% of active large-cap funds underperformed the S&P 500 in 2025; 93% over 20 years.
[16] Ahmed Husain, “The Weirdness Premium,” The Curious Mind, Research Letter No. 2, 26 August 2026 — thecuriousmind.org/p/the-weirdness-premium.

