There is a belief so widely held in the technology industry that it no longer needs to be argued. It goes like this: whoever has the most data, the most chips and the largest data centre will have the best model; whoever crawls the most pages will have the best ranking study; whoever runs the most prompts will have the best visibility score. More is better. More is always better. Scale is the strategy, and the only open question is who can afford it.
This essay takes the belief at its strongest — where it looks least like a belief and most like arithmetic: in the size of the samples the measurement industry now works with.
A hundred and twenty-six million
Now turn the same lens on the tools that sell measurement to marketing departments. The genre is familiar, and the numbers are real. In June 2026 Semrush announced that its AI Visibility Index had scaled "from an initial analysis of 2,500 prompts to 126 million U.S. AI search prompts", analysed over four months across twenty-two industries. 1 Profound describes "our dataset of 680 million citations", collected across three platforms over roughly ten months. 2 Backlinko, now owned by Semrush, "analyzed 11.8 million Google search results" for its ranking-factor study — and, to its credit, adds in the same breath that "as this is a correlation study, it's impossible to determine the underlying reason behind this relationship from our data alone." 3 The number of prompts, pages or citations is the headline. It is meant to be. The larger the crawl, the more the study looks like a census, and a census does not need an argument.
What such a study actually is, in the language of statistics, is a cross-section: one observation per unit, taken at one moment, across a population that was pre-sorted by the very system whose behaviour the study claims to explain. Three things are structurally wrong with it, and no amount of scale repairs any of them. The deepest of the three is not a flaw of the studies. It is a property of numbers, and it was worked out by a statistician before anyone had counted a prompt.
The number that lies
Quality times quantity times difficulty. In 2018 the statistician Xiao-Li Meng published what he called the big-data paradox: the error of an estimate from a non-random sample decomposes into three factors — how correlated the chance of being in the sample is with the thing being measured (data quality), how much of the population is missing (data quantity), and how variable the thing is (problem difficulty). The first factor is the killer. Even a tiny correlation between selection and outcome, multiplied by a large population, destroys the value of size. Meng's worked example is now famous: self-reported voting preferences from one per cent of the US electorate — about 2.3 million people — with a selection correlation of roughly −0.005 have "the same mean squared error as the corresponding sample proportion from a genuine simple random sample of size n≈400". Population inferences with big data, he wrote, are subject to a paradox: "the more the data, the surer we fool ourselves." 4 Three years later Meng and colleagues published the empirical demonstration in Nature. During the US vaccine roll-out, a Facebook-distributed survey collecting about 250,000 responses a week overstated first-dose uptake by 17 percentage points, the Census Bureau's Household Pulse survey by 14, while a conventional online panel of about a thousand responses a week "provided reliable estimates and uncertainty quantification". Their calculation: a survey of 250,000 respondents "can produce an estimate … no more accurate than an estimate from a simple random sample of size 10." 5 The large samples did not reduce the error. They made it confident.
A.:The arithmetic is right. I no longer trust the conclusion.
T.:Then distrust it precisely. Which conclusion?
A.:I am not sure anymore. Ten minutes ago I had a hundred and twenty-six million observations and a standard error that fell with the square root of n. That is not a belief. It is arithmetic.
T.:Then do the arithmetic. Take a napkin. One per cent of the electorate. A defect of minus five in a thousand. What does the fraction over its complement come to?
A.:One over ninety-nine. About a hundredth.
T.:Divide by rho squared.
A.:Twenty-five millionths. … Four hundred and four.
T.:Four hundred. Notice what dropped out.
A.:The size of the population. All that survived was the fraction I hold and the defect in how I got it.
T.:You do not know its sign. You do not know its size. You only know that scale does not cancel it.
A.:You have wanted to say it since the second paragraph. Say it. Small is better.
T.:Yes. Small is better.
A.:No. Read the same source. The thousand tracked the truth because they were drawn at random and weighted, not because they were few. Four hundred at random beat two point three million with a defect. Four hundred drawn badly would still be four hundred drawn badly. You have swapped one superstition for its mirror image.
T.:That is fair. Then precisely, since I owe you precision. It was never large against small. It is: what fraction, of what population, entered by what mechanism. A census has no missing fraction; that term goes to zero, and the defect has nothing left to multiply. Below a census, the defect rules and the count barely matters.
A.:So if I had all of it —
T.:— you would be right. The question is whether a hundred and twenty-six million is all of something, or a lot of something else.
A.:What was the mechanism, then. For my prompts.
T.:Whoever ran the tool. People who use that tool, asking what people who use that tool ask. In 1936 the Literary Digest had two point four million ballots and a mailing list made of telephone directories and car registrations, in a depression. It predicted Landon. Roosevelt took forty-six states. A young man called Gallup, with a few tens of thousands of respondents, called it. 15 Your mailing list is longer. It is still a mailing list.
A.:Then I will fix the defect.
T.:With what? You would need to know who did not enter and how they differ from those who did. That is not in the data. It is what the data is missing. So let me ask the only question I came to ask. What would have to happen for you to stop believing this?
A.:Not yet.
T.:Then finish it, Alex. The argument is not finished with you.
Ergodicity. A cross-section computes an ensemble average — what many units do at one moment. A time series computes a time average — what one unit does over many moments. Ole Peters' 2019 paper in Nature Physics asked the question that economics had spent a century assuming away — "is the time average of an observable equal to its expectation value?" — and showed that for the growth processes that matter, it is not: the expectation value "effectively averages over an ensemble of copies of myself", and none of those copies is me. 6 We treat web performance as a non-ergodic process of exactly this kind, and any crawl shows why. The average page in a crawl of ten million tells you nothing about the path your page will take, because the average is dominated by units you will never be — the marketplaces, the encyclopaedias, the viral one-offs. The only average that predicts your trajectory is your own, taken over enough time.

Intervention, not association. Judea Pearl's do-calculus formalised the distinction the whole industry keeps eliding. An associational quantity is anything that can be read off a joint distribution of observed variables; a causal concept, in Pearl's definition, is "any relationship that cannot be defined from the distribution alone". P(Y | do(X)) — the probability that Y "would occur if treatment condition X = x were enforced uniformly over the population" — is a different object from P(Y | X), and the second level of Pearl's hierarchy "ranks higher than Association because it involves not just seeing what is, but changing what we see." 7 A ranking study reports seeing and sells it as doing. Our own Compendium states the consequence in one line: "A number without a counterfactual is an anecdote with confidence." 8 The counterfactual has to be built — a control group frozen before the intervention, a pool of untreated peers, a placebo period — or it has to be admitted that it is missing.
A number without a counterfactual is an anecdote with confidence.
Multiplicity. Test a hundred traits against an outcome across ten million pages and, at the conventional five per cent threshold, five of them will be "significant" by chance alone — and with a sample that large, effects far too small to matter will clear the threshold as well. Benjamini and Hochberg's 1995 procedure exists precisely for this: it "calls for controlling the expected proportion of falsely rejected hypotheses — the false discovery rate". 9 Calude and Longo proved in 2017 that the problem gets worse with size, not better: "very large databases have to contain arbitrary correlations. These correlations appear only due to the size, not the nature, of data." Their summary: "Too much information tends to behave like very little information." 10 Ioannidis put the consequence for an entire discipline in a title in 2005 — "Why Most Published Research Findings Are False" — and in its first sentence: "It can be proven that most claimed research findings are false." 11 The ranking-factor study is that finding in commercial form.
Parsimony. Add variables to a model and its fit improves; that is arithmetic, not insight. Akaike's information criterion — in his own 1974 formulation, "(-2)log-(maximum likelihood) + 2(number of independently adjusted parameters within the model)" — penalises each added parameter against the likelihood it buys, and the model that wins is the one that explains the most with the least. 12 The fifty-factor study loses that contest to the five-factor one by construction, because forty-five of its factors are fitting the noise of the past and will not survive the next update.
Non-linearity and unequal variance. A linear correlation coefficient assumes the relationship is straight and the scatter is constant; web data are neither. The variance of outcomes grows with the size of the site, and effects saturate, reverse and interact. Rank-based measures, tests for heteroskedasticity and information-theoretic dependence measures exist for exactly this, and none of the vendor studies cited above reports a test of unequal variance or a dependence measure beyond a correlation coefficient. Semrush's 2017 study at least reported that its plain correlations were weak and unstable before reaching for a black-box model. That candour is rarer than it should be.
Chris Anderson wrote in Wired in 2008 that "petabytes allow us to say: 'Correlation is enough.' We can stop looking for models." 13 Six years later, in Science, Lazer and colleagues examined Google Flu Trends — the flagship of that claim — and found that it "was predicting more than double the proportion of doctor visits for influenza-like illness" that the public-health system was recording. They named the pattern: "'Big data hubris' is the often implicit assumption that big data are a substitute for, rather than a supplement to, traditional data collection and analysis." 14 That sentence is the thesis of this essay, published twelve years ago in Science. The industry read Anderson and did not read Lazer.
Fortsetzung folgt.
To be continued.
Sources
Quotations are verbatim from the cited documents. Where a primary text could not be retrieved and a quotation rests on a secondary source, the entry says so. Numbering is local to this chapter.
- Semrush, "Semrush Releases Expanded 2026 AI Visibility Index, Analyzing 126 Million AI Search Prompts", press release, 26 June 2026. https://www.semrush.com/news/463141-semrush-releases-expanded-2026-ai-visibility-index-analyzing-126-million-ai-search-prompts/
- Nick Lafferty, "AI Platform Citation Patterns", Profound, 5 June 2025. https://www.tryprofound.com/blog/ai-platform-citation-patterns
- Brian Dean, "We Analyzed 11.8 Million Google Search Results", Backlinko (updated 14 April 2025). https://backlinko.com/search-engine-ranking
- Xiao-Li Meng, "Statistical paradises and paradoxes in big data (I): Law of large populations, big data paradox, and the 2016 US presidential election", Annals of Applied Statistics 12(2), 2018, 685–726. https://projecteuclid.org/journals/annals-of-applied-statistics/volume-12/issue-2/Statistical-paradises-and-paradoxes-in-big-data-I--Law/10.1214/18-AOAS1161SF.full
- Valerie C. Bradley, Shiro Kuriwaki, Michael Isakov, Dino Sejdinovic, Xiao-Li Meng, Seth Flaxman, "Unrepresentative big surveys significantly overestimated US vaccine uptake", Nature 600, 2021, 695–700. https://pmc.ncbi.nlm.nih.gov/articles/PMC8653636/ · preprint https://arxiv.org/abs/2106.05818
- Ole Peters, "The ergodicity problem in economics", Nature Physics 15, 2019, 1216–1221, doi:10.1038/s41567-019-0732-0. Quotations via Lars P. Syll, 6 Dec 2019 (primary text paywalled). https://larspsyll.wordpress.com/2019/12/06/the-ergodicity-problem-in-economics-wonkish/
- Judea Pearl, "Causal inference in statistics: An overview", Statistics Surveys 3, 2009, 96–146 (Technical Report R-350). https://ftp.cs.ucla.edu/pub/stat_ser/r350.pdf · Judea Pearl, "The Seven Tools of Causal Inference with Reflections on Machine Learning", Technical Report R-481, November 2018. https://ftp.cs.ucla.edu/pub/stat_ser/r481.pdf
- rhinegold Compendium, "Counterfactual". https://insights.rhinegold.de/compendium/counterfactual/
- Yoav Benjamini & Yosef Hochberg, "Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing", Journal of the Royal Statistical Society B 57(1), 1995, 289–300. https://academic.oup.com/jrsssb/article/57/1/289/7035855
- Cristian S. Calude & Giuseppe Longo, "The Deluge of Spurious Correlations in Big Data", Foundations of Science 22, 2017, 595–612, doi:10.1007/s10699-016-9489-4. Author manuscript: https://www.di.ens.fr/users/longo/files/BigData-Calude-LongoAug21.pdf
- John P. A. Ioannidis, "Why Most Published Research Findings Are False", PLoS Medicine 2(8): e124, 2005. https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.0020124
- Hirotugu Akaike, "A new look at the statistical model identification", IEEE Transactions on Automatic Control 19, 1974, 716–723, doi:10.1109/TAC.1974.1100705.
- Chris Anderson, "The End of Theory: The Data Deluge Makes the Scientific Method Obsolete", Wired 16.07, 23 June 2008. https://www.wired.com/2008/06/pb-theory/
- David Lazer, Ryan Kennedy, Gary King, Alessandro Vespignani, "The Parable of Google Flu: Traps in Big Data Analysis", Science 343(6176), 14 March 2014, 1203–1205. https://gking.harvard.edu/files/gking/files/0314policyforumff.pdf
- Peverill Squire, "Why the 1936 Literary Digest Poll Failed", Public Opinion Quarterly 52(1), 1988, 125–133.
Notes on form
The dialogues owe an obvious debt to Douglas Hofstadter's Gödel, Escher, Bach — not in argument or conclusion, but in the idea that dialogue can itself be a form of reasoning.
