The channel with a number next to every euro
Two lines in a marketing budget. The first reports a return by Tuesday: spend, clicks, conversions, cost per acquisition, a ratio that can be defended in a meeting. The second reports nothing of the kind. It produces visibility, mentions, positions, answers given to people who never identify themselves, and an effect that arrives, if it arrives, in a quarter nobody is currently discussing.
A rational decision-maker moves money toward the first line. Not because they believe the second does not work, but because the first can be shown to work and the second cannot. Over several such decisions, a marketing budget stops being a judgement about what creates value and becomes a judgement about what can be observed to create value. Those are different things, and the gap between them is where this essay lives.
Daniel Yankelovich described the sequence in 1971, in a speech about market research:
"The first step is to measure whatever can be easily measured. This is okay as far as it goes. The second step is to disregard that which can't be easily measured or give it an arbitrary quantitative value. This is artificial and misleading. The third step is to presume that what can't be measured easily really isn't very important. This is blindness. The fourth step is to say that what can't be easily measured really doesn't exist. This is suicide." 1
It is worth noting, in an essay about measurement error, that this passage is almost always cited to the wrong source. 1
What the number actually counts
Start with the good news about attribution: it is not wrong. The tracked number measures what it measures, correctly. It goes wrong one step later, when someone reads it as an answer to a different question.
A session-based channel report answers where the conversion fell. It does not answer what the channel is worth. Those are two different accounting questions, and no amount of better tracking merges them. This is a misallocation of credit, not a measurement defect — which is why it cannot be fixed by fixing the tracking. It can only be fixed by keeping two sets of books.
What the tracked view sees is the last mile. In a considered purchase, the research that precedes it runs for weeks across devices, people and platforms, and most of it is structurally invisible: consent limits, cookie lifetimes, cross-device gaps, several people deciding together, corporate networks, and now a growing share of research that happens inside an answer engine and never touches a website at all.
This produces an artefact that looks like a finding. When almost every recorded journey shows a single touch, that does not mean buyers had one touch. It means one touch was seen. We keep that in our standing list of traps for exactly this reason, with the note that it cannot be lifted until upstream tracking exists — and it does not.
The direction of the error is systematic, which is what makes it a bias rather than noise: credit flows downstream, to whatever channel was present at the end. Early channels are structurally understated. Harvesting channels are structurally overstated. Or, in the sentence we use internally when the argument has to fit on one line: the last mile gets the credit for the whole journey.
The experiment that cannot fail
Now the uncomfortable part, and it is not a matter of opinion.
In 2012, eBay ran the experiment that most advertisers never run. It stopped bidding on its own brand keywords in a randomised subset of US markets and watched what happened to traffic and sales. The result, published in Econometrica: brand-keyword ads "have no measurable short-term benefits", because "almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search traffic." 2
The authors then did something more useful than reporting their own result. They compared it to what the standard industry method would have said about the same spend:
"In the absence of causal measures, the industry relies on 'attribution' measures which correlate clicks and purchases. By this measure, eBay performed very well." 2
The numbers are worth setting out, because the spread is the entire argument. Using the same data and the same period, a regression of sales on advertising spend without controls implied a return above 4,100 per cent. With time and geography controls, above 1,400 per cent. The experiment implied minus 63 per cent. 2
That is not a small discrepancy. It is a sign reversal, and the confident number is the wrong one.
This is not an isolated finding. When Lewis and Rao asked how large an advertising experiment must be before it can tell a wildly profitable campaign from one that merely broke even, the answer was that the median campaign in their sample "would have to be nine times larger" — and to resolve a ten-percentage-point difference in return, "62 times larger to possess adequate power—nearly impossible for a campaign of any realistic size." For the observational methods that firms actually use, their verdict is blunt: avoiding the relevant bias "appears to be an impossible statistical feat." 3 A separate line of work named the mechanism: people who see an ad are, in that same window, doing more of everything online. Comparing them to people who did not produces overestimates of effect "on the order of 200 times the correct value." 4 And when fifteen large advertising experiments were re-analysed with the observational methods normally used to report on them, half the estimates were "off by a factor of three." In one, the naive comparison said plus 316 per cent; the experiment said plus 73. 5
So the asymmetry at the start of this essay is not what it appeared to be. The measurable channel is not necessarily valued more accurately than the unmeasurable one. It is reported more confidently. Those are not the same thing, and the difference is not a rounding error — it is the difference between plus four thousand and minus sixty-three.
There is a reason for this, and it is structural. The always-on dashboard is not an experiment. It has no state of the world in which it reports that the channel added nothing, because every conversion that passes through its window is counted as evidence for it. A test that can only confirm proves nothing. We keep that as a standing rule after finding one of our own gates in exactly that condition: a check drawn against a threshold taken from the same distribution it was checking, mathematically incapable of failing. A gate that cannot fail creates confidence without inspection, which makes it more dangerous than no gate at all.
The ratchet
If the error were random, it would be an annoyance. It is not random, and that is what turns a measurement problem into a strategic one.
Budget follows credit. Credit accrues downstream. So budget moves toward the harvesting channel, and away from the channel that produced the demand being harvested. The channel that loses funding gets weaker. As it gets weaker, more of the demand has to be bought rather than earned — which makes the harvesting channel look even better in attribution, because it is now genuinely present at more endings. More budget follows. We describe the terminal state of this loop in one sentence: paying rent for demand you used to own.
Paying rent for demand you used to own.
Economists have studied what loops of this shape do to reversibility. Brian Arthur showed in 1989 that adoption processes with increasing returns become "progressively more 'locked in'", and that they are non-ergodic — historical small events "are not averaged away and 'forgotten' by the dynamics — they may decide the outcome." 6 Past a point, the market "becomes increasingly 'locked-in' to an inferior choice", and a subsidy that would have redirected it early can no longer close the gap. 6 Arthur gives us the formal vocabulary for what such a process can become: path-dependent, increasingly difficult to reverse, and capable of locking in an inferior allocation.
And the counterparty in this arrangement is not passive. In August 2024, a United States federal court found, after trial, that Google holds monopoly power in general search text advertising. The findings of fact are specific about what that power was used for. Google "strategically has used pricing knobs to raise text ads prices", and did so incrementally, so that advertisers "would view price increases as within the ordinary price fluctuations, or 'noise,' generated by the auctions." Internal documents discussed a ten per cent increase as "safe" because it fell "within usual WoW noise". The court's summary: "through barely perceptible and rarely announced tweaks to its ad auctions, Google has increased text ads prices without fear of losing advertisers", and "many advertisers do not even realize that Google is responsible for the changes in price." 7
The same opinion records something that should unsettle anyone who chose this channel because it was transparent. The court found the product had degraded in two ways: advertisers "receive less information in search query reports", and can no longer opt out of keyword matching. Google "removed information from SQRs that provided advertisers with insight into low-volume queries, which diminished advertisers' ability to tailor their ad strategy." The court called these "arguably small changes" that nonetheless "reveal Google as a monopolist unconcerned about product changes that have decreased advertisers' autonomy over the auctions." 7
Read that alongside the ratchet. The channel selected for its measurability becomes less measurable over time, at prices that rise below the threshold of perception, and the budget keeps moving toward it — because the dashboard still produces a number, and the number still looks fine.
The panel became the universe
At this point an essay of this kind usually turns to the reader and explains what decision-makers get wrong. That would be dishonest, because we made the same mistake inside a system built specifically to prevent it, and we made it last month.
We value pages individually where we can measure them. The set of pages we can value that way is a panel — a deliberate, bounded, entirely legitimate measurement panel. What happened next is recorded in our own standing list of traps, in the entry that opened on 25 August:
A measurement panel is not a universe of value. In practice it became one: not in the panel, therefore no page-specific search value, therefore small or worthless in any comparison of interventions.
And the second half of that entry is worse, because it identifies the loop rather than the oversight:
Both scales define value through visibility — a circular argument that keeps invisible precisely those pages whose performance an intervention would first have to activate.
That is Yankelovich's third and fourth step, executed by a measurement system whose entire purpose is to catch that kind of error, against its own data, for months, without anyone noticing. The correction was not a better estimate. It was a change of grammar: every object now carries a coverage status alongside its value — fully valued, partially valued, allocated by proxy, exposure only, unobserved — and "no result" is reported as unknown, never as small.
Nothing about this was visible from the output. That is the part worth dwelling on. We have a separate trap for it, opened after a report was delivered with its title, axes, legend, counters and tiles all correct and only the data points missing. The script exited zero. Every file was written. The note we left ourselves reads: an output that looks almost complete is more dangerous than one that is missing.
Which is the precise answer to the objection that the dashboards were always fine. They were. That is not evidence that the measurement was.
Abraham Kaplan put the general form of this in 1964, and it is more uncomfortable than the hammer line he is usually quoted for: "It comes as no particular surprise to discover that a scientist formulates problems in a way which requires for their solution just those techniques in which he himself is especially skilled." 8 A marketing organisation formulates its strategy in a way that requires for its evaluation just those techniques it already has. Measurability stops being a property of the instrument and becomes a selection rule on the strategy. And once a number is the basis of the decision, Strathern's formulation of Goodhart's law applies: "when a measure becomes a target, it ceases to be a good measure." 9
The second experiment
There is an instrument that answers the question the dashboard cannot. Switch the channel off, in a declared way, and see what still arrives.
It is called the second experiment because the first one was never an experiment. The permanent measurement is a receipt for money already spent. A properly designed holdout is allowed to come back with an answer nobody wanted.
It is also easy to do badly, and doing it badly is worse than not doing it, because it produces a number with a dead experiment's authority. These are the constraints we impose on our own, offered as a design rather than as a result:
Declare before you switch. The specification — segment definition, windows, verdict date — is frozen on the day of the switch-off, before the first post-period observation exists. No retrospective fitting. Where the design later has to deviate, the deviation is declared, not written out.
Freeze the definition of the segment, including the noise you know it contains. A term set assembled after the fact is an argument, not a boundary.
Count in whole periods, and only mature ones. A week counts when its data is complete, not when its last day has passed. No verdict before the end of the switch-off window.
Set an abort threshold that fires automatically. Ours reactivates the channel if reach falls below a declared fraction of baseline for two consecutive complete weeks. An experiment without a stopping rule is not an experiment; it is a gamble with a report attached.
Check leakage where leakage happens, at the level of individual terms rather than by asking the people running the channel — and declare the blind zone you cannot check. In our own case one such guard was not in place at the start. That is in the specification, in writing, as an open risk, because the alternative is a clean-looking result resting on an unexamined assumption.
Use actual costs, not planned ones, and compute rates only from mature periods.
Bring a second, independent source. Our own strongest finding in this area does not come from the tracker at all; it comes from a customer-relationship record that confirms the same thing through a different pipe. A finding that survives two unrelated measurement paths is worth more than a precise finding from one.
Discard your control arm if it is unfit. In one design the obvious comparison series was falling on its own, which would have inflated the measured effect. We ran without it and said so, rather than keeping a comparison that flattered the result.
And then accept what the instrument gives you. What a switch-off measures first is substitution, not demand: the same searchers, now clicking the organic result instead of the paid one. That is a real and valuable finding about who was paying for whom. It is not proof of business damage, and it is not an acquittal either. Our own design does not separate changes below roughly a fifth from noise. That limit is in the specification, not in a footnote.
There is a rule in this that generalises past marketing. Unmeasurable is a permitted result. When the comparison pool is too small, when the base is below the floor, our estimators return not a number — never zero. The most expensive errors we have logged were all cases where something missing passed itself off as something measured.
What this does not prove
An essay that argues against convenient evidence does not get to be selective about inconvenient evidence.
Paid brand search is not universally worthless. A large set of experiments across thousands of brands found a positive causal effect of brand ads of one to four per cent where no competitor was bidding — smaller for stronger brands. 10 The same work is explicit that the eBay case, a very strong brand facing no competitor on its own name, "is not the norm." 10 The null result came from the most favourable possible conditions for a null result.
And there is a case where the spend is rational. For brands that face competitors on their own brand search and choose not to bid, competitors take between 18 and 42 per cent of the clicks; defending against that has a "strongly positive" return. 10 At the system level this begins to resemble rent extraction more than value creation: the advertiser is paying to defend access to demand it already created. But the money is real, and telling a company to stop defending itself is not advice, it is a bill.
Non-brand advertising reaches people the organic channel does not. The eBay analysis found that infrequent and new users are genuinely influenced; the negative average came from spending most of the budget on frequent users who would have bought anyway. 2 The finding is about allocation within the channel as much as between channels.
The opposite overcorrection is also wrong, and we rejected it in writing before anyone outside could: abandoning tracking because the best customers are invisible inverts the causality — and without measurement, not one of the findings in this essay would exist.
Both sides of this argument have an interest. The most-cited evidence that long-term brand effects are systematically underweighted comes from an advertising industry body, drawn from campaigns its own members submitted to an effectiveness competition. That is a self-selected sample published by a party that benefits from the conclusion. The conflict runs in the opposite direction to the one this essay has been describing, and it is not smaller. We have not used it.
And the organic channel is not sovereign either. A university study of the staggered rollout of AI summaries in search found a reduction in external search referrals to English Wikipedia articles of around five per cent, measured against the same articles in two other languages. 11 Nobody at Wikipedia did anything. Both channels depend on design decisions made by a third party. The difference is only that in one of them, you can see it happening.
The question
The point of an incrementality laboratory is not to prove that one channel is better than another. It is to make the real drivers visible — where channels cannibalise each other, and where value is genuinely added — and to be capable of returning an answer the person who commissioned it did not want.
Which is why the honest ending here is not a recommendation about budget. It is a question about instruments.
In our own canon, among the assets we track, there is a line for resilience of distribution — independence from paid media. Next to it, where a figure would go, it says: not carried as a holding; no metric. We named the dependency and we do not measure it. We wrote that down instead of quietly leaving it out, which is the only part of this we would defend.
The first experiment is running in your account right now, and it cannot fail. The second one is the one you have to schedule.
Sources
Quotations are verbatim from the cited documents. Where a primary text could not be retrieved and a quotation rests on a secondary source, the entry says so. Conflicts of interest are stated, including those that cut against this essay's argument.
- Daniel Yankelovich, speech "The New Odds", 15 October 1971; transcript held in the Daniel Yankelovich Papers, UC San Diego Library. The passage is commonly cited to Corporate Priorities (1972) and popularised by Charles Handy, The Empty Raincoat (1994); the abridged version published in Sales Management in November 1971 does not contain the four evaluative clauses. Quoted here from a secondary source that documents the archival record; the archive transcript itself was not inspected.
- Thomas Blake, Chris Nosko, Steven Tadelis, "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica 83(1), 2015, 155–174. https://faculty.haas.berkeley.edu/stadelis/BNT_ECMA_rev.pdf — Conflict: two authors were with eBay Research Labs and the third held an eBay affiliation; the data are eBay's. This is an advertiser reporting on its own spending. Return figures are short-term, and the revenue and cost basis was reconstructed from public filings because the underlying figures are proprietary.
- Randall A. Lewis, Justin M. Rao, "The Unfavorable Economics of Measuring the Returns to Advertising", Quarterly Journal of Economics 130(4), 2015, 1941–1973. — Conflict: both authors were previously employed by Yahoo!, where the twenty-five experiments were run. The finding runs against the platform's interest.
- Randall A. Lewis, Justin M. Rao, David H. Reiley, "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW 2011, 157–166. http://www.davidreiley.com/papers/HereThereEverywhere.pdf — Conflict: all three authors were at Yahoo! Research. The authors also report a case in which observational methods underestimated rather than overestimated.
- Brett R. Gordon, Florian Zettelmeyer, Neha Bhargava, Dan Chapsky, "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook", Marketing Science 38(2), 2019, 193–225. — Conflict, as disclosed in the paper: two co-authors were Facebook employees, all data are Facebook's, and the two academic authors were contingent Facebook employees receiving a token salary of fifteen dollars per week, donated to charity. The finding devalues the platform's own attribution measurement.
- W. Brian Arthur, "Competing Technologies, Increasing Returns, and Lock-In by Historical Events", The Economic Journal 99(394), 1989, 116–131.
- United States v. Google LLC, No. 1:20-cv-03010-APM, Memorandum Opinion, U.S. District Court for the District of Columbia, 5 August 2024 (Mehta, J.). The court found monopoly power in general search services and general search text advertising; it did not find monopoly power in the broader search advertising market. The remedies opinion of 2 September 2025 declined divestiture of Chrome and imposed conduct remedies, including advertiser transparency provisions not analysed here.
- Abraham Kaplan, The Conduct of Inquiry: Methodology for Behavioral Science (Chandler, 1964), p. 28.
- Marilyn Strathern, "'Improving ratings': audit in the British University system", European Review 5(3), 1997, 305–321, at p. 308. Compare Donald T. Campbell, "Assessing the impact of planned social change", Evaluation and Program Planning 2(1), 1979 — quoted here from secondary sources; original pages not inspected.
- Andrey Simonov, Chris Nosko, Justin M. Rao, "Competition and Crowd-Out for Brand Keywords in Sponsored Search", Marketing Science 37(2), 2018, 200–215. — Conflict: the experiments ran on a commercial search platform and the authors have platform employment histories. The findings cut both ways: small positive effects without competition, strongly positive returns to defensive bidding, and crowd-out of roughly half of organic clicks when a brand ad is added.
- Marc Khosravi, Hema Yoganarasimhan, "Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia", University of Washington, arXiv:2602.18455, September 2026.
Internal rules and findings quoted without a number — the standing traps, the coverage states, the holdout design and the canon entry on resilience of distribution — are from rhinegold's own measurement governance and are not public documents. No client, sector, or client figure appears in this essay.
Related in the Compendium:* Attribution · Self-Reported Attribution · Counterfactual · Control Group · Placebo Test · Difference-in-Differences · Synthetic Control · Spillover & Contamination · Touchpoints · Zero-Click Search · Vendor Concentration Risk.
