Difference-in-Differences
- DiD = (treated change) − (untreated change): the subtraction removes what hit everyone.
- Parallel pre-trends are the admission ticket — check them before trusting any estimate.
- Pools of comparable untreated units beat single matched comparators on stability.
- Against a declining market path, flat is a win — report the comparison trajectory, not just the delta.
Consensus definition
DiD is one of the most widely used quasi-experimental designs in applied economics.1 The logic: take the before/after change in the treated group, take the before/after change in a comparable untreated group, and subtract. The untreated change estimates the counterfactual path — what the treated group would have done anyway. The design's canonical demonstration compared employment changes across two neighbouring states after one raised its minimum wage,2 and its load-bearing assumption is parallel trends: absent the intervention, both groups would have moved alike. The assumption is untestable for the treatment window itself, but its plausibility is checkable — groups that moved in parallel before the intervention are credible counterfactual carriers.
rhinegold operator refinement
Rhinegold treats DiD as the workhorse of visibility measurement, because the conditions that break before/after reading are permanent in this field: search demand is seasonal, platforms update continuously, and AI surfaces are redistributing clicks market-wide. All of it hits treated and untreated pages alike — which is exactly what the second difference removes. Two hard-won practice rules. First, the comparison group carries the whole design: it must be structurally similar, untouched by the intervention, and checked for parallel movement in the pre-period. Second, prefer a pool of comparable untreated units over a single matched unit — one comparator carries idiosyncratic noise; the median of a structurally similar pool is a far more stable counterfactual.
Operational use
Define treatment and comparison units before the intervention goes live, snapshot the pre-period for both, and verify parallel pre-trends. After the intervention, read the effect as (treated change) minus (comparison change) — in growth rates when levels differ. Report the comparison group's own trajectory alongside the effect: against a declining market path, an unchanged treated metric is a positive result, and the readout should say so.
Measurement boundary
DiD stands or falls with parallel trends — when the groups were already diverging before the intervention, the second difference measures that divergence, not the effect. It also assumes the intervention does not leak into the comparison group (no spillover), which site-wide signals can violate. DiD yields the incremental effect of an intervention; it does not allocate credit across channels — that remains Attribution.
Distinct from
Against before/after measurement, which has no second difference and absorbs all market drift into the effect estimate. Against a randomized experiment, which creates comparability by assignment — DiD substitutes design and assumption-checking where randomization is impossible, which is the normal case for organic and AI-visibility work. Against the Counterfactual itself: the counterfactual is the question, DiD is one disciplined way of answering it. Against Synthetic Control, which constructs the counterfactual as a weighted combination of donor units rather than using a single observed control — stronger when no individual peer is a clean match, but more demanding on pre-period data.
Observed pattern in practice
Common mistakes
- Skipping the pre-trend check — parallel movement before the intervention is the evidence that the comparison group can carry the counterfactual.
- Letting the intervention contaminate the comparison group — internal links, site-wide template changes, or shared rankings leak treatment into the control.
- Using a single matched comparator and reading its noise as effect — pools of structurally similar units produce stabler counterfactuals.
- Comparing absolute levels between groups of different size — DiD on growth rates or shares, not raw counts, when bases differ.
Where consensus is missing
The method is textbook-settled; its application to visibility data is not. There is no published standard for comparison-pool construction in SEO/GEO contexts, no agreed window lengths for pre- and post-periods against platform update cycles, and no convention for handling interventions that arrive in overlapping waves.
Sources & deeper reading
- rhinegold pool-based DiD practice for organic and AI-visibility interventions — comparison-pool construction and pre-trend verification, shared with engaged clients and partners.
