Every attribution model is an assumption about causation wearing the costume of a measurement. Incrementality testing is how you check the assumption.
The usual objection is that it needs scale, specialists and budget. Two of those three are wrong. What it actually needs is a design decision made before the test starts, and the discipline not to stop it early.
What is incrementality?
Incrementality is the share of conversions that happened because of a marketing activity and would not have happened without it.
The distinction from attributed conversions is the whole point. An attributed conversion is one a model assigned to a touchpoint. An incremental conversion is one the touchpoint caused. Retargeting illustrates the gap: it reaches an audience selected specifically for already being interested, and then receives last-click credit when they do what they were going to do anyway.
Which test design fits your situation?
Three designs cover most cases, and the right one depends on what you can control.
| Design | Requires | Best for | Main weakness |
|---|---|---|---|
| Audience holdout | Platform-level exclusion of a random share | Platforms with built-in holdout support | Only measures within one platform |
| Matched market | Several comparable regions | Geographically spread demand | Markets are never perfectly matched |
| Switchback | Nothing except patience | Small businesses, single market | Confounded by anything else that changes over time |
Audience holdout
The cleanest design. A random share of your target audience is excluded from seeing the advertising, and conversion rates are compared between the exposed and held-out groups.
Randomisation is what makes it clean: the two groups differ only in whether they saw the ads. Several platforms offer this natively as a conversion lift study, which removes most of the implementation work.
Its limit is scope. A platform-run holdout measures that platform's incrementality against no advertising from that platform — not against your other channels, which are still running and may be picking up the slack.
Matched market test
Pick regions with similar historical performance. Keep advertising running in one set, pause it in the other, compare total outcomes.
The design works when demand is geographically spread and regions are genuinely comparable. Matching is the hard part and it is done before the test, not after: select pairs on several months of historical sales patterns, not on population or a guess about similarity.
Its weakness is that markets are never perfectly matched, and anything local — weather, a competitor's promotion, a regional news story — lands inside one side of the comparison.
Switchback
The design available to everyone. Alternate a channel on and off over successive periods — a week on, a week off, repeated several times — and compare.
Repetition is what makes it credible. A single on-off comparison is confounded by everything else that changed between the two periods. Six alternating periods average much of that out, because unrelated events are unlikely to align with your schedule consistently.
It needs patience: a meaningful switchback test runs for months, not weeks. For a single-market business with no way to split audiences, it is frequently the only honest option available.
How do you size the test before running it?
This is the step that gets skipped, and skipping it is why most incrementality tests end inconclusive.
Decide the minimum detectable effect first: the smallest difference that would actually change a decision. Not the smallest difference you could detect with infinite data — the smallest one you would act on.
If discovering that a channel is 10% less incremental than reported would not change your budget, do not design a test to detect 10%. Design it to detect the difference that would. A test powered to find a 40% shortfall needs far less volume than one powered to find 5%, and in most small accounts only the former is achievable.
A worked example
Illustrative figures, chosen to show the arithmetic rather than to suggest benchmarks.
A business spends 20,000 units a month on retargeting. The platform reports 1,000 attributed conversions from it. They suspect a large share would have happened anyway.
They run a 20% audience holdout for six weeks.
| Exposed group (80%) | Held-out group (20%) | |
|---|---|---|
| Audience size | 80,000 | 20,000 |
| Conversions in period | 3,600 | 810 |
| Conversion rate | 4.50% | 4.05% |
| Rate difference | — | 0.45 percentage points |
Scaling the held-out group's rate to the exposed group gives 3,240 conversions expected without advertising, against 3,600 observed — an incremental effect of roughly 360 conversions, not the 1,000 attributed. On these invented numbers the channel is genuinely working, and it is doing about a third of what its reporting claims.
That changes the decision. The true cost per incremental conversion is about three times the reported figure, and whether the channel still clears the acquisition ceiling at that price is now an answerable question rather than an argument.
How much does it cost?
Less than the intuition suggests, and the cost is self-limiting in a useful way.
The cost is the incremental revenue you forgo from the held-out group for the test duration. If the channel is highly incremental, the test is expensive — and you have learned the channel is genuinely valuable. If it is barely incremental, the test costs almost nothing — and you have learned something that will change your budget permanently.
The cost scales with the value of the finding. That is an unusually good property for a piece of research.
Common mistakes
Stopping early because it looks good. Every day of a running test is a sample from a distribution. Stopping at the first favourable reading systematically overstates effects, and this is by far the most common error.
Testing during an unrepresentative period. A holdout across a major sale or seasonal peak measures incrementality during a sale, which may not generalise.
Holding out too small a share. A 5% holdout in a low-volume account produces a group too small to say anything. Better a 30% holdout for six weeks than 5% for six months.
Running one test and treating it as permanent truth. Incrementality changes with saturation, creative, competition and seasonality. A result is accurate for the conditions it was measured in, and worth re-checking when those conditions change materially.
Where this fits
Incrementality testing is the independent check that keeps the rest of the measurement stack honest. Platform reporting optimises; modelling allocates; experiments tell you which of those to believe.
One well-designed holdout per quarter on the largest channel is enough to calibrate everything else — which is a modest habit for the amount of budget it protects. The wider picture is in marketing measurement in a privacy-first world, and the layers it checks are described in the modern attribution stack.