Cookie preferences

Choose what we may use. Strictly necessary cookies keep you signed in and can't be turned off; everything else is your call and changes nothing about how the site works. You can change your mind any time from the footer.

A rack of blue-capped plastic test tubes on a pale blue surface.

Attribution tells you who. Only a holdout tells you whether.

Photo by Tara Winstead (opens in a new tab) on Pexels

Attribution answers which source a customer came through. It cannot answer whether that customer would have bought anyway, and for branded search, retargeting and every channel that reaches people already on their way to you, that is the question that decides the budget. Here is what an incrementality test is, when one is worth running, how a geographic holdout works with your own sales data, and how its answer fits beside an attribution report rather than replacing it.

Contents
  1. What attribution cannot tell you
  2. The cautionary tale every marketer should know
  3. Why lift tests are harder than they look
  4. When a holdout test is worth running
  5. How a geographic holdout works with your own sales
  6. How incrementality and attribution fit together
  7. A short checklist before you pause anything

Attribution tells you which source a customer came through; it cannot tell you whether they would have bought anyway. A holdout test can: withhold a channel from comparable regions or people and compare sales. Run one on your largest or most suspect channel, measure it in closed sales, and read the result beside your attribution report, never instead of it.

Every attribution method, however careful, answers the same question: which source did this customer come through? Matched to closed sales on phone and email, that question has a factual answer, and a business that knows it is in a far better position than one reading platform dashboards.

It is not the only question, and for some channels it is not the important one. A customer who typed your company name into Google and clicked your ad really did come through branded search. Whether that ad did anything — whether the customer would have clicked the free listing directly beneath it — is a question no amount of matching can answer. It needs an experiment.

What attribution cannot tell you

Attribution records association: this customer, this source. Incrementality measures causation: sales that would not have happened without the channel. The two diverge most for channels that reach people who are already on their way to you, and those channels usually look best in attribution reports.

The gap is not a flaw in any particular model. It is built into the question. An attribution report looks only at customers who bought and asks where they came from; it never sees the counterfactual customer who would have bought without the ad. Multi-touch models, modeled conversions and matched revenue all share this limit, because all of them start from the sales that happened.

What each kind of measurement can and cannot answer
MethodQuestion it answersWhat it observesWhat it cannot answer
Platform-reported conversionsWhich events followed our ads within the window?Events the platform seesSales offline; anything outside the window; causation
Matched revenueWhich closed sales came through which source?Your own sales and lead recordsWhether the sale would have happened without the source
Mix modelingHow do sales move with spend across channels over time?Aggregate spend and salesIndividual sales; causation without an experiment to anchor it
Holdout or lift testHow many sales did this channel cause, at this spend, in this period?The difference between treated and withheld groupsWhich individual sales were caused; other channels, periods or spend levels

We compared the first three at length in MTA vs MMM. The holdout belongs in a different column because it is the only one that changes the world on purpose and watches what happens. Everything else observes the world as it was and argues about why.

The cautionary tale every marketer should know

The best-known incrementality result comes from eBay, which switched off paid search to see what it was buying. For brand keywords, almost all of the paid clicks it gave up came back through the free search results. The ads had been collecting customers, not creating them.

eBay is an unusual advertiser — a household name with enormous organic visibility — and a small business bidding on its own name against competitors who also bid on it faces a different situation. The lesson is not that brand search is worthless. It is that an attribution report showing brand search as your best channel is entirely consistent with brand search causing almost nothing, and only a test can tell the two apart.

A channel that reaches people already looking for you will always look excellent in attribution, whether or not it changes anything. Branded search, retargeting and offers to existing customers are the first candidates for a holdout.

Why lift tests are harder than they look

Incrementality tests are simple to describe and hard to make conclusive. Sales vary a lot from week to week and person to person, while the effect of most advertising is small by comparison, so a test needs a large sample or a long run before the signal clears the noise.

That paper is often quoted as an argument against testing. It is better read as an argument for testing only where the answer is worth the cost and the volume can support it. The same authors note that the observational methods marketers use instead suffer from selection bias, because advertising is targeted. A noisy experiment is uncomfortable; a confident observational estimate that is biased is worse.

Platform lift tools face the same arithmetic, and their documentation is candid about it. The feasibility check, the budget recommendation and the minimum detectable effect all exist because a test that is too small returns a confidence interval rather than an answer. Google's geographic version is the closest thing to what an offline business can run for itself.

When a holdout test is worth running

Run a test when a decision of real size hangs on whether a channel is incremental, and when your volume can detect the effect you care about. That usually means your largest channel, a channel you suspect of collecting customers rather than creating them, or a budget change big enough to matter.

  • Your largest line item. If a third of the budget goes to one channel, a test that confirms or questions it once a year is cheap insurance.
  • Channels that reach existing demand. Branded search, retargeting, SMS to existing subscribers and Performance Max campaigns that lean on remarketing.
  • Channels attribution cannot see. Video, billboards, direct mail and podcasts produce few traceable responses, so a test is the only route to their full effect.
  • Before a big change. Doubling a channel or cutting it to zero is a decision worth a few weeks of evidence first.

And do not run one when the answer cannot change anything, when the channel is too small for the effect to be visible in your sales, or when something else in the business — a price change, a new location, a seasonal peak — will move sales during the test by more than the channel could.

How a geographic holdout works with your own sales

Split the areas you serve into two comparable groups, measure how their closed sales compare over a baseline period, pause the channel in one group, and compare again. The difference from the baseline relationship, measured in closed revenue, is the channel's incremental effect at that spend.

  1. Choose the regions. Group ZIP codes, cities or service areas into two sets with similar sales history. Matching on history matters more than matching on size.
  2. Measure the baseline. Take at least eight weeks of closed sales for both groups before the test and compute the ratio between them. If that ratio swings a lot from week to week, note by how much: an effect smaller than the swing will not be visible.
  3. Pause the channel in one group only. Keep everything else the same in both, and write down the start and end dates before you begin.
  4. Measure closed sales, not platform conversions. The point is to measure the business. Use the same sales export you use for attribution, split by location.
  5. Allow for the sales cycle. Keep counting sales from leads created during the test for at least one typical cycle after it ends.

Here is the arithmetic for a business that pauses a paid social channel in half its regions for eight weeks. In the eight weeks before the test, the regions that will keep the ads sold $400,000 and the regions that will lose them sold $380,000, a ratio of 1.0526. During the test, the regions with ads sold $430,000 and the holdout regions sold $372,000. The channel spent $15,000 in the regions where it ran.

Reading an eight-week geographic holdout
LineFigureArithmetic
Baseline ratio, test regions to holdout regions1.0526$400,000 ÷ $380,000
Expected sales in test regions without the channel$391,579$372,000 × $400,000 ÷ $380,000
Actual sales in test regions with the channel$430,000From the sales export
Incremental revenue$38,421$430,000 − $391,579
Spend in test regions$15,000From the invoice
Incremental return per dollar2.56$38,421 ÷ $15,000
Matched revenue credited to the channel in the same regions$61,000From the attribution report
Share of matched revenue that was incremental63%$38,421 ÷ $61,000

Two numbers come out of this, and both are useful. The channel produced about $2.56 of revenue for every dollar it cost, which is a causal estimate and the one to compare with your margin. And about 63% of the revenue your attribution report credited to it would not have happened without it, which tells you how to read that report for this channel until the next test.

Treat both with the humility the method deserves. The estimate holds for that channel, at that spend, in that season. If weekly sales in the baseline varied by ten percent and the effect you measured is nine, you have a hint rather than a result, and the honest report says so.

A holdout gives you a causal number for one channel, at one spend level, in one period. That is narrower than an attribution report and far more trustworthy. Use it to calibrate how you read the report, not to replace it.

How incrementality and attribution fit together

Attribution is the operating view: every month, every channel, every sale traced to a source. Incrementality is the audit: occasional, expensive, narrow and causal. A business that only attributes is trusting its most flattered channels; one that only tests has no idea what happened last month.

The practical sequence is the one we argued for in MTA vs MMM. Start with matched revenue, because it is cheap, monthly and grounded in your own sales. Use it to see which channels matter enough to test. Test the largest and the most suspect, one at a time, and record the design and result. Then read each month's attribution report with those results beside it: a channel whose test showed 63% incrementality is still credited with its matched revenue, and everybody in the room knows roughly how much of it to believe.

What you should not do is multiply the attribution report by the test result and publish the product as revenue. The matched figure is a fact about specific sales. The incremental share is an estimate about a period. Mixing them produces a number that is neither, and it will not survive the first question from finance.

CloseRev is the attribution half of this, not the experiment. It matches closed sales to lead sources on email and phone and reports revenue by channel and campaign, with unmatched sales left in Direct / Unknown; it does not run holdouts or measure lift. If your sales file has a Location column, it can break the same revenue down by region, which is the closed-sales view a geographic holdout reads from.

A short checklist before you pause anything

A holdout costs real revenue in the held-out regions, so the design is worth getting right before the first ad is switched off. Five questions settle most of it.

  1. Which decision will this result change, and by how much money?
  2. Is the channel's likely effect larger than the normal week-to-week swing in sales between the two groups?
  3. Can the two groups be kept apart — is the channel's targeting geographic, and do customers cross the boundary?
  4. Will the test run long enough to cover the sales cycle, and will you count sales that close after it ends?
  5. Have you written down the regions, dates and success measure before starting, so nobody can redraw them after seeing the result?

If the answers hold, run it. One well-designed test a year on your biggest channel will teach you more about where your money goes than any amount of dashboard, and it will make every attribution report you read afterward more honest.

Questions people actually ask

What is an incrementality test?
An experiment that withholds advertising from a comparable group — of people or of regions — and compares their sales with a group that kept seeing it. The difference is the incremental effect: the sales the advertising caused rather than merely accompanied. It is the only common marketing measurement that observes causation instead of inferring it.
What is the difference between attribution and incrementality?
Attribution assigns each sale to a source the customer came through. Incrementality estimates how many sales would not have happened without a channel. A channel can have plenty of attributed revenue and little incremental revenue, if it mostly reaches people who were going to buy anyway. Branded search is the classic example.
What is a geo holdout test?
A test that pauses a channel in some geographic areas and keeps it running in comparable others, then compares closed sales between the two groups against how they compared before the test. It needs no tracking of individuals, which makes it the most practical lift test for a business that sells offline or in several regions.
How long should a holdout test run?
Long enough to cover your sales cycle and for the difference to stand out from normal week-to-week variation. For a business whose sales close within days, four to six weeks is often enough. For one with a multi-month cycle, a short test only measures the leads, and the revenue has to be followed up afterward.
Is incrementality testing worth it for a small business?
Usually only for its largest channel, and only if sales volume is high enough to detect a difference. In one study of twenty-five large experiments, the median confidence interval on return was more than 100 percentage points wide. If your weekly sales swing more than the effect you hope to find, the test cannot answer, and it is better to know that before pausing anything.

See it on your own numbers.

Two exports and a few minutes. 14 days free, no card, nothing to install.