Incrementality testing answers the question attribution can’t: would these sales have happened without the ads? You switch ads off for a controlled group of people or regions, compare them with a group that still sees ads, and the difference is what the spend actually caused. On a mid-size budget, geo holdouts, platform lift studies and simple on/off tests are enough to catch the channels that are claiming credit for sales they didn’t create.
Attribution vs incrementality
Attribution divides credit for conversions that already happened among the touchpoints before them. Incrementality asks a counterfactual question: how many of those conversions would have happened anyway?
| Attribution | Incrementality | |
|---|---|---|
| Question | Which touchpoints were present before the sale? | How many sales did the ads cause? |
| Method | Rules or models applied to tracked journeys | Controlled experiment with a holdout group |
| Frequency | Daily, always on | A few tests a year |
| Typical bias | Over-credits channels close to purchase | Noisy if the test is too small |
| Best use | Day-to-day optimization | Setting budgets and calibrating attribution |
The two work together. Attribution runs the daily optimization; incrementality tells you how much to discount what it reports. If you’re still deciding between modeled approaches, MMM vs multi-touch attribution covers that choice, and experiments are what keep either model honest.
When you need an incrementality test
Run a test when a real budget decision depends on the answer:
- Platform-reported conversions add up to more than your actual orders or pipeline. Several channels are claiming the same customers.
- A large share of spend sits in retargeting, brand search, PMax or existing-customer audiences. These reach people already on their way to buying, so they look efficient almost by definition.
- You’re about to scale a channel meaningfully, or cut one, and want to know what happens to total sales.
- You’re adding a channel that tracks poorly, such as YouTube, connected TV or podcasts.
- Blended numbers disagree with platform numbers. Platform ROAS rises while total revenue stays flat.
Don’t test when your volume is too small for the test to see the effect you expect. For example, if weekly revenue swings 15% on its own and a channel might add 3%, a small, short test will likely end inconclusive and you’ll have paid for the holdout for nothing.
Test types: geo holdouts, lift studies and on/off tests
| Method | How it works | Best for | Main weakness |
|---|---|---|---|
| Geo holdout | Turn ads off in matched regions, compare with regions still running | Large channels, cross-channel questions, anything measured in your own sales data | Needs enough regions and volume; weeks of lost exposure in holdout areas |
| Platform lift study | The platform randomly withholds ads from a control group of users | Meta campaigns, YouTube and video | Counts only what the platform can track; eligibility varies |
| On/off test | Pause a campaign for a period, compare with before and after | Brand search, retargeting, small campaigns | Seasonality and other changes muddy the comparison |
Geo holdouts
Split your market into regions (US states or media markets, UK regions, cities elsewhere) and match treatment and holdout regions on how their sales moved together in the past. Turn the channel off in the holdout regions, then compare actual sales in the treatment regions with what the holdout regions predict they would have sold without the ads.
Measure the outcome in your own systems: Shopify orders by shipping region, CRM pipeline by company location, total revenue. Open-source tools help: Meta’s GeoLift handles market selection and analysis, and Google’s CausalImpact estimates the counterfactual from comparison regions.
Platform lift studies
Meta’s Conversion Lift and Google’s Conversion Lift (mainly for YouTube and video) randomly split people into an exposed group and a control group that can’t see your ads, which removes attribution’s biggest bias. Availability and minimum budgets vary, so check Meta’s Experiments tool, Google Ads’ measurement tools or your Google rep. The catch: the platform measures its own conversions, so validate big results against your own data.
On/off and brand search tests
The simplest version pauses a campaign for two to four weeks and watches total results, not platform results. For brand search, track total brand clicks (paid clicks plus organic clicks from Search Console on brand queries) and total conversions. If organic picks up most of what paid used to take, brand spend is largely not incremental. Check first whether competitors bid on your name; if they do, pausing hands them the top spot, and that should be part of the result. Where you can, pause brand search in holdout regions only, which turns a before-and-after comparison into a controlled one.
Testing Performance Max and retargeting
PMax and retargeting deserve a test because their reported ROAS is built on people who were already close to buying. For PMax, a geo holdout measured in your own sales data is the most independent read. Google Ads’ built-in experiments can compare running with and without PMax, but they grade the result with Google’s own conversion tracking. Apply brand exclusions before testing, or you’ll mostly measure brand demand; the Google Ads audit checklist covers that setup.
For retargeting, use a lift study on the retargeting campaign or pause it in holdout regions. If retargeting runs inside a broad automated campaign such as Meta’s Advantage+ sales campaigns (formerly Advantage+ Shopping) or PMax, you can’t cleanly isolate it; test the whole campaign instead.
Designing a clean test
Most failed tests fail at design, not analysis. Before launch:
- Write one question and one decision: “If iROAS is above 2.0, we keep Meta prospecting at current spend; below 1.2, we cut it by half; in between, we hold and retest.”
- Pick one primary metric measured outside the ad platform: orders, revenue, new customers or qualified pipeline.
- Check that pre-period sales in treatment and holdout regions moved together for at least several months.
- Estimate the smallest effect the test can detect. If it’s bigger than the effect you expect, make the holdout bigger or the test longer.
- Run long enough to cover your purchase cycle, typically three to six weeks for ecommerce, plus a one- to two-week cooldown for delayed conversions. With long B2B sales cycles, read qualified pipeline instead of waiting for closed revenue.
- Freeze everything else in the test regions: no new promotions, price changes or big creative launches.
- Log the start date, end date, spend and every exception.
The decision rule matters most. Written after the results, it bends to whatever the numbers say.
Reading results honestly
Turn the result into three numbers: incremental conversions, incremental cost per acquisition (iCPA) and incremental ROAS (iROAS).
Worked example with made-up round numbers: a DTC brand runs a four-week geo holdout on Meta.
- Treatment regions’ actual revenue: $500,000
- Predicted treatment revenue without Meta, based on the holdout regions: $440,000
- Incremental revenue: $60,000
- Meta spend in treatment regions: $30,000
- iROAS: 2.0
- Meta-reported revenue in the same regions: $120,000, a reported ROAS of 4.0
Incremental revenue divided by reported revenue gives a calibration factor of 0.5. Until the next test, treat every Meta-reported dollar as roughly 50 cents of real revenue.
Then check the result against these traps:
- Ranges, not points. Report the confidence interval. If iROAS falls between 0.8 and 3.2 and your breakeven is 1.5, the test is inconclusive, not a win.
- No detectable lift isn’t zero lift. It means any effect was smaller than the test could reliably detect.
- Stopping early. Checking daily and ending the test when it looks good inflates results. Run the planned duration.
- Spillover. People travel, national campaigns reach holdout regions and marketplace sales may rise elsewhere. Note what the test can’t see.
- One test isn’t permanent truth. Creative, audiences and competition change, so treat a result as current for six to twelve months at most.
Turning results into budget decisions
A test earns its cost only when it changes spend.
| Result | Decision |
|---|---|
| iROAS clearly above target | Scale in steps of 20–30%, then retest at the higher spend |
| Near breakeven | Hold spend; fix creative, audience or offer; retest |
| Clearly below breakeven | Cut to a floor, move budget to the best alternative, watch blended revenue |
| Inconclusive | Don’t treat it as a win or a loss; rerun bigger or longer |
Two rules keep these decisions honest. First, the test measures average incrementality at current spend, and the next dollar usually does worse, so retest after scaling. Second, apply the calibration factor to platform reporting so daily optimization reflects it; a platform showing 4.0 ROAS at a 0.5 factor is really delivering about 2.0.
When I review ad accounts, the pattern I see most is money moving toward whatever reports the highest ROAS, which is usually retargeting and brand. Calibrated numbers often reverse that. Building the measurement this depends on (clean conversion data, regional sales reporting, calibrated dashboards) is the core of my marketing attribution work.
Building a testing calendar
One-off tests go stale. A calendar keeps calibration current without holding back spend all year.
| Quarter | Test | Why then |
|---|---|---|
| Q1 | Geo holdout on the largest paid channel | Normal demand, biggest budget question first |
| Q2 | Brand search on/off test | Cheap to run, often frees budget quickly |
| Q3 | Lift study or holdout on retargeting or PMax | Calibrate before peak-season budgets are set |
| Q4 | No new holdouts; apply calibrations to peak budgets | Holdouts are most expensive in peak weeks |
Adjust the quarters to your seasonality. Run one major test at a time in the same regions so results don’t contaminate each other. Retest a channel after big changes in spend, creative strategy or campaign type, and at least once a year regardless.
Keep a simple test log with the hypothesis, dates, regions, spend, result range, calibration factor and the decision made. After a year, it shows which channels have earned their budget and how far to discount each platform’s reporting.
Get it built
If you want to know which of your channels are actually driving sales, I can design the tests, run the analysis and turn the results into a budget plan. Most engagements start with the $1,500 fixed-price Growth Audit, credited if we keep working together. See pricing or get in touch.