Skip to content
Can Elmas

Attribution · 8 min read

Incrementality Testing: How to Prove Your Ads Actually Drive Sales

TL;DR

Incrementality testing compares sales with ads against a holdout without them, so you measure what the ads caused, not what they claimed. On mid-size budgets, use geo holdouts for big channels, platform lift studies where you qualify, and on/off tests for brand search and retargeting. Set the decision rule first, read ranges rather than single numbers, then reset budgets.

· Published · Updated

Incrementality testing answers the question attribution can’t: would these sales have happened without the ads? You switch ads off for a controlled group of people or regions, compare them with a group that still sees ads, and the difference is what the spend actually caused. On a mid-size budget, geo holdouts, platform lift studies and simple on/off tests are enough to catch the channels that are claiming credit for sales they didn’t create.

Attribution vs incrementality

Attribution divides credit for conversions that already happened among the touchpoints before them. Incrementality asks a counterfactual question: how many of those conversions would have happened anyway?

AttributionIncrementality
QuestionWhich touchpoints were present before the sale?How many sales did the ads cause?
MethodRules or models applied to tracked journeysControlled experiment with a holdout group
FrequencyDaily, always onA few tests a year
Typical biasOver-credits channels close to purchaseNoisy if the test is too small
Best useDay-to-day optimizationSetting budgets and calibrating attribution

The two work together. Attribution runs the daily optimization; incrementality tells you how much to discount what it reports. If you’re still deciding between modeled approaches, MMM vs multi-touch attribution covers that choice, and experiments are what keep either model honest.

When you need an incrementality test

Run a test when a real budget decision depends on the answer:

  • Platform-reported conversions add up to more than your actual orders or pipeline. Several channels are claiming the same customers.
  • A large share of spend sits in retargeting, brand search, PMax or existing-customer audiences. These reach people already on their way to buying, so they look efficient almost by definition.
  • You’re about to scale a channel meaningfully, or cut one, and want to know what happens to total sales.
  • You’re adding a channel that tracks poorly, such as YouTube, connected TV or podcasts.
  • Blended numbers disagree with platform numbers. Platform ROAS rises while total revenue stays flat.

Don’t test when your volume is too small for the test to see the effect you expect. For example, if weekly revenue swings 15% on its own and a channel might add 3%, a small, short test will likely end inconclusive and you’ll have paid for the holdout for nothing.

Test types: geo holdouts, lift studies and on/off tests

MethodHow it worksBest forMain weakness
Geo holdoutTurn ads off in matched regions, compare with regions still runningLarge channels, cross-channel questions, anything measured in your own sales dataNeeds enough regions and volume; weeks of lost exposure in holdout areas
Platform lift studyThe platform randomly withholds ads from a control group of usersMeta campaigns, YouTube and videoCounts only what the platform can track; eligibility varies
On/off testPause a campaign for a period, compare with before and afterBrand search, retargeting, small campaignsSeasonality and other changes muddy the comparison

Geo holdouts

Split your market into regions (US states or media markets, UK regions, cities elsewhere) and match treatment and holdout regions on how their sales moved together in the past. Turn the channel off in the holdout regions, then compare actual sales in the treatment regions with what the holdout regions predict they would have sold without the ads.

Measure the outcome in your own systems: Shopify orders by shipping region, CRM pipeline by company location, total revenue. Open-source tools help: Meta’s GeoLift handles market selection and analysis, and Google’s CausalImpact estimates the counterfactual from comparison regions.

Platform lift studies

Meta’s Conversion Lift and Google’s Conversion Lift (mainly for YouTube and video) randomly split people into an exposed group and a control group that can’t see your ads, which removes attribution’s biggest bias. Availability and minimum budgets vary, so check Meta’s Experiments tool, Google Ads’ measurement tools or your Google rep. The catch: the platform measures its own conversions, so validate big results against your own data.

On/off and brand search tests

The simplest version pauses a campaign for two to four weeks and watches total results, not platform results. For brand search, track total brand clicks (paid clicks plus organic clicks from Search Console on brand queries) and total conversions. If organic picks up most of what paid used to take, brand spend is largely not incremental. Check first whether competitors bid on your name; if they do, pausing hands them the top spot, and that should be part of the result. Where you can, pause brand search in holdout regions only, which turns a before-and-after comparison into a controlled one.

Testing Performance Max and retargeting

PMax and retargeting deserve a test because their reported ROAS is built on people who were already close to buying. For PMax, a geo holdout measured in your own sales data is the most independent read. Google Ads’ built-in experiments can compare running with and without PMax, but they grade the result with Google’s own conversion tracking. Apply brand exclusions before testing, or you’ll mostly measure brand demand; the Google Ads audit checklist covers that setup.

For retargeting, use a lift study on the retargeting campaign or pause it in holdout regions. If retargeting runs inside a broad automated campaign such as Meta’s Advantage+ sales campaigns (formerly Advantage+ Shopping) or PMax, you can’t cleanly isolate it; test the whole campaign instead.

Designing a clean test

Most failed tests fail at design, not analysis. Before launch:

  • Write one question and one decision: “If iROAS is above 2.0, we keep Meta prospecting at current spend; below 1.2, we cut it by half; in between, we hold and retest.”
  • Pick one primary metric measured outside the ad platform: orders, revenue, new customers or qualified pipeline.
  • Check that pre-period sales in treatment and holdout regions moved together for at least several months.
  • Estimate the smallest effect the test can detect. If it’s bigger than the effect you expect, make the holdout bigger or the test longer.
  • Run long enough to cover your purchase cycle, typically three to six weeks for ecommerce, plus a one- to two-week cooldown for delayed conversions. With long B2B sales cycles, read qualified pipeline instead of waiting for closed revenue.
  • Freeze everything else in the test regions: no new promotions, price changes or big creative launches.
  • Log the start date, end date, spend and every exception.

The decision rule matters most. Written after the results, it bends to whatever the numbers say.

Reading results honestly

Turn the result into three numbers: incremental conversions, incremental cost per acquisition (iCPA) and incremental ROAS (iROAS).

Worked example with made-up round numbers: a DTC brand runs a four-week geo holdout on Meta.

  • Treatment regions’ actual revenue: $500,000
  • Predicted treatment revenue without Meta, based on the holdout regions: $440,000
  • Incremental revenue: $60,000
  • Meta spend in treatment regions: $30,000
  • iROAS: 2.0
  • Meta-reported revenue in the same regions: $120,000, a reported ROAS of 4.0

Incremental revenue divided by reported revenue gives a calibration factor of 0.5. Until the next test, treat every Meta-reported dollar as roughly 50 cents of real revenue.

Then check the result against these traps:

  • Ranges, not points. Report the confidence interval. If iROAS falls between 0.8 and 3.2 and your breakeven is 1.5, the test is inconclusive, not a win.
  • No detectable lift isn’t zero lift. It means any effect was smaller than the test could reliably detect.
  • Stopping early. Checking daily and ending the test when it looks good inflates results. Run the planned duration.
  • Spillover. People travel, national campaigns reach holdout regions and marketplace sales may rise elsewhere. Note what the test can’t see.
  • One test isn’t permanent truth. Creative, audiences and competition change, so treat a result as current for six to twelve months at most.

Turning results into budget decisions

A test earns its cost only when it changes spend.

ResultDecision
iROAS clearly above targetScale in steps of 20–30%, then retest at the higher spend
Near breakevenHold spend; fix creative, audience or offer; retest
Clearly below breakevenCut to a floor, move budget to the best alternative, watch blended revenue
InconclusiveDon’t treat it as a win or a loss; rerun bigger or longer

Two rules keep these decisions honest. First, the test measures average incrementality at current spend, and the next dollar usually does worse, so retest after scaling. Second, apply the calibration factor to platform reporting so daily optimization reflects it; a platform showing 4.0 ROAS at a 0.5 factor is really delivering about 2.0.

When I review ad accounts, the pattern I see most is money moving toward whatever reports the highest ROAS, which is usually retargeting and brand. Calibrated numbers often reverse that. Building the measurement this depends on (clean conversion data, regional sales reporting, calibrated dashboards) is the core of my marketing attribution work.

Building a testing calendar

One-off tests go stale. A calendar keeps calibration current without holding back spend all year.

QuarterTestWhy then
Q1Geo holdout on the largest paid channelNormal demand, biggest budget question first
Q2Brand search on/off testCheap to run, often frees budget quickly
Q3Lift study or holdout on retargeting or PMaxCalibrate before peak-season budgets are set
Q4No new holdouts; apply calibrations to peak budgetsHoldouts are most expensive in peak weeks

Adjust the quarters to your seasonality. Run one major test at a time in the same regions so results don’t contaminate each other. Retest a channel after big changes in spend, creative strategy or campaign type, and at least once a year regardless.

Keep a simple test log with the hypothesis, dates, regions, spend, result range, calibration factor and the decision made. After a year, it shows which channels have earned their budget and how far to discount each platform’s reporting.

Get it built

If you want to know which of your channels are actually driving sales, I can design the tests, run the analysis and turn the results into a budget plan. Most engagements start with the $1,500 fixed-price Growth Audit, credited if we keep working together. See pricing or get in touch.

FAQ

Frequently Asked Questions

How much ad spend do you need before incrementality testing is worth it?

Enough that the channel's expected effect stands out from the normal week-to-week swing in your sales. As a working rule, I rarely run a geo holdout on a channel spending less than the low five figures a month, because smaller effects disappear into noise. Brand search and retargeting on/off tests can still be worth running below that.

Can I trust a lift study the ad platform runs itself?

Mostly, because the test and control groups are randomized, which removes the biggest bias in attribution. The limits are that it only counts conversions the platform can see and measures one channel in isolation, so confirm big decisions against your own sales data or a geo holdout.

What is iROAS?

Incremental return on ad spend: the extra revenue the test shows the ads caused, divided by what you spent on them during the test. It is usually lower than platform-reported ROAS, and it is the number to compare against your breakeven.

Should I run incrementality tests during peak season?

Avoid it for your first tests. Holding out regions during your biggest weeks is expensive, and unusual demand makes the pre-period a poor guide to what would have happened, so schedule tests in normal months and apply the results to peak.

Work with me

Let’s find your biggest growth lever

Tell me about your growth challenge. I’ll tell you honestly if I can help — and if I can’t, who can.

  • ✓ No obligation
  • ✓ No sales script
  • ✓ Honest feedback
  • ✓ Clear next steps