Data POEM

Incrementality testing: What geo and holdout lift tests measure, and where they fall short

What geo lift and holdout incrementality tests actually measure, where each falls short, and where always-on causal measurement fills the gap.

According to the IAB's State of Data 2026 report, 75% of US buy-side leaders say their core measurement methods, including attribution, incrementality testing, and MMM, are underperforming on speed, accuracy, or trust.

These are the methods meant to answer “how much of a business result did marketing actually cause after accounting for what would have happened anyway?”

That distinction matters because a campaign can coincide with higher sales without causing them. Reporting the sales that occurred tells you what happened; incrementality testing estimates how much marketing caused.

But the term covers different experimental designs. A geo lift test and a user-level holdout can both estimate causal lift while seeing very different parts of the system.

According to eMarketer and TransUnion, 52% of US brand and agency marketers now run incrementality tests, up from niche adoption two years earlier. That makes the design choice matter more when a test result becomes a budget decision.

Incrementality testing at a glance

  • Incrementality testing estimates what an investment caused by comparing an observed outcome with a credible counterfactual: what would have happened without the intervention.
  • Geo lift tests split geographic markets into treatment and control groups, making them useful when user-level randomization is impractical or impossible.
  • Holdout tests split eligible users, households, or other addressable units into treatment and control groups, giving a more direct randomized comparison when identity can be maintained.
  • Geo tests are limited by market comparability, spillover, aggregation, and the measurement window; holdouts are limited by identity, walled gardens, cross-platform exposure, and outcome visibility.
  • Periodic experiments answer bounded questions at a point in time. They do not create a continuous view of cross-channel, portfolio, and longer-lag causal effects.

What is an incrementality test?

An incrementality test is a controlled experiment designed to estimate the causal effect of marketing. It compares a treatment group exposed to an intervention with a control group that is not, then measures the difference in outcomes. The aim is to estimate what happened because of the investment, not simply what happened after it.

The logic is counterfactual. You observe what happened with the campaign, then construct the best available estimate of what would have happened without it. The gap between those outcomes is the incremental effect.

That makes incrementality different from attribution. Attribution assigns credit across observed touchpoints; incrementality asks whether the outcome would have happened without marketing. 

The experiment can be built at different levels: some split markets, others randomize users, households, stores, or another eligible unit. The design determines the question you can answer, the bias you need to manage, and how far you can generalize the result.

How do geo lift tests work?

A geo lift test measures incremental impact by changing marketing activity in selected geographic markets and comparing the resulting outcome with control markets. Treatment and control geographies are chosen or matched so their pre-test behavior is sufficiently similar to support a credible estimate of what would have happened without the change.

In practice, a brand might increase spend in one set of markets while holding the comparison set steady. The analysis then estimates whether sales, conversions, store visits, or another business outcome moved beyond the counterfactual expected from the control markets.

Geo tests are especially useful when individual exposure cannot be randomized cleanly. Television, out-of-home, radio, some retail activity, and broad media campaigns often operate at a market level instead of a known-user level.

The trade-off is aggregation. A geo test can still produce a strong causal read on a market-level intervention, but its validity depends on comparable markets, limited contamination and enough signal to separate lift from normal geographic variation.

How do holdout tests work?

A holdout test measures incrementality by withholding an intervention from a randomly selected control group while an otherwise eligible treatment group receives it. Because assignment is randomized, differences in outcomes can be attributed more confidently to the intervention when the treatment and control groups remain comparable.

Digital advertising platforms commonly make this design practical because users, households, devices, or accounts can be assigned before exposure. The outcome might be a conversion, purchase, subscription, revenue event, or another measurable business result.

Randomization is the strength of a holdout. If both groups would have behaved similarly without the campaign, the observed difference after treatment is a direct estimate of incremental lift.

But the experimental unit and study power both matter. A holdout also works cleanly only when the same person does not leak across devices, accounts, platforms or channels in ways that break assignment.

What can't geo lift tests measure?

We use a simple diagnostic called the Measurement Boundary Map before interpreting any lift study. Its inputs are the:

  • Test type
  • Campaign or channel under test
  • Measurement window
  • Geographic or user-level split

It produces a map of what the design can observe and what sits outside the experiment, so you can see which questions the result can support before making a budget decision.

For geo lift, the first boundary is geography itself. Google Ads warns that cross-market exposure reduces the measured difference between treatment and control and therefore reduces reported incrementality. Media can spill across market lines through commuting, travel, broadcast overlap, ecommerce, social sharing or platform delivery, contaminating the control group.

The second boundary is aggregation. A market-level result can tell you whether the intervention lifted the measured outcome across those geographies, but it does not automatically explain which audience, creative, device, or customer segment caused the lift.

Geo tests also need enough comparable markets and enough signal to distinguish the treatment effect from local noise. When several channels change at once, the test may estimate the effect of the bundle without isolating each channel's contribution or interaction.

Finally, the test window sets a time boundary. Effects that emerge after measurement ends are not captured by the experiment.

What can't holdout tests measure?

The Measurement Boundary Map looks different for a holdout because its boundary is defined by identity and data access, not geography. A randomized control group can provide a clean estimate inside that boundary, but it cannot observe outcomes or exposures the experiment cannot identify.

Cross-device behavior is one weakness. If the same customer is assigned to control on one device but receives the campaign on another, the holdout is contaminated. Cross-platform exposure creates the same problem when an excluded user still encounters the message elsewhere.

Walled gardens add a second limit. A platform-native holdout may measure incrementality inside the outcomes that platform can see, while missing purchases, offline behavior, or media interactions elsewhere unless those outcomes can be matched back reliably. The industry is alert to the gap: according to the IAB, advertiser focus on cross-platform measurement climbed to 72% in 2026, up from 64% a year earlier.

Holdouts also estimate an effect for the eligible randomized population under the tested conditions. That result does not automatically generalize to every customer, channel, geography, spend level, or future campaign.

As with geo lift, time matters. Brand demand and other delayed effects can continue after the holdout window closes.

Why aren't periodic tests enough?

Periodic tests answer a causal question with an explicit control. For a defined campaign, audience, geography, spend level, and time period, that is often the strongest evidence available.

The problem starts when a bounded experiment is treated as a permanent truth. A lift estimate from one quarter does not automatically hold after the media mix changes, competitors increase spend, pricing moves, distribution shifts, or the same audience sees a different creative sequence.

And the stakes are rising. According to Gartner's 2026 CMO Spend Survey, 62% of CMOs say that failing to meet growth expectations would trigger cuts to their marketing budget, which makes a stale lift estimate an expensive thing to rely on.

Enterprise marketing also creates more causal questions than you can realistically test one at a time. You cannot continuously withhold every channel, market, audience and campaign just to keep every estimate current.

That is the gap between experiments and a measurement system. Experiments are periodic observations. Decision-making is continuous.

For a broader view of that problem, our guide to how enterprise CMOs measure marketing ROI explains why channel-level answers can still conflict when the wider business is measured through separate models.

How does Enterprise Decision AI fill the gap?

Enterprise Decision AI moves incrementality from isolated validation toward continuous causal measurement. Instead of treating each campaign as a separate event, a causal model can estimate how marketing operates alongside pricing, promotions, distribution, competitive activity, seasonality, and other growth drivers as those conditions change.

At DATA POEM, FOUNT is the Large Causal Architecture underneath POEM365. POEM365 uses that architecture to model causal relationships across marketing and the surrounding business context, rather than limiting measurement to the boundaries of a single platform or experiment.

That does not make geo lift or holdout tests obsolete. Experiments remain useful validation points because they create controlled evidence. The difference is that the evidence can inform an ongoing causal model instead of becoming the only moment when incrementality is measured.

That is the role of causal AI: estimate cause and effect from the wider system, update the view as conditions change, and support decisions that cannot wait for the next test cycle.

Frequently asked questions

What is an incrementality test?

An incrementality test is an experiment that estimates how much additional outcome a marketing intervention caused. The control can be built from users, households, stores, or geographic markets, depending on how the campaign can be isolated. The important requirement is a credible counterfactual that represents what would have happened without the intervention.

What is the difference between attribution and incrementality?

The difference between attribution and incrementality is the question each method answers. Attribution distributes credit for an observed conversion across marketing touchpoints, while incrementality estimates whether marketing caused an outcome that would not otherwise have occurred. Attribution is useful for describing customer paths and platform performance; incrementality is more useful when the decision is whether additional spend created additional business value.

How do you calculate incrementality?

You calculate incrementality by comparing the observed treatment outcome with the estimated counterfactual outcome. A simple lift calculation is: (treatment outcome - counterfactual outcome) / counterfactual outcome × 100. For investment decisions, you can also calculate incremental revenue or profit and divide it by the spend required to create that incremental result.

What is the purpose of incrementality testing in Google Ads?

The purpose of incrementality testing in Google Ads is to estimate whether advertising caused additional conversions or business outcomes beyond what would have happened without the ads. A controlled treatment-and-holdout design can reveal when attributed conversions are genuinely incremental, helping teams judge budget, bidding, audience, and campaign decisions on causal impact rather than credited conversions alone.

How do you measure the effectiveness of ads?

You measure the effectiveness of ads by starting with a business outcome and a credible counterfactual, not exposure or clicks alone. Use a geo or holdout experiment where possible to estimate incremental lift, then evaluate incremental revenue, profit, or another decision-relevant KPI. For ongoing planning, combine experimental evidence with continuous causal measurement across the wider marketing mix.

Move measurement from event to state

A strong incrementality program does not ask one experiment to answer more than its design allows. Before commissioning a test, use the Measurement Boundary Map to define the intervention, control, outcome, time horizon, and blind spots. That makes the result easier to interpret and harder to overextend.

For a narrow campaign decision, a geo lift or holdout test may be exactly the right method. For an enterprise budget decision, the question is wider: how does that intervention interact with every other channel and business driver, and how does the answer change as conditions move?

That requires measurement to become a state, not an event. Experiments can anchor the causal evidence, while a continuously updated model carries that evidence into the decisions between tests.

DATA POEM's unified marketing measurement is built around that shift. It measures incremental business outcomes across the marketing mix, accounts for interactions with the wider business, and keeps the causal view current enough to guide the next allocation decision. Book a call to see what it surfaces in your own mix.


See what Data Poem can do for you.

Let's talk about how we can help you grow your business.