Incrementality testing ads measures whether your advertising actually caused a sale, rather than simply claiming credit for one that would have happened anyway. Search Engine Land notes that a properly run test often finds true lift between 0% and 25%, a figure that regularly surprises advertisers used to trusting platform-reported ROAS at face value. Oxedent runs these tests for established ecommerce brands to protect budgets from unnecessary cuts and to expose spend that isn’t earning its place.
Incrementality testing earns its budget when you need to:
- Defend ad spend during a budget review, with causal evidence rather than a dashboard screenshot.
- Validate retail media and marketplace ads, where platforms have every incentive to over-credit themselves.
- Stress-test high-spend channels like Performance Max or branded search, where organic demand often gets mistaken for ad-driven demand.
If none of those apply to you yet, attribution and MMM will probably serve you fine for now.
Key Takeaways
Incrementality testing proves causal ad impact through a controlled counterfactual, while attribution and MMM answer related but different measurement questions.
| Point | Details |
|---|---|
| Definition | Incrementality testing compares an exposed group against a withheld control group to isolate true ad-driven lift. |
| Realistic lift range | Genuine incremental lift often falls between 0% and 25%, well below what attribution reports. |
| Best starting method | Geo holdouts across matched regions run for three to four weeks suit most ecommerce brands. |
| Decision rule first | Agree the exact result that triggers scaling, pausing, or re-testing before launching the test. |
| Where Oxedent fits | Oxedent designs and runs incrementality tests alongside PPC management, feeding results directly into budget reallocation. |
Table of Contents
- What incrementality testing measures (and how it differs from attribution and MMM)
- Why measure incrementality: the business cases that justify the effort
- How incrementality tests work: holdouts, geo experiments, and synthetic controls
- How to calculate incrementality: the formula and a worked example
- Which channels and campaigns deserve incrementality testing first
- Designing a credible test: sample size, duration, holdout share, and budget
- Interpreting your results: combining incrementality with attribution and MMM
- Your first incrementality test: a step-by-step checklist
- What agencies see clients get wrong (and when to bring in help)
- How Oxedent turns your test results into a bigger ad budget
- Frequently asked questions about incrementality testing ads
- Sources
What incrementality testing measures (and how it differs from attribution and MMM)
Incrementality testing answers one question: what happened because of the ad, versus what would have happened anyway? That “what would have happened anyway” is the counterfactual, and building a credible one is the entire discipline. You do this by comparing a treatment group exposed to your advertising against a control group deliberately withheld from it. The gap between the two groups is your lift, and lift is the only honest measure of causation you have.
Attribution, incrementality, and marketing mix modelling (MMM) each answer a different question:
- Attribution answers “which touchpoint gets credit for this conversion?” It’s fast, granular, and useful daily, but it can’t tell you what would have happened without the ad.
- Incrementality answers “did this ad cause a net-new conversion?” It’s the only method built on a genuine counterfactual.
- MMM answers “how should I allocate budget across channels over time?” It works at a macro level using historical spend and revenue data.
Use attribution for day-to-day bid and budget tweaks. Use MMM when you’re planning quarterly allocation across five or more channels. Use incrementality when a decision is expensive enough that you need proof, not a proxy.
Why measure incrementality: the business cases that justify the effort
Running an incrementality test costs time, a temporary revenue dip, and analytical rigour. It earns that cost back in three ways.
First, budget defence. When finance asks why paid social deserves another £20,000 a month, “our ROAS is 4x” is a weaker argument than “we proved a 22% incremental lift versus a held-out control group.”
Second, finding wasted spend. Every ecommerce account has some campaigns capturing demand that was coming anyway, brand search chief among them. Incrementality testing quantifies exactly how much of that spend is genuinely working.
Third, recalibrating ROAS expectations. A campaign showing 6x attributed ROAS might deliver an incremental ROAS closer to 2x once you strip out the customers who’d have bought regardless. That’s not a failure. It’s the number you should actually be optimising towards, and it changes which campaigns deserve priority.
How incrementality tests work: holdouts, geo experiments, and synthetic controls
Three test designs cover most ecommerce use cases, and each suits a different level of control and budget.
-
Randomised audience holdouts. You split your audience into exposed and withheld groups, typically through a platform’s native tool. Google’s Conversion Lift supports both user-based and geo-based designs, and this route gives you the cleanest causal read when you have enough scale to hit statistical significance.
-
Geo or matched market experiments. You pause advertising across a set of matched regions while running as normal elsewhere. This is often the most practical starting point for D2C and ecommerce brands, since it sidesteps user-level tracking limitations and works even under privacy restrictions that block audience holdouts.
-
Synthetic controls. When you can’t withhold ads from any real audience or region, you model a statistical “control” from historical and comparable data instead. It’s the least clean method, useful mainly when operational constraints rule out a true holdout.
Trade-offs run through all three: speed versus confidence, scale versus feasibility, and the risk of contamination if a “held out” group still sees your ads through another channel or a competitor’s overlapping audience.
Pro Tip: Start with a geo holdout before attempting an audience-level test. It’s cheaper to set up, easier to explain to stakeholders, and avoids most of the tracking headaches that plague user-based designs.
How to calculate incrementality: the formula and a worked example
The core formula is straightforward:
Incremental conversions = Conversions (exposed group) − Conversions (control group)
From there, incremental ROAS = Incremental revenue ÷ Ad spend.
Here’s a worked example. Say you run a four-week geo holdout across matched regions. The exposed regions generate 1,000 conversions at an average order value of £60, totalling £60,000 in revenue. The control regions, with no ads running, generate 850 conversions, or £51,000 in revenue.
- Incremental conversions: 1,000 − 850 = 150
- Incremental revenue: £60,000 − £51,000 = £9,000
- If ad spend across the exposed regions was £3,000, incremental ROAS = £9,000 ÷ £3,000 = 3x
Compare that to a blended ROAS calculated from all 1,000 conversions, which would show a much higher (and misleading) number. That gap between attributed and incremental performance is explained in more depth in Oxedent’s guide to blended ROAS.
Watch for two traps: conversions that land after your attribution window closes, and double-counting customers who convert across both test arms through cross-device behaviour.
Which channels and campaigns deserve incrementality testing first
Not every campaign justifies the effort. Prioritise by where the risk of overstated performance is highest.
- Performance Max and other AI-driven, multi-channel campaigns need testing precisely because automation obscures which impression or click actually drove the sale. Oxedent’s notes on Performance Max strategy are worth reading alongside any test plan here.
- Paid social and video often show delayed lift, with audience overlap between platforms muddying a clean read, so build in a longer measurement window than you would for search.
- Branded search and retargeting typically show the lowest incremental lift of any channel, since you’re largely reaching people already planning to buy.
Designing a credible test: sample size, duration, holdout share, and budget
Get the mechanics right before you get the maths right.
- Set the decision rule before launch. Decide what result will trigger action. A useful format: increase budget if incremental ROAS exceeds 3x with a confidence interval that excludes break-even. Agreeing this upfront stops post-hoc rationalising of an inconvenient result.
- Choose your holdout share and duration. Most ecommerce tests hold out a minority share of the audience or geography, and industry guidance recommends running for multiple weeks to ensure reliable results to properly capture delayed conversions and reach statistical significance.
- Run a pre-period where practical. A short baseline before the test starts helps confirm your treatment and control groups behave similarly absent any intervention, strengthening confidence in the result.
- Prepare stakeholders for the revenue dip. Holding out spend means a temporary, controlled drop in sales from the withheld group. Flag this to finance and leadership before the test starts, not after someone notices the numbers.
Minimum detectable lift matters more than most marketers realise: a test too small to detect a 10% lift will report “no significant difference” even when a real lift exists, which is a design failure, not a null result.
Pro Tip: If your budget can’t support a full holdout test, Bayesian methods have lowered the minimum spend required for some platform-run experiments, making smaller-scale tests viable where they weren’t before.
Interpreting your results: combining incrementality with attribution and MMM
Expect incremental ROAS to come in lower than your attributed ROAS. That’s not the test failing. It’s the test doing its job by stripping out conversions that were never caused by the ad in the first place.
Use your pre-agreed decision rule to act:
- Scale the channel if incremental ROAS clears your threshold with a confidence interval that doesn’t touch break-even.
- Pause or reduce if lift is negligible or the interval spans zero.
- Re-test if results sit in an ambiguous middle ground, ideally after fixing tracking issues, which is covered in Oxedent’s guide to fixing conversion tracking problems.
The strongest measurement programmes don’t pick one method. They use attribution for daily optimisation, MMM for cross-channel allocation, and incrementality to validate the causal story behind both. Document your assumptions, report confidence intervals rather than a single point estimate, and always end your test write-up with a recommended next action, not just a number.
Your first incrementality test: a step-by-step checklist
Move from theory to execution with this sequence:
- Pick one high-spend campaign or channel and define your KPI and decision rule.
- Choose your method: geo holdout for most ecommerce brands, audience holdout if you have platform-level scale.
- Confirm tracking is clean before launch, using an account audit checklist if you haven’t reviewed setup recently.
- Run the test for three to four weeks minimum, watching for contamination between groups.
- Calculate lift and incremental ROAS, document every assumption you made along the way.
- Execute the decision: scale, pause, or schedule a re-test.
What agencies see clients get wrong (and when to bring in help)
The most common mistake is running a test for two weeks, seeing an inconclusive number, and abandoning the method entirely rather than fixing the design. The second is skipping the decision rule, so a clean result still triggers an argument about what it means.
Run tests internally once your team has clean tracking and the patience for a multi-week commitment. Bring in specialist support when you need a channel like Performance Max stress-tested without pausing your own optimisation work, which is where an agency’s structured process pays for itself.
How Oxedent turns your test results into a bigger ad budget
Running a credible incrementality test alongside daily campaign management is hard to do properly when you’re also the person managing bids, feeds, and creative refreshes. Oxedent designs and runs incrementality tests for established ecommerce brands as part of its Google Ads, Meta, and Performance Max management, so the result feeds straight back into budget decisions rather than sitting in a report nobody reads.
Oxedent’s eCommerce PPC management service is built around this loop: test, measure real lift, reallocate spend towards what’s genuinely working, and repeat, with no long-term contract locking you in if priorities shift. If you’re carrying a Performance Max budget you’ve never stress-tested, or a retail media line you suspect is capturing demand rather than creating it, request an audit and Oxedent will scope a test design around your actual account before you commit a penny of extra spend.
Frequently asked questions about incrementality testing ads
What is incrementality testing in advertising?
It’s a controlled experiment that compares customers exposed to your ads against a withheld control group, isolating conversions the advertising actually caused rather than ones that would have happened anyway.
How is incrementality testing different from attribution?
Attribution assigns credit across touchpoints in a single customer journey; incrementality measures what wouldn’t have happened without the ad at all, using a genuine control group rather than a modelled credit split.
How long should an incrementality test run?
Most guidance recommends a minimum of three to four weeks, long enough to capture delayed conversions and reach statistical significance without reacting to short-term noise.
Which channels benefit most from incrementality testing?
Retail media, marketplaces, and Performance Max campaigns benefit most, since automation and platform incentives make it hardest to separate genuine lift from captured organic demand.
Can small ecommerce brands afford to run incrementality tests?
Yes. Bayesian testing methods have lowered minimum budget requirements for some platform-run experiments, and geo holdouts remain accessible without large audience scale.
Sources
For deeper technical grounding beyond this guide, these sources cover the method from different angles:
