Measurement Methodology
Incrementality Testing
Attribution tells you who gets credit. Incrementality tells you what actually worked. We measure both, because the difference between 'attributed' and 'incremental' is where most wasted spend hides.
The uncomfortable question every brand should ask:
"How much of what Google claims to have driven would have happened anyway?"
The Attribution Problem
Google Ads claims credit for conversions using attribution models - such as data-driven or last click. These models answer: "Which ad touchpoint should get credit for this sale?"
But they never answer the more important question: "Would this sale have happened without the ad?"
Brand searches are the clearest example. Some people who search your brand name would have bought through an organic listing anyway; others might have gone to a competitor bidding on your name. The platform credits the brand ad either way, so attributed brand ROAS cannot tell you which share was caused by the ad. Only a controlled test can.
This isn't fraud. It's how attribution works. But if you're making budget decisions based on attributed performance alone, you're almost certainly overspending in some areas and underspending in others.
Illustrative example: invented numbers, not a client result
What the platform reported
4.2x
Blended ROAS
£340k
Attributed Revenue
After incrementality testing
1.8x
Incremental ROAS
£145k
Incremental Revenue
Synthetic figures for illustration only. In this invented scenario a controlled test estimated that 57% of attributed revenue would have occurred without ads; a real result would carry an uncertainty range.
How We Test Incrementality
Four methodologies, each suited to different account structures and commercial questions.
Geographic Holdout
Pause all paid activity in a matched region. Compare sales trends against an active region with similar demographics and seasonality.
Best For
Brands with national coverage and consistent regional demand
Duration
Set from your own pre-period volatility and conversion lag
Measures
The difference in total sales between treatment and control regions, against a pre-registered outcome
Campaign-Level Holdout
Pause a specific campaign type (e.g., Brand, PMax, Generic) while keeping everything else running. Isolate the true lift of that campaign.
Best For
Measuring the real value of brand campaigns or Performance Max
Duration
Set from your own pre-period volatility and conversion lag
Measures
Whether pausing a campaign reduces total sales or just shifts attribution
Budget Scaling Test
Increase or decrease spend by a fixed percentage in a controlled period. Measure whether the marginal spend produces marginal profit.
Best For
Testing whether 'scaling' actually improves outcomes
Duration
Set from your own pre-period volatility and conversion lag
Measures
Marginal POAS - the profit generated by the last £1,000 of spend
Channel Isolation
Change one channel in a randomised or matched design while holding others steady, to estimate that channel's contribution.
Best For
Multi-channel brands unsure which platform drives genuine demand
Duration
Set from your own pre-period volatility and conversion lag
Measures
Estimated channel contribution, with an uncertainty range, compared with platform-claimed attribution
The 5-Phase Process
Phase lengths depend on your conversion volume, conversion lag and the effect size the test needs to detect. They are agreed in the pre-registered design, not fixed in advance.
Length set in the test design
Commercial Baseline
- Establish true P&L baseline (not platform P&L)
- Map current attribution claims vs. bank-reconciled revenue
- Identify the gap between reported and actual
- Define the commercial question the test will answer
Length set in the test design
Test Design
- Select test methodology based on account structure
- Define control and test groups with statistical rigour
- Set minimum detectable effect thresholds
- Align test duration with business cycles (avoid peak/promo periods)
Length set in the test design
Execution & Monitoring
- Implement test with clean controls
- Monitor for contamination (e.g., organic changes, competitor activity)
- Track both platform metrics AND commercial metrics simultaneously
- Weekly check-ins to ensure test integrity
Length set in the test design
Commercial Analysis
- Reconcile platform data against actual revenue and margin
- Calculate true incremental contribution (not platform-attributed)
- Build diminishing returns curve for marginal spend
- Produce CFO-ready summary with P&L impact
Length set in the test design
Reallocation
- Redirect budget from non-incremental to genuinely incremental activity
- Set new POAS targets based on proven incrementality
- Agree the triggers for re-testing (material changes in spend, mix, competition or range)
- Document findings for board/finance review
Reference
A branded-search test protocol you can use
Brand Search is where attributed and incremental results diverge most. This is the sequence we use to test it. Google's own custom experiments and Conversion Lift documentation covers the platform tools.
Three kinds of evidence, not one
Attributed conversions
What the platform credits to ads. Useful for optimisation, but it says nothing about what would have happened without the ad.
Pre/post observation
Sales before and after a change. Seasonality, promotions and demand shifts are mixed into the difference.
Causal lift
The difference between treatment and a comparable control over the same period, with uncertainty stated. A causal reading needs valid assignment (randomised or pre-matched units), a genuine control and stated assumptions (no spillover between units, no other change hitting one group); a contemporaneous difference alone is not enough.
1. Hypothesis
Write one falsifiable sentence before anything changes, e.g. "Pausing brand Search in treatment regions loses less contribution before advertising than the media cost it saves."
2. Eligible test units
Choose units you can switch independently and measure separately: usually geographic regions (Google Ads location targeting, orders matched by delivery postcode). Exclude units with launches, store openings or unusual promotions planned.
3. Treatment and control
Assign units before looking at results, ideally at random within matched pairs. Treatment: brand Search paused or reduced. Control: unchanged. Record the assignment list.
4. Contamination and overlap
List every other campaign that can serve on brand queries: Performance Max, Shopping, dynamic search ads, competitors, affiliates bidding on brand. Decide how each is handled in treatment, or the test measures a mix.
5. Primary business outcome
Pick one before the test: total orders or contribution after advertising across all channels in each unit. Not platform-attributed conversions, which can fall or shift attribution even when total business sales do not change.
6. Power and duration
Pre-specify the smallest effect worth acting on and use your own pre-period volatility to estimate the duration needed to detect it. Cover at least one full purchase cycle plus conversion lag. There is no universal duration.
7. Stopping rules
Agree in advance what ends the test early (for example, a loss beyond an agreed contribution threshold, a stock-out or a site outage) and that nobody stops it because an early reading looks good.
8. Uncertainty and limits
Report the estimate with a range, not a single number. If the range includes zero, the result is inconclusive, not a win. Results hold for that period, those regions and that competitive landscape.
Synthetic worked example: invented numbers, not a client or research result
Pre-period (four weeks): treatment regions 1,000 orders; control regions 1,200 orders.
Test period (four weeks, brand Search paused in treatment only): control rose to 1,236 orders, a 3% change (1,236 ÷ 1,200 = 1.03).
Expected treatment orders without the pause: 1,000 × 1.03 = 1,030. Observed: 950. Estimated orders caused by brand Search: 1,030 − 950 = 80.
At £30 contribution before advertising per order, that is 80 × £30 = £2,400 of contribution lost. Brand spend saved in treatment: £3,000. Net effect of pausing: £3,000 − £2,400 = £600 better off, with advertising subtracted once. This adjusted pre/post arithmetic illustrates the calculation; it is not itself proof, which depends on the assignment, control and assumptions above.
Before acting, check the uncertainty range. If the plausible range for lost orders runs from, say, 20 to 140, the net effect runs from +£2,400 to −£1,200, and the honest verdict is inconclusive. Do not extend the test after seeing that result unless a sequential design was pre-specified; otherwise accept the decision risk explicitly or write a new protocol for a fresh test.
Contribution follows our POAS convention.
Pre-launch worksheet
Why Incentives Matter
Incrementality testing is uncomfortable. It can show that some attributed conversions would have happened without the ad. Brand campaigns with high reported ROAS can include many conversions that would have happened anyway, and Performance Max can take credit for demand other campaigns would have captured. How much varies by account; only a test tells you.
For an agency whose fee is justified by attributed ROAS, this is a structural threat. Proving that part of the spend isn't incremental means recommending a budget cut, and under a percentage-of-spend model, a corresponding fee reduction.
We charge fixed fees specifically so this incentive conflict doesn't exist. When we find waste, we recommend cutting it. When we find incrementality, we recommend scaling it. The commercial alignment is clean.