Skip to main content
    European Search Awards 2026 · Best Small PPC Agency

    Measurement Methodology

    Incrementality Testing

    Attribution tells you who gets credit. Incrementality tells you what actually worked. We measure both, because the difference between 'attributed' and 'incremental' is where most wasted spend hides.

    The uncomfortable question every brand should ask:

    "How much of what Google claims to have driven would have happened anyway?"

    The Attribution Problem

    Google Ads claims credit for conversions using attribution models - such as data-driven or last click. These models answer: "Which ad touchpoint should get credit for this sale?"

    But they never answer the more important question: "Would this sale have happened without the ad?"

    Brand searches are the clearest example. Some people who search your brand name would have bought through an organic listing anyway; others might have gone to a competitor bidding on your name. The platform credits the brand ad either way, so attributed brand ROAS cannot tell you which share was caused by the ad. Only a controlled test can.

    This isn't fraud. It's how attribution works. But if you're making budget decisions based on attributed performance alone, you're almost certainly overspending in some areas and underspending in others.

    Illustrative example: invented numbers, not a client result

    What the platform reported

    4.2x

    Blended ROAS

    £340k

    Attributed Revenue

    After incrementality testing

    1.8x

    Incremental ROAS

    £145k

    Incremental Revenue

    Synthetic figures for illustration only. In this invented scenario a controlled test estimated that 57% of attributed revenue would have occurred without ads; a real result would carry an uncertainty range.

    How We Test Incrementality

    Four methodologies, each suited to different account structures and commercial questions.

    Geographic Holdout

    Pause all paid activity in a matched region. Compare sales trends against an active region with similar demographics and seasonality.

    Best For

    Brands with national coverage and consistent regional demand

    Duration

    Set from your own pre-period volatility and conversion lag

    Measures

    The difference in total sales between treatment and control regions, against a pre-registered outcome

    Campaign-Level Holdout

    Pause a specific campaign type (e.g., Brand, PMax, Generic) while keeping everything else running. Isolate the true lift of that campaign.

    Best For

    Measuring the real value of brand campaigns or Performance Max

    Duration

    Set from your own pre-period volatility and conversion lag

    Measures

    Whether pausing a campaign reduces total sales or just shifts attribution

    Budget Scaling Test

    Increase or decrease spend by a fixed percentage in a controlled period. Measure whether the marginal spend produces marginal profit.

    Best For

    Testing whether 'scaling' actually improves outcomes

    Duration

    Set from your own pre-period volatility and conversion lag

    Measures

    Marginal POAS - the profit generated by the last £1,000 of spend

    Channel Isolation

    Change one channel in a randomised or matched design while holding others steady, to estimate that channel's contribution.

    Best For

    Multi-channel brands unsure which platform drives genuine demand

    Duration

    Set from your own pre-period volatility and conversion lag

    Measures

    Estimated channel contribution, with an uncertainty range, compared with platform-claimed attribution

    The 5-Phase Process

    Phase lengths depend on your conversion volume, conversion lag and the effect size the test needs to detect. They are agreed in the pre-registered design, not fixed in advance.

    1

    Length set in the test design

    Commercial Baseline

    • Establish true P&L baseline (not platform P&L)
    • Map current attribution claims vs. bank-reconciled revenue
    • Identify the gap between reported and actual
    • Define the commercial question the test will answer
    2

    Length set in the test design

    Test Design

    • Select test methodology based on account structure
    • Define control and test groups with statistical rigour
    • Set minimum detectable effect thresholds
    • Align test duration with business cycles (avoid peak/promo periods)
    3

    Length set in the test design

    Execution & Monitoring

    • Implement test with clean controls
    • Monitor for contamination (e.g., organic changes, competitor activity)
    • Track both platform metrics AND commercial metrics simultaneously
    • Weekly check-ins to ensure test integrity
    4

    Length set in the test design

    Commercial Analysis

    • Reconcile platform data against actual revenue and margin
    • Calculate true incremental contribution (not platform-attributed)
    • Build diminishing returns curve for marginal spend
    • Produce CFO-ready summary with P&L impact
    5

    Length set in the test design

    Reallocation

    • Redirect budget from non-incremental to genuinely incremental activity
    • Set new POAS targets based on proven incrementality
    • Agree the triggers for re-testing (material changes in spend, mix, competition or range)
    • Document findings for board/finance review

    Reference

    A branded-search test protocol you can use

    Brand Search is where attributed and incremental results diverge most. This is the sequence we use to test it. Google's own custom experiments and Conversion Lift documentation covers the platform tools.

    Three kinds of evidence, not one

    Attributed conversions

    What the platform credits to ads. Useful for optimisation, but it says nothing about what would have happened without the ad.

    Pre/post observation

    Sales before and after a change. Seasonality, promotions and demand shifts are mixed into the difference.

    Causal lift

    The difference between treatment and a comparable control over the same period, with uncertainty stated. A causal reading needs valid assignment (randomised or pre-matched units), a genuine control and stated assumptions (no spillover between units, no other change hitting one group); a contemporaneous difference alone is not enough.

    1. 1. Hypothesis

      Write one falsifiable sentence before anything changes, e.g. "Pausing brand Search in treatment regions loses less contribution before advertising than the media cost it saves."

    2. 2. Eligible test units

      Choose units you can switch independently and measure separately: usually geographic regions (Google Ads location targeting, orders matched by delivery postcode). Exclude units with launches, store openings or unusual promotions planned.

    3. 3. Treatment and control

      Assign units before looking at results, ideally at random within matched pairs. Treatment: brand Search paused or reduced. Control: unchanged. Record the assignment list.

    4. 4. Contamination and overlap

      List every other campaign that can serve on brand queries: Performance Max, Shopping, dynamic search ads, competitors, affiliates bidding on brand. Decide how each is handled in treatment, or the test measures a mix.

    5. 5. Primary business outcome

      Pick one before the test: total orders or contribution after advertising across all channels in each unit. Not platform-attributed conversions, which can fall or shift attribution even when total business sales do not change.

    6. 6. Power and duration

      Pre-specify the smallest effect worth acting on and use your own pre-period volatility to estimate the duration needed to detect it. Cover at least one full purchase cycle plus conversion lag. There is no universal duration.

    7. 7. Stopping rules

      Agree in advance what ends the test early (for example, a loss beyond an agreed contribution threshold, a stock-out or a site outage) and that nobody stops it because an early reading looks good.

    8. 8. Uncertainty and limits

      Report the estimate with a range, not a single number. If the range includes zero, the result is inconclusive, not a win. Results hold for that period, those regions and that competitive landscape.

    Synthetic worked example: invented numbers, not a client or research result

    Pre-period (four weeks): treatment regions 1,000 orders; control regions 1,200 orders.

    Test period (four weeks, brand Search paused in treatment only): control rose to 1,236 orders, a 3% change (1,236 ÷ 1,200 = 1.03).

    Expected treatment orders without the pause: 1,000 × 1.03 = 1,030. Observed: 950. Estimated orders caused by brand Search: 1,030 − 950 = 80.

    At £30 contribution before advertising per order, that is 80 × £30 = £2,400 of contribution lost. Brand spend saved in treatment: £3,000. Net effect of pausing: £3,000 − £2,400 = £600 better off, with advertising subtracted once. This adjusted pre/post arithmetic illustrates the calculation; it is not itself proof, which depends on the assignment, control and assumptions above.

    Before acting, check the uncertainty range. If the plausible range for lost orders runs from, say, 20 to 140, the net effect runs from +£2,400 to −£1,200, and the honest verdict is inconclusive. Do not extend the test after seeing that result unless a sequential design was pre-specified; otherwise accept the decision risk explicitly or write a new protocol for a fresh test.

    Contribution follows our POAS convention.

    Pre-launch worksheet

    Why Incentives Matter

    Incrementality testing is uncomfortable. It can show that some attributed conversions would have happened without the ad. Brand campaigns with high reported ROAS can include many conversions that would have happened anyway, and Performance Max can take credit for demand other campaigns would have captured. How much varies by account; only a test tells you.

    For an agency whose fee is justified by attributed ROAS, this is a structural threat. Proving that part of the spend isn't incremental means recommending a budget cut, and under a percentage-of-spend model, a corresponding fee reduction.

    We charge fixed fees specifically so this incentive conflict doesn't exist. When we find waste, we recommend cutting it. When we find incrementality, we recommend scaling it. The commercial alignment is clean.

    Common Questions

    Incrementality testing measures the true causal impact of your advertising by comparing outcomes with and without ads running. Unlike attribution modelling (which assigns credit after the fact), well-designed incrementality tests estimate whether ads caused sales that would not have happened otherwise. The estimate depends on valid assignment, a genuine control and stated assumptions.

    The primary cost is the revenue you forgo during the holdout period. For geographic holdouts, that depends on how much of your market is in treatment and for how long, both set in advance from your own data. There is no universal duration or guaranteed payback; the value is a decision based on causal evidence rather than attribution.

    Controlled tests take planning, conversion volume and a willingness to pause or reduce spend. Under a percentage-of-spend fee, a result that recommends lower spend also lowers the fee, which is one reason we charge fixed fees.

    There is no fixed schedule; re-test when spend, mix, competition or the product range changes materially. Market conditions, competition, and consumer behaviour change - incrementality is not a fixed number. What was incremental 6 months ago may not be today.

    Attribution tells you which touchpoints a customer interacted with before converting - it assigns credit after the fact. Incrementality tells you whether the conversion would have happened without the ad at all. Attribution answers 'who gets credit.' Incrementality answers 'did it actually work.' They are fundamentally different questions.

    We use cookies to improve your experience. Privacy Policy