Measurement · Incrementality

Incrementality testing

Incrementality testing answers one question: how many of these sales would have happened anyway. You hold a group back, run the ads at everyone else, and read the gap. Google recommends 4 to 6 weeks for an experiment. The button that ends it early is available on day three, and ending is the only action here you cannot undo.

See it read your own accounts, free

Every guide on this topic teaches you how to design the test.

Almost none teach you how to keep it alive for six weeks.

That is the part that fails. Not the maths.

Two operators said it out loud, and neither named it

Same week last summer, on r/analytics. One asked it as a question, addressed to nobody in particular:

“Curious if you’ve had geo tests survive contact with a real budget or if they die in the org before they finish.”

u/WickedReports, r/analytics, July 2026. Collected in our own ICP voice bank.

The other answered it without meaning to, in a different thread:

“I wholeheartedly disagree with this, but that’s what the management wants and I cannot fight them (I tried).”

u/trp_wip, r/analytics, July 2026. Same voice bank.

Nobody in either thread proposed a fix

The subject came up and moved on, because there is no accepted name for this failure and therefore nothing to look up.

So this page is in two halves. The first is the standard material, kept short. The second is the part nobody publishes: what you write down on day zero so that a bad Tuesday in week two stays a bad Tuesday.

What is incrementality testing?

Incrementality testing measures how many conversions the advertising actually caused, by holding a group back on purpose. One group sees the ads. A comparable group does not. The gap between them is the incremental result, and everything else in your reporting is a guess about that gap.

The arithmetic is not the hard part:

Incremental conversions and lift

incremental = conversions(exposed) - conversions(control) * scale scale = size(exposed) / size(control) lift % = incremental / (conversions(control) * scale) * 100 incr. CPA = spend / incremental

The last line is the one worth staring at. Your incremental CPA is always worse than your reported CPA, sometimes by a lot, because the reported one is being credited with sales that were coming anyway. That is not a tracking error. It is the difference between two questions.

There are four ways to hold a group back, and they are not interchangeable.

Method What is held back What it costs you Best for
Platform experiment
(Google Ads experiments)
A split of the same campaign’s traffic, control against treatment Nothing extra. The budget is already being spent Comparing two settings. Bidding, match types, a landing page. It answers “which of these two”, never “is any of this working”
Geo holdout Whole regions, switched off for the duration Real revenue in the dark regions, for weeks The honest answer to “is this channel worth it”. Slow, expensive, and the only version most finance teams will believe
Platform lift study
(conversion or brand lift)
A share of the audience the platform selects and manages Usually a minimum spend, plus the control group’s lost sales Speed. Easiest to launch. Hardest to audit, because the platform builds the control group and then grades the result
Switch it off and watch Nothing. Time is the only variable Feels free. It is not Almost nothing, and it is what most people actually do. Seasonality, pricing and every other change ride along inside the result

Note what the first row cannot do. A campaign experiment splits traffic that was going to be bought either way, so it compares two versions of spending the money. It never compares spending to not spending. Plenty of reports labelled incrementality are running that first row.

Not sure which of these four your account has the volume for? We read it and tell you, free →

Why can’t attribution answer this?

Because every conversion in your attribution data was exposed to the ads. There is no unexposed group in there to compare against. Attribution distributes credit across the touchpoints it can see, using a model and a lookback window that you chose. Change either and the same week produces a different number, without a single thing happening in the account. That is the whole argument in marketing attribution.

So attribution can tell you which channel to credit. It cannot tell you which channel to keep.

This gap got wider in 2026, not narrower. Broad, keywordless matching means the platform decides more of what you buy, and the reporting on that decision is split across views the platform tells you not to add up. We wrote that out for the September migration in the AI Max auto-upgrade. When the report cannot be reconciled, holding a group back is what is left.

How long does the test have to run?

Longer than anyone expects, and the platform says so itself.

“It’s recommended that you run the experiment for at least 4-6 weeks or longer if you have a long conversion delay.”

Google Ads Help, Experiments FAQs (retrieved 26 August 2026). The same page recommends waiting for one to two conversion cycles.

Can your account resolve the effect at all?

Google publishes a score for that, before you start. Experiment Power is calculated from historical spend, conversions, expected uplift and duration, and the bands are blunt:

A horizontal scale from 0 to 99 percent split into three bands. Low, 0 to 49 percent, takes up half the width and is struck out. Medium, 50 to 79 percent, takes up three tenths. High, 80 to 99 percent, takes up the last fifth in solid orange.
The widths are the part the numbers hide: “low power” covers half of everything the score can say, and it means no answer rather than a weak one.

Underneath that sits the same arithmetic as any other measurement question. The smallest difference you can distinguish shrinks with the square root of your conversion count, so below roughly 50 conversions per side you can only see very large effects. The table of thresholds is in what is a good ROAS, and it applies here unchanged.

Which produces the uncomfortable first step. Before designing anything, work out whether your account can produce a readable answer in the time you have. Sometimes it cannot, and knowing that on day zero is worth more than six weeks of held-back revenue.

The bar it gets graded against

When the verdict arrives, Google grades it against a fixed one. Its methodology page puts it in one line: “Two-tailed significance testing is then run using the 95% confidence interval”. That is from The statistical methodology behind experiments, retrieved 26 August 2026. The bar does not move because your quarter ends on Friday.

Find out on day zero, not on day forty-two: we count what your accounts can resolve, free →

The one control in an experiment you cannot take back

Here is the thing I have not seen written down anywhere, and it is sitting in the help docs. Two controls, side by side, and only one of them can be taken back.

The control What Google’s documentation says Can you undo it?
Extend the end date The end date can be pushed “at any time during the experiment, up to a maximum of 12 weeks from the current date” Yes, as often as you like. Patience has unlimited undo
End now “Your experiment will finish at the end date, or you can finish the experiment manually before the end date when you select End now” No. “After the experiment ends, you can’t resume it or edit the end date”

Both quotes are from Find and edit your experiments, retrieved 26 August 2026.

Two buttons that look the same

The irreversible one is the one an anxious stakeholder asks for in week two. Nothing marks it as the point of no return, because to the product it is just a status change.

Why week two is when the pressure arrives

This is not psychology. It is the ramp-up. The early days of an experiment are volatile by construction, so that is precisely when the chart looks alarming and somebody starts watching it daily. The statistics literature has a name for what daily watching does:

“A/B tests are typically analyzed via frequentist p-values and confidence intervals; but these inferences are wholly unreliable if users endogenously choose samples sizes by continuously monitoring their tests.”

Johari, Pekelis and Walsh, Always Valid Inference: Bringing Sequential Analysis to A/B Testing, arXiv:1512.04922 (v3, 16 July 2019).

Nobody in this story is the villain

Read that as an operator, not as a statistician. It does not say your organisation is undisciplined. It says that letting the watcher decide when to stop breaks the test mathematically, however well intentioned the watcher is. The person asking for a daily update is not the villain here. The missing decision rule is. Which is the whole fix, and it fits on one screen. Write it before the test starts. Paste it where the team can see it, and hand it to the next person who asks how the test is going.

What to write down before the test starts

Test charter. Filled in before the test starts, not after

QUESTION Does <change> produce incremental <metric>? HOLDOUT <what is held back, where, and how much> STARTS <date> ENDS <date> not "when we know" MINIMUM READ <N> conversions per side OR <N> days, whichever comes later IGNORE UNTIL <date, start + 14 days> ramp-up, not signal DECISION RULE if lift >= <X>% -> <action> if lift < <X>% -> <action> if inconclusive -> <action> write this one too EARLY STOP only if <spend or safety condition> NOT because the chart looks flat WHO CAN END IT <one name> LOGGED WHERE <where the start and every mid-test change is recorded>

The branch everyone leaves out

It is the third one in the decision rule. In a small account, inconclusive is the most likely outcome. If nobody agreed in advance what happens then, the argument you were trying to avoid arrives anyway, with six weeks of sunk cost attached. And the last line of the charter matters for a reason that took me a while to understand. It is the next section.

Want the baseline pulled before you start the clock? Free read of your Meta and Google accounts →

What this cost me to learn

In 2025 I spent more than 10,000 euros of my own money on ad management platforms. Not a client’s money. Mine.

What they returned was the commentary Google and Meta already show you inside their own interfaces, plus generic rules applied identically to every account. Not one of them ever asked whether the account could answer the question I was asking it. None of them kept a record of what I had changed while I was asking.

That second gap is the one I ended up building for, and it is where I hit something I did not expect.

Two horizontal tracks on the same 42-day axis. The top track is a six-week experiment: days 1 to 14 striped as volatile ramp-up, days 15 to 42 in orange as the readable part. The bottom track is the change_event API window asked on day 42: days 1 to 12 struck out as not retrievable, days 13 to 42 lit in orange as still queryable.
Run a test for the six weeks Google recommends, and on the day you conclude, the API has already forgotten the first twelve days of it.

Six weeks is forty-two days. The API remembers thirty.

Google recommends 4 to 6 weeks for an experiment. The Google Ads API resource that tells you what changed in the account, change_event, must be filtered to the past 30 days and capped at 10,000 rows per query. There is no backfill. The interface keeps two years of change history, the API keeps thirty days, and the gap is not recoverable by any tool you connect.

Six weeks is forty-two days. Forty-two minus thirty is twelve. On the day you sit down to defend the result, nothing can tell you what else was touched in the first twelve days of your own test. We wrote the mechanics of that window up separately in Google Ads change history.

How I found out, which was not cheap

Building the change-history import. The first run pulled four identical rows, and the before-and-after impact sat on “measuring” forever because the same campaign arrived with two different internal references across two imports. The join returned nothing and said nothing. The obvious diagnosis, stale data, was wrong.

The lesson generalises past our tooling. A test is a claim about a period, and the account’s own memory of that period expires while the test is still running. So the charter has a line for where the start gets logged, and the answer cannot be “the platform remembers”. It does not. This is the same argument as deciding from memory, one level up: the account forgets too.

What to do when the organisation is the variable

Five things, in the order they actually help.

The honest counter-argument

Leaving this out would be selling.

Most operators do not want any of this. When Google announced a bidding change in August 2026, the top-voted reply in r/PPC was to set an ideal CPA or ROAS, let it run, and stop worrying. Our own ICP research has logged that pattern nine separate times across four subreddits. Faced with a measurement that is expensive or awkward, the popular move is to lower what you ask of it. Not to fix it. In r/agency the same reflex shows up as dropping the metric entirely.

That is a rational trade, and it works for the pain most people actually feel, which is a client asking why a number moved. It stops working in exactly two situations. When the number is wired to something that spends money on its own, which is what not to automate in your ad accounts. And when somebody is deciding whether to keep funding the channel at all. Incrementality is for the second one. If nobody is asking that question, do not spend six weeks of held-back revenue answering it.

Start with the readability check: your conversion counts per side, on your real accounts, free →

What I still do not know

Related reading. Marketing attribution: why every tool gives you a different number for the same week. What is a good ROAS: the conversion volume that decides whether any comparison is readable. Google Ads change history: the thirty-day window, and why there is no backfill. AI Max auto-upgrade: the reporting split that pushed us here. Deciding from memory: writing the expectation down before the result arrives.

What is incrementality testing?

Incrementality testing measures how many conversions the advertising actually caused, by holding a group back on purpose. One group sees the ads, a comparable group does not, and the gap between them is the incremental result. It is the only method that separates sales you bought from sales you would have had anyway.

How is incrementality different from attribution?

Attribution assigns credit for conversions that already happened, using a model and a lookback window you chose. Incrementality asks what would have happened with no ads at all. Attribution can never answer that, because every conversion in the data was exposed. Only a held-back group creates the missing comparison.

How long should an incrementality test run?

Google Ads Help recommends at least 4 to 6 weeks, longer if your conversion delay is long, and suggests waiting for one to two conversion cycles. The first one to two weeks are ramp-up and usually too volatile to read. Plan the read date at the start rather than deciding it later.

What is a holdout group?

A holdout is the share of your audience, geography or budget deliberately excluded from the ads for the length of the test. It is the control. Holding it back costs real revenue, and that cost is the price of the answer. A test with no holdout is a before-and-after comparison wearing a better name.

Can I run an incrementality test on a small account?

Often no, and it is better to know that first. The smallest difference you can distinguish shrinks with the square root of your conversion count. Under roughly 50 conversions per side you can only see very large effects, so a small account can run a perfect test and still end with no readable answer.

Why do incrementality tests get stopped early?

Because somebody with authority watches the daily numbers during the volatile phase and reacts. Continuous monitoring is a known statistical problem, not a discipline problem. The fix is procedural: write the end date, the decision rule and the person allowed to stop it before the test starts.

Do I need a platform lift study or can I run this myself?

A platform lift study is easier to launch and harder to audit, because the platform builds the control group and grades its own result. A geo holdout or a campaign experiment you configure yourself is more work and gives you the raw numbers. For a test that has to survive an argument, run the one you can show your working on.

Last updated:

Who wrote this

I am Manu. I have been buying media for eight years, and I got tired of platforms that hand back the same commentary the ad platform already shows you. So I started building an open-source console that reads my Meta and Google data and never changes anything I have not approved. It is Apache-2.0, with a runnable demo on synthetic data. If you want the conversion counts and the change log run against your own accounts before you design a test, the audit below does exactly that and changes nothing.