📈 THE INCREMENTALITY MEASUREMENT FRAMEWORK
Incrementality is the measure of what changed because of your marketing: not what got credited to it. Every method on this sheet exists to answer one question honestly: would this outcome have happened anyway? The ladder below ranks them by how rigorously they answer it.
Not to Be Confused With
🎯 Attribution
Attribution splits credit between touchpoints; incrementality asks whether the outcome would have happened anyway. Use attribution for steering, incrementality for truth.
📉 Pre/post without a control
“Sales went up after the campaign” is a story, not a measurement. Without a counterfactual, you cannot know what would have happened anyway. That is exactly what the higher rungs control for.
🧮 Media mix modelling
MMM allocates spend at scale; it does not prove cause. Feed incrementality findings into your MMM assumptions.
North Star Setup
🎯 Define One Outcome
Pick one business-linked metric for decisions (e.g., qualified trials, retained revenue), not a vanity proxy.
One metric beats five conflicting ones⏱️ Set Read Windows
Define exactly when outcomes are counted so tests are comparable across channels and cycles.
🧮 Make It Decision-Usable
Your metric must tell you whether to scale, hold, or cut spend this week.
💼 Finance-Ready
If finance wouldn’t care when it moves, it isn’t your North Star metric.
Incrementality Ladder
| Level | Use when | Strength | Weakness |
|---|---|---|---|
| 1. Attribution | Need rapid directional steering | Fast and cheap | Weakest causal truth |
| 2. Pre/Post | Controls limited, medium-stakes decision | Quick calibration | High confounding risk |
| 3. Matched Markets/Cohorts | Need better quasi-experimental confidence | Stronger comparability | More setup and monitoring |
| 4. Randomised Holdout | High-stakes budget or strategy calls | Best causal evidence | Hardest operationally |
A concrete version of rung 3, from my own work: hold out India, keep Pakistan as the control market. The control tells you what would have happened in the holdout, which is the whole game.
Two more rungs practitioners ask about. Ghost ads (or PSA holdouts) hold out the ad itself at serving time: the control group sees a public-service ad instead, so you are testing the creative rather than the audience. It is the cleanest digital holdout I know. Brand lift studies survey exposed versus unexposed groups for the cases where the outcome is perception, not behaviour.
How to Run Tests
- Pre-register everything. Hypothesis, KPI, duration, stop rules, and decision thresholds.
- Pick highest feasible rigor. Don’t pretend attribution is experimentation.
- Protect test integrity. No mid-test KPI changes, no peeking and early stopping.
- Read business + statistical significance. Tiny but significant effects usually don’t move the company.
- Reallocate by marginal lift. Keep coverage where needed, cut weak marginal spend first.
- Back-test by switching off. Periodically turn off marketing entirely in a region to validate your models. At Facebook we ran a fully-off test on paid direct response marketing for one of our products in India. MAU began shrinking, and recovered when we reversed the test. It proved the methodology and defended the budget.
What Matters vs What Doesn’t
✅ This Matters
- Sustained lift over pre-registered windows
- Result survives replication across cohorts/regions
- Magnitude changes budget decisions
- Findings feed MMM/planning assumptions
🚫 This Doesn’t
- Microscope-level gains with heavy interpretation
- Winner-only reporting after many tests
- Attribution-only proof for budget defense
- One-day significance spikes
Failure Modes to Avoid
🧷 P-Hacking
Stopping tests once significance appears or changing metrics midstream inflates false wins.
🔎 Last-Click Overstatement
High-intent channels (e.g., branded search/app store search) often absorb credit for demand created elsewhere.
🪫 Underpowered Tests
Small samples create noise disguised as confidence.
🧠 No Learning Loop
If insights are not documented and reused, each quarter resets to zero learning.
🧭 Operating Rule
Use attribution for steering, incrementality for truth, and MMM for scale allocation. One system, three roles.
The Ground Is Moving
Two regime changes made everything on this sheet more valuable, and I expect more to come. Apple’s App Tracking Transparency broke deterministic attribution on iOS, and the loss of third-party cookies did the same on the web. Modelled measurement got weaker across the board. That is exactly why experiments got more valuable: when you can no longer track the user, you test the market instead.
Know Its Limits
Match the rigour to the stakes. A randomised holdout is the gold standard, but it is also the hardest thing on this sheet to run operationally. Don’t gold-plate a small decision with a six-week geo experiment. The ladder exists so you can climb as high as the decision warrants, not so every call starts at the top.
And remember what incrementality can’t tell you: whether the money worked, yes. Not what to do next. The test proves the lift; the creative leap that produces the next lift still comes from people, not holdouts.
My position: the test is never the expensive part. Not knowing is. The most rigorous test on this sheet requires deliberately not marketing to people you could reach, and I have seen teams skip the holdout that would have saved them millions because the test itself felt expensive, then spend the millions anyway.
