Influencer Marketing Incrementality: How to Run a Holdout Test
· 6 min read
Discount codes and tracked links tell you which sales touched a creator. They do not tell you which sales would not have happened otherwise, and those are different numbers — sometimes by a factor of two. The only way to separate them is to withhold creator spend somewhere and compare what happens there against where you spent.
That is an incrementality test, and creator marketing is genuinely harder to test than paid media: you cannot flip creators off mid-flight without breaking contracts, the content keeps working after the campaign ends, and a code posted to a national audience leaks into whatever region you were trying to hold out. All three are workable, if you design around them rather than discover them at read-out.
Why coded sales overstate
Three leaks, in rough order of size:
- Cannibalised demand. A customer already heading to your site finds a creator's code first. The sale is coded but it was coming anyway. This is largest for brands with strong branded search and for codes that leak to deal sites.
- Code-hunting. The checkout field itself creates the search. A share of "creator sales" are people who opened a new tab, searched "brand + discount code", and took whichever one ranked.
- Under-counting in the other direction. Long-form and gift-guide style content converts weeks later through branded search, which attribution reads as organic. This one pushes the true number up, which is why the honest answer is a test rather than a discount factor you guessed.
Attribution is still worth doing well — tracking codes and links properly is the input to everything here. A test tells you what to multiply that number by.
Three designs that work for creator campaigns
| Design | How it works | Needs | Read time | Main failure |
|---|---|---|---|---|
| Geo holdout | Run creator spend in matched markets, withhold in others, compare revenue | Regional sales data; ability to target or restrict by geo | 4–8 weeks plus a pre-period | National codes and organic reach contaminate the holdout |
| Time-based on/off | Alternate active and dark periods, compare like weeks | Enough weeks and stable seasonality | 8–12 weeks minimum | Seasonality and other channels moving underneath you |
| Staggered creator rollout | Launch cohorts in waves; each unstarted cohort's audience is a temporary control | Enough creators to make cohorts (roughly 20+) | Length of the rollout | Cohorts differ by audience, not just by timing |
Geo holdout is the cleanest, if you have regional data. Time-based on/off is what you use when you sell one product nationally with no regional split — weaker, because everything else in the business also changes between November and January. Staggered rollout suits always-on programmes: you were onboarding creators in waves anyway, so the test costs nothing extra.
Sizing the test before you run it
The failure that wastes a quarter is running a test too small to detect the effect you were looking for, then reading the noise as an answer.
A rough sizing check, before committing: if creator spend in the test region is less than about 2–3% of that region's revenue, a normal-length test will not separate its effect from week-to-week variance. You have three levers — spend more in fewer places, run longer, or accept a wider confidence interval and say so.
A worked example, with the assumptions stated because they are assumptions, not measurements:
- Test markets do roughly $400k of revenue over an 8-week window, with weekly revenue varying by about 12% week to week.
- You put $60k of creator spend into those markets — 15% of revenue, a genuinely detectable share.
- At that spend and that variance, an 8-week test can distinguish a lift of roughly 5% or more from noise. Below that, you will get a point estimate you should not act on.
- 5% of $400k is $20k of incremental revenue against $60k of spend. If your contribution margin is 60%, that is $12k of margin — a clear loss, and a real result.
The point of doing this arithmetic first is that it tells you what the test can and cannot answer. If the minimum effect you can detect is bigger than the effect that would change your decision, do not run it — increase the spend concentration or change the question.
Two more design rules. Always include a pre-period of at least four weeks and check that test and control markets moved together in it; if they did not, your matching is wrong and the post-period difference is not the campaign. And hold everything else constant — no regional promo, no paid social geo-shift, no retail push into the test markets — or you have measured your own marketing calendar.
Reading the result without fooling yourself
Report the range, not the point. A test that returns "creator marketing drove 8% lift, 95% CI 1% to 15%" means the effect is probably positive and you cannot yet size it. That is a legitimate finding and a reason to run a bigger second test, not a headline number for a deck. The same discipline the ROI post applies to attributed numbers applies harder here, because a lift test carries an air of authority a spreadsheet does not.
Then check contamination directly. Pull the share of code redemptions in the test that came from outside the test geography — if it is meaningful, your holdout was not held out and the true lift is larger than measured. Prefer geo-restricted incentives or no incentive at all in a lift test, and where creators must post nationally, accept that geo tests measure your lower bound.
Finally, take the ratio you got and apply it as a discount factor to attributed sales going forward — "coded sales are running at roughly 0.6× incremental" — then re-test every two or three quarters. That factor is what makes ordinary campaign reporting honest between tests, and it drifts as your brand awareness and creator mix change.
The prerequisite nobody budgets for
Every design above assumes you can answer, precisely: which creators posted, on which dates, in which markets, under which codes, at what cost. Most brands cannot. The campaign lives across an inbox, a spreadsheet, a payments export and a shared drive, and reconstructing a clean flight window six weeks later takes longer than the analysis does.
Keeping the deal terms, posting dates, codes and payments on one record per creator is the unglamorous half of measurement — it is what CreatorCast is for, and it turns a lift test from a project into a query. Set the record up before the test starts; a test you have to reconstruct is a test you will not repeat.
One last caution: test a programme you intend to keep running. A one-off burst tells you about a burst, and creator marketing's returns — saved posts, branded search, the second purchase — show up on a horizon longer than any holdout you can afford.
Frequently asked questions
What is the difference between attribution and incrementality? Attribution assigns a sale to whichever touchpoint it can see — a code or a link. Incrementality asks whether the sale would have happened without that touchpoint. Attributed sales are usually larger than incremental ones, because some buyers were already coming.
How long should an influencer holdout test run? Four to eight weeks of live campaign, plus a pre-period of at least four weeks to confirm your test and control markets track each other. Shorter than that and week-to-week variance swamps the effect; much longer and other business changes creep in.
Can I test incrementality without regional sales data? Yes, with a time-based on/off design or a staggered creator rollout, both of which are weaker because seasonality and other channels move underneath you. A survey question at checkout — "where did you hear about us" — is a useful cross-check, not a substitute.
What if the test says creator marketing is not incremental? Check the design before the conclusion: national codes leaking into holdout markets, a promo running in test regions, or a spend level too small to detect. If the design holds, the honest read is that this creator mix at this spend was not incremental — which is an argument for changing the mix, not necessarily for cutting the channel.
Run your creator program without the spreadsheet
Find creators, run outreach from your own inbox, approve content and send payouts — all in CreatorCast.
Get started