π Your MMM Overstates Paid Search: Why a 4.2x Channel Reads as 10.6x, and How to Take the Parameters, Not Just the Lift, From a Geo Test
A Zalando preprint: a standard MMM reads a 4.2x paid-search channel as 10.6x. Better controls don't fix it; the geo-test time series does. The protocol.
π THE EXECUTIVE SUMMARY
The Definition: A marketing mix model (MMM) is a regression on weekly aggregates that assigns sales to channels through three parameters per channel: adstock decay (how long an ad keeps working), saturation and effectiveness. A geo experiment withholds a channel in some regions and reads the sales gap. On 21 August 2026 a preprint by Niklas Heusch of Zalando proposed estimating all three parameters directly from the weekly time series of such experiments, instead of reducing each test to one lift number.
The Core Insight: On synthetic data with a known truth, a standard MMM reported a paid-search return of 10.61x against a true 4.20x, and its 90 percent credible interval (6.56 to 14.36) did not contain the truth. Perfect controls only brought it down to 8.41x. The structural estimate from two geo tests landed at 4.31x, from four at 4.14x, both with intervals around the truth. This is a budget-setting problem, not a data-quality one: spend follows demand, so an observational model reads demand as effect. The fix is variation you created on purpose, kept as a time series.
Why observational models flatter paid search
The paper's words: "Marketing budgets are not randomly assigned. Companies increase advertising during periods of expected high demand, algorithmic bidding systems chase performance signals, and strategic decisions coordinate marketing with promotional calendars."
Paid search is the sharpest case because the auction does the confounding for you. Google's own Meridian documentation is blunt: query volume "is often an important confounder between media and sales", and "failing to control for GQV can lead to overestimation of the causal effect of paid search." A 2018 Google paper by Chen and colleagues built a correction for exactly this. The result that should worry anyone running an MMM is the oracle case: with every true driver of demand handed to the model, the estimate still came in at twice the truth. Two mechanisms survive: the bidder responds to realised weekly performance "including its random component, which no covariate an analyst could hold constant can absorb", and the baseline "combines its components multiplicatively while the controls enter linearly". The algorithm chases noise, and noise looks like return.
What the paper does instead
The setup is 156 weeks of a simulated online retailer with three channels (paid search, social, TV), seasonality, trend, endogenous budgets and algorithmic bidding. Four holdout tests run on paid search, each with four weeks of pre-period, four weeks with spend set to zero in the treatment regions and eight weeks of cooldown, at four different spend levels. Instead of summing each test to one lift, the method differences treatment and control outcomes week by week and fits adstock, saturation and effectiveness to how that gap opens and closes. Together the tests cover adstocked spend from roughly EUR 43K to EUR 194K a week.

The compliance angle: aggregate by design
Neither side of this workflow touches a person: an MMM consumes weekly spend and sales, a geo test regional sales by day. No cookie, no identifier, no lawful-basis question under GDPR, nothing for a consent banner to gate, which is why, as our geo-incrementality issue argued, this measurement family survives falling consent rates. The obligation moves: once the model decides budgets, the tests behind it need a paper trail a new CMO can follow.
Model, test, or both
| Route | What you get | Cost | Where it breaks |
|---|---|---|---|
| Observational MMM alone (Meridian, Apache-2.0; Robyn, open source) | Channel returns and response curves from weekly aggregates; a budget optimiser | No licence fee; analyst and data-engineering time | Endogenous spend; paid search reads high even with perfect controls |
| One-number lift test (Google Conversion Lift, or a geo test summed to a lift) | A causal read for one channel, one window, one spend level | Conversion Lift from USD 5,000 since November 2025; 7 to 14 days | One point on the curve; no adstock, no saturation; a different estimand from the MMM's |
| Structural calibration (geo tests kept as weekly series) | Adstock, saturation and effectiveness with intervals; priors the MMM can take directly | 16-week cycles including cooldown, at several spend levels; a statistician | Assumes the functional form; no spillovers; valid only over the spend range tested |
The house view: never run the first row without at least the second. If you already pay for geo tests, the third row costs storage discipline, not budget.
The Expert Perspective
Calibration has been the industry's answer for years. Robyn "implements the MMM calibration as an objective function in the multi-objective optimization" against lift results, and Meridian's guidance is that "one common approach is to use an experiment's point estimate as the prior mean and its standard error as the prior standard deviation." Both take the experiment as one number, and Meridian admits the discomfort: "The ROI measured by an experiment rarely aligns perfectly with the ROI measured by MMM", because the two have different estimands. Heusch closes the gap from the experiment side: fit the same three parameters the MMM uses, from the experiment's own weeks, then hand them over as priors, as a joint likelihood, or by subtracting the channel's contribution and modelling the residual. The caveats are the paper's own: synthetic data only, an assumed geometric adstock and logistic saturation, no spillover between regions, and tests long enough (two-week tests "may not generate enough adstock variation") and varied enough in spend to trace the curve. Dew, Padilla and Shchetkina's 2024 "Your MMM is Broken" reached the same conclusion: the test design decides the model form, not the fit.
Conclusion & Next Steps
The installable artefact is a geo-test protocol written for calibration, not for a headline:
- Length rule: pre-period, test, cooldown. The paper's design is 4, 4 and 8 weeks; the cooldown is where adstock decay is measured, so do not cut it.
- Level rule: run tests at different spend levels across the year; saturation is identified only over the range you probe.
- Storage rule: keep the weekly treatment-minus-control series, by region, in the warehouse. A lift percentage in a deck cannot be re-analysed.
- Control rule: put population-scaled query volume in the observational model anyway, split brand and generic; Meridian supports it natively.
- Handover rule: feed the estimated parameters, with intervals, into the MMM as priors and record which test produced them. Our MMM blueprint covers the model from there.
FAQ
Does this mean my MMM's paid-search ROAS is 2.5 times too high?
No. The 2.5 factor belongs to one synthetic retailer with one bidding rule; the direction generalises, the size does not.
Can I use Google's Conversion Lift instead of a geo test?
It returns one lift number for one period, not the weekly regional series the structural method needs, and Search, Shopping and Performance Max studies still go through an account representative.
Is this only a paid-search problem?
Any channel whose budget responds to results, including algorithmically bid social and retargeting, carries the same bias in smaller measure.
References & Sources Cited
- Niklas Heusch, Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments, arXiv:2608.21128 (21 August 2026).
- PPC Land, MMM overstates paid search ROAS by 2.5 times (31 August 2026); Google lowers incrementality testing threshold to $5,000 (11 November 2025).
- Chen et al., Bias Correction For Paid Search In Media Mix Modeling (2018).
- Dew, Padilla and Shchetkina, Your MMM is Broken (2024).
- Meridian docs: Paid search modeling; ROI priors and calibration; GitHub.
- Meta, Robyn features: calibration.
Forwarded this by a friend? Subscribe here. Found it useful? Forward it to one person who'd want it.
See you soon,
Team Data Measured
Data Measured is researched and fact-checked by our Editorial Team. We explain measurement; nothing here is legal advice.