πŸ“ Your MMM Overstates Paid Search: Why a 4.2x Channel Reads as 10.6x, and How to Take the Parameters, Not Just the Lift, From a Geo Test

A Zalando preprint: a standard MMM reads a 4.2x paid-search channel as 10.6x. Better controls don't fix it; the geo-test time series does. The protocol.

Share
A tall navy column with a grey wave inside it beside a short blue column, a dashed line marking the true height

πŸš€ THE EXECUTIVE SUMMARY

The Definition: A marketing mix model (MMM) is a regression on weekly aggregates that assigns sales to channels through three parameters per channel: adstock decay (how long an ad keeps working), saturation and effectiveness. A geo experiment withholds a channel in some regions and reads the sales gap. On 21 August 2026 a preprint by Niklas Heusch of Zalando proposed estimating all three parameters directly from the weekly time series of such experiments, instead of reducing each test to one lift number.

The Core Insight: On synthetic data with a known truth, a standard MMM reported a paid-search return of 10.61x against a true 4.20x, and its 90 percent credible interval (6.56 to 14.36) did not contain the truth. Perfect controls only brought it down to 8.41x. The structural estimate from two geo tests landed at 4.31x, from four at 4.14x, both with intervals around the truth. This is a budget-setting problem, not a data-quality one: spend follows demand, so an observational model reads demand as effect. The fix is variation you created on purpose, kept as a time series.

The paper's words: "Marketing budgets are not randomly assigned. Companies increase advertising during periods of expected high demand, algorithmic bidding systems chase performance signals, and strategic decisions coordinate marketing with promotional calendars."

Paid search is the sharpest case because the auction does the confounding for you. Google's own Meridian documentation is blunt: query volume "is often an important confounder between media and sales", and "failing to control for GQV can lead to overestimation of the causal effect of paid search." A 2018 Google paper by Chen and colleagues built a correction for exactly this. The result that should worry anyone running an MMM is the oracle case: with every true driver of demand handed to the model, the estimate still came in at twice the truth. Two mechanisms survive: the bidder responds to realised weekly performance "including its random component, which no covariate an analyst could hold constant can absorb", and the baseline "combines its components multiplicatively while the controls enter linearly". The algorithm chases noise, and noise looks like return.

What the paper does instead

The setup is 156 weeks of a simulated online retailer with three channels (paid search, social, TV), seasonality, trend, endogenous budgets and algorithmic bidding. Four holdout tests run on paid search, each with four weeks of pre-period, four weeks with spend set to zero in the treatment regions and eight weeks of cooldown, at four different spend levels. Instead of summing each test to one lift, the method differences treatment and control outcomes week by week and fits adstock, saturation and effectiveness to how that gap opens and closes. Together the tests cover adstocked spend from roughly EUR 43K to EUR 194K a week.

Bar chart: paid-search ROAS estimates against a true value of 4.20x. Standard MMM with realistic controls 10.61x, with oracle controls 8.41x; structural estimation from 2 geo tests 4.31x, from 4 geo tests 4.14x, each with 90 percent credible intervals.
Same channel, same truth, four estimates. The observational model misses even with perfect controls; two geo tests kept as time series land within a tenth.

The compliance angle: aggregate by design

Neither side of this workflow touches a person: an MMM consumes weekly spend and sales, a geo test regional sales by day. No cookie, no identifier, no lawful-basis question under GDPR, nothing for a consent banner to gate, which is why, as our geo-incrementality issue argued, this measurement family survives falling consent rates. The obligation moves: once the model decides budgets, the tests behind it need a paper trail a new CMO can follow.

Model, test, or both

RouteWhat you getCostWhere it breaks
Observational MMM alone (Meridian, Apache-2.0; Robyn, open source)Channel returns and response curves from weekly aggregates; a budget optimiserNo licence fee; analyst and data-engineering timeEndogenous spend; paid search reads high even with perfect controls
One-number lift test (Google Conversion Lift, or a geo test summed to a lift)A causal read for one channel, one window, one spend levelConversion Lift from USD 5,000 since November 2025; 7 to 14 daysOne point on the curve; no adstock, no saturation; a different estimand from the MMM's
Structural calibration (geo tests kept as weekly series)Adstock, saturation and effectiveness with intervals; priors the MMM can take directly16-week cycles including cooldown, at several spend levels; a statisticianAssumes the functional form; no spillovers; valid only over the spend range tested

The house view: never run the first row without at least the second. If you already pay for geo tests, the third row costs storage discipline, not budget.

The Expert Perspective

Calibration has been the industry's answer for years. Robyn "implements the MMM calibration as an objective function in the multi-objective optimization" against lift results, and Meridian's guidance is that "one common approach is to use an experiment's point estimate as the prior mean and its standard error as the prior standard deviation." Both take the experiment as one number, and Meridian admits the discomfort: "The ROI measured by an experiment rarely aligns perfectly with the ROI measured by MMM", because the two have different estimands. Heusch closes the gap from the experiment side: fit the same three parameters the MMM uses, from the experiment's own weeks, then hand them over as priors, as a joint likelihood, or by subtracting the channel's contribution and modelling the residual. The caveats are the paper's own: synthetic data only, an assumed geometric adstock and logistic saturation, no spillover between regions, and tests long enough (two-week tests "may not generate enough adstock variation") and varied enough in spend to trace the curve. Dew, Padilla and Shchetkina's 2024 "Your MMM is Broken" reached the same conclusion: the test design decides the model form, not the fit.

Conclusion & Next Steps

The installable artefact is a geo-test protocol written for calibration, not for a headline:

  • Length rule: pre-period, test, cooldown. The paper's design is 4, 4 and 8 weeks; the cooldown is where adstock decay is measured, so do not cut it.
  • Level rule: run tests at different spend levels across the year; saturation is identified only over the range you probe.
  • Storage rule: keep the weekly treatment-minus-control series, by region, in the warehouse. A lift percentage in a deck cannot be re-analysed.
  • Control rule: put population-scaled query volume in the observational model anyway, split brand and generic; Meridian supports it natively.
  • Handover rule: feed the estimated parameters, with intervals, into the MMM as priors and record which test produced them. Our MMM blueprint covers the model from there.

FAQ

Does this mean my MMM's paid-search ROAS is 2.5 times too high?
No. The 2.5 factor belongs to one synthetic retailer with one bidding rule; the direction generalises, the size does not.

Can I use Google's Conversion Lift instead of a geo test?
It returns one lift number for one period, not the weekly regional series the structural method needs, and Search, Shopping and Performance Max studies still go through an account representative.

Is this only a paid-search problem?
Any channel whose budget responds to results, including algorithmically bid social and retargeting, carries the same bias in smaller measure.

References & Sources Cited


Forwarded this by a friend? Subscribe here. Found it useful? Forward it to one person who'd want it.

See you soon,
Team Data Measured

Data Measured is researched and fact-checked by our Editorial Team. We explain measurement; nothing here is legal advice.