Found in AI answers?Get a free review
Theory Road.
← PerspectivesPaid Media

Marketing Mix Modeling for Mid-Size Brands: When It's Worth It.

Marketing mix modeling has gone from a consultancy product to free open-source software, but the math was never the hard part. Here is what MMM needs, what it costs, when it is not worth it, and the geo holdout test we run instead for mid-size budgets.

By Theory RoadSeptember 21, 202612 min read

Marketing mix modeling (MMM) is a statistical method that estimates how much each marketing channel contributed to sales by relating weekly spend, promotions and outside factors to revenue over time. For mid-size brands it is worth it when you spend meaningfully in several channels, have roughly two years of clean weekly data, and face real budget decisions it can inform. When those conditions are missing, a well-designed geo lift test usually answers the question you actually have faster and for less money.

This guide explains what marketing mix modeling does and does not do, how it compares with multi-touch attribution and incrementality testing, the data a model needs, what the open-source tools from the major ad platforms offer, what it really costs, when it is not worth it, and how we run geo holdout tests for brands whose budgets are too small for a full model to be reliable.

What marketing mix modeling is.

A marketing mix model is a regression, usually Bayesian in modern tools, that treats sales as the sum of a baseline (what you would sell with no marketing at all) plus the contribution of each channel, plus the effect of price, promotions, seasonality, holidays and external factors. It works on aggregated data, typically weekly totals by channel, so it needs no user-level tracking, cookies or click paths. That matters more every year as privacy changes erode user-level tracking.

Two ideas make MMM more than a simple regression. The first is adstock, or carryover: a television or video flight keeps working for weeks after it airs, so the model spreads its effect over time. The second is saturation, or diminishing returns: the tenth thousand dollars in a channel buys less than the first. Together they let the model answer the practical question, which is not only which channel worked but where the next dollar should go.

MMM vs multi-touch attribution vs incrementality testing.

Three families of methods get lumped together as measurement. They answer different questions, they need different data, and they fail in different ways. Knowing which one you are using matters more than which vendor sells it.

Marketing mix modeling, multi-touch attribution and incrementality testing compared
MethodQuestion it answersData it needsWhere it is strongBlind spots
Marketing mix modelingHow much did each channel contribute to total sales, and where should the next dollar go?Weekly spend and outcomes by channel over roughly two years or more, plus promotions, pricing and seasonalityCovers offline, upper-funnel and walled-garden channels; needs no user trackingCorrelational; struggles with channels whose spend never varies; slow to update
Multi-touch attributionWhich tracked touchpoints preceded each conversion?User-level click and impression paths, usually from cookies, pixels and platform APIsGranular and fast; useful for in-channel optimizationMisses offline and untracked exposure; credits touchpoints rather than measuring cause; weakened by privacy changes
Incrementality testing (geo lift or holdout)Did this channel or campaign cause sales that would not have happened otherwise?Outcomes by geography or randomized group before and during a controlled testClosest thing to proof of cause; simple to explainTests one question at a time; needs enough volume per market; costs the sales you give up in holdout areas

Multi-touch attribution tells you who touched what. It does not tell you whether the touch mattered. A retargeting ad shown to someone already on the way to checkout gets credit in almost every attribution model, even when the sale would have happened anyway. Incrementality testing is the corrective: it withholds the ad from a comparable group and measures the difference. MMM sits between them, estimating causal contribution across the whole mix from historical variation, and it is strongest when calibrated with a few real experiments.

Attribution tells you who was in the room. Incrementality tells you who changed the outcome. MMM tries to estimate the second from history, and it gets much better when you feed it the results of real tests.

The data a marketing mix model needs.

Most MMM projects that disappoint fail on data, not math. Before you pay for a model or assign an analyst to an open-source one, check that you can assemble the following at a weekly grain, going back as far as possible. Two years is a common working minimum because the model needs to see each season at least twice, and more history helps.

  • Spend by channel per week, split the way you would actually make decisions: branded search separate from non-brand, prospecting social separate from retargeting, and each retail media or connected TV buy on its own line.
  • The outcome you care about, ideally revenue or new customers from your own system of record rather than platform-reported conversions.
  • Promotions, discounts and price changes, with dates and depth.
  • Seasonality and holiday markers, plus known one-off events such as a stockout, a site outage, a product launch or a press moment.

Open-source MMM tools from the major ad platforms.

A few years ago, marketing mix modeling meant hiring a specialist consultancy. Today the major ad platforms publish open-source modeling frameworks that any analyst can run. At the time of writing, Meta maintains Robyn, an open-source MMM package, and GeoLift, an open-source library for designing and reading geo experiments. Google maintains Meridian, an open-source Bayesian MMM framework that can incorporate reach and frequency data and calibrate against experiment results. All three are free to use, well documented and actively developed, and their features change, so check the current documentation before you plan around any one capability.

Free software is not a free model. These frameworks handle the statistics; they do not collect your data, clean it, choose sensible priors, or decide whether the output makes business sense. A model's conclusions depend heavily on the analyst's choices, and those should be made by someone accountable to your results rather than to a media seller. Commercial MMM platforms add data connectors, dashboards and support on top of similar methods, and they charge for it.

What marketing mix modeling really costs.

The line items are the same whether you buy a platform, hire a consultancy or run an open-source framework in house. What changes is who pays for each one.

  • Data engineering: pulling spend from every ad platform, outcomes from your commerce or CRM system, and promotion calendars from whoever keeps them, then reconciling it into one weekly table. For most brands this is the largest single cost of the first model.
  • Modeling time: specifying the model, setting priors, running it, checking diagnostics and rerunning when results do not make sense.
  • Calibration experiments: geo tests or platform lift studies that anchor the model to measured cause. These cost analyst time and the sales you give up in holdout markets.
  • Refresh and maintenance: a model is a snapshot. Most brands that rely on MMM refresh it quarterly or monthly, and each refresh repeats part of the data work.

When marketing mix modeling is not worth it.

MMM is a good tool with a narrow sweet spot for mid-size brands. It is usually not worth building yet when:

  • Most of your spend sits in one or two channels. A model with two inputs is a geo test that has not been run yet.
  • You have less than about two years of consistent data, or the business changed so much in that time that older data describes a different company.
  • Spend in your main channels barely varied, so the model has little signal to learn from.
  • Your outcome data lives in several systems that do not agree, and nobody owns reconciling them.
  • The question you really have is narrow, such as whether branded search or a new connected TV test is incremental. A targeted experiment answers that faster.

Incrementality testing: the other half of the answer.

Incrementality testing measures cause directly. You split an audience or a set of markets into test and control, change marketing in one and not the other, and compare outcomes. Platform lift studies do this inside a single ad platform using randomized user groups, and they are worth running when you qualify, but they are graded by the platform being tested and cannot compare channels against each other.

Geo lift tests do the same thing with geography. You pick markets, withhold or increase spend in some of them, and compare sales against matched markets where nothing changed. Because the outcome is measured in your own systems, the result does not depend on any platform's tracking. The trade-off is volume: each market needs enough sales for a change to be visible above normal noise, which is why test design matters more than the software that reads the result. We have written more about the same logic applied to partner programs in our guide to affiliate marketing incrementality.

How Theory Road runs geo holdout tests for mid-size budgets.

Most of the brands we work with spend enough to need real measurement but not enough to support a full model refreshed every month. For them, a disciplined program of geo holdout tests is the most useful measurement investment. This is how we run one inside our paid media and media buying work.

Pick one question with money behind it.
We start with the channel where the budget is largest and the evidence is weakest. Common examples are branded search, prospecting social, connected TV and retail media. The question is written down with the decision it will drive, such as cut, hold or scale.
Measure the outcome in your own systems.
We pull orders, leads or booked jobs by ZIP code or designated market area from your commerce platform or CRM, not from the ad platform. If the business cannot map outcomes to geography yet, fixing that comes first.
Build matched markets from history.
Using at least several months of pre-period data, we group markets whose sales moved together in the past, then assign test and control groups so they start out as similar as possible. Open-source tools help with this matching step.
Size the test before launch.
We run a power analysis to estimate the smallest lift the test could detect with your volume and the planned duration. If that detectable lift is larger than any realistic effect, we change the design, extend the test or pick a different question rather than run a test that cannot answer.
Run the change cleanly.
Spend is paused or increased only in the test markets, with other channels held steady and no promotions that hit test and control differently, usually for several weeks plus a short cooldown.
Read the lift against a synthetic control.
We compare actual sales in test markets with what the matched controls predict they would have been, and report the lift with a range, not a single number, along with the implied incremental cost per acquisition.
Turn the result into a budget decision.
The finding goes back to the owner of the decision named in step one, and into the plan. Where a brand later builds a marketing mix model, each completed test becomes a calibration point for it.

The same design works for channels that are hard to measure any other way. Connected TV and programmatic video rarely produce clicks worth attributing, so geo tests are often the only credible read; our CTV and programmatic work is built around them. Retail media is similar: ads on a retailer's site influence purchases in stores and on other marketplaces that the retailer's own reporting never sees, which is why we test retail media budgets by geography where the retailer's data allows it.

Putting MMM, attribution and experiments together.

The brands that measure well do not pick one method. They use each for what it does best and let the others check it.

  • Attribution and platform reporting for daily, in-channel optimization: which ad, keyword or audience is working better than its neighbors.
  • Geo lift and holdout tests for the big causal questions, run a few times a year on the channels with the most money and the least certainty.
  • A marketing mix model, once the data and spend justify one, for the annual and quarterly allocation across channels, calibrated with the experiments above.
  • One reconciled report that ties all three back to revenue in your own systems, reviewed on a fixed schedule.

If your reporting today stops at platform-reported conversions, start there before any of this. Our guide to paid media reporting covers the four numbers every leadership team should see each month.

Where to start.

For most mid-size brands the sequence is the same: clean up conversion tracking so outcomes are measured in your own systems, make sure outcomes can be broken out by geography, run one well-designed geo holdout test on your largest uncertain channel, and only then decide whether a full marketing mix model will change enough decisions to pay for itself. If you want a second opinion on where your measurement stands, tell us what you are spending and where and we will give you a straight read.

What is marketing mix modeling?

Marketing mix modeling is a statistical method that estimates how much each marketing channel, promotion and outside factor contributed to sales, using aggregated weekly data over time. Because it does not need user-level tracking, it can measure offline and upper-funnel channels and is not weakened by cookie and privacy changes.

What is the difference between MMM and incrementality testing?

MMM estimates the contribution of every channel at once from historical variation in spend and sales. Incrementality testing runs a controlled experiment on one channel or campaign, withholding it from a comparable group, to measure what it caused. Tests are closer to proof but answer one question at a time, so the strongest programs use tests to calibrate the model.

How much data do you need for marketing mix modeling?

A common working minimum is about two years of weekly data by channel, so the model sees each season at least twice, along with promotions, pricing and major events. More history helps, and so does spend that varied over time. Brands with a short or inconsistent history usually get more from geo lift tests first.

What is a geo lift test?

A geo lift test measures a channel's incremental impact by changing marketing in some geographic markets and not in others that historically behaved the same way, then comparing sales. Because outcomes are measured in your own systems by ZIP code or market, the result does not depend on any ad platform's tracking.

Is marketing mix modeling worth it for a mid-size brand?

It is worth it when you spend meaningfully across several channels, have clean multi-year weekly data, and have budget decisions the model can change. If most spend is in one or two channels, or the data is thin, a geo holdout test on your largest uncertain channel will usually answer the question faster and for less.

Work with us

Let’s talk about what’s next.

A short note on where the business is and where it needs to go. A senior partner replies within one business day.

t@theoryroad.com

Step 1 of 2

How can we help you get found?

What do you need help with?

Next: name, email and budget. That’s it.