Blog

Why Open-Source MMMs Undervalue Direct Mail — And How to Fix the Model Before You Lock In Your Media Plan

9 Min Read
by Allison Nick

Open-source marketing mix models — Meridian, Robyn, LightweightMMM — have become the default planning tool for a growing segment of mid-market and DTC brands. They’re free, well-documented, and backed by some of the largest companies in advertising. They also systematically undercount direct mail’s contribution to revenue. Not because the frameworks are flawed in principle, but because their default configurations were built for digital channels. Feed a physical, household-level medium through a model calibrated for impression-based digital touchpoints, and the model gets the answer wrong in predictable, correctable ways.

The stakes are real. Teams are using these outputs to set H2 and FY2027 budgets right now. If direct mail’s coefficient is suppressed by a modeling artifact — not by actual underperformance — the channel loses the budget it earned. We’ve identified three specific mechanisms by which open-source MMMs undervalue direct mail, how to correct them, and how holdout-based incrementality results can serve as ground truth for calibration.

The goal isn’t to argue that direct mail deserves more credit than it earns. It’s to make sure the model reflects what actually happened.

Why These Models Default to Digital Logic

Google publicly announced Meridian in March 2024 and made it generally available in January 2025 — it’s a fully Bayesian open-source framework built on Python, designed to integrate with Google’s own ecosystem and intended to replace LightweightMMM as Google’s primary MMM offering. Meta’s Robyn, which has been through multiple significant releases since its original launch, is the de facto starting point for teams building their first marketing mix model. LightweightMMM remains embedded in dozens of active planning workflows.

All three are capable frameworks. None of them ship with defaults that work well for direct mail.

The reason is structural. These tools were designed by teams whose primary channels (paid social, paid search, display, video) share a common data architecture: impression-level event logs, near-instant response curves, and individual-level click or view tracking. Meridian uses geo-level hierarchical modeling with Bayesian inference, which is genuinely powerful. But geo-level priors are initialized around digital response timelines, and none of these frameworks include out-of-the-box configurations for a physical channel that takes days to deliver and weeks to convert.

Three specific default behaviors create the undercount.

Problem 1: Adstock Decay Functions Run Too Fast

Adstock captures the idea that advertising has a carry-over effect — exposure today can drive conversions tomorrow, next week, or later. The question is how long that effect lasts.

Robyn offers geometric adstock, which uses a single decay parameter and assumes the effect peaks the week the ad runs, then fades at a fixed rate — and Weibull PDF adstock, a two-parameter option that enables lagged effects where the peak can arrive weeks after spend. The problem isn’t that Robyn can’t model long response curves. It’s that many implementations — especially first-time builds by teams without direct mail experience — default to geometric adstock with conservative decay settings that assume the bulk of the response occurs within one to two weeks.

For direct mail, that’s far too fast. In-home delivery alone takes three to eight business days depending on geography and mail class. Matchback conversion data from Postie campaigns consistently shows response curves extending 30 to 60 days post-drop, with meaningful conversion activity still registering in weeks five and six for certain verticals.

When the model’s decay function cuts off response attribution at day 14, it doesn’t just miss late converters — it assigns their conversions to whatever digital channel was running concurrently. The mail piece drove the action. The model credits the retargeting ad.

Problem 2: Saturation Curves Are Built for Impression Frequency, Not Household Reach

After adstock, Robyn passes each channel through the Hill function, where the alpha parameter controls the curve shape and gamma sets the inflection point — the math behind “the first dollar in a channel works harder than the hundred-thousandth.”

The default saturation priors assume each incremental impression to the same user delivers less marginal value. That’s reasonable when you’re serving the same display ad for the fifteenth time. Direct mail saturation works differently. It operates at the household level, and a second or third piece to the same address within a 60-day window often increases response rates rather than diminishing them — up to a frequency threshold that varies by vertical. The default Hill function priors penalize direct mail spend at volumes where the response is actually still climbing, prematurely flattening the channel’s modeled return curve and understating its value at current spend levels.

Problem 3: Mail-Plus-Digital Interaction Effects Are Ignored

None of the three open-source MMMs model channel interaction terms out of the box. For all-digital media mixes, the omission is suboptimal but manageable — the channels share similar response timelines and the misattribution tends to average out. For direct mail, the omission has real consequences.

Campaign data from Postie programs consistently shows that households exposed to both a mail piece and a digital retargeting sequence convert at meaningfully higher rates than those exposed to either channel alone. Without an interaction term, the model attributes the full lift to whichever channel’s adstock function captures the conversion timestamp. Given the decay defaults described above, that’s almost always the digital channel.

The net effect across all three problems: a model built to be channel-agnostic in theory becomes systematically biased against offline media in practice.

How to Correct the Model to Include Your Direct Mail Efforts

Fixing the undercount doesn’t require building a custom MMM from scratch. It requires three targeted modifications that any data scientist with access to the model code and direct mail campaign data can implement.

Modify 1: Adjust adstock decay to reflect direct mail’s actual response curve.

For Robyn, switch from geometric adstock to Weibull PDF — the framework supports it natively. Weibull PDF enables the peak response to arrive after the first period, which is the right shape for direct mail. Set the shape parameter to represent a peak response window of 10–21 days post-drop, with a tail extending to 45–60 days. For Meridian, adjust the carryover prior to reflect a longer half-life than the digital-calibrated default. For LightweightMMM, use the delayed adstock option with a peak lag of 12–18 days.

These ranges reflect what Postie matchback data shows across DTC, financial services, insurance, and home services verticals. Your own campaign data should inform the exact parameters, but the directional adjustment applies broadly: direct mail response curves are several times longer than digital defaults assume.

Modify 2: Reshape saturation priors for household-level frequency dynamics.

The Hill function’s inflection point parameter governs when diminishing returns kick in. Initialize it with a higher value for direct mail than you’d use for a display or paid social channel with equivalent spend. Direct mail response scales more linearly with volume than digital impression delivery, particularly when volume increases represent reaching new households rather than adding frequency to the same ones. This prevents the model from prematurely flattening direct mail’s response curve.

Modify 3: Add an explicit interaction term for mail-plus-digital exposure.

This takes the most effort but delivers the largest correction. Create a variable capturing overlapping exposure windows — specifically, households that received a mail piece and were also in the addressable audience for a concurrent digital campaign within 7–30 days of the mail drop. This can be constructed from CRM match data and digital campaign audience logs. Include it as an additional regressor in the model.

Without this term, the model doesn’t just undervalue direct mail, it artificially inflates the apparent contribution of concurrent digital channels, compounding the budget misallocation in both directions.

Calibrate Against Holdout-Based Ground Truth

The modifications above correct structural biases. To verify the model is now accurate, you need ground truth — and in direct mail, ground truth comes from holdout-based incrementality measurement.

Use holdout groups as calibration inputs. Postie campaigns can be configured with randomized holdout groups: households that meet all targeting criteria but are withheld from the mail drop. The conversion rate differential between the mailed group and the holdout is the closest thing available to a true incremental lift measurement. After fitting the model, compare its estimated contribution for direct mail against holdout-measured lift. If the model’s estimate is materially lower — which with default configurations, it almost always is — use the holdout result to inform a prior update and refit.

Calibrate at the campaign level, not the channel level. A single annual coefficient for “direct mail” blends prospecting campaigns and reactivation campaigns, which can show substantially different ROAS profiles. At minimum, break direct mail spend into two inputs (prospecting and retention) with separate adstock and saturation parameters for each.

Run a pre-budget validation. Before model outputs become allocations, translate the model’s predicted direct mail contribution into an implied incremental CPA and compare it to what you’ve observed in holdout-measured campaigns. If the model says direct mail’s incremental CPA is $95 but your holdout data shows $52, the model is wrong — and any reallocation based on that output moves budget away from a channel that’s outperforming what the model believes.

Give your modeling team channel-specific priors. Many mid-market teams outsource MMM builds to analytics consultants or rely on data scientists who’ve never run a direct mail campaign. They have no reason to question the defaults. The performance marketing team — the people who see matchback data and holdout results — need to provide explicit prior guidance. This isn’t overriding the model. It’s informing it with first-party data the modeler doesn’t have.

The Fix Is Straightforward — If You Make It Before Budgets Lock

Open-source MMMs aren’t biased against direct mail by design. They’re biased against it by default, which is a more solvable problem — but only if you address it before the model’s outputs harden into a media plan.

The three mechanisms of undercount are each individually correctable. Together, they can materially suppress direct mail’s modeled contribution relative to what holdout-based incrementality testing confirms the channel actually delivered. A channel running strong incremental CPAs in holdout-validated testing shouldn’t lose budget because the model’s default decay parameter was calibrated for paid search response times.

Adjust the decay. Reshape the saturation prior. Add the interaction term. Calibrate against holdout-measured ground truth. The frameworks support every one of these modifications. The only question is whether the team building the model has the channel-specific inputs to configure it correctly.

See how Postie’s matchback attribution and holdout methodology provide the ground truth your MMM needs →