Direct Mail Attribution: Promo Codes Lie, Holdouts Don't

Promo codes measure redemption, not causation. See the geo holdout design BlueAlpha uses to get a real cost per incremental order on direct mail.

Incrementality Testing

Measurement Blind Spots

Direct mail attribution built on promo codes measures redemption, not causation. A code tells you which customers used it. It cannot tell you whether the mailer created the order, because the people most likely to redeem a code are the people already closest to buying. The fix is a geo holdout: withhold mail from a matched set of markets, model what those markets would have done with mail, and measure the gap. This guide covers why code-based measurement misreads offline spend, how to design a direct mail holdout, and how to read the result.


What Promo Code Attribution Actually Measures

Promo code attribution is last-click attribution with an extra step. When a customer enters MAIL25 at checkout, your system records a direct mail conversion and assigns the full order to the mailer. That is a redemption event, and redemption is a coupon-usage metric rather than a causal one.

Redemption and causation come apart in a specific direction. Coupon codes are redeemed mostly by people who were already going to buy and found a discount waiting. A high redemption count is therefore consistent with two very different worlds: a mailer that generated demand, and a mailer that discounted demand you already had. The code cannot distinguish them, and neither can any dashboard built on top of it.

This produces a familiar pattern inside growth teams. Direct mail shows up as the most efficient channel in the reporting, gets treated as lower funnel because its conversions look immediate and cheap, and becomes politically difficult to cut. Meanwhile the digital channels that did the work of building consideration receive nothing, because the customer's fifteen digital touchpoints leave no trace at the moment the code is typed.


The Three Ways a Promo Code Misreads Direct Mail

It overstates direct mail by absorbing credit from everything upstream. A customer who saw four paid social ads, searched your brand twice, and then received a mailer with an offer is recorded as a direct mail acquisition. Every channel that moved that customer toward the purchase is recorded as having contributed nothing. The error compounds in mixed-media programs, because the more digital pressure you apply in a market, the more redemptions your mail appears to generate there.

It understates direct mail by missing everyone who ignored the code. Recipients who saw the mailer, remembered the brand, and later bought at full price through search or direct traffic are invisible to code-based measurement. In programs where the offer is weak or the audience is already brand-aware, this can be the larger share of the mailer's real effect. The result is a channel whose reported performance depends heavily on offer strength rather than on demand creation.

It cannot price the counterfactual. That is the only number that matters for a budget decision. The question a CFO is asking is not how many codes came back. It is what would have happened to revenue in those markets if the mail had never landed. A redemption count contains no information about that, so it cannot support a decision to scale, hold, or cut. Teams end up arguing from priors: one side points at cost per piece, the other points at redemption volume, and neither has evidence that answers the question.


Geo Holdouts Answer the Question Promo Codes Cannot

A geo holdout test constructs the missing counterfactual. You withhold direct mail from a set of markets, keep mailing everywhere else, and use the withheld markets to model what the mailed markets would have done without the drop. This is the same matched market testing design used to measure paid channels, applied to an offline one. The difference between modeled and actual performance is the incremental effect of the mail.

Direct mail is well suited to this design, better suited than most digital channels. Mail is bought and delivered by geography already, so a holdout requires suppressing a list segment rather than fighting a platform's targeting controls. Drop dates are known in advance, which gives you a clean pre-period and a defined in-home window to measure against. Your list tells you exactly who was mailed and when. Most of the friction that makes geo testing awkward on social platforms does not exist here.

The measurement target should be your own transaction data rather than any channel-side report. That distinction is the difference between asking the mailer whether it worked and checking the cash register. BlueAlpha ran a geo holdout on a paid channel for the beverage brand Cann. The design deliberately measured against Shopify transactions rather than the platform's dashboard or a third-party attribution tool. Both are modeled estimates layered on top of each other. The same principle applies to mail: measure orders, revenue, and new customers as your systems actually recorded them.


How to Design a Direct Mail Geo Holdout in Six Steps

Define the outcome before the design. Pick the business metric the mail is supposed to move, at the level your finance team uses. Gross order volume, net new customers, and revenue are defensible. Code redemptions are not, because they are the thing being tested.

Choose holdout markets that match on pre-period behavior, not on size alone. Matched markets need similar baseline trajectories in the outcome metric over the weeks before the drop. Two DMAs (designated market areas) with identical population and very different seasonality make a poor pair. Where your footprint is narrow, zip clusters can substitute for DMAs, provided each cluster carries enough weekly volume to detect the effect you expect.

Run a power analysis before committing budget. Power analysis is the step that separates a usable test from an expensive one, and it is standard practice in incrementality testing across any channel. The analysis tells you whether a drop of the size you are planning can produce a detectable signal at all. Running an underpowered test is the most expensive mistake available here, because it consumes a full mail cycle and returns an ambiguous result that changes no one's mind.

Suppress the holdout at the list level and document it. The holdout markets receive no mail for the full test window, including evergreen and reactivation drops that would otherwise contaminate the read. Cross-channel activity should be held steady across both groups; a concurrent promotion running only in mailed markets will show up as mail lift.

Set the measurement window around the in-home date, not the drop date. Mail has a delivery lag and a response tail. The window opens when pieces land and stays open long enough to capture delayed response, which for considered purchases runs well past the first week.

Fix the decision rule in writing before results arrive. State in advance what result scales the program, what result holds it flat, and what result cuts it. Pre-committing removes the temptation to reinterpret an inconvenient interval after the fact.


What a Holdout Found on a $300K Offline Campaign

BlueAlpha ran the first statistically validated measurement of beehiiv's $300,000 New York City subway campaign using a geo holdout design and Bayesian structural time series modeling. New York was the treatment geography, comparable cities served as controls, and the model measured impact at each funnel stage. The test detected roughly 100,000 incremental new website users at about $4 per incremental visitor, at 95% confidence.

The more useful part of that result is what happened further down the funnel. The same test measured just over 100 incremental free signups at roughly $2,700 each, and 15 to 20 incremental paid purchases at roughly $17,000 each. Against beehiiv's lifetime value of $1,250 per net new purchase, the campaign returned an estimated $19,000 to $25,000 in customer value on a $300,000 investment. The channel was genuinely working at the top of the funnel and genuinely uneconomic at the bottom, and only a stage-by-stage causal read could show both at once.

That campaign was out-of-home rather than direct mail, and the numbers belong to subway advertising. The design is what transfers. An offline medium delivered by geography, with a known exposure window and a synthetic control built from matched markets, is one measurement problem. That does not change whether the medium is a subway platform or a mailbox. beehiiv's own case study names direct mail among the channels the approach covers, and our playbook on measuring out-of-home advertising walks through the full workflow step by step.


Reading the Result: Lift, Intervals, and What to Do

A causal test returns a distribution rather than a single number, and the interval is where the decision lives. The rule is simple. If the credible interval sits entirely above zero, the mail is producing lift. If it sits entirely below zero, the mail is destroying value. If it spans zero, the measured effect cannot be distinguished from no effect, regardless of how encouraging the point estimate looks in isolation.

Cann's AppLovin test is the clearest illustration of why the interval matters more than the headline. The point estimate for incremental gross sales came back at positive $90,000, which reads as a win. The 95% credible interval ran from negative $152,000 to positive $310,000, which meant the result was no different from zero. The team paused the channel and reallocated roughly $480,000 a year into channels where performance could be proven. The 19-day test cost about $30,000 in media and returned roughly 16x in year one.

Direct mail adds one more layer to the read. Because the point of the exercise is usually a reallocation decision, the number to compute is cost per incremental order, not cost per redemption. Suppose a drop generates 1,000 redemptions and the holdout shows the mail caused 300 net new orders. Your true cost per incremental order is then more than three times what the redemption-based report displays. Those figures are arithmetic illustration rather than a benchmark. The ratio for your program is whatever your own test returns, and that is the only version worth putting in front of finance.


Keep the Promo Code, Change What You Conclude From It

None of this is an argument for removing codes from your creative. Codes drive response, they let you vary offers by segment, and their redemption data is genuinely useful for understanding which audiences engage with which offer. The change is in what you allow the redemption number to decide.

Treat redemption as a response metric that tells you how an offer performed. Treat incremental lift, measured against a holdout, as the performance metric that decides budget. Once a holdout has run, the ratio between the two becomes a calibration factor you can apply to subsequent drops without retesting every cycle. Periodic retests keep that factor honest as your list, offer, and market mix change.

This is the same discipline serious teams apply to every channel that claims credit for itself. A platform reports the conversions it can see and has an incentive to see generously. A promo code reports the redemptions it can see and has no incentive at all, which somehow makes it feel more trustworthy. Both are answering a question about visibility. Only a holdout answers the question about causation, which is the one your budget actually turns on.


Get a Direct Mail Test Design for Your Next Drop

You may be running direct mail at meaningful scale with redemption volume as your only read on it. If so, one geo holdout on the next drop will tell you more than another year of code-based reporting. BlueAlpha designs the holdout, runs the power analysis against your own volumes, and produces the market split and decision rule before the mail goes out. The result is then measured against your transaction data and fed back into the causal measurement model that sets your budget. Book a strategy call to scope a test on your next mail cycle.


Methodology and sources:

The beehiiv figures come from BlueAlpha's OOH measurement case study, a geo holdout with Bayesian structural time series modeling and synthetic control validation, reported at 95% confidence. The Cann figures come from BlueAlpha's AppLovin geo holdout case study, a 19-day test measured against Shopify transaction data at a 95% credible interval. Both are paid and out-of-home channels rather than direct mail; they are cited here for test design and result interpretation, not as direct mail results. The 1,000-redemption example is arithmetic illustration, not a benchmark. BlueAlpha does not publish response-rate or cost-per-piece norms for direct mail, because those vary too widely by list, offer, and vertical to be useful as planning inputs.


FAQ

Why are promo codes unreliable for direct mail attribution?

Promo codes measure redemption, which is a coupon-usage metric, not a causal one. A code records that a customer used it, not that the mailer caused the purchase, and codes are redeemed disproportionately by customers who were already close to buying. This overstates direct mail by absorbing credit from upstream digital touchpoints and simultaneously understates it by missing recipients who bought later at full price without the code.

How do I accurately measure direct mail ROI?

Measure direct mail with a geo holdout test. Withhold mail from a set of markets matched to your mailed markets on pre-period performance. Model what the mailed markets would have done without the drop, then measure the difference in orders or revenue as your own systems recorded it. The output is cost per incremental order, which is the figure that supports a budget decision, unlike cost per redemption.

What is a geo holdout test for direct mail?

A geo holdout test for direct mail suppresses mail in a defined set of geographies while continuing to mail everywhere else. The suppressed markets are then used to build a statistical counterfactual for the mailed markets. Direct mail suits this design well because it is already bought by geography, drop dates are known in advance, and the mailed list defines exposure precisely. The measured gap between actual and counterfactual performance is the mail's incremental contribution.

Can I run a direct mail holdout if I only operate in a few markets?

Yes, provided each unit of geography carries enough weekly conversion volume for the test to detect the effect you expect. Businesses operating in a handful of cities can run holdouts at the zip cluster level instead of the DMA level. Zips are grouped into treatment and control sets that match on pre-period trajectory. A power analysis run before the test determines whether your volumes support the design or whether you need a longer window or a larger drop.

How long does a direct mail incrementality test take?

The measurement window should open at the in-home date rather than the drop date. It stays open long enough to capture the response tail, which for considered purchases extends well beyond the first week. Add a pre-period of several weeks so the model can learn the relationship between treatment and control markets before the mail lands. BlueAlpha's 19-day geo holdout for Cann shows that a well-powered test on a high-volume channel can resolve in under three weeks.

Should I stop putting promo codes on direct mail?

No. Codes drive response and their redemption data is useful for comparing offers across segments. What changes is the conclusion you draw from them. Redemption becomes a response metric that tells you how an offer performed. Incremental lift, measured against a holdout, becomes the performance metric that sets budget.

What is the difference between direct mail attribution and direct mail incrementality?

Direct mail attribution assigns credit for orders that have already been recorded, usually through promo code redemption or a matchback against the mailed list. Direct mail incrementality measures how many of those orders would not have happened without the mail, which requires a control group that never received it. Attribution can be computed from data you already have; incrementality requires a deliberately designed experiment.

See which of your marketing dollars are actually working.

See which of your marketing dollars are actually working.

See which of your marketing dollars are actually working.