MMMM Studio

The headline, in one sentence

Media produced about $153.2M of your $259.3M — roughly 59 cents in every dollar. Of your seven measurable channels, none has yet shown it paid for itself and seven genuinely cannot be called yet.

That second half is not a hedge and it is not a fault in the model. Two years of weekly history is simply not enough evidence to separate nine channels that mostly ran at the same time. Below, the channels are sorted by how much we actually know about them — not by how good their number looks.

We can't tell “Search — Brand” apart from “Search — Non-brand”

Their spend went up and down together almost every week of this period — that is the two lines below, and they are very nearly one line. When two channels move as one, nothing in the data can say which of them made the sale, so we are leaving both unscored rather than showing you a number we would have to take back. Those channels account for 82% of everything media did in this period, so read the headline below as the most media could have done, not as what it did.

Search — BrandSearch — Non-brand

Run them apart for about six weeks — keep one at its normal level and pause or cut the other — then fit again. Once they stop moving together we can tell you what each one is worth. Everything else on this page is unaffected.

Show me the statistics

Withheld because the sampler’s chains never agreed: Search — Brand, Search — Non-brand. The counts and the return ranking below are computed over the 7 that did converge. The contribution split still lists the withheld rows, because that table has to add up to the observed outcome, but those rows are as arbitrary as the returns withheld alongside them.

They carry 82% of media contribution, so the media share and the blended return are largely a restatement of them — an upper bound on what this fit knows, not a result. The validation tab carries the per-channel R-hat and effective sample size.

How do I read this?

Two years of revenue, and where it came from

Each band is one channel's contribution, stacked on the business you'd have had without any media at all. The white line is what you actually sold.

$259.3M

actually sold

$4.0M/WK13 WEEKS WE HID FROM THE MODEL →20222023

The pale mass underneath is the business you’d have had anyway: $110.5M, two fifths of everything.

Search — Non-brand40.8%
Search — Brand6.9%
Social — Meta2.5%
TV — Broadcast1.9%
TV — Cable1.3%
Video — CTV1.3%
Social — TikTok1.3%
Programmatic Display1.2%
Out of Home0.7%
Everything else41.9%
±media total 47.0%–70.2%
How do I read this?
Nothing is settled yet

Not one of your seven channels has a range that sits entirely above break-even.

That is a real result, not a missing one: it says your data cannot yet separate what each channel did from what the others did at the same time. Nothing below should be cut on this evidence, and nothing below should be scaled on it either.

What would settle it: one channel deliberately switched off, or moved out of step with the others, for six weeks or more. The evidence tab shows how little each range moved away from what you assumed going in.

What each dollar came back as

Sorted by how much your data actually pinned down. The dotted line is break-even — a band crossing it means we can't tell you whether it won or lost.

TV — BroadcastCan't call it$0.86

$0.27 to $2.07 · best-measured channel you have, and it still straddles break-even

TV — CableCan't call it$1.06

$0.29 to $2.99

Out of HomeCan't call it$1.21

$0.28 to $4.04

Video — CTVCan't call it$1.16

$0.25 to $4.60

Social — MetaCan't call it$1.27

$0.28 to $5.45

Programmatic DisplayCan't call it$1.22

$0.28 to $5.46 · widest range on the screen relative to its size

Social — TikTokCan't call it$1.22

$0.29 to $5.36

Search — Non-brand

withheldwe got a different answer every time we tried

Search — Brand

withheldwe got a different answer every time we tried

Not this

“TV — Broadcast loses money.” Its best guess is $0.86, but the range runs to $2.07. The screen is not allowed to convert a median below break-even into a verdict — the same rule that stops it calling the seven maybes wins.

How do I read this?

Do this first

Don't cut anything on this evidence.

Nothing here has shown it lost money — the ranges simply cross break-even in both directions. A cut made on this page would be a cut made on noise, and next year would look like proof that the cut worked.

Do this next

Run one lift test on a Social — Meta burst. It is worth more than another year of data.

Social — Meta and Social — TikTok each switch on once across the whole window, and that is all the model has to learn from. One deliberate on/off test would narrow both faster than waiting.

Don't do this

Don't rank the seven uncertain channels against each other.

Their ranges overlap almost completely. Any ordering you read off the medians is noise, and it will reverse itself next quarter. If someone asks for a ranking, show them this panel.

13 things would make next quarter’s answer sharper. 8 of them are ours.

Everything the run flagged, in one place and sorted by who actually does the work. Nothing in the first two columns is generic — each item is the engine’s own words about a specific finding in your data. The third column is opinion, and is drawn so you can tell at a glance.

Email this list

Your media team

changes to how you buy, not to the data

Two channels are near-duplicates

These two were bought as one thing, so model them as one thing: combine the columns. If you need them split, the only way to get there is an experiment that moves one without the other.

FROM · collinearity · 1 vs 0.9

TV — Broadcast is off air most weeks

TV — Cable is off air most weeks

Out of Home is off air most weeks

Nothing to fix in the data — this is a media plan, not a defect. Expect a wide range on this channel and weight it accordingly; a lift test on a single burst would pin it faster than more history would.

FROM · burst_flighting · TV — Broadcast, TV — Cable, Out of Home

Severe variance inflation

Usually the same cause as near-duplicate channels: this one is reconstructable from the others. Combine it with whichever channel it shadows, or drop a control that is standing in for it. A channel that launched partway through the window hits this too, because the baseline absorbs the period it was dark.

FROM · vif · TV — Cable · 5170.9 vs 10

Us

nothing for you to do — one click and we rerun

Failed checkChain convergence (R-hat)

Worst R-hat 1.2264 at roi. Above 1.01 the chains have not mixed and the posterior is not yet the posterior.

FROM · scorecard · rhat · 1.23 vs 1.01

Failed checkBulk effective sample size

Lowest bulk ESS 13 at beta. Under 400 the point estimates carry visible Monte Carlo error.

FROM · scorecard · ess_bulk · 13.04 vs 400

Failed checkTail effective sample size

Lowest tail ESS 16 at beta. Under 400 the interval endpoints carry visible Monte Carlo error.

FROM · scorecard · ess_tail · 16.41 vs 400

Runs the same fit with twice the draws (1000 → 2000), and a more careful search (target_accept_prob 0.9 → 0.95). Nothing else moves — same data, same priors, same holdout — so the two are comparable. If the answer shifts, that shift is itself the finding.

WarningIn-sample accuracy

MAPE 6.8% (target <10%), R-squared 0.879 (target >0.9). In-sample fit is necessary, not sufficient — a model with enough knots fits anything.

FROM · scorecard · in_sample · 6.79 vs 10

WarningHoldout accuracy

MAPE 15.5% (target <15%) and MASE 0.62 (target <1) on 13 held-out cells the likelihood never saw. MASE compares the error against forecasting the same period last year; above 1 the model has not earned its complexity out of sample.

FROM · scorecard · holdout · 15.5 vs 15

Failed checkVariance inflation (post-transform)

Highest VIF 5170.9 on TV — Cable (warn 5, hard 10), measured after adstock and saturation because that is what the model sees.

FROM · scorecard · vif · 5170.9 vs 5

Failed checkDesign conditioning

Condition number 216.3 (target <30) on the media block with the baseline and controls projected out.

FROM · scorecard · condition_number · 216.32 vs 30

WarningObservations per parameter

104 observations, 38 parameters, 2.7x (target >4).

FROM · scorecard · degrees_of_freedom · 2.74 vs 4

WarningData moved the prior

Indistinguishable from the prior: Social — TikTok (contraction 2%, KS p=0.50), Search — Brand (contraction -12816%, KS p=0.24), Programmatic Display (contraction 3%, KS p=0.20). A channel whose posterior is its prior has been assumed, not estimated.

FROM · scorecard · prior_shift · -128.16 vs 0.2

WarningNo prior/data conflict

Outside the prior's 1st-99th percentile: Search — Brand (10.92 vs [0.15, 9.91]), Search — Non-brand (42.94 vs [0.15, 9.91]). Either the prior is wrong or the data are being over-read; the fit is a compromise between two claims that disagree.

FROM · scorecard · prior_conflict · 2 vs 0

Those are a property of the data and the priors rather than of the search, so running it again will not move them. Listed here so nobody goes looking for a fix that does not exist.

Your file

Thin degrees-of-freedom budget

Merging two closely related channels would buy back headroom. Otherwise expect wide intervals and treat them as the honest answer rather than a defect.

FROM · dof · 2.74 vs 4

Your file

Ill-conditioned design matrix

The design as a whole is near-singular, so no single column is to blame. Simplify: fewer channels, fewer controls, or a longer window.

FROM · condition · 216.32 vs 30

General advice — not from your data

true of most MMMs, not measured on this one

Keep a channel dark somewhere, on purpose

The single cheapest thing you can do for next year's model is leave one region or one fortnight without a channel that normally runs everywhere. Continuous channels are the hardest to measure precisely because they never stop.

Record promotions and price changes weekly

A column of promotion weeks and a column of average price cost nothing to keep, and their absence is the usual reason media ends up with the credit for a price cut. Most uploads arrive without them.

This column is separated on purpose. Everything in the other two is traceable to a finding in your run; nothing in this one is.

Stamped so this can be reproduced

Every fit is written to disk with the data it read and the version that read it, so a number in last quarter’s deck can be traced to the run that produced it — and that run can be produced again. Nothing on these pages is computed when you open them.

run
20260831-122256-16895b
data
d1639c05b048
engine
0.1.0
holdout
13 periods
dataset
hard_case
fitted
31 Aug 2026
took
1 min
name
Demo run

Past runs

Every fit is kept, with its verdict and its date, so a number in last quarter’s deck can always be traced back to the run that produced it.

This is the only fit on record. The second one is where this list starts earning its space.

How do I read this?