The headline, in one sentence
Media produced about $153.2M of your $259.3M — roughly 59 cents in every dollar. Of your seven measurable channels, none has yet shown it paid for itself and seven genuinely cannot be called yet.
That second half is not a hedge and it is not a fault in the model. Two years of weekly history is simply not enough evidence to separate nine channels that mostly ran at the same time. Below, the channels are sorted by how much we actually know about them — not by how good their number looks.
We can't tell “Search — Brand” apart from “Search — Non-brand”
Their spend went up and down together almost every week of this period — that is the two lines below, and they are very nearly one line. When two channels move as one, nothing in the data can say which of them made the sale, so we are leaving both unscored rather than showing you a number we would have to take back. Those channels account for 82% of everything media did in this period, so read the headline below as the most media could have done, not as what it did.
Run them apart for about six weeks — keep one at its normal level and pause or cut the other — then fit again. Once they stop moving together we can tell you what each one is worth. Everything else on this page is unaffected.
Show me the statistics
Withheld because the sampler’s chains never agreed: Search — Brand, Search — Non-brand. The counts and the return ranking below are computed over the 7 that did converge. The contribution split still lists the withheld rows, because that table has to add up to the observed outcome, but those rows are as arbitrary as the returns withheld alongside them.
They carry 82% of media contribution, so the media share and the blended return are largely a restatement of them — an upper bound on what this fit knows, not a result. The validation tab carries the per-channel R-hat and effective sample size.
Two years of revenue, and where it came from
Each band is one channel's contribution, stacked on the business you'd have had without any media at all. The white line is what you actually sold.
$259.3M
actually sold
The pale mass underneath is the business you’d have had anyway: $110.5M, two fifths of everything.
Not one of your seven channels has a range that sits entirely above break-even.
That is a real result, not a missing one: it says your data cannot yet separate what each channel did from what the others did at the same time. Nothing below should be cut on this evidence, and nothing below should be scaled on it either.
What would settle it: one channel deliberately switched off, or moved out of step with the others, for six weeks or more. The evidence tab shows how little each range moved away from what you assumed going in.
What each dollar came back as
Sorted by how much your data actually pinned down. The dotted line is break-even — a band crossing it means we can't tell you whether it won or lost.
$0.27 to $2.07 · best-measured channel you have, and it still straddles break-even
$0.29 to $2.99
$0.28 to $4.04
$0.25 to $4.60
$0.28 to $5.45
$0.28 to $5.46 · widest range on the screen relative to its size
$0.29 to $5.36
withheld — we got a different answer every time we tried
withheld — we got a different answer every time we tried
“TV — Broadcast loses money.” Its best guess is $0.86, but the range runs to $2.07. The screen is not allowed to convert a median below break-even into a verdict — the same rule that stops it calling the seven maybes wins.
Do this first
Don't cut anything on this evidence.
Nothing here has shown it lost money — the ranges simply cross break-even in both directions. A cut made on this page would be a cut made on noise, and next year would look like proof that the cut worked.
Do this next
Run one lift test on a Social — Meta burst. It is worth more than another year of data.
Social — Meta and Social — TikTok each switch on once across the whole window, and that is all the model has to learn from. One deliberate on/off test would narrow both faster than waiting.
Don't do this
Don't rank the seven uncertain channels against each other.
Their ranges overlap almost completely. Any ordering you read off the medians is noise, and it will reverse itself next quarter. If someone asks for a ranking, show them this panel.
13 things would make next quarter’s answer sharper. 8 of them are ours.
Everything the run flagged, in one place and sorted by who actually does the work. Nothing in the first two columns is generic — each item is the engine’s own words about a specific finding in your data. The third column is opinion, and is drawn so you can tell at a glance.
Your media team
changes to how you buy, not to the data
Two channels are near-duplicates
These two were bought as one thing, so model them as one thing: combine the columns. If you need them split, the only way to get there is an experiment that moves one without the other.
FROM · collinearity · 1 vs 0.9
TV — Broadcast is off air most weeks
TV — Cable is off air most weeks
Out of Home is off air most weeks
Nothing to fix in the data — this is a media plan, not a defect. Expect a wide range on this channel and weight it accordingly; a lift test on a single burst would pin it faster than more history would.
FROM · burst_flighting · TV — Broadcast, TV — Cable, Out of Home
Severe variance inflation
Usually the same cause as near-duplicate channels: this one is reconstructable from the others. Combine it with whichever channel it shadows, or drop a control that is standing in for it. A channel that launched partway through the window hits this too, because the baseline absorbs the period it was dark.
FROM · vif · TV — Cable · 5170.9 vs 10
Us
nothing for you to do — one click and we rerun
Worst R-hat 1.2264 at roi. Above 1.01 the chains have not mixed and the posterior is not yet the posterior.
FROM · scorecard · rhat · 1.23 vs 1.01
Lowest bulk ESS 13 at beta. Under 400 the point estimates carry visible Monte Carlo error.
FROM · scorecard · ess_bulk · 13.04 vs 400
Lowest tail ESS 16 at beta. Under 400 the interval endpoints carry visible Monte Carlo error.
FROM · scorecard · ess_tail · 16.41 vs 400
Runs the same fit with twice the draws (1000 → 2000), and a more careful search (target_accept_prob 0.9 → 0.95). Nothing else moves — same data, same priors, same holdout — so the two are comparable. If the answer shifts, that shift is itself the finding.
MAPE 6.8% (target <10%), R-squared 0.879 (target >0.9). In-sample fit is necessary, not sufficient — a model with enough knots fits anything.
FROM · scorecard · in_sample · 6.79 vs 10
MAPE 15.5% (target <15%) and MASE 0.62 (target <1) on 13 held-out cells the likelihood never saw. MASE compares the error against forecasting the same period last year; above 1 the model has not earned its complexity out of sample.
FROM · scorecard · holdout · 15.5 vs 15
Highest VIF 5170.9 on TV — Cable (warn 5, hard 10), measured after adstock and saturation because that is what the model sees.
FROM · scorecard · vif · 5170.9 vs 5
Condition number 216.3 (target <30) on the media block with the baseline and controls projected out.
FROM · scorecard · condition_number · 216.32 vs 30
104 observations, 38 parameters, 2.7x (target >4).
FROM · scorecard · degrees_of_freedom · 2.74 vs 4
Indistinguishable from the prior: Social — TikTok (contraction 2%, KS p=0.50), Search — Brand (contraction -12816%, KS p=0.24), Programmatic Display (contraction 3%, KS p=0.20). A channel whose posterior is its prior has been assumed, not estimated.
FROM · scorecard · prior_shift · -128.16 vs 0.2
Outside the prior's 1st-99th percentile: Search — Brand (10.92 vs [0.15, 9.91]), Search — Non-brand (42.94 vs [0.15, 9.91]). Either the prior is wrong or the data are being over-read; the fit is a compromise between two claims that disagree.
FROM · scorecard · prior_conflict · 2 vs 0
Those are a property of the data and the priors rather than of the search, so running it again will not move them. Listed here so nobody goes looking for a fix that does not exist.
Your file
Thin degrees-of-freedom budget
Merging two closely related channels would buy back headroom. Otherwise expect wide intervals and treat them as the honest answer rather than a defect.
FROM · dof · 2.74 vs 4
Your file
Ill-conditioned design matrix
The design as a whole is near-singular, so no single column is to blame. Simplify: fewer channels, fewer controls, or a longer window.
FROM · condition · 216.32 vs 30
General advice — not from your data
true of most MMMs, not measured on this one
Keep a channel dark somewhere, on purpose
The single cheapest thing you can do for next year's model is leave one region or one fortnight without a channel that normally runs everywhere. Continuous channels are the hardest to measure precisely because they never stop.
Record promotions and price changes weekly
A column of promotion weeks and a column of average price cost nothing to keep, and their absence is the usual reason media ends up with the credit for a price cut. Most uploads arrive without them.
This column is separated on purpose. Everything in the other two is traceable to a finding in your run; nothing in this one is.
Stamped so this can be reproduced
Every fit is written to disk with the data it read and the version that read it, so a number in last quarter’s deck can be traced to the run that produced it — and that run can be produced again. Nothing on these pages is computed when you open them.
- run
- 20260831-122256-16895b
- data
- d1639c05b048
- engine
- 0.1.0
- holdout
- 13 periods
- dataset
- hard_case
- fitted
- 31 Aug 2026
- took
- 1 min
- name
- Demo run
Past runs
Every fit is kept, with its verdict and its date, so a number in last quarter’s deck can always be traced back to the run that produced it.
This is the only fit on record. The second one is where this list starts earning its space.