MMMM Studio

The verdict on the model itself

Believe the direction. Don't quote the decimal.

14 checks ran on this fit. 11 passed, two flagged, one failed. The failure is real, and it is why this page does not say “validated” — but none of it changes which channel worked.

11

passed

2

flagged

1

failed

The one that failed

The search hit rough ground.

Measured 0.0173 against a line at 0.0010. Fit this again — it usually clears. If it keeps happening you are asking the model for more detail than your data holds, and merging the channels that overlap is what fixes it.

The two that flagged

All of them are borderline rather than broken — worth knowing before someone argues with a slide, not worth refitting for.

The four searches did not quite agree. Measured 1.0104 against a line at 1.0100.

Some of your channels ran too closely together to be told apart. Measured 6.8043 against a line at 5.0000.

We are telling you this because nobody else will

On the 13 weeks we hid from the model, simply assuming “same week as last year” was more accurate than our forecast.

Our error was 11.1%. The free guess was 10.4%. That gap is small, and by the stricter week-on-week measure in the holdout check below the model is comfortably ahead — but on this particular test the simple answer won, and burying that would make everything else on this page worth less.

What it means for you: this is an attribution tool, not a forecasting tool. It is built to answer which channel caused what, and that answer does not depend on next quarter’s revenue being predicted correctly. If someone asks you to forecast with it, this is the number to show them.

The test: 13 weeks the model never saw

We removed them before fitting, then asked the finished model to predict them.

11.1%

our error

10.4%

free guess

5.5%

on weeks it did see

$1.6M$3.3M$5.0M13 WEEKS THE MODEL NEVER SAW →2022-01-032025-09-292025-12-22
What you actually soldWhat the model saidThe range it was prepared to be wrong byJust repeating the same week last yearThe 13 withheld weeks are drawn wider than the rest — they are 6% of the history and a third of the picture, because they are the whole test.

Holdout r² 0.234 · on weeks it did see 0.906 · free guess is “the same week 52 weeks ago

What passed, in plain words

The model isn't over-claiming. There is no chance it invented a negative baseline to hand media extra credit.
It isn't hiding behind noise. Its own estimate of how wrong it might be matches the errors it actually makes.
Your data moved every guess. No channel came back as the assumption you fed in, dressed up as a finding.
Nothing you believed was contradicted outright. Every result sits inside the range you would have called plausible before you started.
It predicts weeks it was never shown. Measured on data withheld before fitting, against the free answer of repeating last year.
It explains the history it was shown. Necessary but not sufficient — a flexible enough model traces any past. Only weeks withheld from it can settle that.
The answer doesn't wobble. Changing one week of spend a little does not move the whole result a lot.
It has room to work. There are several weeks of evidence for every separate quantity being worked out.
It looked long enough. Enough independent samples of the answer to make the middle of every range stable rather than lucky.
The edges are stable, not just the middles. The ends of every range are sampled as well as the centre — which is what makes a range a range.
It found returns it was never shown. This dataset was built from returns chosen in advance. The model was fitted without them and then marked against them — the one check that can call a marketing model right or wrong outright, and one no real dataset can offer.

The full list, for whoever asks

14 checks in 5 groups — did the maths settle down, does it predict, could your data tell the channels apart, did your data actually teach the model anything, on made-up data, did it find the answer we hid. Each row further down carries the measured value, the threshold, and the engine’s own one-line explanation.

11 passed · 2 warned · 1 failed

In the engine’s order, always all 14, never a subset.

Open the full list ↓

Weeks of evidence

208

Every week in every region the model was allowed to learn from.

Unknowns to solve for

34

Every separate quantity the model had to work out from your data.

Data per unknown

6.1×

How many weeks of evidence there are for each thing being worked out. Below about 4× you are asking more questions than your data can answer.

How wobbly the answer is

2

Above 30, changing one week of spend by a little moves the whole answer by a lot. Overlapping channels are what drive this up.

Did the maths settle down?

The model does not solve your data in one pass — it searches for the answer from several starting points at once, and it is only finished when they all arrive at the same place. If they did not, nothing further down this page means anything.

  • WarnChain convergence (R-hat)

    Worst R-hat 1.0104 at baseline_knots. Above 1.01 the chains have not mixed and the posterior is not yet the posterior.

    measured 1.010 · threshold 1.010

  • PassBulk effective sample size

    Lowest bulk ESS 450 at baseline_knots. Under 400 the point estimates carry visible Monte Carlo error.

    measured 450.303 · threshold 400.000

  • PassTail effective sample size

    Lowest tail ESS 902 at baseline_knots. Under 400 the interval endpoints carry visible Monte Carlo error.

    measured 901.990 · threshold 400.000

  • FailDivergent transitions

    69 divergences in 4000 draws (1.73%). Divergences mean the sampler could not follow the posterior's curvature, so the region it failed in is under-represented. Raise target_accept_prob (spec section 6) before trusting the numbers.

    measured 0.017 · threshold 0.001

Does it predict?

Matching history proves very little: a flexible enough model can trace any past. What counts is the stretch of weeks we hid from it — and whether it beat the free answer of assuming this year looks like last year.

  • PassIn-sample accuracy

    MAPE 5.5% (target <10%), R-squared 0.906 (target >0.9). In-sample fit is necessary, not sufficient — a model with enough knots fits anything.

    measured 5.453 · threshold 10.000

  • PassHoldout accuracy

    MAPE 11.1% (target <15%) and MASE 0.63 (target <1) on 13 held-out cells the likelihood never saw. MASE compares the error against forecasting the same period last year; above 1 the model has not earned its complexity out of sample.

    measured 11.090 · threshold 15.000

  • PassPosterior-predictive check

    Bayesian p-value 0.520 (target 0.05-0.95). High values mean the estimated noise is larger than the residuals — the model is not being held to account by the data.

    measured 0.520 · threshold 0.050

  • PassBaseline stays positive

    P(baseline < 0) = 0.00 (target <0.2). A negative baseline is the classic symptom of over-attribution: the model has given media credit for more than the business would sell with no marketing at all.

    measured 0.000 · threshold 0.200

Could your data tell the channels apart?

When two channels go up and down together, no amount of modelling can say which one did the work — the split you get back is the one we assumed going in. These checks measure that rather than hope about it.

  • WarnVariance inflation (post-transform)

    Highest VIF 6.8 on Paid Search (warn 5, hard 10), measured after adstock and saturation because that is what the model sees.

    measured 6.804 · threshold 5.000

  • PassDesign conditioning

    Condition number 2.0 (target <30) on the media block with the baseline and controls projected out.

    measured 1.999 · threshold 30.000

  • PassObservations per parameter

    208 observations, 34 parameters, 6.1x (target >4).

    measured 6.118 · threshold 4.000

Did your data actually teach the model anything?

The classic way marketing models mislead is by handing back the assumption they started with, dressed up as a finding. A channel is flagged here only when two separate signals agree the data never budged it.

  • PassData moved the prior

    Every channel's ROI posterior moved off its prior; the least informed narrowed by -396%. A channel whose posterior is its prior has been assumed, not estimated.

    measured -3.964 · threshold 0.200

  • PassNo prior/data conflict

    Every posterior mean ROI sits inside its prior's 1st-99th percentile.

    measured 0.000 · threshold 0.000

On made-up data, did it find the answer we hid?

Only possible on synthetic data, where we know the right answer because we chose it. Passing means the ranges came out wide enough to contain the truth on this dataset.

  • PassRecovers known ROI on synthetic data

    6/6 true ROIs inside the 90% interval (need 6), the median interval spanning 2.2x the channel's true return.

    measured 1.000 · threshold 0.900

How do I read this?

Channel overlap

How closely every pair of channels moved together. Measured after we account for advertising that keeps working after it stops running, and for the fact that the tenth million spent does less than the first — so this is the overlap the model has to live with, not the overlap you would see in a spend report.

TV — Broadcast0.060.180.190.040.08
Video — CTV0.060.170.120.220.11
Paid Social0.180.170.780.51-0.04
Paid Search0.190.120.780.580.09
Programmatic Display0.040.220.510.58-0.02
Out of Home0.080.11-0.040.09-0.02

Read a cell as: when this channel went up, did that one go up too? 1.00 means always, 0.00 means never, −1.00 means the opposite. Anything above 0.80 and the two are one channel as far as your data is concerned — you can trust what the pair did together, but not the split between them.

How do I read this?

How each channel spends

A channel that is always on at the same level gives the model almost nothing to learn from, however big its budget — it is the ups and downs that reveal what the spend does. This table shows which of your channels have them.

ChannelSpendPatternWeeks darkVariationVIF
TV — Broadcast28%bursty62%1.371.2
Video — CTV16%variable0%0.603.2
Paid Social19%continuous0%0.333.0
Paid Search17%continuous0%0.426.8
Programmatic Display11%continuous0%0.321.9
Out of Home9%bursty56%1.171.2