The verdict on the model itself
Believe the direction. Don't quote the decimal.
12 checks ran on this fit. Six passed, five flagged, one failed. The failure is real, and it is why this page does not say “validated” — but none of it changes which channel worked.
6
passed
5
flagged
1
failed
The one that failed
Some of your channels ran too closely together to be told apart.
Measured 34.0081 against a line at 5.0000. Channel overlap, further down this page, names the pairs and says which one to deal with first.
The five that flagged
All of them are borderline rather than broken — worth knowing before someone argues with a slide, not worth refitting for.
It did not look long enough. Measured 184.7059 against a line at 400.0000.
The edges of the ranges are not settled. Measured 249.9867 against a line at 400.0000.
The search hit rough ground. Measured 0.0030 against a line at 0.0010.
It cannot even explain the history it was shown. Measured 6.0886 against a line at 10.0000.
Your data flatly contradicts something you told the model to expect. Measured 1.0000 against a line at 0.0000.
No out-of-sample test on this fit
Nothing was hidden from this model, so nothing on this page proves it can predict.
Every week of your data was used to fit it. The checks below still apply, but the one that matters most — does it work on weeks it has never seen — was not run. Refit with a holdout to get it.
What passed, in plain words
The full list, for whoever asks
13 checks in 4 groups — did the maths settle down, does it predict, could your data tell the channels apart, did your data actually teach the model anything. Each row further down carries the measured value, the threshold, and the engine’s own one-line explanation.
6 passed · 5 warned · 1 failed · 1 not run
In the engine’s order, always all 13, never a subset.
Open the full list ↓Weeks of evidence
156
Every week in every region the model was allowed to learn from.
Unknowns to solve for
32
Every separate quantity the model had to work out from your data.
Data per unknown
4.9×
How many weeks of evidence there are for each thing being worked out. Below about 4× you are asking more questions than your data can answer.
How wobbly the answer is
2
Above 30, changing one week of spend by a little moves the whole answer by a lot. Overlapping channels are what drive this up.
What we changed about this fit
Your data tripped a check, so the model was estimated differently. Your numbers were not touched — no imputing, no winsorising, no dropped rows. Each entry says what it cost.
Some channels held at their current budget
Appliedtriggered by vifRetail MediaAffiliateThese channels' spend does not vary enough, did not run for enough of the window, or moves too closely with the rest of the model for the data to say what they returned. Their numbers come mostly from the starting assumption. They stay in the model, so they do not distort the others, but the budget planner will not move them.
What it costs you
You will get no recommendation for these channels — deliberately. A plan that reallocated budget on the strength of an assumption would look exactly like one built on evidence.
Did the maths settle down?
The model does not solve your data in one pass — it searches for the answer from several starting points at once, and it is only finished when they all arrive at the same place. If they did not, nothing further down this page means anything.
- PassChain convergence (R-hat)
Worst R-hat 1.0096 at roi. Above 1.01 the chains have not mixed and the posterior is not yet the posterior.
measured 1.010 · threshold 1.010
- WarnBulk effective sample size
Lowest bulk ESS 185 at baseline_knots. Under 400 the point estimates carry visible Monte Carlo error.
measured 184.706 · threshold 400.000
- WarnTail effective sample size
Lowest tail ESS 250 at roi. Under 400 the interval endpoints carry visible Monte Carlo error.
measured 249.987 · threshold 400.000
- WarnDivergent transitions
3 divergences in 1000 draws (0.30%). Divergences mean the sampler could not follow the posterior's curvature, so the region it failed in is under-represented. Raise target_accept_prob (spec section 6) before trusting the numbers.
measured 0.003 · threshold 0.001
Does it predict?
Matching history proves very little: a flexible enough model can trace any past. What counts is the stretch of weeks we hid from it — and whether it beat the free answer of assuming this year looks like last year.
- WarnIn-sample accuracy
MAPE 6.1% (target <10%), R-squared 0.895 (target >0.9). In-sample fit is necessary, not sufficient — a model with enough knots fits anything.
measured 6.089 · threshold 10.000
- SkippedHoldout accuracy
No holdout was reserved. Fit with `data.with_holdout(n)` to get an out-of-sample number; in-sample fit alone cannot detect overfitting.
- PassPosterior-predictive check
Bayesian p-value 0.520 (target 0.05-0.95). High values mean the estimated noise is larger than the residuals — the model is not being held to account by the data.
measured 0.520 · threshold 0.050
- PassBaseline stays positive
P(baseline < 0) = 0.00 (target <0.2). A negative baseline is the classic symptom of over-attribution: the model has given media credit for more than the business would sell with no marketing at all.
measured 0.000 · threshold 0.200
Could your data tell the channels apart?
When two channels go up and down together, no amount of modelling can say which one did the work — the split you get back is the one we assumed going in. These checks measure that rather than hope about it.
- FailVariance inflation (post-transform)
Highest VIF 34.0 on Retail Media (warn 5, hard 10), measured after adstock and saturation because that is what the model sees.
measured 34.008 · threshold 5.000
- PassDesign conditioning
Condition number 2.3 (target <30) on the media block with the baseline and controls projected out.
measured 2.319 · threshold 30.000
- PassObservations per parameter
156 observations, 32 parameters, 4.9x (target >4).
measured 4.875 · threshold 4.000
Did your data actually teach the model anything?
The classic way marketing models mislead is by handing back the assumption they started with, dressed up as a finding. A channel is flagged here only when two separate signals agree the data never budged it.
- PassData moved the prior
Every channel's ROI posterior moved off its prior; the least informed narrowed by -571%. A channel whose posterior is its prior has been assumed, not estimated.
measured -5.713 · threshold 0.200
- WarnNo prior/data conflict
Outside the prior's 1st-99th percentile: Paid Search (13.66 vs [0.15, 9.91]). Either the prior is wrong or the data are being over-read; the fit is a compromise between two claims that disagree.
measured 1.000 · threshold 0.000
Channel overlap
How closely every pair of channels moved together. Measured after we account for advertising that keeps working after it stops running, and for the fact that the tenth million spent does less than the first — so this is the overlap the model has to live with, not the overlap you would see in a spend report.
| TV — Broadcast | — | 0.18 | 0.21 | 0.13 | -0.08 | -0.09 |
|---|---|---|---|---|---|---|
| Paid Social | 0.18 | — | 0.82 | 0.70 | 0.15 | -0.02 |
| Paid Search | 0.21 | 0.82 | — | 0.71 | 0.16 | 0.03 |
| Programmatic Display | 0.13 | 0.70 | 0.71 | — | 0.36 | -0.11 |
| Affiliate | -0.08 | 0.15 | 0.16 | 0.36 | — | 0.11 |
| Retail Media | -0.09 | -0.02 | 0.03 | -0.11 | 0.11 | — |
Read a cell as: when this channel went up, did that one go up too? 1.00 means always, 0.00 means never, −1.00 means the opposite. Anything above 0.80 and the two are one channel as far as your data is concerned — you can trust what the pair did together, but not the split between them.
How each channel spends
A channel that is always on at the same level gives the model almost nothing to learn from, however big its budget — it is the ups and downs that reveal what the spend does. This table shows which of your channels have them.
| Channel | Spend | Pattern | Weeks dark | Variation | VIF |
|---|---|---|---|---|---|
| TV — Broadcast | 33% | bursty | 60% | 1.38 | 1.1 |
| Paid Social | 23% | continuous | 0% | 0.34 | 3.8 |
| Paid Search | 20% | continuous | 0% | 0.44 | 6.8 |
| Programmatic Display | 13% | continuous | 0% | 0.32 | 3.2 |
| Affiliate | 7% | continuous | 0% | 0.01 | 2.4 |
| Retail Media | 4% | continuous since launch | 63% | 1.44 | 34.0 |