The verdict on the model itself
Nothing on this page is safe to read yet.
13 checks ran on this fit. Three passed, five flagged, five failed. Every other number in this report is measured on top of the one that failed, so there is not much point reading on until it is fixed.
3
passed
5
flagged
5
failed
Chain convergence (R-hat)
The four searches did not quite agree.
Measured 1.2264 against a line at 1.0100. Fit this again and let it run longer. That is almost always all it needs.
Bulk effective sample size
It did not look long enough.
Measured 13.0385 against a line at 400.0000. Fit this again and let it run longer. That is almost always all it needs.
Tail effective sample size
The edges of the ranges are not settled.
Measured 16.4101 against a line at 400.0000. Fit this again and let it run longer. That is almost always all it needs.
Variance inflation (post-transform)
Some of your channels ran too closely together to be told apart.
Measured 5170.8980 against a line at 5.0000. Channel overlap, further down this page, names the pairs and says which one to deal with first.
Design conditioning
The answer wobbles — one week of spend moves the whole result.
Measured 216.3155 against a line at 30.0000. Channel overlap, further down this page, names the pairs and says which one to deal with first.
The five that flagged
All of them are borderline rather than broken — worth knowing before someone argues with a slide, not worth refitting for.
It cannot even explain the history it was shown. Measured 6.7887 against a line at 10.0000.
It does not predict weeks it was never shown. Measured 15.4981 against a line at 15.0000.
You are asking more questions than your data can answer. Measured 2.7368 against a line at 4.0000.
Your data barely moved the model off what it already expected. Measured -128.1594 against a line at 0.2000.
Your data flatly contradicts something you told the model to expect. Measured 2.0000 against a line at 0.0000.
On the 13 weeks we hid from the model, our forecast beat simply assuming “same week as last year”.
Our error was 15.5%. The free guess was 18.0% — a gap of 2.5 percentage points on weeks the model was never allowed to see. That is the version of this comparison worth quoting, and it is the only one that lets the rest of this page be quoted too.
What it means for you: this is an attribution tool, not a forecasting tool. It is built to answer which channel caused what, and that answer does not depend on next quarter’s revenue being predicted correctly. If someone asks you to forecast with it, this is the number to show them.
The test: 13 weeks the model never saw
We removed them before fitting, then asked the finished model to predict them.
15.5%
our error
18.0%
free guess
6.8%
on weeks it did see
Holdout r² -0.965 · on weeks it did see 0.879 · free guess is “the same week 52 weeks ago”
What passed, in plain words
The full list, for whoever asks
14 checks in 5 groups — did the maths settle down, does it predict, could your data tell the channels apart, did your data actually teach the model anything, on made-up data, did it find the answer we hid. Each row further down carries the measured value, the threshold, and the engine’s own one-line explanation.
3 passed · 5 warned · 5 failed · 1 not run
In the engine’s order, always all 14, never a subset.
Open the full list ↓Weeks of evidence
104
Every week in every region the model was allowed to learn from.
Unknowns to solve for
38
Every separate quantity the model had to work out from your data.
Data per unknown
2.7×
How many weeks of evidence there are for each thing being worked out. Below about 4× you are asking more questions than your data can answer.
How wobbly the answer is
216
Above 30, changing one week of spend by a little moves the whole answer by a lot. Overlapping channels are what drive this up.
What we changed about this fit
Your data tripped a check, so the model was estimated differently. Your numbers were not touched — no imputing, no winsorising, no dropped rows. Each entry says what it cost.
Some channels held at their current budget
Appliedtriggered by vifTV — CableThese channels' spend does not vary enough, did not run for enough of the window, or moves too closely with the rest of the model for the data to say what they returned. Their numbers come mostly from the starting assumption. They stay in the model, so they do not distort the others, but the budget planner will not move them.
What it costs you
You will get no recommendation for these channels — deliberately. A plan that reallocated budget on the strength of an assumption would look exactly like one built on evidence.
Did the maths settle down?
The model does not solve your data in one pass — it searches for the answer from several starting points at once, and it is only finished when they all arrive at the same place. If they did not, nothing further down this page means anything.
- FailChain convergence (R-hat)
Worst R-hat 1.2264 at roi. Above 1.01 the chains have not mixed and the posterior is not yet the posterior.
measured 1.226 · threshold 1.010
- FailBulk effective sample size
Lowest bulk ESS 13 at beta. Under 400 the point estimates carry visible Monte Carlo error.
measured 13.039 · threshold 400.000
- FailTail effective sample size
Lowest tail ESS 16 at beta. Under 400 the interval endpoints carry visible Monte Carlo error.
measured 16.410 · threshold 400.000
- PassDivergent transitions
2 divergences in 4000 draws (0.05%). Divergences mean the sampler could not follow the posterior's curvature, so the region it failed in is under-represented. Raise target_accept_prob (spec section 6) before trusting the numbers.
measured 0.001 · threshold 0.001
Does it predict?
Matching history proves very little: a flexible enough model can trace any past. What counts is the stretch of weeks we hid from it — and whether it beat the free answer of assuming this year looks like last year.
- WarnIn-sample accuracy
MAPE 6.8% (target <10%), R-squared 0.879 (target >0.9). In-sample fit is necessary, not sufficient — a model with enough knots fits anything.
measured 6.789 · threshold 10.000
- WarnHoldout accuracy
MAPE 15.5% (target <15%) and MASE 0.62 (target <1) on 13 held-out cells the likelihood never saw. MASE compares the error against forecasting the same period last year; above 1 the model has not earned its complexity out of sample.
measured 15.498 · threshold 15.000
- PassPosterior-predictive check
Bayesian p-value 0.530 (target 0.05-0.95). High values mean the estimated noise is larger than the residuals — the model is not being held to account by the data.
measured 0.530 · threshold 0.050
- PassBaseline stays positive
P(baseline < 0) = 0.00 (target <0.2). A negative baseline is the classic symptom of over-attribution: the model has given media credit for more than the business would sell with no marketing at all.
measured 0.000 · threshold 0.200
Could your data tell the channels apart?
When two channels go up and down together, no amount of modelling can say which one did the work — the split you get back is the one we assumed going in. These checks measure that rather than hope about it.
- FailVariance inflation (post-transform)
Highest VIF 5170.9 on TV — Cable (warn 5, hard 10), measured after adstock and saturation because that is what the model sees.
measured 5170.898 · threshold 5.000
- FailDesign conditioning
Condition number 216.3 (target <30) on the media block with the baseline and controls projected out.
measured 216.315 · threshold 30.000
- WarnObservations per parameter
104 observations, 38 parameters, 2.7x (target >4).
measured 2.737 · threshold 4.000
Did your data actually teach the model anything?
The classic way marketing models mislead is by handing back the assumption they started with, dressed up as a finding. A channel is flagged here only when two separate signals agree the data never budged it.
- WarnData moved the prior
Indistinguishable from the prior: Social — TikTok (contraction 2%, KS p=0.50), Search — Brand (contraction -12816%, KS p=0.24), Programmatic Display (contraction 3%, KS p=0.20). A channel whose posterior is its prior has been assumed, not estimated.
measured -128.159 · threshold 0.200
- WarnNo prior/data conflict
Outside the prior's 1st-99th percentile: Search — Brand (10.92 vs [0.15, 9.91]), Search — Non-brand (42.94 vs [0.15, 9.91]). Either the prior is wrong or the data are being over-read; the fit is a compromise between two claims that disagree.
measured 2.000 · threshold 0.000
On made-up data, did it find the answer we hid?
Only possible on synthetic data, where we know the right answer because we chose it. Passing means the ranges came out wide enough to contain the truth on this dataset.
- SkippedRecovers known ROI on synthetic data
Not assessed: Chain convergence (R-hat), Bulk effective sample size, Tail effective sample size did not converge, so these draws are not the posterior. A wider interval covers more truths, so grading coverage here would reward the sampling failure rather than the model.
Channel overlap
How closely every pair of channels moved together. Measured after we account for advertising that keeps working after it stops running, and for the fact that the tenth million spent does less than the first — so this is the overlap the model has to live with, not the overlap you would see in a spend report.
| TV — Broadcast | — | 1.00 | 0.18 | 0.19 | 0.21 | 0.20 | 0.14 | -0.02 | 0.08 |
|---|---|---|---|---|---|---|---|---|---|
| TV — Cable | 1.00 | — | 0.18 | 0.18 | 0.21 | 0.20 | 0.14 | -0.02 | 0.08 |
| Social — Meta | 0.18 | 0.18 | — | 1.00 | 0.96 | 0.96 | 0.96 | 0.24 | 0.05 |
| Social — TikTok | 0.19 | 0.18 | 1.00 | — | 0.96 | 0.96 | 0.95 | 0.24 | 0.07 |
| Search — Brand | 0.21 | 0.21 | 0.96 | 0.96 | — | 1.00 | 0.93 | 0.23 | 0.07 |
| Search — Non-brand | 0.20 | 0.20 | 0.96 | 0.96 | 1.00 | — | 0.93 | 0.23 | 0.08 |
| Programmatic Display | 0.14 | 0.14 | 0.96 | 0.95 | 0.93 | 0.93 | — | 0.19 | 0.02 |
| Video — CTV | -0.02 | -0.02 | 0.24 | 0.24 | 0.23 | 0.23 | 0.19 | — | -0.04 |
| Out of Home | 0.08 | 0.08 | 0.05 | 0.07 | 0.07 | 0.08 | 0.02 | -0.04 | — |
Read a cell as: when this channel went up, did that one go up too? 1.00 means always, 0.00 means never, −1.00 means the opposite. Anything above 0.80 and the two are one channel as far as your data is concerned — you can trust what the pair did together, but not the split between them.
3 pairs of your channels are impossible to separate
The closest is TV — Broadcast and TV — Cable. They went up and down together so consistently that nothing in your data distinguishes them. You can still trust what the pair achieved between them — what cannot be recovered, by any model, is how much of it belonged to each.
Start with TV — Broadcast and TV — Cable — report them as one line and the model can tell you what they are worth together. Then work down the list. Fixing only the first is the usual mistake: the rest go on widening every range in this report. To get the split back you have to run them on different schedules for a while — there is no way to recover it from the data you already have.
Show me the statistics
Pairs at or above 0.99 correlation, sorted by variance inflation rather than by correlation — correlation saturates near 1.00 and VIF does not, so the list order is the only thing distinguishing a pair that costs a little from one that costs ten times as much.
- TV — Broadcast + TV — Cabler 0.9999 · VIF 5171
- Search — Brand + Search — Non-brandr 0.9985 · VIF 505
- Social — Meta + Social — TikTokr 0.9978 · VIF 311
How each channel spends
A channel that is always on at the same level gives the model almost nothing to learn from, however big its budget — it is the ups and downs that reveal what the spend does. This table shows which of your channels have them.
| Channel | Spend | Pattern | Weeks dark | Variation | VIF |
|---|---|---|---|---|---|
| TV — Broadcast | 23% | bursty | 61% | 1.34 | 5169.3 |
| TV — Cable | 12% | bursty | 61% | 1.34 | 5170.9 |
| Social — Meta | 16% | continuous | 0% | 0.25 | 311.4 |
| Social — TikTok | 8% | continuous | 0% | 0.25 | 310.4 |
| Search — Brand | 7% | continuous | 0% | 0.45 | 505.1 |
| Search — Non-brand | 11% | continuous | 0% | 0.42 | 486.9 |
| Programmatic Display | 8% | continuous | 0% | 0.19 | 15.6 |
| Video — CTV | 9% | variable | 0% | 0.60 | 9.1 |
| Out of Home | 6% | bursty | 60% | 1.23 | 1.2 |