Evidence
10 Aug 2026, 13:01 UTCOver 10 paired rounds, the model no difference demonstrated (last rally order).
| Baseline | Race | Rounds | Model | Baseline | Gain | Verdict |
|---|---|---|---|---|---|---|
| Last rally order | rally | 10 | 8.12 | 8.65 | +0.54 | no difference demonstrated |
| Standings order | rally | 10 | 8.12 | 7.81 | -0.31 | no difference demonstrated |
Paired bootstrap on the round-by-round difference: 0.535, 95% CI [-0.946, 2.739] in positions gained. Positive means the model is closer to the real finishing order than last rally order is; the interval covers zero, so no difference has been demonstrated.
forward evaluation — each round was forecast before it ran and scored afterwards; not a backtest. Historical replays and the forward record are computed separately and never merged.
- Some comparisons straddle zero: a difference has not been demonstrated, in either direction.
- Metrics are only comparable within this series. A position error over this field size means nothing next to another series' number.
WRC · 2026
Model accuracy
How the WRC model’s leakage-safe pre-rally forecasts have scored against the actual results, over 10 completed rounds of 2026. Every number is scored finishers-only, using only data available before each rally.
Winner hit rate
20%
Podium hit rate
37%
Mean position error
8.12
NDCG@5
0.67
Model health at a glance
WatchA self-check on the live model: whether recent forecasts are still landing as well as they should, and whether the numbers feeding the model have drifted from what it was tuned on.
Forecast quality
▲ 13%
higher error than the season benchmark
Rounds monitored
10
recent rounds in the rolling check
Input drift
5/5
model inputs that moved vs baseline
Input drift by feature
Some model inputs have drifted from their reference range — normal early in a season with a small sample of rounds so far. It is flagged here for transparency, but the rolling forecast-quality check shows the predictions themselves are being watched closely.
Per-round accuracy
Podium-weighted rally accuracy per round. Tap a cell for the breakdown.
Per round
| Round | Winner | Podium hits | Mean error | NDCG@5 | Win Brier |
|---|---|---|---|---|---|
| R1 · Rallye Monte Carlo | miss | 1/3 | 6.231 | 0.92 | 0.0702 |
| R2 · Rally Sweden | miss | 1/3 | 6.417 | 0.87 | 0.0636 |
| R3 · Safari Rally Kenya | miss | 1/3 | 8.429 | 0.57 | 0.0632 |
| R4 · Croatia Rally | ✓ hit | 2/3 | 15.067 | -0.22 | 0.0437 |
| R5 · Rally Islas Canarias | miss | 1/3 | 7.909 | 0.86 | 0.0970 |
| R6 · Vodafone Rally de Portugal | miss | 0/3 | 7.909 | 0.69 | 0.0998 |
| R7 · FORUM8 Rally Japan | miss | 2/3 | 5.545 | 0.90 | 0.0773 |
| R8 · EKO Acropolis Rally Greece | miss | 1/3 | 6.077 | 0.87 | 0.0747 |
| R9 · Rally Estonia | miss | 1/3 | 8.5 | 0.25 | 0.0710 |
| R10 · Secto Rally Finland | ✓ hit | 1/3 | 9.083 | 0.92 | 0.0576 |
Win Brier scores the model’s win probabilities against who actually won — lower is sharper and better calibrated.
Vs the baselines
Does the forecast beat championship form?
A rally has no qualifying to lean on, so the fair test is whether the forecast is sharper than simply ordering crews by their championship standings. Over 10 completed rounds the model is sharper than championship form on both the win and podium probability scores. A last-rally momentum baseline is a touch sharper on the win score this season — shown here honestly rather than hidden.
| Measure | Our forecast | Championship form | Last rally |
|---|---|---|---|
| Win probability score(lower is sharper) | 0.875 | 0.876 | 0.844 |
| Podium probability score(lower is sharper) | 0.734 | 0.765 | 0.818 |
| Winner called(higher is better) | 20% | 10% | 20% |
The skill model alone scores 0.901 on the win score and 0.792 on podium — weaker than championship form on its own. Blending it with championship form is what edges the combined forecast ahead.
Green marks where the forecast beats championship-form order.
Walk-forward validation
Model vs the championship-form baseline
Every completed round is re-forecast using only earlier rounds, then scored against a baseline that simply ranks crews by their championship standings. Gold marks the better side. Beating championship form on the ranking metrics is the bar the model has to clear.
Rally
10 rounds · model vs championship form| Metric | Model | Championship form |
|---|---|---|
| Mean position error | 8.12 | 7.81 |
| Top-5 ranking | 0.665 | 0.608 |
| Order agreement | 0.530 | 0.501 |
| Podium hits / round | 1.10 | 1.30 |
Probability calibration
How trustworthy the probabilities are
A well-calibrated model assigns probabilities that match how often things actually happen. WRC’s forecasts are tuned against the real classified rally results so a stated 30% podium chance means roughly 3-in-10 over the long run.
Training rounds
10
real completed rounds
Status
Calibrated on real WRC rally results (one classification per round).
Generated 8/30/2026, 12:01:23 PM
Calibration samples per market
Win
124
observations
Podium
124
observations
Top 6
124
observations
Top 10
124
observations
Each figure is how many prior crew-outcomes fed the calibrator for that market. More samples means a steadier probability estimate.
Model health
Win-market Brier trend
Lower is better · 10 rounds
Diagnostics
- ⚠ pWin: PSI 0.405 (significant drift vs baseline)
- ⚠ pPodium: PSI 0.671 (significant drift vs baseline)
- ⚠ finishRangeHigh: PSI 0.447 (significant drift vs baseline)
- • predictedValue: PSI 0.169 (moderate drift vs baseline)
- • meanFinish: PSI 0.208 (moderate drift vs baseline)
- • rolling Brier regression +12.7%
Feature drift and rolling-Brier are tracked round-to-round; a spike flags where the field behaved unlike the rounds the model learned from.