Evidence
30 Aug 2026, 11:58 UTCOver 12 paired rounds, the model no difference demonstrated (grid order).
| Baseline | Race | Rounds | Model | Baseline | Gain | Verdict |
|---|---|---|---|---|---|---|
| Grid order | feature | 12 | 4.02 | 3.88 | -0.14 | no difference demonstrated |
| Grid order | sprint | 12 | 3.66 | 3.42 | -0.24 | no difference demonstrated |
| lastStandings | feature | 11 | 4.08 | 4.24 | +0.17 | no difference demonstrated |
| lastStandings | sprint | 11 | 3.67 | 4.22 | +0.55 | beats the baseline |
Paired bootstrap on the round-by-round difference: -0.136, 95% CI [-0.375, 0.107] in positions gained. Positive means the model is closer to the real finishing order than grid order is; the interval covers zero, so no difference has been demonstrated.
forward evaluation — each round was forecast before it ran and scored afterwards; not a backtest. Historical replays and the forward record are computed separately and never merged.
- Some comparisons straddle zero: a difference has not been demonstrated, in either direction.
- Metrics are only comparable within this series. A position error over this field size means nothing next to another series' number.
MotoGP · 2026
Model accuracy
How the MotoGP model’s leakage-safe pre-race forecasts have scored against the actual results, over 12 completed rounds of 2026. Every number is scored finishers-only, using only data available before each race.
Winner hit rate
25%
Podium hit rate
50%
Mean position error
4.02
NDCG@5
0.91
Model health at a glance
HealthyA self-check on the live model: whether recent forecasts are still landing as well as they should, and whether the numbers feeding the model have drifted from what it was tuned on.
Forecast quality
▼ 4%
lower error than the season benchmark
Rounds monitored
12
recent rounds in the rolling check
Input drift
3/5
model inputs that moved vs baseline
Input drift by feature
Some model inputs have drifted from their reference range — normal early in a season with a small sample of rounds so far. It is flagged here for transparency, but the rolling forecast-quality check shows the predictions themselves are still holding up.
Per-round accuracy
Podium-weighted Grand Prix accuracy per round. Tap a cell for the breakdown.
Per round (Grand Prix)
| Round | Winner | Podium hits | Mean error | NDCG@5 | Win Brier |
|---|---|---|---|---|---|
| R1 · Buriram | miss | 1/3 | 3.421 | 0.94 | 0.0354 |
| R2 · Goiania | miss | 2/3 | 2.833 | 0.97 | 0.0380 |
| R3 · Austin | miss | 2/3 | 3 | 0.90 | 0.0356 |
| R4 · Jerez de la Frontera | miss | 2/3 | 2.85 | 0.97 | 0.0484 |
| R5 · Le Mans | miss | 1/3 | 4.062 | 0.97 | 0.0632 |
| R6 · Barcelona | miss | 0/3 | 7.235 | 0.51 | 0.0526 |
| R7 · Mugello | ✓ hit | 2/3 | 2.842 | 0.93 | 0.0230 |
| R8 · Balatonfokajár | ✓ hit | 2/3 | 5.625 | 0.91 | 0.0303 |
| R9 · Brno | miss | 1/3 | 3.824 | 0.96 | 0.0540 |
| R10 · Assen | miss | 2/3 | 4.562 | 0.96 | 0.0513 |
| R11 · Sachsenring | ✓ hit | 1/3 | 3.867 | 0.97 | 0.0269 |
| R12 · Silverstone | miss | 2/3 | 4.125 | 0.95 | 0.0508 |
Win Brier scores the model’s win probabilities against who actually won — lower is sharper and better calibrated.
Vs the baselines
Does the forecast beat the grid?
Our Grand Prix forecast is made after qualifying, so it starts from the real grid. Over 12 completed rounds it is sharper than the grid order alone on win and podium probabilities. A pre-qualifying, form-only forecast is weaker than the grid — so we don’t claim to call the grid from nothing.
| Measure | Our forecast | Grid order | Form only |
|---|---|---|---|
| Win probability score(lower is sharper) | 0.042 | 0.054 | 0.051 |
| Podium probability score(lower is sharper) | 0.079 | 0.106 | 0.115 |
| Winner called(higher is better) | 25% | 33% | 0% |
Green marks where the grid-conditioned forecast beats the bare grid order.
Walk-forward validation
Model vs the “last race repeats” baseline
Every completed round is re-forecast using only earlier rounds, then scored against a trivial predictor that just replays the previous result. Gold marks the better side. Beating this baseline is the bar the model has to clear.
Sprint race
12 rounds · model vs last-race| Metric | Model | Last-race |
|---|---|---|
| Mean position error | 3.66 | — |
| Top-5 ranking | 0.893 | — |
| Order agreement | 0.791 | — |
| Podium hits / round | 1.50 | — |
Feature race
12 rounds · model vs last-race| Metric | Model | Last-race |
|---|---|---|
| Mean position error | 4.02 | — |
| Top-5 ranking | 0.911 | — |
| Order agreement | 0.844 | — |
| Podium hits / round | 1.50 | — |
Probability calibration
How trustworthy the probabilities are
A well-calibrated model assigns probabilities that match how often things actually happen. MotoGP’s forecasts are tuned against the real classified results so a stated 30% podium chance means roughly 3-in-10 over the long run.
Training rounds
12
real completed rounds
Status
Calibrated on real MotoGP results (Sprint + Grand Prix, per race-type stratum).
Generated 8/30/2026, 12:01:13 PM
Calibration samples per market
Win
429
observations
Podium
429
observations
Top 6
429
observations
Top 10
429
observations
Each figure is how many prior rider-outcomes fed the calibrator for that market. More samples means a steadier probability estimate.
Model health
Win-market Brier trend
Lower is better · 12 rounds
Diagnostics
- • pWin: PSI 0.195 (moderate drift vs baseline)
- • pPodium: PSI 0.128 (moderate drift vs baseline)
- • meanFinish: PSI 0.118 (moderate drift vs baseline)
Feature drift and rolling-Brier are tracked round-to-round; a spike flags where the field behaved unlike the rounds the model learned from.