Evidence
8 Jul 2026, 07:14 UTCOver 5 paired rounds, the model no difference demonstrated (last race order).
| Baseline | Race | Rounds | Model | Baseline | Gain | Verdict |
|---|---|---|---|---|---|---|
| Last race order | feature | 5 | 5.51 | 4.89 | -0.62 | no difference demonstrated |
| Last race order | sprint | 5 | 5.69 | 5.25 | -0.44 | does NOT beat the baseline |
Paired bootstrap on the round-by-round difference: -0.617, 95% CI [-2.190, 0.957] in positions gained. Positive means the model is closer to the real finishing order than last race order is; the interval covers zero, so no difference has been demonstrated.
forward evaluation — each round was forecast before it ran and scored afterwards; not a backtest. Historical replays and the forward record are computed separately and never merged.
- The model does not beat last race order on every measured comparison. That is stated rather than hidden.
- Some comparisons straddle zero: a difference has not been demonstrated, in either direction.
- Metrics are only comparable within this series. A position error over this field size means nothing next to another series' number.
Formula 2 · 2026
Model accuracy
How the F2 model’s leakage-safe pre-race forecasts have scored against the actual results, over 6 completed rounds of 2026. Every number is scored finishers-only, using only data available before each race.
Winner hit rate
0%
Podium hit rate
50%
Mean position error
5.54
NDCG@5
0.68
Per-round accuracy
Podium-weighted feature-race accuracy per round. Tap a cell for the breakdown.
Per round (feature race)
| Round | Winner | Podium hits | Mean error | NDCG@5 | Win Brier |
|---|---|---|---|---|---|
| R1 · Australia | miss | 1/3 | 5.684 | 0.60 | 0.0496 |
| R2 · Miami | miss | 1/3 | 6.8 | 0.57 | 0.0624 |
| R3 · Canada | miss | 2/3 | 7.286 | 0.63 | 0.0757 |
| R4 · Monaco | miss | 1/3 | 4.905 | 0.70 | 0.0560 |
| R5 · Spain | miss | 2/3 | 4.714 | 0.67 | 0.0502 |
| R6 · Austria | miss | 2/3 | 3.842 | 0.93 | 0.0477 |
Win Brier scores the model’s win probabilities against who actually won — lower is sharper and better calibrated.
Walk-forward validation
Model vs the “last race repeats” baseline
Every completed round is re-forecast using only earlier rounds, then scored against a trivial predictor that just replays the previous result. Gold marks the better side. Beating this baseline is the bar the model has to clear.
Sprint race
6 rounds · model vs last-race| Metric | Model | Last-race |
|---|---|---|
| Mean position error | 5.70 | 5.25 |
| Top-5 ranking | 0.623 | 0.579 |
| Order agreement | 0.317 | 0.257 |
| Podium hits / round | 0.50 | 0.80 |
Feature race
6 rounds · model vs last-race| Metric | Model | Last-race |
|---|---|---|
| Mean position error | 5.54 | 4.89 |
| Top-5 ranking | 0.683 | 0.628 |
| Order agreement | 0.348 | 0.256 |
| Podium hits / round | 1.50 | 1.00 |
Candidate model
A shadow model runs alongside the live one
position-head candidate · gated behind F2_USE_POSITION_HEAD
Comparison basis: pooled (sprint+feature) mean_position_error
Production error
5.46
mean positions off
Candidate error
6.27
mean positions off
Gap
+0.81
candidate minus production
insufficient overlap (4 common rounds; need >= 5). The candidate only gets promoted once it beats the live model on enough real rounds — until then the site keeps serving the production forecast.
Probability calibration
How trustworthy the probabilities are
A well-calibrated model assigns probabilities that match how often things actually happen. F2’s forecasts are tuned against the real classified results so a stated 30% podium chance means roughly 3-in-10 over the long run.
Training rounds
6
real completed rounds
Status
Calibrated on real F2 results.
Generated 8/30/2026, 12:00:15 PM
Calibration samples per market
Win
223
observations
Podium
223
observations
Top 6
223
observations
Top 10
223
observations
Each figure is how many prior driver-outcomes fed the calibrator for that market. More samples means a steadier probability estimate.
Historical performance
Backtest across 109 driver-rounds
Predicted finishing order scored against the official classification for the 2026 season — every round, every driver. Each round was replayed using only signals available before lights-out, so nothing here is hindsight.
Rounds Evaluated
6
Mean Position Error
5.54 pos
Within 3 Positions
41.3%
Within 5 Positions
58.7%
Podium Hit Rate
50.0%
Winner Hit Rate
0.0%
Order Agreement
0.348
Top-5 Ranking
0.683
Model health
Win-market Brier trend
Lower is better · 6 rounds
Diagnostics
- ⚠ predictedValue: PSI 0.858 (significant drift vs baseline)
- ⚠ pWin: PSI 1.861 (significant drift vs baseline)
- ⚠ pPodium: PSI 1.281 (significant drift vs baseline)
- ⚠ meanFinish: PSI 1.069 (significant drift vs baseline)
- ⚠ finishRangeHigh: PSI 1.814 (significant drift vs baseline)
Feature drift and rolling-Brier are tracked round-to-round; a spike flags where the field behaved unlike the rounds the model learned from.