Evidence
7 Jul 2026, 15:02 UTCOver 4 paired rounds, the model too few rounds to say (last race order).
| Baseline | Race | Rounds | Model | Baseline | Gain | Verdict |
|---|---|---|---|---|---|---|
| Last race order | feature | 4 | 6.35 | 6.44 | +0.09 | too few rounds to say |
| Last race order | sprint | 4 | 7.22 | 8.20 | +0.98 | too few rounds to say |
Paired bootstrap on the round-by-round difference: 0.090, 95% CI [-0.588, 0.570] in positions gained. Positive means the model is closer to the real finishing order than last race order is; the interval covers zero, so no difference has been demonstrated.
forward evaluation — each round was forecast before it ran and scored afterwards; not a backtest. Historical replays and the forward record are computed separately and never merged.
- Metrics are only comparable within this series. A position error over this field size means nothing next to another series' number.
Formula 3 · 2026
Model accuracy
How the F3 model’s leakage-safe pre-race forecasts have scored against the actual results, over 5 completed rounds of 2026. Every number is scored finishers-only, using only data available before each race.
Winner hit rate
0%
Podium hit rate
13%
Mean position error
7.16
NDCG@5
0.69
Model health at a glance
HealthyA self-check on the live model: whether recent forecasts are still landing as well as they should, and whether the numbers feeding the model have drifted from what it was tuned on.
Forecast quality
Not enough graded rounds yet.
Rounds monitored
5
recent rounds in the rolling check
Input drift
5/5
model inputs that moved vs baseline
Input drift by feature
Some model inputs have drifted from their reference range — normal early in a season with a small sample of rounds so far. It is flagged here for transparency, but the rolling forecast-quality check shows the predictions themselves are still holding up.
Per-round accuracy
Podium-weighted feature-race accuracy per round. Tap a cell for the breakdown.
Per round (feature race)
| Round | Winner | Podium hits | Mean error | NDCG@5 | Win Brier |
|---|---|---|---|---|---|
| R1 · Australia | miss | 0/3 | 10.391 | 0.61 | 0.0414 |
| R2 · Monaco | miss | 0/3 | 7.077 | 0.67 | 0.0388 |
| R3 · Spain | miss | 1/3 | 4.429 | 0.84 | 0.0389 |
| R4 · Austria | miss | 1/3 | 6.769 | 0.65 | 0.0398 |
| R5 · Great Britain | miss | 0/3 | 7.111 | 0.70 | 0.0399 |
Win Brier scores the model’s win probabilities against who actually won — lower is sharper and better calibrated.
Walk-forward validation
Model vs the “last race repeats” baseline
Every completed round is re-forecast using only earlier rounds, then scored against a trivial predictor that just replays the previous result. Gold marks the better side. Beating this baseline is the bar the model has to clear.
Sprint race
5 rounds · model vs last-race| Metric | Model | Last-race |
|---|---|---|
| Mean position error | 6.90 | 8.20 |
| Top-5 ranking | 0.660 | 0.661 |
| Order agreement | 0.470 | 0.225 |
| Podium hits / round | 0.60 | 0.50 |
Feature race
5 rounds · model vs last-race| Metric | Model | Last-race |
|---|---|---|
| Mean position error | 7.16 | 6.44 |
| Top-5 ranking | 0.694 | 0.728 |
| Order agreement | 0.370 | 0.456 |
| Podium hits / round | 0.40 | 0.75 |
Candidate model
A shadow model runs alongside the live one
position-head candidate · gated behind F3_USE_POSITION_HEAD
Comparison basis: pooled (sprint+feature) mean_position_error
Production error
6.25
mean positions off
Candidate error
6.37
mean positions off
Gap
+0.13
candidate minus production
insufficient overlap (3 common rounds; need >= 5). The candidate only gets promoted once it beats the live model on enough real rounds — until then the site keeps serving the production forecast.
Probability calibration
How trustworthy the probabilities are
A well-calibrated model assigns probabilities that match how often things actually happen. F3’s forecasts are tuned against the real classified results so a stated 30% podium chance means roughly 3-in-10 over the long run.
Training rounds
5
real completed rounds
Status
Calibrated on real F3 results.
Generated 8/30/2026, 12:00:22 PM
Calibration samples per market
Win
265
observations
Podium
265
observations
Top 6
265
observations
Top 10
265
observations
Each figure is how many prior driver-outcomes fed the calibrator for that market. More samples means a steadier probability estimate.
Historical performance
Backtest across 130 driver-rounds
Predicted finishing order scored against the official classification for the 2026 season — every round, every driver. Each round was replayed using only signals available before lights-out, so nothing here is hindsight.
Rounds Evaluated
5
Mean Position Error
7.16 pos
Within 3 Positions
33.1%
Within 5 Positions
50.8%
Podium Hit Rate
13.3%
Winner Hit Rate
0.0%
Order Agreement
0.370
Top-5 Ranking
0.694
Model health
Win-market Brier trend
Lower is better · 5 rounds
Diagnostics
- ⚠ predictedValue: PSI 2.233 (significant drift vs baseline)
- ⚠ pWin: PSI 2.044 (significant drift vs baseline)
- ⚠ pPodium: PSI 1.819 (significant drift vs baseline)
- ⚠ meanFinish: PSI 1.983 (significant drift vs baseline)
- ⚠ finishRangeHigh: PSI 1.519 (significant drift vs baseline)
Feature drift and rolling-Brier are tracked round-to-round; a spike flags where the field behaved unlike the rounds the model learned from.