Evidence
16 Aug 2026, 15:20 UTCOver 17 paired rounds, the model no difference demonstrated (grid order).
| Baseline | Race | Rounds | Model | Baseline | Gain | Verdict |
|---|---|---|---|---|---|---|
| Grid order | race | 17 | 5.55 | 5.46 | -0.08 | no difference demonstrated |
| Last race order | race | 16 | 5.58 | 5.53 | -0.05 | no difference demonstrated |
Paired bootstrap on the round-by-round difference: -0.082, 95% CI [-0.630, 0.519] in positions gained. Positive means the model is closer to the real finishing order than grid order is; the interval covers zero, so no difference has been demonstrated.
forward evaluation — each round was forecast before it ran and scored afterwards; not a backtest. Historical replays and the forward record are computed separately and never merged.
- Some comparisons straddle zero: a difference has not been demonstrated, in either direction.
- Metrics are only comparable within this series. A position error over this field size means nothing next to another series' number.
Formula E · Season 2025-26
Model accuracy
How the Formula E model’s pre-race forecasts have scored against the actual results, over 17 completed rounds of the season. Every number is scored finishers-only, using only data available before each race.
Winner hit rate
12%
Podium hit rate
22%
Mean position error
5.55
NDCG@5
0.66
Per-round accuracy
Podium-weighted race accuracy per round. Tap a cell for the breakdown.
Per round
| Round | Winner | Podium hits | Mean error | NDCG@5 | Win Brier |
|---|---|---|---|---|---|
| R1 · São Paulo | miss | 1/3 | 5.077 | 0.78 | 0.0736 |
| R2 · Mexico City | ✓ hit | 2/3 | 3.765 | 0.96 | 0.0450 |
| R3 · Miami | miss | 0/3 | 6.947 | 0.49 | 0.0607 |
| R4 · Jeddah | ✓ hit | 2/3 | 4 | 0.96 | 0.0393 |
| R5 · Jeddah II | miss | 0/3 | 6 | 0.62 | 0.0554 |
| R6 · Madrid | miss | 2/3 | 4.1 | 0.92 | 0.0480 |
| R7 · Berlin | miss | 0/3 | 5 | 0.55 | 0.0555 |
| R8 · Berlin II | miss | 1/3 | 5.778 | 0.56 | 0.0442 |
| R9 · Monaco | miss | 1/3 | 7.389 | 0.45 | 0.0639 |
| R10 · Monaco II | miss | 0/3 | 5.556 | 0.68 | 0.0679 |
| R11 · Sanya | miss | 0/3 | 5.267 | 0.80 | 0.0635 |
| R12 · Shanghai | miss | 0/3 | 5.9 | 0.56 | 0.0603 |
| R13 · Shanghai II | miss | 0/3 | 6.556 | 0.30 | 0.0583 |
| R14 · Tokyo | miss | 0/3 | 5.421 | 0.67 | 0.0655 |
| R15 · Tokyo II | miss | 1/3 | 6.875 | 0.63 | 0.0631 |
| R16 · London | miss | 1/3 | 4.438 | 0.72 | 0.0787 |
| R17 · London II | miss | 0/3 | 6.222 | 0.51 | 0.0661 |
Win Brier scores the model’s win probabilities against who actually won — lower is sharper and better calibrated.
Walk-forward validation
Model vs the naive baselines
Every completed round is re-forecast using only earlier rounds, then scored against two trivial predictors — one that replays the previous result, one that assumes the grid finishes in order. Blue marks the best column. Beating these baselines is the bar the model has to clear.
E-Prix
17 rounds · model vs naive baselines| Metric | Model | Last-race | Grid order |
|---|---|---|---|
| Mean position error | 5.55 | 5.53 | 5.46 |
| Top-5 ranking | 0.657 | 0.603 | 0.676 |
| Order agreement | 0.244 | 0.148 | 0.256 |
| Podium hits / round | 0.65 | 0.69 | 1.06 |
Candidate model
A shadow model runs alongside the live one
position-head candidate · gated behind FE_USE_POSITION_HEAD
Comparison basis: race mean_position_error
Production error
5.55
mean positions off
Candidate error
5.85
mean positions off
Gap
+0.29
candidate minus production
per-round guard tripped (worst regression 48.8% > 20%). The candidate only gets promoted once it beats the live model on enough real rounds — until then the site keeps serving the production forecast.
Probability calibration
How trustworthy the probabilities are
A well-calibrated model assigns probabilities that match how often things actually happen. The forecasts are tuned against the season’s real classified results — separately for street and permanent circuits, because the two race very differently — so a stated 30% podium chance means roughly 3-in-10 over the long run.
Training rounds
17
real completed rounds
Status
Street stratumCircuit stratum
Calibrated on real Formula E results, stratified street vs permanent circuit.
Generated 8/30/2026, 12:00:36 PM
Calibration samples per market
Win
302
observations
Podium
302
observations
Top 6
302
observations
Top 10
302
observations
Each figure is how many prior driver-outcomes fed the calibrator for that market. More samples means a steadier probability estimate.
Historical performance
Backtest across 1,121 driver-rounds
Predicted finishing order scored against the official classification for the 2026 season — every round, every driver. Each round was replayed using only signals available before lights-out, so nothing here is hindsight.
Rounds Evaluated
16
Mean Position Error
5.29 pos
Within 3 Positions
40.6%
Within 5 Positions
57.7%
Podium Hit Rate
43.8%
Winner Hit Rate
25.0%
Order Agreement
0.381
Top-5 Ranking
0.745
Model health
WatchA self-check on the live model: whether recent forecasts are still landing as well as they should, and whether the numbers feeding the model have drifted from what it was tuned on.
Forecast quality
▲ 11%
higher error than the season benchmark
Rounds monitored
17
recent rounds in the rolling check
Input drift
5/5
model inputs that moved vs baseline
Input drift by feature
Average Probability Error, Round by Round
How far the model's probabilities sat from what actually happened — lower is a better-calibrated forecast.
Some model inputs have drifted from their reference range — normal across a season and with a small sample of races. It is flagged here for transparency, but the rolling forecast-quality check above shows the predictions themselves are being watched closely.