Skip to main content

Evidence

16 Aug 2026, 15:20 UTC

Over 17 paired rounds, the model no difference demonstrated (grid order).

BaselineRaceRoundsModelBaselineGainVerdict
Grid orderrace175.555.46-0.08no difference demonstrated
Last race orderrace165.585.53-0.05no difference demonstrated

Paired bootstrap on the round-by-round difference: -0.082, 95% CI [-0.630, 0.519] in positions gained. Positive means the model is closer to the real finishing order than grid order is; the interval covers zero, so no difference has been demonstrated.

forward evaluation — each round was forecast before it ran and scored afterwards; not a backtest. Historical replays and the forward record are computed separately and never merged.

  • Some comparisons straddle zero: a difference has not been demonstrated, in either direction.
  • Metrics are only comparable within this series. A position error over this field size means nothing next to another series' number.

Formula E · Season 2025-26

Model accuracy

How the Formula E model’s pre-race forecasts have scored against the actual results, over 17 completed rounds of the season. Every number is scored finishers-only, using only data available before each race.

Winner hit rate

12%

Podium hit rate

22%

Mean position error

5.55

NDCG@5

0.66

Per-round accuracy

Podium-weighted race accuracy per round. Tap a cell for the breakdown.

Per round

RoundWinnerPodium hitsMean error
R1 · São Paulomiss1/35.077
R2 · Mexico City✓ hit2/33.765
R3 · Miamimiss0/36.947
R4 · Jeddah✓ hit2/34
R5 · Jeddah IImiss0/36
R6 · Madridmiss2/34.1
R7 · Berlinmiss0/35
R8 · Berlin IImiss1/35.778
R9 · Monacomiss1/37.389
R10 · Monaco IImiss0/35.556
R11 · Sanyamiss0/35.267
R12 · Shanghaimiss0/35.9
R13 · Shanghai IImiss0/36.556
R14 · Tokyomiss0/35.421
R15 · Tokyo IImiss1/36.875
R16 · Londonmiss1/34.438
R17 · London IImiss0/36.222

Win Brier scores the model’s win probabilities against who actually won — lower is sharper and better calibrated.

Walk-forward validation

Model vs the naive baselines

Every completed round is re-forecast using only earlier rounds, then scored against two trivial predictors — one that replays the previous result, one that assumes the grid finishes in order. Blue marks the best column. Beating these baselines is the bar the model has to clear.

E-Prix

17 rounds · model vs naive baselines
MetricModelLast-race
Mean position error5.555.53
Top-5 ranking0.6570.603
Order agreement0.2440.148
Podium hits / round0.650.69

Candidate model

A shadow model runs alongside the live one

position-head candidate · gated behind FE_USE_POSITION_HEAD

Comparison basis: race mean_position_error

Production model still ahead

Production error

5.55

mean positions off

Candidate error

5.85

mean positions off

Gap

+0.29

candidate minus production

per-round guard tripped (worst regression 48.8% > 20%). The candidate only gets promoted once it beats the live model on enough real rounds — until then the site keeps serving the production forecast.

Probability calibration

How trustworthy the probabilities are

Calibration applied

A well-calibrated model assigns probabilities that match how often things actually happen. The forecasts are tuned against the season’s real classified results — separately for street and permanent circuits, because the two race very differently — so a stated 30% podium chance means roughly 3-in-10 over the long run.

Training rounds

17

real completed rounds

Status

Street stratumCircuit stratum

Calibrated on real Formula E results, stratified street vs permanent circuit.

Generated 8/30/2026, 12:00:36 PM

Calibration samples per market

Win

302

observations

Podium

302

observations

Top 6

302

observations

Top 10

302

observations

Each figure is how many prior driver-outcomes fed the calibrator for that market. More samples means a steadier probability estimate.

Historical performance

Backtest across 1,121 driver-rounds

Predicted finishing order scored against the official classification for the 2026 season — every round, every driver. Each round was replayed using only signals available before lights-out, so nothing here is hindsight.

Rounds Evaluated

16

Mean Position Error

5.29 pos

Within 3 Positions

40.6%

Within 5 Positions

57.7%

Podium Hit Rate

43.8%

Winner Hit Rate

25.0%

Order Agreement

0.381

Top-5 Ranking

0.745

Model health

Watch

A self-check on the live model: whether recent forecasts are still landing as well as they should, and whether the numbers feeding the model have drifted from what it was tuned on.

Forecast quality

11%

higher error than the season benchmark

Rounds monitored

17

recent rounds in the rolling check

Input drift

5/5

model inputs that moved vs baseline

Input drift by feature

Predicted merit· ShiftedWin probability· ShiftingPodium probability· ShiftingMean finish· ShiftedFinish range (high)· Shifted

Average Probability Error, Round by Round

How far the model's probabilities sat from what actually happened — lower is a better-calibrated forecast.

Some model inputs have drifted from their reference range — normal across a season and with a small sample of races. It is flagged here for transparency, but the rolling forecast-quality check above shows the predictions themselves are being watched closely.