Skip to main content

Evidence

8 Jul 2026, 07:14 UTC

Over 5 paired rounds, the model no difference demonstrated (last race order).

BaselineRaceRoundsModelBaselineGainVerdict
Last race orderfeature55.514.89-0.62no difference demonstrated
Last race ordersprint55.695.25-0.44does NOT beat the baseline

Paired bootstrap on the round-by-round difference: -0.617, 95% CI [-2.190, 0.957] in positions gained. Positive means the model is closer to the real finishing order than last race order is; the interval covers zero, so no difference has been demonstrated.

forward evaluation — each round was forecast before it ran and scored afterwards; not a backtest. Historical replays and the forward record are computed separately and never merged.

  • The model does not beat last race order on every measured comparison. That is stated rather than hidden.
  • Some comparisons straddle zero: a difference has not been demonstrated, in either direction.
  • Metrics are only comparable within this series. A position error over this field size means nothing next to another series' number.

Formula 2 · 2026

Model accuracy

How the F2 model’s leakage-safe pre-race forecasts have scored against the actual results, over 6 completed rounds of 2026. Every number is scored finishers-only, using only data available before each race.

Winner hit rate

0%

Podium hit rate

50%

Mean position error

5.54

NDCG@5

0.68

Per-round accuracy

Podium-weighted feature-race accuracy per round. Tap a cell for the breakdown.

Per round (feature race)

RoundWinnerPodium hitsMean error
R1 · Australiamiss1/35.684
R2 · Miamimiss1/36.8
R3 · Canadamiss2/37.286
R4 · Monacomiss1/34.905
R5 · Spainmiss2/34.714
R6 · Austriamiss2/33.842

Win Brier scores the model’s win probabilities against who actually won — lower is sharper and better calibrated.

Walk-forward validation

Model vs the “last race repeats” baseline

Every completed round is re-forecast using only earlier rounds, then scored against a trivial predictor that just replays the previous result. Gold marks the better side. Beating this baseline is the bar the model has to clear.

Sprint race

6 rounds · model vs last-race
MetricModelLast-race
Mean position error5.705.25
Top-5 ranking0.6230.579
Order agreement0.3170.257
Podium hits / round0.500.80

Feature race

6 rounds · model vs last-race
MetricModelLast-race
Mean position error5.544.89
Top-5 ranking0.6830.628
Order agreement0.3480.256
Podium hits / round1.501.00

Candidate model

A shadow model runs alongside the live one

position-head candidate · gated behind F2_USE_POSITION_HEAD

Comparison basis: pooled (sprint+feature) mean_position_error

Production model still ahead

Production error

5.46

mean positions off

Candidate error

6.27

mean positions off

Gap

+0.81

candidate minus production

insufficient overlap (4 common rounds; need >= 5). The candidate only gets promoted once it beats the live model on enough real rounds — until then the site keeps serving the production forecast.

Probability calibration

How trustworthy the probabilities are

Calibration applied

A well-calibrated model assigns probabilities that match how often things actually happen. F2’s forecasts are tuned against the real classified results so a stated 30% podium chance means roughly 3-in-10 over the long run.

Training rounds

6

real completed rounds

Status

Calibrated on real F2 results.

Generated 8/30/2026, 12:00:15 PM

Calibration samples per market

Win

223

observations

Podium

223

observations

Top 6

223

observations

Top 10

223

observations

Each figure is how many prior driver-outcomes fed the calibrator for that market. More samples means a steadier probability estimate.

Historical performance

Backtest across 109 driver-rounds

Predicted finishing order scored against the official classification for the 2026 season — every round, every driver. Each round was replayed using only signals available before lights-out, so nothing here is hindsight.

Rounds Evaluated

6

Mean Position Error

5.54 pos

Within 3 Positions

41.3%

Within 5 Positions

58.7%

Podium Hit Rate

50.0%

Winner Hit Rate

0.0%

Order Agreement

0.348

Top-5 Ranking

0.683

Model health

Win-market Brier trend

Lower is better · 6 rounds

Diagnostics

  • predictedValue: PSI 0.858 (significant drift vs baseline)
  • pWin: PSI 1.861 (significant drift vs baseline)
  • pPodium: PSI 1.281 (significant drift vs baseline)
  • meanFinish: PSI 1.069 (significant drift vs baseline)
  • finishRangeHigh: PSI 1.814 (significant drift vs baseline)

Feature drift and rolling-Brier are tracked round-to-round; a spike flags where the field behaved unlike the rounds the model learned from.