Live Calibration
Live scoring begins as 2026 races resolve. No 2026 race outcomes are recorded yet. The model is freezing its predictions weekly into forecast_snapshots so each race can be scored the moment its winner is recorded in race_outcomes.
Reliability plot
Each dot is a 10-percentage-point bucket of forecasts. The x-axis is the model's mean predicted P(D) in that bucket; the y-axis is the empirical D-win rate of those races. A perfectly calibrated model puts every dot on the diagonal. Dot size is the resolved-race count in the bucket. The plot is empty until 2026 races resolve.
Snapshots
The forecast snapshotter runs weekly (Monday 13:00 ET) and writes today's projected P(D) for every forward 2026 race into forecast_snapshots. So far: 526 snapshot rows total across all races; 0 snapshot rows are for races already resolved and contribute to the live Brier above.
Model vs market — in progress
Maine Senate. On the June 12, 2026 snapshot the model priced Maine at 6% D while the betting market priced it at 62% D. That is one of the largest model-vs-market gaps on the board this cycle. Since then, Graham Platner (D) won the Maine Democratic primary on June 9, 2026 — a nomination outcome consistent with the model’s contrarian read: the model has been leaning on Maine’s personal-vote spread for Susan Collins (R), which is what the market appears to have been under-weighting. This race is not yet resolved and is not in the live Brier above. We are publishing the snapshot now so the eventual result can be scored against both readings when Maine votes on November 3, 2026.
Framing note: this is model vs. market in progress, not a claim of a resolved race. If Collins wins by her usual personal-brand margin, the model looks right and the market looked expensive; if Platner wins, the market looked right and our personal-vote feature was over-weighted. Either way the outcome will be scored on this page in the live-Brier table below.
Backtest Brier (2024)
The published model lr-2026-06-10-chal-recal-twospeed reports a 2024 held-out test Brier of 0.1191. This is a backtest number, not live. This model serves on [2022,2024] (the two most recent labeled cycles), so it cannot hold out 2024 itself. The reported Brier is the genuine out-of-sample score of the identical narrow-winsor15 configuration trained on [2020,2022] and tested on the held-out 2024 cycle (the prior-window policy). Serving window and skill-measurement window are reported separately and never conflated. It is shown here for context only and is never combined with the live Brier above.