Model accuracy
Does the number hold up?
Every 2025 game replayed with the rating each team carried before that game — the same calculation the site shows for Week 1, with no hindsight. Then the 2026 season, tracked game by game as results land.
Tuning
Fitted to 2025, not to a textbook
The model started with the usual Elo defaults. Each constant was swept against the 272 stored 2025 games, fit on Weeks 1-13 and validated on Weeks 14-18 — a setting only survives if it also helps on games it was not fit to. Log loss is the objective; anything worth less than 0.002 was left at its round-number default.
Straight-up
61.8% → 62.9%
held-out weeks 64.1% → 66.7%
Brier score
0.226 → 0.224
held-out weeks 0.225 → 0.222
Log loss
0.644 → 0.639
held-out weeks 0.637 → 0.628
| Constant | Was | Now | Held-out LL |
|---|---|---|---|
K-factor Ratings were moving too slowly; 24-28 all beat 20 on both halves of the season. | 20 | 26 | -0.0015 |
Home field Home teams went 142-123 (.536) in 2025. A 55-point bonus was pricing an edge the season did not show. | 55 pts | 35 pts | -0.0088 |
MOV dampener Flat from 1.6 to 3.0 — the blowout dampener barely moves the grade. Left alone. | 2.2 | 2.2 | kept |
Spread divisor 22 cut margin error from 10.265 to 10.255 points per game. Noise, so the round number stays. | 25 | 25 | kept |
Offseason regression 30-35% graded a hair better on the in-season proxy, inside the noise guard. | 25% | 25% | kept |
Blend weights Elo already prices margin of victory, so the scoring terms were double-counting it. | 55 / 20 / 15 / 10 | 80 / 5 / 15 / 0 | -0.0059 |
Honest read: the two changes that mattered were a faster K and a smaller home-field bonus — home teams went 142-123 (.536) in 2025, so 55 points was overpricing the host. Together they are worth about one extra winner per 100 games. Over 272 games that is real but modest; NFL Elo does not have much more room than this.
Straight-up
63.2%
172 of 272 winners called in 2025
Brier score
0.224
Lower is better. 0.25 = always guessing 50/50
Avg spread error
10.3 pts
44% of games within a touchdown
Log loss
0.639
Penalises confident wrong calls hardest
Accuracy by confidence band
The bands on every matchup card, judged against reality. A "Strong" call should win far more often than a "Lean".
Coin flip
58%
43-31 · claimed 53%
Lean
58%
63-46 · claimed 59%
Confident
65%
35-19 · claimed 70%
Strong
89%
31-4 · claimed 81%
Calibration
When the model says 70%, does it win 70% of the time?
Grey bar: what the model claimed. Coloured bar: what actually happened. A well-calibrated model has them roughly equal in every band.
2026 season tracker
No 2026 games have final scores yet. 16 Week 1 predictions are on the board and this table fills in as they are played. See the slate →
Worst calls of 2025
The games the model was most sure about and still got wrong.
| Week | Game | Final | Model pick | Result |
|---|---|---|---|---|
| Wk 18 | WSH at PHI | 24–17 | PHI 83% | Miss |
| Wk 13 | CIN at BAL | 32–14 | BAL 79% | Miss |
| Wk 13 | LAR at CAR | 28–31 | LAR 77% | Miss |
| Wk 17 | LAR at ATL | 24–27 | LAR 76% | Miss |
| Wk 14 | NO at TB | 24–20 | TB 74% | Miss |
Sharpest calls of 2025
Right winner, and the model spread landed nearly on the final margin.
| Week | Game | Final | Model pick | Result |
|---|---|---|---|---|
| Wk 18 | NO at ATL | 17–19 | ATL 57% | Hit |
| Wk 3 | DEN at LAC | 20–23 | LAC 61% | Hit |
| Wk 15 | TEN at SF | 24–37 | SF 86% | Hit |
| Wk 16 | NE at BAL | 28–24 | NE 63% | Hit |
| Wk 10 | ATL at IND | 25–31 | IND 71% | Hit |
Week by week, 2025
Honest caveat: the 2025 backtest is in-sample in one sense — ratings were built from the same season — but each prediction only uses ratings as they stood before that kickoff, so no game informs its own forecast. Early-season weeks are the weakest, because every team starts the year near the league average.