Methodology
How VotePredictor turns polls into calibrated forecasts — and an honest account of what the model does and doesn't add.
The model
Forecasts come from TabPFN, a tabular foundation model. Election forecasting is a small-data problem — a few hundred comparable races per cycle, not millions — which is exactly where this class of model is strong. For each race we build a feature row from its polling (a recency-weighted average, the trend, how much pollsters disagree, how many polls) plus the seat's prior result, and the model predicts the win probability and final margin.
Trained on 25 years, tested out-of-sample
The model trains on every polled Senate, Governor, House and presidential race from 1998–2025. We validated it with a strict walk-forward backtest: to forecast any past cycle, the model only ever saw earlier cycles — never the outcome it was being graded on. No result data leaks into the inputs.
What we found (and what we don't claim)
- Aggregating polls is the big win. Going from a single poll to a race-level average lifts winner accuracy from ~80% to ~88%.
- The model's edge is sharper margins and calibration, not the winner call (the poll average already calls most races right). Its predicted margins beat a naive average, and its probabilities are well-calibrated.
- Calibration matters more than bravado. Over many races, our 70% calls win about 70% of the time. We'd rather be honestly uncertain than falsely confident.
Does the fancy model actually beat a simple average — and beat other models?
A fair question, asked two ways. We raced TabPFN against a 14-model field — random forests, XGBoost, LightGBM, CatBoost, a Gaussian process, a neural net, kernel and nearest-neighbor methods, regularized linears — on the same races and the same walk-forward rules. Nothing beats it, and the plain bias-corrected poll average — no machine learning at all — beats every model except TabPFN. TabPFN is built for exactly this small-data regime, which is why it's the one learner that adds value over averaging — the aggregation does the heavy lifting, and the model is a small, real bonus on top. See the full bake-off on the AI Models page.
Pollster ratings
Each pollster is rated on its final poll before each election — the call that actually matters — against the result (see the leaderboard). The score is sample-size-adjusted, so a lucky run over a few races doesn't outrank a steady record over hundreds. Ratings are walk-forward inside the model — a pollster's rating on any date uses only its earlier polls. We use them to de-bias polls (subtract a pollster's lean) rather than to weight by accuracy, which our backtest showed doesn't help.
We also de-bias by voter screen. On full-campaign data, registered-voter and adult-sample polls overstate the Democrat by ~2 points versus likely-voter polls of the same races late in the cycle, and correcting it cut race error ~6% out-of-sample. So we shrink the Democratic margin on RV/adult polls toward their likely-voter equivalent — but only as the election nears (where the effect is real and measured), not months out where it isn't.
Polls and fundamentals: what each race is forecast from
Senate and Governor. Polled races are forecast from polls by the race model above, then shrunk toward a fundamentals forecast by inverse variance: the fundamentals model (state presidential lean, incumbency, the seat's prior result, the national environment) is validated walk-forward against naive baselines and carries its own spread, about three to four times a polled race's, so it takes roughly 7% of the weight. Races nobody polls are forecast from the fundamentals model alone and marked as such. Chamber control comes from a Monte-Carlo that draws one correlated national swing across all races each run, so it reflects the real risk that polls miss in the same direction everywhere (independent draws look far too confident).
House.The published margin for each of the 435 districts is the district's 2024 presidential margin on its enacted 2026 lines, relative to the national vote, plus one fitted intercept, plus the national generic ballot as a uniform swing — a two-parameter rule that beat a tree on lean, incumbency, money and expert ratings on every metric in every walk-forward cycle. The win probability adds a small incumbency premium, re-estimated each cycle with a two-cycle half-life because the premium keeps shrinking. Where a district has polls, the poll average is blended in by inverse variance, with about half the weight on a single poll and more as polls accumulate; the same number is what every page, the API and the seat simulation use. The House simulation carries its own national term for generic-ballot error at a six-week horizon.
Each of these rules was preregistered — frozen with its gates before it shipped and before any 2026 outcome — and its walk-forward record is in the repository's ledger. On the rankings board VotePredictor is scored on the frozen race model's calibrated probabilities, with no House head blended in: the board is the model's own out-of-sample record, not a tuned version of it.
How much is an early poll worth?
A poll six months out is not a poll on election eve. Using the full-campaign poll feeds across five cycle-office series (President, Senate and Governor races in 2020, 2022 and 2024 — every poll with its field dates), we bucketed polls by how far they were from the election and measured how far the average missed the final margin.
Two lessons, and they pull in opposite directions. In a normal cycle (2024), polls converge: the average miss grows from ~4 points on election eve to ~6 points three-plus months out, so earlier polls genuinely carry less information. But lead time only shrinks noise, not bias: in 2020 the polls were off by 5–7 points at every horizon — the late polls were actually the worst — because a systematic miss can't be outrun by waiting. That is exactly why our live forecast widens its uncertainty the further it is from November, and why being early is a smaller risk than the chance the whole cycle is tilted.
How we stack up against the pros
Pollster ratings grade the people who take polls. This grades VotePredictor's own model against the other forecasters — FiveThirtyEight, Cook, the Economist, Sabato, DDHQ, Silver Bulletin and a dozen more — using each outfit's final pre-election win probability per race (2016–2024), archived by JHK Forecasts. Everyone is scored identically: Brier score — the squared error of the win probability — on the same races and the same outcomes, against VotePredictor Elections' strictly out-of-sample (walk-forward) calls.
392 races where the final margin was within 15 points. Lower Brier is better. Only forecasters that rated every race in the set are ranked; the rest are listed unranked, with VotePredictor's Brier on exactly the races they rated.
| # | Forecaster | Races | Brier | Called right | VP, same races |
|---|---|---|---|---|---|
| 1 | FiveThirtyEight | 392 | 0.113 | 86% | 0.137 |
| 2 | Cook Political | 392 | 0.123 | 95% | 0.137 |
| 3 | Sabato's Crystal Ball | 392 | 0.129 | 83% | 0.137 |
| 4 | VotePredictor Elections | 392 | 0.137 | 82% | |
| — | Race to the WH | 172 | 0.116 | 81% | 0.127 |
| — | Inside Elections | 222 | 0.122 | 89% | 0.124 |
| — | JHK Forecasts | 255 | 0.123 | 81% | 0.134 |
| — | Fox News | 123 | 0.123 | 93% | 0.146 |
| — | Split Ticket | 137 | 0.127 | 82% | 0.139 |
| — | The Economist | 342 | 0.128 | 80% | 0.140 |
| — | CNalysis | 172 | 0.136 | 79% | 0.127 |
| — | DDHQ/Decision Desk | 357 | 0.137 | 83% | 0.137 |
| — | Elections Daily | 158 | 0.141 | 80% | 0.131 |
| — | RealClearPolitics | 133 | 0.199 | 79% | 0.151 |
The honest read: VotePredictor Elections lands in the pack but behind the leaders, and the gap is widest on the hard, genuinely competitive races. That tracks — it is the frozen race model scored on its own — no fundamentals shrink, no hand-adjustment — competing here with its walk-forward probability rather than the production version. It's the clearest evidence we have that priors buy something exactly where polls are thin and noisy, which is what the fundamentals shrink and the House poll blend above now act on.
Data & limits
- Live polls from VoteHub (CC-BY 4.0); historical polls from FiveThirtyEight and the Silver Bulletin; results + partisan lean from FiveThirtyEight / MIT.
- VoteHub omits candidate party, so we attach it from a curated crosswalk that we cross-check against Wikipedia.
- House district lean is the 2024 presidential result on each district's enacted 2026 lines, sourced per district for the ten states that redrew; a few borrowed figures are flagged on the district page.
- Race results are the certified candidate margins summed across fusion ballot lines; the poll archive's own outcome field had 14 New York and Connecticut winners wrong.
- This is an independent research project, not affiliated with any campaign, and not a betting product.