Hunting for Weather-Influenced Upsets
The last post found that weather doesn't change who wins a game, on average. That got us wondering about a sneakier possibility: what if it doesn't move the average, but it still causes the occasional upset — just scattered in both directions, canceling out when you zoom out? We went looking for those games. What we found is a good lesson in not fooling yourself.
How do you even measure a "surprise"?
First we needed a stand-in for "what should have happened." We trained a simple model to guess a game's final margin using only team strength (SP+ ratings) — nothing about weather at all — and had it guess every game in our database the honest way: by only ever looking at games it hadn't already seen the answer to, the same way you'd have to predict a game before it's played. The difference between what actually happened and what the model expected is that game's surprise, measured in points.
Even with zero weather involved, the typical game still surprises the model by about 11 points. Football has a lot of randomness in it on a normal, sunny Saturday. That number matters because it's the bar any "weather-caused" upset has to clear to mean anything — an 11-point surprise isn't a story, it's a Tuesday.
Does bad weather make games harder to predict at all?
Before hunting for individual games, we checked the more basic question directly: are surprises simply bigger, on average, when the weather's bad — even if they go in both directions and cancel out? If weather were a "great equalizer," this is where it would show up.
None of those differences are big enough to trust as real (a statistician would say they're not statistically significant, and we checked). If anything, snow games lean slightly more predictable, which fits something we found in the last post: teams play it safer in the snow, leaning on the run instead of risking a windblown pass — a more conservative game plan producing a more predictable game, not a wilder one.
So here's the honest conclusion: weather doesn't make results less predictable, at any level of "bad" we tried. Which means the list below is exactly what it looks like — the biggest upsets that happened to occur on a rough-weather day — not proof that rough-weather days produce more upsets. We're publishing it anyway because it's a genuinely fun list of games, not because we think the weather caused any of them.
The list
The 30 biggest gaps between "what the model expected from team strength alone" and "what actually happened," among games with measurable rain, any snow, or wind at 15+ mph. An Upset badge means the team the model favored actually lost outright — not just won by less than expected.
| Date | Matchup | Score | Expected | Surprise | Conditions |
|---|---|---|---|---|---|
| Loading… | |||||
What would actually change our mind
One list of upsets proves nothing by itself — that's the whole reason we checked the predictability question first, before looking at individual games. To turn "curiosity" into "real finding," we'd want to see the same pattern show up consistently across independent slices of data: different seasons tested separately, a genuine pre-game weather forecast instead of the actual weather that happened, different levels of competition. We haven't seen that yet. If a future re-run of this, on more data, finds a real effect, that'll be the next post.