Skip to content

Stubborn or jumpy? Bias and variance

A model can be too rigid to learn, or so sensitive that one result rewrites everything it believes. Those are bias and variance, and twenty-five years of SPFL results show what each costs.

Beginner Part 5 of Machine Learning Through Football

Contents

The football question

Celtic lose at home to Rangers: their first defeat of the season. What should a prediction model make of it?

One model shrugs: Celtic are still top, so it backs them again next week. Another tears up its notes and decides Rangers are winning the league by twenty points. Neither reaction is right, and the two ways of getting it wrong have names: bias and variance.

The concept

A model can be too simple and miss the pattern, or so complicated it starts learning the noise. Bias and variance are the reasons why.

Bias: consistently missing the same thing

Say I build a prediction model using only league position. Celtic are top, Rangers are third, so the model backs Celtic. It knows nothing about xG, injuries, home advantage, form or who's suspended.

That's high bias. The model's predictions are consistent, but they're consistently missing the same parts of the picture. It's why high bias goes hand in hand with underfitting, from underfitting vs overfitting.

Variance: changing its mind at the slightest nudge

Give a second model everything: recent results, xG, xGA, shots, possession, rest days, referee, previous meetings, the weather, what the manager had for breakfast. This one becomes incredibly sensitive to the exact matches it was trained on. Train it on one set of games, get one lot of predictions. Nudge the training data slightly and the predictions swing all over the place.

That sensitivity is variance. A high-variance model is saying "change my training data a little and I might change my mind a lot". It's why high variance goes with overfitting, as in overfitting.

The trade-off

So the two extremes are a model too rigid to learn the useful patterns, and a model so twitchy it reacts to every quirk in the data. What we want is somewhere in between: flexible enough to learn real football patterns, stable enough that one daft weekend doesn't rewrite everything it believes.

Make the model more flexible and bias comes down. Keep adding complexity and eventually variance takes over.

A football example

One setting controls how twitchy a model is: its memory. Rate each team by its points per game over its last k league matches, and back whichever side rates higher. A short memory reacts to every result. A long memory barely notices one.

How far can a single result move the rating? A result is worth 0, 1 or 3 points, so one match can shift the average by up to

$$\text{biggest move} = \frac{3}{k} \text{ points a game}$$

In plain football

  • k is how many past matches the rating remembers.
  • 3 is the gap between a win and a defeat.
  • With k = 1, one result moves the rating by up to 3 points a game: the whole scale. With k = 38, a season's worth, the most one result can do is 0.08.

So when Celtic lose at home to Rangers, the one-match model now rates Rangers at 3 points a game and Celtic at 0, and backs Rangers against anyone until Celtic win again. The 38-match model knocks Celtic down by 3/38 = 0.08 and carries on. The first has torn up its notes. The second has learned a little from the result without deciding everything it knew about football was wrong.

Now test every memory length on real Scottish Premiership results from 2000/01 to 2025/26. To keep it fair, use the same 4,097 matches for every model: those where both teams already had at least 152 Premiership matches behind them.

The share of results each memory length called right. The twitchiest model, remembering one match, is barely better than always picking a home win. Accuracy climbs as the memory lengthens, then levels off at around a season.
Remembers Average move per result Right
1 match 1.29 46.1%
3 matches 0.42 47.7%
10 matches 0.13 50.1%
38 matches 0.03 51.5%
152 matches 0.01 51.8%

The one-match model is all variance. Every result rewrites its opinion, by 1.29 points a game on average, and it scores 46.1%, barely clear of always picking a home win (44.0%). Lengthen the memory and it steadies: by 20 to 38 matches the rating moves a few hundredths of a point per result, and accuracy settles at about 51.5–51.8%.

Where's the bias?

At the long end, you'd expect bias to creep back in: a rating that remembers four seasons is slow to notice a team getting better or worse. In the Scottish Premiership, that barely costs anything. Remembering 152 matches does as well as remembering 38, because the strength of most sides changes slowly, and the top two change hardly at all.

The bias in these models is elsewhere. Every one of them uses a single feature, points per game. None knows about home advantage, injuries or form in the way a bookmaker does, and all of them level off around 52%, where the bookmakers reach 56% (underfitting vs overfitting). A longer memory can't fix that. Only better features can.

Show the mathsHow bias and variance add up to the total error. Optional.

For a model predicting a number, such as goals, the average squared error on a new match splits into three parts:

$$\begin{aligned} &E\big[(y - \hat{f}(x))^2\big] \\ &= \text{Bias}^2 + \text{Variance} + \sigma^2 \end{aligned}$$

  • Bias is how far the model's average prediction, over all the training sets it might have seen, is from the truth.
  • Variance is how much its prediction jumps around from one training set to another.
  • σ² is the noise: the part of football no model can predict. Deflections, penalties, a keeper's bad day.

Flexibility trades the first two against each other. The third never goes away, which is why even the bookmakers get nearly half of all results wrong.

Why it matters

  • Two different failures need two different fixes. High bias needs more or better features; high variance needs a simpler model or more data.
  • Stability is a feature. A model whose predictions swing on one result will swing the wrong way as often as the right way.
  • More data calms variance. A 38-match memory beats a 3-match one because it averages over more football, not because it's cleverer.
  • Nobody cares how well your model predicts games that have already happened. The question isn't "how closely does this fit the matches we've already seen?" It's "will the patterns it learned still hold when the next fixtures arrive?"

Limitations

  • Memory is one dial among many. Real models have lots of settings that trade bias against variance: how many features, how deep a tree, how strong a penalty.
  • The SPFL is unusually stable. With two clubs so far ahead, a long memory costs little here. In a league where fortunes change faster, it may cost more.
  • Accuracy hides the probabilities. A sensible model would move its probabilities a little after a shock result, not flip its pick. That's where the real difference between a steady and a twitchy model shows.
  • Weighting recent matches more is the usual compromise. Many ratings, such as Elo, update a little after every result rather than using a fixed window.

Because Saturday is the exam.

Try it yourself

After your team's next shock result, write down what each of three pundits would predict for the following match: one who only remembers last week, one who remembers the last ten, and one who remembers the whole season. Which sounds most like the reaction on social media, and which turns out to be right?

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:

import csv
from collections import defaultdict
from datetime import datetime

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
MEMORY = [1, 2, 3, 5, 10, 20, 38, 76, 152]  # how many past league matches a team's rating remembers

games = []
for y in range(2000, 2026):
    with open(f"SC0_{y % 100:02d}{(y + 1) % 100:02d}.csv", encoding="latin-1") as f:
        games += [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))

rating = lambda past, k: sum(past[-k:]) / k  # points per game over the last k matches
history, right, moved, n, updates, home = defaultdict(list), defaultdict(int), defaultdict(float), 0, 0, 0
for r in games:
    h, a = r["HomeTeam"], r["AwayTeam"]
    if min(len(history[h]), len(history[a])) >= max(MEMORY):  # same matches for every memory length
        n += 1
        home += r["FTR"] == "H"
        for k in MEMORY:
            right[k] += ("H" if rating(history[h], k) >= rating(history[a], k) else "A") == r["FTR"]
    for team, p in zip((h, a), POINTS[r["FTR"]]):
        if len(history[team]) >= max(MEMORY):
            updates += 1
            for k in MEMORY:
                moved[k] += abs(rating(history[team] + [p], k) - rating(history[team], k))
        history[team].append(p)

print(f"{n} matches; always a home win {home / n:.1%}")
for k in MEMORY:
    print(f"last {k:3} matches: right {right[k] / n:.1%}, rating moves {moved[k] / updates:.2f} points a game per result")

Further reading

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.