Skip to content

Brilliant in training, gone on matchday. Overfitting

A model can learn its training matches too well, rules and all, and then fall apart on matches it hasn't seen. Real SPFL seasons show how it happens, how to spot it, and what to do about it.

Beginner Part 3 of Machine Learning Through Football

Contents

The football question

A model predicts the result of 99% of its training matches correctly. Is it a good model?

It sounds like it must be. But there's a player like this at every club: unplayable in training, then vanishes on matchday. What counts is Saturday.

The concept

Overfitting is when a model learns the training data too well. Instead of the underlying patterns, it learns the quirks: the coincidences and the noise that happened to be in those particular matches.

What we actually want is a model that picks up the underlying patterns well enough to do a decent job on matches it's never seen. Recent xG and xGA, home advantage, opponent strength, injuries, rest days and recent form all carry real information about the next match. An overfitted model goes further, and finds "rules" like these:

  • This team always wins after two draws and a Tuesday kick-off.
  • This team struggles when they had exactly 7 shots on target the week before.
  • This team always beats that opponent in the rain in October.

Each might be true of the training data. None is football. It's the team that's rehearsed one set piece a hundred times and has no idea what to do once the game gets scrappy.

The warning sign is a gap:

  • Training performance: excellent.
  • Test performance: a lot worse.

That gap, between matches the model learned from and matches it hasn't seen, is how you spot overfitting. It's why training data and test data are kept apart.

A football example

Let's build an overfitted model on purpose. The model is a lookup table: group the training matches by some features, and for each group, predict whichever result happened most often. Then make the groups more and more specific, one feature at a time, using only things known before kick-off:

  1. Nothing: one group, so it always picks a home win.
  2. + form: the gap between the two sides' points per game over their last five matches, in five bands, as in features and targets.
  3. + last 2: the home side's last two results, such as "two draws".
  4. + day: the day of the week it's played.
  5. + month: the month.
  6. + teams: which two teams are playing.

Train on the Scottish Premiership from 2000/01 to 2020/21 (4,096 matches) and test on the five seasons since, 2021/22 to 2025/26 (990 matches). Each season's opening matches are skipped until both sides have played five, as there's no form to measure yet.

Every extra detail pushes training accuracy up. Test accuracy only improves while the rules stay broad, then falls away. The shaded gap is the warning sign.
Rule uses Groups Training Test
Nothing 1 43.8% 47.2%
+ form 5 49.7% 51.8%
+ last 2 45 50.1% 53.0%
+ day 252 53.4% 51.2%
+ month 1,148 64.3% 46.5%
+ teams 3,995 99.0% 47.8%

Training accuracy climbs all the way from 43.8% to 99%. Test accuracy tells a different story: it improves with form, levels off, then drops back to about what always picking a home win scores.

The Groups column shows why. Split 4,096 matches into 1,148 groups and most groups are tiny: 473 hold a single match. A group of one always "predicts" its own result perfectly. That's not learning, it's the memoriser from training data and test data in disguise. By the last rule there are 3,995 groups for 4,096 matches, so nearly every match has a group to itself. Hence 99%.

The October rule

One of those groups is almost exactly the kind of rule at the top of this page. A home side in better form, that has won its last two, playing on a Saturday in October. In the 21 training seasons, that happened four times, and the away side won all four.

The model learned it: in-form home side, two wins, October Saturday, back the away team. In the five test seasons the same situation came up three times. The home side won two and drew one. The rule was wrong every time. Four matches was never evidence of anything; it was a coincidence the model mistook for football.

Tackling overfitting

There are ways to tackle overfitting:

  • Simpler models. Fewer features, fewer groups, fewer chances to memorise.
  • A validation set, or cross-validation. Hold back part of the training data, or take turns holding back each part, and choose the model on that, never on the test set.
  • Regularisation. Penalise a model for being complicated, so it only uses detail that really earns its place.
  • Pruning. Grow a detailed decision tree, then cut back the branches that only fit a handful of matches.

Here's the validation set at work. Learn each rule from 2000/01 to 2017/18, and score it on 2018/19 to 2020/21. The test seasons stay untouched. On validation, form alone and form plus the last two results tie exactly, 275 of 542 right each, and the more detailed rules all score lower. On a tie, take the simpler model. Refitted on all the training seasons, the form rule scores 51.8% on the test seasons: close to the best any rule managed, found without ever peeking at the test.

Why it matters

  • A model that's brilliant on training data isn't automatically a good model. What matters is how it does on data it hasn't seen.
  • Detail isn't free. Every extra feature splits the data thinner, until the model is fitting single matches.
  • Football invites overfitting. A season is only a few hundred matches, and there are endless things to slice it by: days, months, referees, weather, kits. Some slice will always look like a pattern.
  • The gap is the diagnosis. Look at training and test scores side by side. A big gap means the model has learned the training data, not the game.

Limitations

  • These are deliberately crude models. Lookup tables overfit quickly; proper models overfit more subtly, but the same gap gives them away.
  • Five test seasons still wobble. With 990 test matches, each accuracy is give or take about 3 points, so form (51.8%) and form plus the last two results (53.0%) can't really be told apart.
  • Some specific rules are real. A team that genuinely struggles away in midweek after European trips might be a real pattern. The test is whether it holds on matches the model hasn't seen.
  • The opposite problem exists too. A model that's too simple to learn the real patterns is underfitting, and the art is landing between the two.

It looked like prime Glasgow Celtic in training. Then Saturday came.

Try it yourself

Think of a "rule" you've heard about your team: we never win on Boxing Day, we always lose after an international break, that striker only scores in cup ties. How many matches is it based on? Write down what it predicts for the next few times it applies, then check. Does it hold, or was it the October rule?

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:

import csv
from collections import Counter, defaultdict
from datetime import datetime

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}

def season(s):  # one row per match, with features known before kick-off
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    for r in games:
        r["day"] = datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y")
    games.sort(key=lambda r: r["day"])
    pts, res, rows = defaultdict(list), defaultdict(list), []
    for r in games:
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(pts[h]) >= 5 and len(pts[a]) >= 5:
            gap = (sum(pts[h][-5:]) - sum(pts[a][-5:])) / 5
            r["form"] = 4 - ((gap <= -1) + (gap < -0.4) + (gap <= 0.4) + (gap < 1))  # 0 away far better ... 4 home far better
            r["last2"] = "".join(res[h][-2:])  # home side's last two results
            r["weekday"], r["month"] = r["day"].strftime("%a"), r["day"].month
            rows.append(r)
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            pts[team].append(p)
            res[team].append({3: "W", 1: "D", 0: "L"}[p])
    return rows

RULES = [(), ("form",), ("form", "last2"), ("form", "last2", "weekday"),
         ("form", "last2", "weekday", "month"), ("form", "last2", "weekday", "month", "HomeTeam", "AwayTeam")]

def fit(train, keys):  # most common result for each combination seen in training; otherwise a home win
    seen = defaultdict(Counter)
    for r in train:
        seen[tuple(r[k] for k in keys)][r["FTR"]] += 1
    return lambda r: seen[k].most_common(1)[0][0] if (k := tuple(r[x] for x in keys)) in seen else "H"

def accuracy(predict, rows):
    return sum(predict(r) == r["FTR"] for r in rows) / len(rows)

names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
data = {s: season(s) for s in names}
rows = lambda first, last: [r for s in names[first:last] for r in data[s]]
train, test = rows(0, 21), rows(21, 26)  # 2000/01-2020/21, 2021/22-2025/26
for keys in RULES:
    predict = fit(train, keys)
    print(f"{' + '.join(keys) or 'nothing':50} training {accuracy(predict, train):.1%}  test {accuracy(predict, test):.1%}")

# choose the rule on a validation set (2018/19-2020/21), never the test set; on a tie, the simpler rule
fit_on, valid = rows(0, 18), rows(18, 21)
best = max(RULES, key=lambda k: (accuracy(fit(fit_on, k), valid), -len(k)))
print("chosen:", best, f"test {accuracy(fit(train, best), test):.1%}")

# the October rule: home side better (form 3), won its last two, Saturday in October
october = lambda r: (r["form"], r["last2"], r["weekday"], r["month"]) == (3, "WW", "Sat", 10)
print(Counter(r["FTR"] for r in train if october(r)), Counter(r["FTR"] for r in test if october(r)))

Further reading

  • Overfitting, Google for Developers. The training-against-validation gap, with loss curves, and what causes overfitting.
  • Overfitting, Wikipedia. Overfitting and underfitting, with the classic picture of a curve threaded through every point.
  • Underfitting vs. Overfitting, scikit-learn. A short worked example that uses cross-validation to pick the right amount of complexity.

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.