Too simple or too clever? Underfitting vs overfitting
A model can fail by being too simple to see the patterns or too complicated to ignore the noise. Real SPFL seasons, and the bookmakers, show what each looks like and where the sweet spot sits.
Beginner Part 4 of Machine Learning Through Football
Contents
The football question
Two coaches. One says "just give it to the best player." The other has a 47-step plan for every throw-in, wind direction and blade of grass.
Which one has the better game plan?
Probably neither. The first is too simple to cope with anything the opponent does. The second is so detailed that it falls apart the moment the match doesn't go to plan. Machine learning models go wrong in exactly the same two ways.
The concept
A model can miss the target from either side:
- Underfitting: the model's too simple. It misses important patterns and does poorly even on the data it was trained on.
- Overfitting: the model's too complicated. It learns the training data far too closely, noise and quirks included, then struggles on matches it hasn't seen. That's the model in overfitting: brilliant in training, gone on matchday.
The coach who says "give it to the best player" is probably underfitting. The one with the 47-step throw-in plan is possibly overfitting.
What you want sits in between: complex enough to pick up the patterns that matter, but not so complex it memorises every strange little thing that happened in the past.
How to tell which one you've got
Put the training score next to the test score, as in training data and test data, and compare both with a naive baseline:
| Training | Test | Probably |
|---|---|---|
| Poor | Poor, close to the baseline | Underfitting |
| Good | Close to training | About right |
| Excellent | Drops sharply | Overfitting |
For football results the naive baseline is easy: just predicting a home win every time gets you right around 45% of the time. A decent model should be pushing into the low-to-mid 50s, which is about where the bookmakers land.
A football example
Real Scottish Premiership results again. Train on 2000/01 to 2020/21 and test on the five seasons since, 2021/22 to 2025/26 (990 matches), skipping each season's opening matches until both sides have played five. Four models, from too simple to too clever, plus the bookmakers:
- Always a home win. No features at all.
- Higher in the table. One feature: pick whichever side is higher in the league table on the morning of the match.
- Form + table position. Two features, each in five bands: the gap in points per game over the last five matches, as in features and targets, and the gap in league position. 25 groups, and for each the most common result in training.
- Everything, down to the teams. Form, the last two results, the day, the month and the two teams: the 3,995-group lookup table from overfitting.
- The bookmaker's favourite. Whichever result Bet365 priced shortest before kick-off. It doesn't learn from our training seasons, so it only has a test score.
| Model | Probably |
|---|---|
| Always a home win | Underfitting |
| Higher in the table | Simple, but decent |
| Form + table position | About right |
| Everything, down to the teams | Overfitting |
| Bookmaker's favourite | The benchmark |
Always a home win is underfitting in its purest form: poor on training, poor on test, and it is the baseline.
Everything, down to the teams is the opposite: 99% on the matches it memorised, then back to the home-win baseline on seasons it hadn't seen.
Form + table position sits in the sweet spot: its test score (53.5%) is as good as its training score, so what it learned was real. That's low-to-mid 50s, right where a decent model should be.
League position alone
Predicting matches using just one feature, league position, sounds like a textbook underfit. Tested, it's better than that: 52.6%, five points clear of the home-win baseline and only a point behind form + table position. League position already sums up a lot of football: goals, results, the quality of the squad.
But one feature can only go so far. It never predicts a draw, knows nothing about form or injuries, and falls short of the bookmakers by 3.6 points.
The bookmakers
The bookmaker's favourite came in 56.2% of the time. Bookmakers aren't doing magic. They have more and better features than our lookup tables: team news, injuries, suspensions, and a market of bettors who punish any price that's wrong. Beating them consistently is very hard, which is why they make a good bar to aim at.
With 990 test matches, each score here is give or take about 3 points, so the middle three can't really be split. The two ends can: the home-win baseline and the overfitted table at one end, the bookmakers at the other.
Why it matters
- There are two ways to fail. A model that does poorly on test data may be too simple or too complicated, and the fix is opposite: add features, or take them away. Bias and variance explains why.
- Training and test side by side diagnose it. Both poor means underfitting; a big gap means overfitting.
- Always compare with a baseline. A model that scores 47% sounds fine until you learn that always backing the home side does the same.
- Simple goes a long way. One well-chosen feature got within a point of the best of our models. Start simple, and add only what earns its place on unseen matches.
Limitations
- Accuracy only counts the favourite. The bookmakers' real strength is in their probabilities, not just who they make favourite. Scores that grade probabilities are in evaluating prediction models.
- These models are crude. Better methods sit between the extremes more gracefully; the same diagnosis still applies to them.
- "About right" moves. With more data, a model can afford more detail before it overfits. The sweet spot depends on how much football you've got.
- Draws are the hard part. None of these models predicts draws usefully, and they're nearly a quarter of results.
What we really want is a model that performs well on data it's never seen. We're not trying to explain last weekend perfectly. We want something that still works when Saturday comes round.
Try it yourself
Before your team's next five matches, predict each result two ways: first with one rule ("the team higher in the table wins"), then with everything you know (form, injuries, who's rested, the weather). Afterwards, count how many each got right. Did all that extra detail actually help?
Reproduce the analysis
The results and odds files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:
import csv
from collections import Counter, defaultdict
from datetime import datetime
POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
def season(s): # one row per match, with features known before kick-off
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
for r in games:
r["day"] = datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y")
games.sort(key=lambda r: r["day"])
pts, res, gd, gf, rows = defaultdict(list), defaultdict(list), Counter(), Counter(), []
for r in games:
h, a = r["HomeTeam"], r["AwayTeam"]
if len(pts[h]) >= 5 and len(pts[a]) >= 5:
gap = (sum(pts[h][-5:]) - sum(pts[a][-5:])) / 5
r["form"] = 4 - ((gap <= -1) + (gap < -0.4) + (gap <= 0.4) + (gap < 1)) # 0 away far better ... 4 home far better
table = sorted(pts, key=lambda t: (sum(pts[t]), gd[t], gf[t]), reverse=True)
r["pos_h"], r["pos_a"] = table.index(h) + 1, table.index(a) + 1
d = r["pos_a"] - r["pos_h"] # places the home side is above the away side
r["pos"] = (d >= -6) + (d >= -2) + (d > 2) + (d > 6) # 0 away far higher ... 4 home far higher
r["last2"] = "".join(res[h][-2:])
r["weekday"], r["month"] = r["day"].strftime("%a"), r["day"].month
rows.append(r)
hg, ag = int(r["FTHG"]), int(r["FTAG"])
gd[h] += hg - ag; gd[a] += ag - hg; gf[h] += hg; gf[a] += ag
for team, p in zip((h, a), POINTS[r["FTR"]]):
pts[team].append(p)
res[team].append({3: "W", 1: "D", 0: "L"}[p])
return rows
def fit(train, keys): # most common result for each combination seen in training; otherwise a home win
seen = defaultdict(Counter)
for r in train:
seen[tuple(r[k] for k in keys)][r["FTR"]] += 1
return lambda r: seen[k].most_common(1)[0][0] if (k := tuple(r[x] for x in keys)) in seen else "H"
def accuracy(predict, rows):
rows = [r for r in rows if predict(r)]
return sum(predict(r) == r["FTR"] for r in rows) / len(rows)
def favourite(r): # the bookmaker's shortest price; None where there are no odds
try:
odds = {k: float(r["B365" + k]) for k in "HDA"}
except (KeyError, ValueError):
return None
return min(odds, key=odds.get)
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
train = [r for s in names[:21] for r in season(s)] # 2000/01-2020/21
test = [r for s in names[21:] for r in season(s)] # 2021/22-2025/26
models = {
"always a home win": fit(train, ()),
"higher in the table": lambda r: "H" if r["pos_h"] < r["pos_a"] else "A",
"form + table position": fit(train, ("form", "pos")),
"everything, down to the teams": fit(train, ("form", "last2", "weekday", "month", "HomeTeam", "AwayTeam")),
}
for name, predict in models.items():
print(f"{name:30} training {accuracy(predict, train):.1%} test {accuracy(predict, test):.1%}")
print(f"{'bookmaker favourite':30} test {accuracy(favourite, test):.1%}")
Further reading
- Underfitting vs. Overfitting, scikit-learn. Three curves, too simple, about right and too complex, fitted to the same points.
- Overfitting, Wikipedia. Both failures side by side, with the section on underfitting.
- Bias–variance tradeoff, Wikipedia. The theory underneath: why too simple and too complex fail in different ways.