Learning from its mistakes. Gradient boosting
Start with a rough guess, look at what it got wrong, and fix a little of it with a small tree. Repeat a few hundred times. On five SPFL test seasons gradient boosting draws level with logistic regression, and shows where three numbers run out.
Intermediate Part 11 of Machine Learning Through Football
Contents
The football question
A manager watches the first half, spots what's going wrong, and makes one change at the break. Not a new team, one adjustment. Next match, another. Over a season, lots of small fixes add up to a side that's much better than where it started.
Can a model learn the same way: start with a rough guess, look at what it got wrong, and fix a little at a time? That's gradient boosting, and it's the last of the big tree methods in this series, after decision trees and random forests.
The concept
A random forest grows its trees side by side, independently, and averages them. Boosting grows them one after another, and each new tree has one job: fix what the model so far is getting wrong.
- Start simple. Before any trees, every match gets the base rates: how often home wins, draws and away wins happen in the training seasons.
- Measure the misses. For every training match and every result, the miss is what happened minus what the model said.
- Fit a small tree to the misses. Here, a tree that asks one question. It finds the cut that best separates the matches the model is underrating from the ones it's overrating.
- Take a small step. Nudge each result's score by a tenth of what the tree says, and go back to step 2.
Each round:
$$\begin{aligned} &\text{score}_{\text{new}} \\ &\quad = \text{score} + 0.1 \times \text{tree}(\text{misses}) \end{aligned}$$
In plain football
- The score is the same kind as in logistic regression: one for each result, turned into chances that add up to 1.
- The misses are 1 minus the chance the model gave, for the result that happened, and 0 minus the chance for the ones that didn't. A home win the model gave 30% leaves a miss of +0.7 on the home score: it should have been higher.
- 0.1 is the step size. Small steps mean each tree fixes only a little, so no single tree can drag the model off course.
If this looks familiar, it should. The misses are the downhill slope of the log loss, the same slope logistic regression followed downhill. Boosting is gradient descent where each step is a small tree, and that's where the name comes from.
The same three features as the last three parts, each the home side's figure minus the away side's, in points a game: recent form, this season so far and last season. The same 3,900 training matches, 2001/02 to 2020/21, and the same 990 test matches, 2021/22 to 2025/26.
When to stop
Keep boosting long enough and it will start fixing "mistakes" that were just luck: overfitting again. So how many rounds? Not by looking at the test seasons, which would spoil the exam. Instead, as overfitting suggested, hold some training seasons back: learn on 2001/02 to 2015/16 (2,963 matches), and check each round on 2016/17 to 2020/21 (937 matches).
| Trees ask | Best after | Held-back log loss |
|---|---|---|
| One question | 380 rounds | 0.9660 |
| Two questions | 160 rounds | 0.9668 |
Both settle at almost the same place. The two-question version gets there faster, then gets steadily worse, 0.9683 by round 280: each extra round is fitting noise. The one-question version is slower but steadier.
Then, with the number of rounds fixed, boost again on all 3,900 training matches and face the test seasons for the first time.
The result
| Model | Test accuracy | Test log loss |
|---|---|---|
| Base rates | 47.2% | 1.056 |
| Form + table position (lookup) | 53.5% | 0.997 |
| Forest, fully grown trees | 53.3% | 0.980 |
| Tree, two questions | 55.5% | 0.954 |
| Forest, groups of 100+ | 54.3% | 0.953 |
| Boosting, two-question trees | 54.5% | 0.953 |
| Boosting, one-question trees | 54.5% | 0.951 |
| Logistic regression | 54.8% | 0.950 |
| Bookmaker | 56.2% | 0.932 |
Boosting draws level with logistic regression, 0.951 against 0.950, and doesn't pass it.
What it learned
Because each tree asks only one question, the boosted model is a sum of simple pieces, one set for each gap. That means it can be taken apart: add up every step that asked about this season's gap, and you see exactly how the model thinks that gap matters, with no straight line forced on it.
- Through the middle, about the same climb. From −1 to +1 points a game, boosting rises about as far as logistic regression's line, which adds 0.74 per point for this season and 1.00 for last season, but in uneven steps: at +1 it has reached +0.97 on this season's gap and +1.31 on last season's.
- At the ends, boosting goes flat. Beyond +1.5 on this season's gap, or +1 on last season's, being even further ahead adds nothing more. On the minus side the shape is less tidy, with a small rebound at −2 for this season. The flat ends rest on a minority of matches: 267 of the 3,900 have a this-season gap beyond 1.5 either way, 578 a last-season gap beyond 1.
- Recent form hardly figures. Of 1,140 small trees, 210 asked about recent form, and all their steps together move the score by no more than 0.11 either way, against about 1.3 for each season gap. Every model in this series has reached the same verdict on form.
So boosting found a bend that logistic regression can't draw. It just doesn't matter much: the bend is at the extremes, where a minority of matches are, and where a strong favourite already has most of the chances.
Four models, one ceiling
Logistic regression, a single tree, a forest and now boosting all score within 0.004 of each other on the same test matches, 0.950 to 0.954. They work in completely different ways, and they hit the same ceiling.
That's the most useful result in this part of the series. The limit isn't the model; it's the three numbers going in. Recent form, this season and last season have been squeezed dry. The bookmakers, at 0.932, know things these three numbers can't tell us: injuries, suspensions, team news, a manager's first game, a squad that's been rotated. To get closer, a model needs better features, not a cleverer method, and rating every team's attack and defence is one way to get them.
Why it matters
- Boosting is one of the strongest methods there is on tables of data with many columns. Its well-known versions, such as XGBoost and LightGBM, win data science competitions and run inside many real prediction systems.
- It's gradient descent again. The same idea that finds the best dial in calculus and trains logistic regression, with trees as the steps.
- Held-back seasons choose the settings. The number of rounds came from 2016/17 to 2020/21, not from the test, so the test score is honest.
- A simple model that ties is the better choice. Logistic regression is faster, fits in a few numbers and explains itself. Boosting's value here is in what it shows: the straight line was nearly right.
Limitations
- Three features. Same ceiling as every model before it.
- Settings. The step size (0.1), the tree size and the smallest group (50 matches) were set by hand. Tuning them on the held-back seasons might gain a little.
- One held-back split. The rounds were chosen on one block of five seasons. Cross-validation over several blocks would be steadier.
- The flat ends are thin. They rest on a few hundred matches, so they may not hold in another league, or even in the next few seasons.
Try it yourself
Guess the chance of a home win for each of this weekend's matches. After the results, look at your misses: which kind of match did you get most wrong? Big favourites, even games, sides in form? Adjust your guesses for that kind of match next weekend, just a little. Keep doing that for a few weeks: you're boosting by hand.
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. The first half builds the same three features as logistic regression; the second boosts. It runs hundreds of rounds, so allow about half a minute:
import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import exp, log, sqrt
POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
FEATURES = ["recent form", "this season so far", "last season"]
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
def season(s):
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
return games
def points_per_game(games):
pts, n = Counter(), Counter()
for r in games:
for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
pts[team] += p
n[team] += 1
return {t: pts[t] / n[t] for t in n}
# three features for every match, each the home side's figure minus the away side's, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
last = points_per_game(season(s_last))
promoted = 0.85 * sum(last.values()) / len(last)
history = defaultdict(list)
for r in season(s):
h, a = r["HomeTeam"], r["AwayTeam"]
if len(history[h]) >= 5 and len(history[a]) >= 5:
rows.append((s, [
(sum(history[h][-5:]) - sum(history[a][-5:])) / 5, # points a game, last five
sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]), # points a game this season
last.get(h, promoted) - last.get(a, promoted), # points a game last season
], r["FTR"]))
for team, p in zip((h, a), POINTS[r["FTR"]]):
history[team].append(p)
train = [r for r in rows if r[0] < "2122"] # 2001/02-2020/21
test = [r for r in rows if r[0] >= "2122"] # 2021/22-2025/26
def tree(X, misses, rows, depth, smallest=50):
"""A small tree fitted to the misses: each question is the cut that best separates big misses from small ones."""
n, total = len(rows), sum(misses[i] for i in rows)
best = None
if depth > 0:
for j in range(len(FEATURES)):
rows = sorted(rows, key=lambda i: X[i][j])
left = 0.0
for k in range(1, n):
left += misses[rows[k - 1]]
if smallest <= k <= n - smallest and X[rows[k - 1]][j] < X[rows[k]][j]:
gain = left ** 2 / k + (total - left) ** 2 / (n - k)
if gain > (best[0] if best else total ** 2 / n + 1e-12):
best = (gain, j, (X[rows[k - 1]][j] + X[rows[k]][j]) / 2)
if not best: # a group: its average miss
return total / n
_, j, cut = best
return (j, cut, tree(X, misses, [i for i in rows if X[i][j] <= cut], depth - 1, smallest),
tree(X, misses, [i for i in rows if X[i][j] > cut], depth - 1, smallest))
def value(node, x):
while isinstance(node, tuple):
node = node[2] if x[node[0]] <= node[1] else node[3]
return node
def chances(score): # scores to probabilities that add up to 1, as in part 8
top = max(score.values())
e = {c: exp(v - top) for c, v in score.items()}
return {c: v / sum(e.values()) for c, v in e.items()}
def boost(data, rounds, depth, watch=None, rate=0.1):
"""Start from the base rates; each round, fit a tree to each result's misses and take a small step."""
X, Y = [x for _, x, _ in data], [y for _, _, y in data]
start = {c: log(Y.count(c) / len(Y)) for c in "HDA"}
scores = [dict(start) for _ in data]
watched = [dict(start) for _ in watch or []]
steps, history = [], {}
for r in range(1, rounds + 1):
p = [chances(s) for s in scores]
step = {c: tree(X, [(Y[i] == c) - p[i][c] for i in range(len(Y))], list(range(len(Y))), depth) for c in "HDA"}
steps.append(step)
for x, s in zip(X, scores):
for c in "HDA":
s[c] += rate * value(step[c], x)
if watch:
for (_, x, _), s in zip(watch, watched):
for c in "HDA":
s[c] += rate * value(step[c], x)
if r % 20 == 0:
history[r] = sum(-log(chances(s)[y]) for (_, _, y), s in zip(watch, watched)) / len(watch)
return start, steps, history
def forecast(model, x, rate=0.1):
start, steps, _ = model
return chances({c: start[c] + rate * sum(value(step[c], x) for step in steps) for c in "HDA"})
print(f"{len(train)} training matches, {len(test)} test matches")
# 1. how many rounds? Learn on 2001/02-2015/16, check on 2016/17-2020/21; the test seasons stay unseen
early = [r for r in train if r[0] < "1617"]
late = [r for r in train if r[0] >= "1617"]
rounds = {}
for depth, most in ((1, 500), (2, 300)):
history = boost(early, most, depth, watch=late)[2]
rounds[depth] = min(history, key=history.get)
print(f"{depth}-question trees, {len(early)} matches learned, {len(late)} checked: best after {rounds[depth]} rounds, "
f"{history[rounds[depth]]:.4f}; by round:", {r: round(v, 4) for r, v in history.items() if r % 40 == 0 or r == 20})
# 2. learn again on all the training seasons with that many rounds, then face the test seasons
for depth in (1, 2):
model = boost(train, rounds[depth], depth)
ll = sum(-log(forecast(model, x)[y]) for _, x, y in test) / len(test)
acc = sum(max("HDA", key=forecast(model, x).get) == y for _, x, y in test) / len(test)
print(f"{depth}-question trees, {rounds[depth]} rounds: test log loss {ll:.3f}, accuracy {acc:.1%}")
if depth == 1:
stumps = model
# 3. what the one-question model learned: how each gap on its own moves the home score against the away score
start, steps, _ = stumps
print("questions asked:", dict(Counter(FEATURES[s[c][0]] for s in steps for c in "HDA" if isinstance(s[c], tuple))))
sd = [sqrt(sum((x[j] - sum(r[1][j] for r in train) / len(train)) ** 2 for _, x, _ in train) / len(train)) for j in range(3)]
PART8 = {0: (0.04, -0.01), 1: (0.37, -0.21), 2: (0.29, -0.36)} # part 8's home and away weights, per standard deviation
for j in (0, 1, 2):
def home_minus_away(v): # steps about the other gaps give the same at v and at 0, so they cancel below
x = [v if k == j else 0 for k in range(3)]
return 0.1 * sum(value(s["H"], x) - value(s["A"], x) for s in steps)
curve = {g / 4: round(home_minus_away(g / 4) - home_minus_away(0), 2) for g in range(-8, 9)}
print(f"{FEATURES[j]}: boosting {curve}; logistic regression {(PART8[j][0] - PART8[j][1]) / sd[j]:.2f} per point a game")
print("training matches beyond the flat ends:", sum(abs(x[1]) > 1.5 for _, x, _ in train), "with a this-season gap over 1.5,",
sum(abs(x[2]) > 1 for _, x, _ in train), "with a last-season gap over 1")
Further reading
- Gradient boosted decision trees, Google for Developers. How each tree is trained on the previous model's errors, and the step size that keeps it from overfitting.
- Gradient boosting, Wikipedia. The method, its link to gradient descent, and the usual ways of stopping it overfitting.
- Introduction to boosted trees, XGBoost. The maths behind the most widely used boosting library, step by step.
- Ensembles, scikit-learn. Gradient boosting alongside random forests and the other ways of combining models.