Skip to content

Three questions and a prediction. Decision trees

A decision tree predicts a match the way a pundit reasons, with a string of yes-or-no questions, and it learns which questions to ask from the data. On five SPFL test seasons two questions nearly match logistic regression; twelve fall apart.

Intermediate Part 9 of Machine Learning Through Football

Contents

The football question

Ask a pundit to call a match and you'll hear something like: are they better this season? Were they better last season too? By a lot? Each answer narrows it down, until there's a verdict.

Can a model learn to reason like that: pick its own questions from past results, and turn the answers into a probability for home win, draw and away win? That's a decision tree. It's a different kind of model from logistic regression, which adds up weighted numbers, so this part puts the two side by side on the same matches.

The concept

A decision tree is a string of yes-or-no questions. Each question splits the matches in two, each half gets its own next question, and the matches that end up together at the bottom form a group. The prediction for a new match is simply what happened in its group: if 75% of the training matches in the group were home wins, the tree says 75% home win.

The same three features as logistic regression, each the home side's figure minus the away side's, in points a game: recent form (last five matches), this season so far and last season. A question looks like "is the this-season gap 0.14 or less?"

The tree chooses each question by trying every feature and every place to cut it, and keeping the one that leaves the two halves most predictable:

$$\begin{aligned} \text{cost} &= n_{\text{yes}} \times \text{spread}_{\text{yes}} \\ &\quad + n_{\text{no}} \times \text{spread}_{\text{no}} \end{aligned}$$

In plain football

  • n is how many matches fall on each side of the question.
  • Spread is how mixed a side's results are: zero if every match went the same way, highest when home wins, draws and away wins are equally common. It's the log loss you'd get by predicting that side's results back to itself.
  • The best question is the one with the lowest cost: it sorts the matches into two piles whose results are as easy to call as possible. Then the tree does the same again inside each pile.

Two details stop the probabilities going wrong. Each group's shares include one imaginary home win, draw and away win, the fix from is accuracy the right score?, so no group ever says 0%. And the tree needs a limit, or it will keep asking questions until every group is a single match. How far to let it go is the big decision, and it's where this part ends up.

A football example

Every Scottish Premiership match from 2001/02 to 2020/21 where both sides had played at least five games: 3,900 training matches, the same as logistic regression. Here's the tree it grows when allowed three questions:

The tree three questions deep. Each gap is the home side's points a game minus the away side's; a minus means the away side was better. Every path ends in a group of training matches and the share of home wins, draws and away wins in it.

Follow one match down. The home side has taken 0.5 points a game more than the visitors this season (no, it's more than 0.14), and was 0.3 better last season (yes, 0.58 or less), but not by more than 0.81 this season (yes). It lands in a group of 848 matches: home 49%, draw 24%, away 27%.

What it learned

  • The first question is about this season, and the second about last season. Between them they do almost all the work.
  • Recent form isn't asked at all. Trees three and four questions deep never use it; it first appears five questions deep, which, as the next section shows, is just when the tree starts learning noise. Logistic regression came to the same conclusion by a completely different route: its weight on recent form was close to zero.
  • The draw peaks in the middle. The highest draw share, 31%, is in the group where the home side is at most 0.14 points a game better this season but was more than 0.10 better last season: close to evenly matched, all told. Even there a home win is more likely, so, like every model in this series, the tree never picks a draw.
  • The groups are uneven. The biggest holds 1,218 matches, the smallest 133. The tree cuts wherever the results change, not into neat equal bands.

How many questions?

Let the tree go deeper and it keeps finding questions that help on the training matches. The real test is the five seasons it never saw, 2021/22 to 2025/26, the same 990 matches used to test logistic regression. Training and test columns are log loss:

Log loss, lower is better. Every extra question lowers it on the training matches. The test log loss is best at two questions, then gets steadily worse: the tree is learning the training seasons, not football.
Questions (groups) Training Test Test accuracy
0 (1): base rates 1.067 1.056 47.2%
1 (2) 1.017 1.008 49.6%
2 (4) 0.977 0.954 55.5%
3 (8) 0.966 0.960 51.6%
6 (57) 0.925 0.974 53.5%
12 (544) 0.759 1.069 47.2%

Twelve questions deep, the tree has split 3,900 matches into 544 groups, about seven matches each. It calls 67.3% of the training matches right, and 47.2% of the test matches: no better than always picking a home win. Its test log loss, 1.069, is worse than a model that knows nothing but how often home teams win (1.056). Tiny groups give confident probabilities built on a handful of matches, and most of that confidence is wrong.

That's overfitting, exactly as the lookup rules showed earlier in the series, and the depth of the tree is the dial between stubborn and jumpy: too few questions and it can't tell teams apart, too many and it memorises.

Pruning

Choosing the depth is one fix. Another is to let the tree ask as many questions as it likes, but never make a group smaller than 200 matches. That stops it chasing a few odd results. Grown that way, it stops by itself at 15 groups, and scores 0.958 on the test seasons, close to the best fixed depth, without anyone choosing a depth at all. Of its 14 questions, only two are about recent form.

This is pruning, one of the fixes listed in overfitting. How the limit gets picked matters too: set it by looking at the test seasons and the test stops being a fair exam. The honest way is cross-validation on the training seasons.

Tree against logistic regression

The same 990 test matches, the same three features:

Model Test accuracy Test log loss
Base rates 47.2% 1.056
Form + table position (lookup) 53.5% 0.997
Tree, pruned to groups of 200+ 0.958
Tree, two questions 55.5% 0.954
Logistic regression 54.8% 0.950
Bookmaker 56.2% 0.932

The tree two questions deep calls more results right than logistic regression, 55.5% against 54.8%, but its log loss is slightly worse. That's the lesson of is accuracy the right score? again: the tree's four groups give only four different forecasts, so its probabilities are coarse, while logistic regression moves smoothly with every tenth of a point.

On this data, the tree roughly matches logistic regression and doesn't beat it. That's not a failure. Logistic regression assumes the chances change smoothly as the gaps grow; a tree assumes nothing, and it would have found a sharp threshold or an odd interaction if there were one. Finding none tells us the simple, smooth picture was about right.

Try it yourself in the Decision Tree Match Predictor: set the three gaps and watch the match find its way down the tree.

Why it matters

  • You can read it. A tree is a flowchart. A coach, a scout or a fan can follow exactly why it said 75%, which is harder with weights.
  • It picks its own questions. Nobody told it recent form doesn't matter; it never found a reason to ask.
  • Depth is the dial. Every tree-based model has to be stopped somewhere, by depth, by group size or by pruning, and the test data decides whether you stopped in the right place.
  • One tree is the start. A single tree is jumpy: move the cut from 0.14 to 0.16 and matches jump groups, and a slightly different set of seasons can grow a different tree. The fix is to grow many trees and let them vote, a random forest, which is usually how trees beat simpler models.

Limitations

  • Steps, not slopes. A match just either side of a cut gets very different forecasts, although the teams are almost identical. An evenly matched game, every gap 0, lands in the 1,218-match group, which also holds plenty of weaker home sides, so the tree makes the away side favourite, 40% to 34%; logistic regression says home 43%.
  • Only three features. Injuries, team news, managers and money aren't in it, and a tree can only split on what it's given.
  • Unstable. Trees grown on slightly different seasons can look quite different, even when they score about the same.
  • One league. The cuts, 0.14, −0.68 and so on, are learned from the Scottish Premiership, with its two dominant clubs. Another league would grow a different tree.

Try it yourself

Take last season's results in your league. Split the matches by one question, say "was the home side higher in the table?", and count home wins, draws and away wins on each side. Now try a different question. Which split gives the two piles whose results are easier to call? You've just grown the first branch of a decision tree.

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. The first half builds the same three features as logistic regression; the second grows the trees. It runs in about a second:

import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import log

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
FEATURES = ["recent form", "this season so far", "last season"]
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return games

def points_per_game(games):
    pts, n = Counter(), Counter()
    for r in games:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            n[team] += 1
    return {t: pts[t] / n[t] for t in n}

# three features for every match, each the home side's figure minus the away side's, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
    last = points_per_game(season(s_last))
    promoted = 0.85 * sum(last.values()) / len(last)
    history = defaultdict(list)
    for r in season(s):
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(history[h]) >= 5 and len(history[a]) >= 5:
            rows.append((s, [
                (sum(history[h][-5:]) - sum(history[a][-5:])) / 5,                   # points a game, last five
                sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]),  # points a game this season
                last.get(h, promoted) - last.get(a, promoted),                          # points a game last season
            ], r["FTR"]))
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            history[team].append(p)

train = [r for r in rows if r[0] < "2122"]   # 2001/02-2020/21
test = [r for r in rows if r[0] >= "2122"]   # 2021/22-2025/26

def spread(counts):  # the log loss of predicting a group's own results back to it; 0 if they all went the same way
    n = sum(counts.values())
    return -sum(c / n * log(c / n) for c in counts.values() if c)

def grow(data, depth, smallest=1):
    """Ask the question that most lowers the log loss, split, and repeat on each side."""
    counts = Counter(r[2] for r in data)
    best = None
    if depth > 0:
        for j in range(len(FEATURES)):
            data = sorted(data, key=lambda r: r[1][j])
            left, right = Counter(), Counter(counts)
            for i in range(1, len(data)):  # the first i matches answer yes, the rest no
                left[data[i - 1][2]] += 1
                right[data[i - 1][2]] -= 1
                if smallest <= i <= len(data) - smallest and data[i - 1][1][j] < data[i][1][j]:
                    cost = i * spread(left) + (len(data) - i) * spread(right)
                    if cost < (best[0] if best else len(data) * spread(counts) - 1e-9):
                        best = (cost, j, (data[i - 1][1][j] + data[i][1][j]) / 2)
    if not best:  # a leaf: the group's results, plus one imaginary result of each kind so nothing is 0%
        return {c: (counts[c] + 1) / (len(data) + 3) for c in "HDA"}
    _, j, cut = best
    return (j, cut, grow([r for r in data if r[1][j] <= cut], depth - 1, smallest),
            grow([r for r in data if r[1][j] > cut], depth - 1, smallest))

def predict(node, x):
    while isinstance(node, tuple):
        node = node[2] if x[node[0]] <= node[1] else node[3]
    return node

def leaves(node):
    return leaves(node[2]) + leaves(node[3]) if isinstance(node, tuple) else 1

def score(node, data):  # log loss and accuracy
    return (sum(-log(predict(node, x)[y]) for _, x, y in data) / len(data),
            sum(max("HDA", key=predict(node, x).get) == y for _, x, y in data) / len(data))

def show(node, data, indent=""):  # the tree as questions, with how many training matches reach each group
    if not isinstance(node, tuple):
        return print(f"{indent}{len(data)} matches: " + ", ".join(f"{c} {node[c]:.0%}" for c in "HDA"))
    j, cut = node[0], node[1]
    print(f"{indent}{FEATURES[j]} gap <= {cut:+.2f}?")
    show(node[2], [r for r in data if r[1][j] <= cut], indent + "  yes: ")
    show(node[3], [r for r in data if r[1][j] > cut], indent + "  no:  ")

def asked(node):  # how often each feature is used as a question
    return asked(node[2]) + asked(node[3]) + Counter([FEATURES[node[0]]]) if isinstance(node, tuple) else Counter()

print(f"{len(train)} training matches, {len(test)} test matches")
for depth in (0, 1, 2, 3, 4, 5, 6, 8, 10, 12):
    tree = grow(train, depth)
    (tl, ta), (sl, sa) = score(tree, train), score(tree, test)
    print(f"{depth:2} questions deep, {leaves(tree):3} groups: training log loss {tl:.3f} ({ta:.1%}), test {sl:.3f} ({sa:.1%})")
    print("   questions asked:", dict(asked(tree)))
show(grow(train, 3), train)
pruned = grow(train, 12, smallest=200)  # no group smaller than 200 matches
print(f"at least 200 matches a group: {leaves(pruned)} groups, test log loss {score(pruned, test)[0]:.3f}, asked {dict(asked(pruned))}")

Further reading

  • Decision trees, Google for Developers. A short introduction to how a tree makes a prediction, from Google's course on decision forests.
  • Decision tree learning, Wikipedia. How trees choose their splits (information gain, the idea behind this part's "spread", and Gini impurity), with their advantages and limitations.
  • Decision trees, scikit-learn. The standard Python library's guide: splitting criteria including entropy, the minimum group size (min_samples_leaf), and pruning.
  • Random forests, Google for Developers. Many trees, each grown on a different sample of the data: where single trees go next.

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.