Skip to content

Ask a hundred pundits. Random forests

One decision tree is jumpy, so grow a hundred, each on a slightly different set of matches, and average what they say. On five SPFL test seasons the forest fixes most of a single tree's wild guesses, and still finishes just behind logistic regression.

Intermediate Part 10 of Machine Learning Through Football

Contents

The football question

Ask one pundit to call a match and you get one opinion, shaped by whichever games they happened to watch and whatever stuck in their mind. Ask a hundred, each with a slightly different set of games behind them, and average what they say. Is the panel better than any one pundit, and is it better than a single well-built model?

That's a random forest: a crowd of decision trees, each grown a little differently, all voting on every match.

The concept

A single decision tree is jumpy. Change the training matches a little and it can ask a different first question, and every question after that changes with it. That's high variance, the "jumpy" half of stubborn or jumpy. A random forest turns the jumpiness into an advantage, with two bits of deliberate randomness:

  1. A different set of matches for every tree. Each tree is grown on a resample: 3,900 matches drawn at random from the 3,900 training matches, with repeats allowed. Some matches turn up twice or three times, about a third not at all.
  2. A different choice of question. Every time a tree asks a question, it may only choose from features picked at random. Here, with three features, each question gets one at random.

Then the trees vote. The forest's forecast is the average of its trees' forecasts:

$$\begin{aligned} &P_{\text{forest}}(\text{home}) \\ &\quad = \frac{P_1 + P_2 + \dots + P_T}{T} \end{aligned}$$

In plain football

  • T is the number of trees, here up to 100.
  • P₁ is the first tree's chance of a home win, P₂ the second tree's, and so on: the share of home wins in the group the match lands in, as in part 9.
  • Each tree's mistakes are partly its own, from its own matches and its own questions. Averaged across a hundred trees, those separate mistakes cancel out, while the patterns every tree finds survive.

The same three features as logistic regression and decision trees, each the home side's figure minus the away side's, in points a game: recent form, this season so far and last season. The same 3,900 training matches, 2001/02 to 2020/21, and the same 990 test matches, 2021/22 to 2025/26.

How jumpy is one tree?

Grow 20 trees, each on its own resample of the training matches, and let each ask just one question. They come up with 16 different first questions:

One dot for each of 20 trees, placed where it cuts. Most ask about this season, but they disagree about where to cut, and five ask about last season instead. The same data, shuffled a little, gives a different first question.

Part 9's tree asked "is the this-season gap +0.14 or less?" first, and three of the 20 agree. The rest cut anywhere from −0.35 to +0.52, or ask about last season instead. None of them is wrong; the difference between them is mostly luck in which matches each one saw.

Grown all the way, one question after another until the groups are tiny, single trees are worse still. A hundred of them, each on its own resample, score between 1.045 and 1.170 on the test seasons. A model that knows nothing but how often each result happens scores 1.056.

A hundred trees

Now let them vote. Two forests, each of 100 trees: one of fully grown trees, and one where no group may be smaller than 100 matches, the pruning idea from part 9:

Log loss on the test seasons, lower is better. Adding trees helps a lot at first and then levels off; most of the gain comes from the first ten.
Trees Fully grown Groups of 100+
1 1.152 0.976
5 1.002 0.956
10 0.987 0.955
100 0.980 0.953

Averaging does exactly what it promises. The fully grown trees, which on their own score between 1.045 and 1.170, vote their way from 1.152 with one tree to 0.980 with a hundred. That's the whole idea of a random forest: many jumpy models, averaged, make a steady one.

The gentler forest starts from better trees and ends better, at 0.953. Among its hundred trees, the luckiest one on its own scored 0.951 on the test seasons, but there was no way to know in advance which tree that would be. The forest doesn't need to know.

Forest against the rest

The same 990 test matches:

Model Test accuracy Test log loss
Base rates 47.2% 1.056
One tree, 12 questions deep 47.2% 1.069
Forest, fully grown trees 53.3% 0.980
Tree, two questions 55.5% 0.954
Forest, groups of 100+ 54.3% 0.953
Logistic regression 54.8% 0.950
Bookmaker 56.2% 0.932

The forest fixes the single tree's worst habit, and still finishes just behind logistic regression. There are two reasons, and both come from the data rather than the method.

  • Three features aren't much of a forest. Random forests usually do best with many features that combine in awkward ways. Here the pattern is simple and smooth, stronger teams win more, and a straight-line model like logistic regression already captures most of it.
  • The random choice of question forces in recent form. With one feature picked at random for each question, the forest asks about recent form almost as often as the other two (901 times against 941 and 920 in the gentler forest), although part 9 and logistic regression both found it adds almost nothing. With dozens of features, most of them useful, that randomness helps; with three, it costs.

An evenly matched game

What does each model say when every gap is 0?

Model Home Draw Away
Logistic regression 43.5% 25.8% 30.7%
Forest, groups of 100+ 37.8% 23.1% 39.1%
Forest, fully grown 33.8% 20.3% 45.9%
Close matches, training seasons 38.5% 28.5% 33.0%

The last row is what actually happened in the 397 training matches where the sides were within 0.2 points a game of each other on both season gaps. The home side was still slightly more likely to win. Logistic regression's straight lines push the home side too high; the forests, like the single tree in part 9, lean the other way. None of them is spot on, and the 397 matches have their own luck in them too, but it's a reminder that a better score overall doesn't mean every forecast is better.

Why it matters

  • Averaging beats picking. A hundred mediocre, independent opinions can beat the best of them, as long as their mistakes aren't all the same mistakes. It's the idea behind the wisdom of crowds, and behind betting markets.
  • Variance can be fixed without losing flexibility. A fully grown tree can learn almost any pattern but can't be trusted; a forest keeps the flexibility and drops most of the jumpiness.
  • It's one of the most used models there is. On tables of data with many columns, random forests and their cousins, such as gradient boosting, which also combine many trees, are often the first thing a data scientist tries.
  • The simplest model can still win. On three features, logistic regression is hard to beat. Trying the fancier model and finding it doesn't help is a result, not a failure.

Limitations

  • Three features. With so few, the forest's random choice of question is more a handicap than a help, and there's no hidden pattern for it to find.
  • Hard to read. One tree is a flowchart; a hundred are not. The forest trades part 9's readability for steadiness.
  • Random. Grow the forest again with different random choices and the scores move slightly, by a few thousandths in log loss when we tried five different seeds. The snippet fixes the random seed so its numbers repeat.
  • Settings chosen in hindsight. The two forests here are compared on the test seasons to show the idea. Choosing between them properly means cross-validation on the training seasons.

Try it yourself

Before the next round of fixtures, ask five friends to give a chance for each home win, and write down their average as well as their individual answers. After the matches, score them all, as in is accuracy the right score?. Was the average better than most of your friends? Better than the best of them?

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. The first half builds the same three features as logistic regression; the second grows the forests. The random seed is fixed, so it prints the same numbers every time; it runs in under ten seconds:

import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import log
import random

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
FEATURES = ["recent form", "this season so far", "last season"]
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return games

def points_per_game(games):
    pts, n = Counter(), Counter()
    for r in games:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            n[team] += 1
    return {t: pts[t] / n[t] for t in n}

# three features for every match, each the home side's figure minus the away side's, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
    last = points_per_game(season(s_last))
    promoted = 0.85 * sum(last.values()) / len(last)
    history = defaultdict(list)
    for r in season(s):
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(history[h]) >= 5 and len(history[a]) >= 5:
            rows.append((s, [
                (sum(history[h][-5:]) - sum(history[a][-5:])) / 5,                   # points a game, last five
                sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]),  # points a game this season
                last.get(h, promoted) - last.get(a, promoted),                          # points a game last season
            ], r["FTR"]))
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            history[team].append(p)

train = [r for r in rows if r[0] < "2122"]   # 2001/02-2020/21
test = [r for r in rows if r[0] >= "2122"]   # 2021/22-2025/26

def spread(counts):  # the log loss of predicting a group's own results back to it
    n = sum(counts.values())
    return -sum(c / n * log(c / n) for c in counts.values() if c)

def grow(data, depth, smallest, rng, choose=3):
    """A decision tree, as in part 9, except each question may only use `choose` features picked at random."""
    counts = Counter(r[2] for r in data)
    best = None
    if depth > 0:
        for j in rng.sample(range(len(FEATURES)), choose):
            data = sorted(data, key=lambda r: r[1][j])
            left, right = Counter(), Counter(counts)
            for i in range(1, len(data)):
                left[data[i - 1][2]] += 1
                right[data[i - 1][2]] -= 1
                if smallest <= i <= len(data) - smallest and data[i - 1][1][j] < data[i][1][j]:
                    cost = i * spread(left) + (len(data) - i) * spread(right)
                    if cost < (best[0] if best else len(data) * spread(counts) - 1e-9):
                        best = (cost, j, (data[i - 1][1][j] + data[i][1][j]) / 2)
    if not best:  # a group: its results plus one imaginary result of each kind
        return {c: (counts[c] + 1) / (len(data) + 3) for c in "HDA"}
    _, j, cut = best
    return (j, cut, grow([r for r in data if r[1][j] <= cut], depth - 1, smallest, rng, choose),
            grow([r for r in data if r[1][j] > cut], depth - 1, smallest, rng, choose))

def predict(node, x):
    while isinstance(node, tuple):
        node = node[2] if x[node[0]] <= node[1] else node[3]
    return node

def vote(trees, x):  # the forest's forecast: the average of its trees' forecasts
    return {c: sum(predict(t, x)[c] for t in trees) / len(trees) for c in "HDA"}

def score(trees, data):  # log loss and accuracy on data
    forecasts = [(vote(trees, x), y) for _, x, y in data]
    return (sum(-log(p[y]) for p, y in forecasts) / len(data),
            sum(max("HDA", key=p.get) == y for p, y in forecasts) / len(data))

def resample(data, rng):  # the same number of matches, drawn at random with repeats
    return [rng.choice(data) for _ in data]

def asked(node):
    return asked(node[2]) + asked(node[3]) + Counter([FEATURES[node[0]]]) if isinstance(node, tuple) else Counter()

rng = random.Random(1)
print(f"{len(train)} training matches, {len(test)} test matches")

# 1. one question each, grown on 20 different resamples of the training matches
firsts = Counter()
for _ in range(20):
    t = grow(resample(train, rng), 1, 1, rng)
    firsts[f"{FEATURES[t[0]]} gap <= {t[1]:+.2f}"] += 1
print(f"20 resampled trees, {len(firsts)} different first questions:", dict(firsts.most_common()))

# 2. two forests of 100 trees: fully grown, and with no group under 100 matches; one feature to choose from per question
for name, smallest in (("fully grown", 1), ("groups of 100+", 100)):
    trees = [grow(resample(train, rng), 12, smallest, rng, choose=1) for _ in range(100)]
    singles = [score([t], test)[0] for t in trees]
    print(f"{name}: single trees score {min(singles):.3f} to {max(singles):.3f} on test")
    print("   test log loss by number of trees:", {n: round(score(trees[:n], test)[0], 3) for n in (1, 2, 3, 5, 10, 20, 50, 100)})
    ll, acc = score(trees, test)
    p = vote(trees, [0, 0, 0])
    print(f"   100 trees: test log loss {ll:.3f}, accuracy {acc:.1%}; every gap 0: home {p['H']:.1%}, draw {p['D']:.1%}, away {p['A']:.1%}")
    print("   questions asked:", dict(sum((asked(t) for t in trees), Counter())))

# 3. what really happened when the sides were close on both season gaps
close = Counter(y for _, x, y in train if abs(x[1]) <= 0.2 and abs(x[2]) <= 0.2)
n = sum(close.values())
print(f"training matches within 0.2 on both season gaps: {n}, home {close['H'] / n:.1%}, draw {close['D'] / n:.1%}, away {close['A'] / n:.1%}")

Further reading

  • Random forests, Google for Developers. How resampling and random feature choice make trees different, and why they're grown deep.
  • Random forest, Wikipedia. The algorithm, its history, and how feature importance is measured.
  • Ensembles, scikit-learn. The standard Python library's guide to random forests and the other ways of combining models, such as gradient boosting.
  • Wisdom of the crowd, Wikipedia. Why averaging many independent guesses can beat an expert, and when it fails.

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.