Skip to content

How does a model learn? Rolling downhill with gradient descent

Machine learning models learn by measuring how their error changes as they adjust a setting, then stepping downhill. That's the derivative at work. 25 SPFL seasons show it finding how much of last season's scoring carries into this one.

Intermediate Part 6 of Calculus Through Football

Contents

The football question

A side scored 2 goals a game last season. How many should we expect this season?

Two lazy answers: "the same as last season", or "the league average, like everyone else". The truth is somewhere in between, and finding exactly where is a job for a model. But how does a model find the answer?

Almost every machine learning model, from a simple line to the huge models behind chatbots, learns the same way: it measures how wrong it is, works out which way is downhill, and takes a step. That's gradient descent, and it runs on the derivative.

The concept

Give the model one dial, w, and let it predict:

$$\begin{aligned} &\text{this season} = \text{average} \\ &\quad + w \times (\text{last season} - \text{average}) \end{aligned}$$

In plain football

  • Last season − average is how far above or below the league average a team scored last time.
  • w is how much of that gap we expect to carry over. w = 1: all of it, "the same as last season". w = 0: none of it, "everyone's average".
  • The model's whole job is to find the best w.

"Best" means the smallest error: the average squared miss between what the model predicts and what teams actually scored. Every value of w gives an error, and together they make a valley. Gradient descent finds the bottom of the valley without ever seeing the whole of it:

  1. Start anywhere, say w = 0.
  2. Work out the slope of the valley where you're standing: the derivative of the error with respect to w.
  3. Step downhill, a distance proportional to the slope.
  4. Repeat until the slope is flat.

As an update rule:

$$w_{\text{new}} = w - \text{step size} \times \text{slope}$$

In plain football

  • Slope is how much the error changes as you nudge w: the derivative. Positive means turning w up makes things worse, so the minus sign turns it down.
  • Step size (machine learning people call it the learning rate) is how far to move each time.
  • Steep ground, big step; nearly flat, small step. At the bottom the slope is zero and the dial stops moving.

A football example

Every Scottish Premiership side that played two top-flight seasons in a row between 2000/01 and 2025/26: 271 team-seasons, each a pair of numbers, last season's goals per game and this season's. Here's the valley:

The error for every setting of the dial. Starting from w = 0, each gold dot is one step of gradient descent: big steps where the valley is steep, smaller ones as it flattens near the bottom at 0.83.
The dial What it means Error
w = 0 Everyone scores the league average 0.2306
w = 1 Everyone scores the same as last season 0.0875
w = 0.83 The bottom of the valley 0.0809

Gradient descent, with a step size of 1, gets there in about ten steps: from 0 to 0.36, 0.57, 0.68, 0.74, 0.78, 0.80, 0.81, 0.82, then creeping the last few thousandths. The same answer comes out of the textbook formula for a straight line, 0.826, so we can check the rolling ball ended up in the right place.

What 0.83 means

About 83% of a team's gap from average carries into the next season; the other 17% was luck, or changes that don't last. A side that scored 2.0 a game is expected to score 1.87; a side that scored 1.0 is expected to score 1.05. Both are pulled a little towards the league's 1.36.

That's regression to the mean, found by rolling downhill. It's the same pull as in where priors come from and updating a team's scoring rate, which found last season worth about half a season of new evidence.

Getting the step size right

The step size is the one thing gradient descent can't work out for itself:

Too small, and after twelve steps the dial has barely moved. About right, and it glides to 0.83. Too big, and it overshoots from side to side of the valley. Bigger still, step 8, and each overshoot is worse than the last.
Step size After 12 steps
0.1 0.34: creeping, it would take hundreds of steps
1 0.83: there
4 0.80, after zigzagging from 1.45 to 0.36 and back
8 −51,569: flown off to infinity

For this valley, any step size above 4.56 makes each overshoot bigger than the last, and the dial flies off. Real models tune the step size carefully, and often shrink it as they get close.

Show the mathsThe error, its derivative, and the exact answer. Optional.

With x last season's goals per game and y this season's, measured from their averages, the error is

$$\begin{aligned} E(w) &= \frac{1}{n}\sum (y - w x)^2 \\ &= S_{yy} - 2w\,S_{xy} + w^2 S_{xx} \end{aligned}$$

where \(S_{xx}\) is the average of \(x^2\), \(S_{xy}\) of \(xy\) and \(S_{yy}\) of \(y^2\). Its derivative is

$$\frac{dE}{dw} = 2w\,S_{xx} - 2S_{xy}$$

which is zero at \(w = S_{xy} / S_{xx} = 0.826\). Each step multiplies the distance from the bottom by \(1 - 2\eta S_{xx}\), where η is the step size, so the steps shrink only if \(\eta < 1 / S_{xx} \approx 4.56\).

Why it matters

  • This is how machine learning learns. A real model has thousands or billions of dials, not one, and the slope becomes a gradient, one slope per dial. The idea is exactly this: measure the error, find downhill, step.
  • The derivative does the work. Without a way to measure slope, a model would have to try every setting. With it, it can walk straight to the bottom.
  • Step size matters. Too timid wastes time; too bold never arrives.
  • Training is this loop. When the Machine Learning series talks about a model learning from its training data, this is what's happening underneath.

Limitations

  • One dial is a toy. This valley has one bottom and an exact formula; real models have bumpy landscapes with many dips, where gradient descent can settle in the wrong one.
  • Learning the training data too well is still possible. Rolling to the very bottom of the training error can mean overfitting; the test data still has the final say.
  • Goals per game is one number. Squad changes, managers and money all move a team's scoring, and none of them is in this model.

Try it yourself

Pick a number between 0 and 1 for w. Predict this season's goals per game for three teams in your league, using last season's figures and the league average. At the end of the season, work out your average squared miss, then try a different w. Which way should you move it? You've just taken a step of gradient descent by hand.

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Then:

import csv
from collections import Counter

names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def goals_per_game(s):
    gf, n = Counter(), Counter()
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        for r in csv.DictReader(f):
            if r.get("FTR") in ("H", "D", "A"):
                gf[r["HomeTeam"]] += int(r["FTHG"]); n[r["HomeTeam"]] += 1
                gf[r["AwayTeam"]] += int(r["FTAG"]); n[r["AwayTeam"]] += 1
    return {t: gf[t] / n[t] for t in n}

rates = {s: goals_per_game(s) for s in names}
# every team that played two Premiership seasons in a row: (last season, this season)
pairs = [(rates[a][t], rates[b][t]) for a, b in zip(names, names[1:]) for t in rates[b] if t in rates[a]]
n = len(pairs)
avg_last, avg_this = sum(x for x, _ in pairs) / n, sum(y for _, y in pairs) / n

def error(w):  # average squared miss of: this season = average + w x (last season - average)
    return sum((y - (avg_this + w * (x - avg_last))) ** 2 for x, y in pairs) / n

def slope(w):  # the derivative of the error with respect to w
    return sum(-2 * (x - avg_last) * (y - (avg_this + w * (x - avg_last))) for x, y in pairs) / n

print(f"{n} team-seasons; error if w = 0 (everyone average) {error(0):.4f}, w = 1 (same as last season) {error(1):.4f}")
for step_size in (0.1, 1, 4, 8):
    w, path = 0.0, []
    for _ in range(12):
        w -= step_size * slope(w)  # a step downhill
        path.append(round(w, 3))
    print(f"step size {step_size}: w = {path}")

exact = sum((x - avg_last) * (y - avg_this) for x, y in pairs) / sum((x - avg_last) ** 2 for x, _ in pairs)
print(f"exact answer w = {exact:.3f}, error {error(exact):.4f}; steps above {n / sum((x - avg_last) ** 2 for x, _ in pairs):.2f} fly off")
print(f"so a 2.0-goal side is expected to score {avg_this + exact * (2.0 - avg_last):.2f}, a 1.0-goal side {avg_this + exact * (1.0 - avg_last):.2f}")

Further reading

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.