Skip to content

Every team its own attack and defence. Dixon-Coles from scratch

Four machine learning models stalled at the same score on three numbers per match. Rate every team's attack and defence from the goals they score and let in, refit before every matchday, and a classic football model finally gets past them.

Advanced Part 12 of Machine Learning Through Football

Contents

The football question

Logistic regression, a decision tree, a random forest and gradient boosting all finished within 0.004 of each other on the same 990 test matches, stuck at about 0.950. Every one of them saw a match as three numbers: how far apart the two sides were in recent form, this season and last season.

But a fan sees more than that. Celtic score a lot. Hearts are hard to break down. A side can be poor overall and still dangerous going forward. Can a model that rates every team's attack and defence separately, from the goals they score and let in, do better than all of them?

That's the model Mark Dixon and Stuart Coles published in 1997, building on Mike Maher's from 1982, and it's still a standard starting point for football prediction. This part builds it from scratch. It's the most mathematical piece in the series so far, so every formula has its plain-football version.

The model

Every team gets two numbers: an attack rating and a defence rating. There's one more for the whole league, home advantage. A match's expected goals are:

$$\begin{aligned} \lambda &= \text{attack}_{\text{home}} \times \text{defence}_{\text{away}} \times \text{home} \\ \mu &= \text{attack}_{\text{away}} \times \text{defence}_{\text{home}} \end{aligned}$$

In plain football

  • λ (lambda) is how many goals the home side should score, μ (mu) the away side.
  • Attack is a multiplier: 1 is a typical attack, 1.5 scores half as many again against the same defence.
  • Defence is how many goals a typical attack would expect to score when visiting them. Lower is better.
  • Home is how much playing at home multiplies a side's goals.

Each side's goals then follow the Poisson distribution with those averages, which gives a chance for every scoreline, and adding them up gives home win, draw and away win.

The low-score correction

Dixon and Coles noticed that plain Poisson gets the low scores slightly wrong, because it treats the two sides' goals as independent. Their fix multiplies the chance of just four scorelines by a correction that depends on one extra number, ρ (rho):

$$\begin{aligned} \text{0–0:}&\;\; 1 - \lambda\mu\rho \\ \text{1–0:}&\;\; 1 + \mu\rho \\ \text{0–1:}&\;\; 1 + \lambda\rho \\ \text{1–1:}&\;\; 1 - \rho \end{aligned}$$

In plain football

  • Every other scoreline is left alone.
  • A negative ρ makes 0–0 and 1–1 a little more likely and 1–0 and 0–1 a little less: more draws among the low scores.
  • ρ = 0 is plain Poisson. The Dixon-Coles model page shows the correction side by side with plain Poisson.

Fitting it

The ratings are chosen by maximum likelihood: the set of numbers under which the results that actually happened were as likely as possible. Recent matches count for more than old ones, through a weight that fades with time:

$$\begin{aligned} &\text{maximise} \;\; \sum_{\text{matches}} w \times \log P(\text{score}) \\ &w = e^{-\xi \times \text{days ago}} \end{aligned}$$

In plain football

  • P(score) is the chance the model gave to the score that happened. A good set of ratings gives the real scores high chances. Taking logs and adding is the same scoring rule as log loss, turned the right way up.
  • w is how much a match counts: 1 for yesterday's, less for older ones.
  • ξ (xi) is how fast old matches fade. Too slow and the ratings remember last manager's team; too fast and they swing on a couple of results.

In practice the ratings are found by taking turns. Given everyone's defences, each team's best attack rating is simply its weighted goals scored divided by the goals a typical attack would have scored against the same opponents. Then, given the attacks, the same for defences; then home advantage; and round again, 60 times, until nothing moves. Finally, ρ is chosen as the value that makes the low scores most likely. The four years before each matchday are used, and a team needs at least five matches in them to be rated.

The model is refitted before every matchday, using only results that had already happened.

How fast should old results fade?

Not chosen on the test seasons. As in gradient boosting, the decay is chosen on the held-back seasons, 2016/17 to 2020/21, scoring the same kind of matches as the test (both sides at least five games in):

Decay per day Half-life Held-back log loss
0.001 693 days 0.9726
0.002 347 days 0.9686
0.003 231 days 0.9675
0.004 173 days 0.9683
0.005 139 days 0.9700

A match counts half as much after 231 days, about two-thirds of a year. Last season still matters, two seasons ago much less, which is roughly how a fan weighs things too.

The ratings

After the last match of 2025/26, for that season's twelve clubs:

Attack across, defence up: the best sides are top right. Celtic and Rangers, in gold, are in a different place from everyone else going forward; Hearts have the league's best defence.

Home advantage is ×1.32: a side scores about a third more at home than it would away against the same opponents. The fitted correction is small, ρ = −0.03.

Take Hearts at home to Hibernian. Hearts' attack is 1.21, Hibs' defence 0.95, home advantage 1.32, so Hearts should score 1.21 × 0.95 × 1.32 = 1.52. Hibs' attack is 1.16 and Hearts' defence 0.79, so Hibs should score 0.92. Poisson and the correction turn those into:

Hearts v Hibernian 1.52 to 0.9250.7%26.7%22.6%
Celtic v Livingston 3.54 to 0.6588.9%
Home win Draw Away win. Chances from the ratings after the last match of 2025/26. Celtic v Livingston: draw 7.8%, away win 3.3%.

The test

The same 990 test matches as every model in this series, 2021/22 to 2025/26:

Model Test accuracy Test log loss
Base rates 47.2% 1.056
Form + table position (lookup) 53.5% 0.997
Decision tree, two questions 55.5% 0.954
Random forest 54.3% 0.953
Gradient boosting 54.5% 0.951
Logistic regression 54.8% 0.950
Dixon-Coles 55.3% 0.945
Same ratings, no correction 55.3% 0.944
Bookmaker 56.2% 0.932

Team ratings get past the ceiling. From logistic regression's 0.950 to the bookmakers' 0.932 is a gap of 0.018 in log loss; Dixon-Coles closes 0.005 of it, a little over a quarter.

Three reasons, all about what goes in rather than how clever the method is:

  • Every team is itself. The earlier models saw only gaps between two sides. Here a strong attack and a leaky defence are two separate facts, and a match is the meeting of one side's attack with the other's defence.
  • Goals, not points. A 4–0 and a 1–0 are both three points, but they say different things about a team. Goals carry more information per match.
  • Always up to date. The ratings are refitted before every matchday, so a side that improves or falls away starts to show it within weeks.

That last point needs saying plainly: it's also an advantage the machine learning models didn't have. Their weights were learned once, from 2001/02 to 2020/21, and never updated through the test seasons, although their inputs did. Dixon-Coles keeps learning from each week's results. That's how it's meant to be used, but it isn't a like-for-like contest of methods.

The correction adds nothing here

The famous low-score correction makes no difference: 0.945 with it, 0.944 without. The fitted ρ is small, −0.03, so the correction barely moves any forecast, and what it does move doesn't help. It was designed for English football in the 1990s; in these seasons of the Scottish Premiership, plain Poisson on good ratings is just as good.

Why it matters

  • Better inputs beat cleverer methods. Four machine learning models couldn't get below 0.950; a model from 1997, built on team-by-team ratings, can.
  • It explains itself. Two ratings per team, one for home advantage: every forecast can be traced back to numbers a fan can argue about.
  • It gives every scoreline. Not just home, draw or away, but 2–1, 0–0 and 4–0, the same way the Poisson Match Predictor does from expected goals.
  • It's a classic for a reason. Nearly thirty years on, it's still one of the first models people build for football scores, and the one many others are measured against.

Limitations

  • One league. Promoted sides arrive with little or no Premiership history, so their ratings rest on a handful of matches until they've played a few.
  • Goals are noisy. A few lucky wins can inflate a rating. Ratings built on expected goals (xG) are often steadier, but the data here doesn't have it.
  • No team news. Injuries, suspensions, a new manager, a rotated side before a European tie: the bookmakers know, the ratings don't.
  • A simpler fit. The ratings and ρ are found in two steps, the ratings first and then ρ, rather than all together as in the original paper. With a ρ this small the difference is tiny, but it isn't the full method.
  • Choices. The four-year window, the five-match minimum and 60 rounds of fitting were set by hand, not tuned.

Try it yourself

Take your league's table. For each side, divide goals scored per game by the league's average: that's a rough attack rating. Divide goals conceded per game by the league average: a rough defence rating. For this weekend's matches, multiply one side's attack by the other's defence by the league's average goals per side, add a bit for home advantage, and you have expected goals. The Poisson Match Predictor turns them into chances.

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It refits the model before every matchday, hundreds of times, so allow about three minutes:

Show the Python114 lines, ready to copy and run.
import csv
from collections import Counter
from datetime import datetime, timedelta
from math import exp, factorial, log

names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
    for r in games:
        r["date"] = datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y")
    return sorted(games, key=lambda r: r["date"])

# every Premiership match; `tested` marks the ones parts 8-11 were scored on (both sides 5+ games into the season)
matches = []
for s in names:
    played = Counter()
    for r in season(s):
        h, a = r["HomeTeam"], r["AwayTeam"]
        matches.append({"date": r["date"], "season": s, "home": h, "away": a, "hg": int(r["FTHG"]), "ag": int(r["FTAG"]),
                        "result": r["FTR"], "tested": played[h] >= 5 and played[a] >= 5})
        played[h] += 1
        played[a] += 1

def low_score(x, y, l, u, rho):  # Dixon and Coles' correction: only 0-0, 1-0, 0-1 and 1-1 are touched
    return {(0, 0): 1 - l * u * rho, (0, 1): 1 + l * rho, (1, 0): 1 + u * rho, (1, 1): 1 - rho}.get((x, y), 1)

def fit(today, decay, correct=True):
    """Every team's attack and defence, home advantage and the correction, from the four years before today."""
    recent = [m for m in matches if timedelta(0) < today - m["date"] <= timedelta(4 * 365)]
    games = Counter(t for m in recent for t in (m["home"], m["away"]))
    weighted = [(m, exp(-decay * (today - m["date"]).days)) for m in recent if games[m["home"]] >= 5 and games[m["away"]] >= 5]
    teams = {t for m, _ in weighted for t in (m["home"], m["away"])}
    attack, defence, home = dict.fromkeys(teams, 1.0), dict.fromkeys(teams, 1.0), 1.3
    for _ in range(60):  # take turns: best attacks for these defences, best defences for these attacks, then home advantage
        scored, expected = Counter(), Counter()
        for m, w in weighted:
            scored[m["home"]] += w * m["hg"]
            expected[m["home"]] += w * defence[m["away"]] * home
            scored[m["away"]] += w * m["ag"]
            expected[m["away"]] += w * defence[m["home"]]
        attack = {t: scored[t] / expected[t] for t in teams}
        conceded, expected = Counter(), Counter()
        for m, w in weighted:
            conceded[m["away"]] += w * m["hg"]
            expected[m["away"]] += w * attack[m["home"]] * home
            conceded[m["home"]] += w * m["ag"]
            expected[m["home"]] += w * attack[m["away"]]
        defence = {t: conceded[t] / expected[t] for t in teams}
        home = sum(w * m["hg"] for m, w in weighted) / sum(w * attack[m["home"]] * defence[m["away"]] for m, w in weighted)
        middle = exp(sum(log(v) for v in attack.values()) / len(attack))  # so the typical attack is 1
        attack = {t: v / middle for t, v in attack.items()}
        defence = {t: v * middle for t, v in defence.items()}
    def fit_of(rho):  # how likely the results are with this correction; values that make a chance negative are ruled out
        total = 0.0
        for m, w in weighted:
            t = low_score(m["hg"], m["ag"], attack[m["home"]] * defence[m["away"]] * home, attack[m["away"]] * defence[m["home"]], rho)
            if t <= 0:
                return float("-inf")
            total += w * log(t)
        return total
    rho = max((r / 100 for r in range(-25, 11)), key=fit_of) if correct else 0.0
    return attack, defence, home, rho

def poisson(k, rate):
    return exp(-rate) * rate ** k / factorial(k)

def chances(model, h, a):
    attack, defence, home, rho = model
    l, u = attack[h] * defence[a] * home, attack[a] * defence[h]  # expected goals for each side
    p = {"H": 0.0, "D": 0.0, "A": 0.0}
    for x in range(11):
        for y in range(11):
            p["H" if x > y else "D" if x == y else "A"] += low_score(x, y, l, u, rho) * poisson(x, l) * poisson(y, u)
    return {c: v / sum(p.values()) for c, v in p.items()}

def score(targets, decay, correct=True):  # refit before every matchday, using only results already played
    loss = right = 0
    model, day = None, None
    for m in targets:
        if m["date"] != day:
            model, day = fit(m["date"], decay, correct), m["date"]
        p = chances(model, m["home"], m["away"])
        loss -= log(p[m["result"]])
        right += max(p, key=p.get) == m["result"]
    return loss / len(targets), right / len(targets)

held_back = [m for m in matches if m["tested"] and "1617" <= m["season"] < "2122"]
test = [m for m in matches if m["tested"] and m["season"] >= "2122"]
print(f"{len(held_back)} held-back matches (2016/17-2020/21), {len(test)} test matches (2021/22-2025/26)")

# 1. how fast should old matches fade? Chosen on the held-back seasons
by_decay = {d: score(held_back, d)[0] for d in (0.001, 0.002, 0.003, 0.004, 0.005)}
best = min(by_decay, key=by_decay.get)
for d, v in by_decay.items():
    print(f"decay {d} a day (a match counts half as much after {log(2) / d:.0f} days): held-back log loss {v:.4f}")
print("best decay:", best)

# 2. the test seasons, with and without the low-score correction
for correct in (True, False):
    loss, acc = score(test, best, correct)
    print(f"{'Dixon-Coles' if correct else 'Poisson, no correction'}: test log loss {loss:.3f}, accuracy {acc:.1%}")

# 3. the ratings after the last match of 2025/26, for that season's twelve clubs
attack, defence, home, rho = fit(matches[-1]["date"] + timedelta(days=1), best)
clubs = {m["home"] for m in matches if m["season"] == "2526"}
print(f"home advantage x{home:.2f}, correction {rho:+.2f}")
for t in sorted(clubs, key=lambda t: defence[t] / attack[t]):
    print(f"  {t:15} attack {attack[t]:.2f}  defence {defence[t]:.2f}")
for h, a in (("Celtic", "Livingston"), ("Hearts", "Hibernian")):
    p = chances((attack, defence, home, rho), h, a)
    print(f"{h} v {a}: expected goals {attack[h] * defence[a] * home:.2f} to {attack[a] * defence[h]:.2f}; "
          f"home {p['H']:.1%}, draw {p['D']:.1%}, away {p['A']:.1%}")

Further reading

  • Predicting football results with statistical modelling: Dixon-Coles and time-weighting, David Sheehan. A Python walk-through of plain Poisson, the Dixon-Coles correction and time decay, on the English Premier League.
  • Maximum likelihood estimation, Wikipedia. The principle behind the fit: choose the numbers that make what happened most likely.
  • Poisson regression, Wikipedia. The general family this model belongs to: counts, such as goals, whose average depends on other numbers.
  • Dixon, M. J. and Coles, S. G. (1997), "Modelling association football scores and inefficiencies in the football betting market", Journal of the Royal Statistical Society, Series C (Applied Statistics), 46(2), 265–280. The original paper.

Get new pieces by email

An email when something new is published, and the occasional update. Unsubscribe in one click. How your email is used.