When is the title decided?
A title can be sealed in the table months after the numbers already know. Simulating 25 Scottish Premiership seasons from every third round, the model is typically 90% sure of the champion by round 12. The table doesn't settle it until round 34.
Intermediate
Contents
The research question
A title is "mathematically" won when no one can catch the leaders on points. That usually happens late, with a handful of games left. But anyone who watched Celtic in 2016/17 knew who'd win long before then. How early does the evidence settle a title race, and how often does it fool us?
The dataset
Every Scottish Premiership match from 2000/01 to 2025/26, with results only. The races analysed are 2001/02 to 2025/26: each needs the previous season to build its starting beliefs. 2019/20, stopped by COVID after 29 rounds and settled on points per game, is shown but left out of the averages.
The method
The model is the season simulation from who'll win the league?, run not once but at every third round of every season:
- Stop the season after round 3, then round 6, 9 and so on.
- Update beliefs. Each team's attack and defence start from last season and are updated with the matches played so far, as in updating a team's scoring rate.
- Play the rest of the real fixtures 500 times, allowing for doubt about how good each team is, and count how often each side finishes top.
For each season, two moments:
- 90% sure: the first checkpoint from which the model gave the eventual champion at least a 90% chance, and stayed there to the end.
- Sealed: the point in the real season after which no other side could reach the champion's points, even by winning every remaining match.
And one check the other way: the most the model ever gave a side that didn't win.
Results
Three seasons
- 2016/17: Celtic were 93% after three rounds and never below 97% after that. The title was sealed in round 30, but the evidence had settled it 27 rounds earlier.
- 2021/22: Rangers were champions-elect for months; the model gave Celtic between 13% and 22% until round 21. Then Celtic's run turned it around, 97% by round 33.
- 2025/26: Celtic started as 98% favourites, drifted below 50%, were down to 23% after round 33, with Hearts favourites at round 36, and won it on the final day.
Every season
| Typical round (median) | |
|---|---|
| Model 90% sure of the champion | 12 |
| Title sealed in the table | 34 |
Across the 18 seasons with both, the model was typically sure by round 12 and the title was sealed in round 34: the evidence knew about 22 rounds, more than half a season, before the table did. In the most one-sided seasons, 2012/13, 2014/15, 2016/17 and 2017/18, it was sure after just three rounds.
When it went to the wire
Six races were never 90% certain until the end, and all six went to the final round: 2002/03 (settled on goal difference), 2004/05, 2007/08, 2008/09, 2010/11 and 2025/26. That's what a good model should do: stay unsure when the race really is close.
When it was sure of the wrong side
Twice the model was 90% sure of a side that didn't win: Celtic in 2002/03 (90% after round 3) and Celtic in 2004/05 (91% after round 6). Both were in the first six rounds, when little of the season's evidence was in, and both titles went to the final day. From round 9 onwards it was never 90% sure of the wrong team. Its other big misjudgements were Rangers at 87% after six rounds of 2021/22, and Rangers at 87% after twelve rounds of 2011/12, the season they went into administration and were docked ten points.
A warning about small samples
Twenty-five title races is not many, and in most of them the same two clubs finished first and second. The "typical" rounds could move a few places either way with a different set of seasons, and a league with a more open title race would look very different. The checkpoints are every three rounds, so "round 12" means "somewhere between the 10th and 12th round".
Limitations
- Results only. No transfers, injuries, managers or finances: 2011/12 shows the cost.
- No points deductions. The simulated and real tables use results only, so the "sealed" round ignores deductions.
- The split is simplified. The simulation plays the fixtures that actually happened after the split.
- 500 runs per checkpoint. Each chance is good to about ±2.6 points at 90%, enough for these questions but not for fine distinctions.
Conclusion
In the Scottish Premiership, the title is usually decided long before it's won. By around round 12 the evidence has typically settled it, more than half a season before the table catches up, and in the most one-sided years the model knew after three games. When the race is genuinely close it stays unsure, and those races go to the final day. Its two confident mistakes both came in the first six rounds, which is the real lesson: early in a season, even a 90% chance leaves room for a Helicopter Sunday.
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It takes about half a minute:
import csv
import random
from collections import Counter
from datetime import datetime
from math import exp
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
def season(s): # (home, away, home goals, away goals) in date order
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
return [(r["HomeTeam"], r["AwayTeam"], int(r["FTHG"]), int(r["FTAG"])) for r in games]
data = {s: season(s) for s in names}
def totals(games): # goals for, goals against and matches, per team
gf, ga, n = Counter(), Counter(), Counter()
for h, a, x, y in games:
gf[h] += x; ga[h] += y; n[h] += 1
gf[a] += y; ga[a] += x; n[a] += 1
return gf, ga, n
def table(games, teams): # points, then goal difference, then goals scored (no points deductions)
pts, gd, gs = Counter({t: 0 for t in teams}), Counter(), Counter()
for h, a, x, y in games:
pts[h] += 3 * (x > y) + (x == y); pts[a] += 3 * (y > x) + (x == y)
gd[h] += x - y; gd[a] += y - x; gs[h] += x; gs[a] += y
return sorted(teams, key=lambda t: (pts[t], gd[t], gs[t]), reverse=True), pts
def poisson(lam, rng):
limit, k, p = exp(-lam), 0, rng.random()
while p > limit:
k += 1
p *= rng.random()
return k
def title_chances(s, sims, cut, k=20, seed=1):
"""Stop season s after `cut` matches, then play its remaining fixtures sims times."""
rng = random.Random(seed)
games, last = data[s], data[names[names.index(s) - 1]]
teams = sorted({t for g in games for t in g[:2]})
lf, la, ln = totals(last)
mu = sum(x + y for *_, x, y in last) / (2 * len(last)) # goals per team per match last season
home = sum(x for *_, x, _ in last) / len(last) / mu
away = sum(y for *_, y in last) / len(last) / mu
played, rest = games[:cut], games[cut:]
gf, ga, n = totals(played)
# priors from last season, worth k matches; promoted sides score less and concede more than average
pf = {t: lf[t] / ln[t] if ln[t] else 0.85 * mu for t in teams}
pa = {t: la[t] / ln[t] if ln[t] else 1.15 * mu for t in teams}
wins = Counter()
for _ in range(sims):
# we aren't sure how good each team is, so each run draws a plausible attack and defence from the Gamma posterior
att = {t: rng.gammavariate(k * pf[t] + gf[t], 1 / (k + n[t])) for t in teams}
dfn = {t: rng.gammavariate(k * pa[t] + ga[t], 1 / (k + n[t])) for t in teams}
sim = [(h, a, poisson(att[h] * dfn[a] / mu * home, rng), poisson(att[a] * dfn[h] / mu * away, rng)) for h, a, *_ in rest]
wins[table(played + sim, teams)[0][0]] += 1
now, pts = table(played, teams)
return {t: wins[t] / sims for t in teams}, now, pts, table(games, teams)[0][0]
def sealed(games, teams, champion): # the round after which no rival could reach the champion's points
total = Counter(t for g in games for t in g[:2])
pts, played = Counter(), Counter()
for i, (h, a, x, y) in enumerate(games, 1):
pts[h] += 3 * (x > y) + (x == y); pts[a] += 3 * (y > x) + (x == y)
played[h] += 1; played[a] += 1
if all(pts[t] + 3 * (total[t] - played[t]) < pts[champion] for t in teams if t != champion):
return i / (len(teams) // 2)
return None # level on points at the end: settled on goal difference
lines, both, wrong_90 = {}, [], []
for s in names[1:]:
games = data[s]
teams = sorted({t for g in games for t in g[:2]})
champion = table(games, teams)[0][0]
per_round, rounds = len(teams) // 2, len(games) // (len(teams) // 2)
chances = [(r, title_chances(s, 500, r * per_round)[0]) for r in range(3, rounds, 3)]
path = [(r, c[champion]) for r, c in chances]
wrong = max(((p, t, r) for r, c in chances for t, p in c.items() if t != champion), default=(0, "", 0))
sure = next((r for i, (r, _) in enumerate(path) if all(p >= 0.9 for _, p in path[i:])), None)
maths = sealed(games, teams, champion)
lines[s] = path
if sure and maths and rounds == 38:
both.append((sure, maths))
if wrong[0] >= 0.9:
wrong_90.append((s, wrong))
note = " (stopped early, decided on points per game)" if rounds < 38 else ""
print(f"20{s[:2]}/{s[2:]} {champion:8} 90% sure from round {sure or '-':>2}; sealed in round {f'{maths:.1f}' if maths else 'none, goal difference'} of {rounds}{note}"
f"; most it ever gave another side: {wrong[1]} {wrong[0]:.0%} (round {wrong[2]})")
middle = lambda xs: sorted(xs)[len(xs) // 2] if len(xs) % 2 else sum(sorted(xs)[len(xs) // 2 - 1:len(xs) // 2 + 1]) / 2
print(f"{len(both)} seasons with both: 90% sure by round {middle([a for a, _ in both]):g} and sealed in round {middle([b for _, b in both]):.1f}, typically")
print("90% sure of the wrong side:", ", ".join(f"20{s[:2]}/{s[2:]} {t} {p:.0%} in round {r}" for s, (p, t, r) in wrong_90))
for s in ("1617", "2122", "2526"):
print(f"20{s[:2]}/{s[2:]} champion's chance every third round:", " ".join(f"{p:.0%}" for _, p in lines[s]))
Further reading
- Monte Carlo method, Wikipedia. Simulating something many times to estimate a probability.
- Posterior predictive distribution, Wikipedia. Forecasting while allowing for doubt about the underlying strengths, as every simulated season here does.
- Poisson regression, Wikipedia. The standard way to model goals from attack, defence and home advantage.