Which ground is hardest to visit? Partial pooling
Every club's home record says something about its ground, and a lot about luck. Partial pooling lets the data decide how far to trust each club's own figure. Over 26 Scottish Premiership seasons the clubs differ far less than the raw table suggests, and Hearts still come out on top.
Advanced Part 6 of Bayesian Thinking Through Football
New to the notation? The symbols explained
Contents
The football question
Every fan believes their ground is special: the crowd, the pitch, the long trip for the visitors. The results seem to agree for some clubs more than others. Over 26 Scottish Premiership seasons, Hearts took 0.61 more points a game at home than away; St Johnstone just 0.08. So which ground really is the hardest to visit, and how much of those gaps is luck?
The two obvious answers are both wrong. Trusting each club's own figure takes every lucky and unlucky run at face value. Giving every club the league average, as Elo does, says grounds don't differ at all. Partial pooling sits between them, and lets the data decide where.
Three ways to estimate
For each club, the home advantage is its home points a game minus its away points a game. The same squad plays both, so the difference is mostly about where the match is played.
| Method | Each club's estimate | The problem |
|---|---|---|
| No pooling | Its own figure | A few lucky seasons look like a fortress |
| Complete pooling | The league average, +0.33 | Assumes grounds don't differ at all |
| Partial pooling | In between, by an amount the data chooses |
This is the idea from where priors come from and updating a team's scoring rate: start from what's typical, then move towards what this club has shown, as far as the evidence justifies. The difference here is that the league average and how much to trust it both come from the data, rather than being chosen by hand.
The model
A club's pooled estimate pulls its own figure towards the league average:
$$\begin{aligned} &\text{pooled} \\ &\quad = \text{average} + k \times (\text{raw} - \text{average}) \\ &k = \frac{s^2}{s^2 + v} \end{aligned}$$
In plain football
- raw is the club's own home advantage; average is the league's.
- s is how much clubs' true home advantages really differ from one another.
- v is how noisy this club's own figure is. Fewer home games, or more up-and-down results, means more noise.
- k is how much of its own figure the club keeps. If clubs really differ a lot (big s) or the club's figure is precise (small v), k is close to 1 and the club keeps its own figure. If its figure is noisy, k is small and it's pulled most of the way to the average.
The noise v comes straight from the results: the spread of the club's home points divided by its home games, plus the same for its away games. The spread between clubs s is the part that needs estimating.
Letting the data choose s
If clubs' true home advantages spread out by s, and each raw figure adds its own noise v on top, then each raw figure should vary around the league average by about √(s² + v). Choose the s that makes the raw figures we actually have most likely, the same maximum-likelihood idea as Dixon-Coles:
$$\begin{aligned} &\text{maximise over } s: \\ &\sum_{\text{clubs}} \log \text{Normal}\big(\text{raw} \mid \text{average},\; s^2 + v\big) \end{aligned}$$
In plain football
- If the clubs' raw figures are no more spread out than their noise alone would produce, s comes out near zero: grounds don't really differ.
- If they're far more spread out than noise explains, s is large: grounds really do differ, and each club keeps more of its own figure.
Does it work?
A fair test predicts something the method hasn't seen. Split the 26 seasons into odd and even. Estimate every club's home advantage from one half, three ways, and see which best predicts its home advantage in the other half. Then swap the halves round.
| Method | Odd → even seasons | Even → odd seasons |
|---|---|---|
| No pooling | 0.0575 | 0.0575 |
| Complete pooling | 0.0309 | 0.0336 |
| Partial pooling | 0.0300 | 0.0326 |
The numbers are mean squared errors, so lower is better, for the 18 clubs with matches in both halves. Trusting each club's own figure is by far the worst: its errors are nearly twice as big. Partial pooling beats the league average both ways round, but only narrowly, which says something on its own: the clubs don't differ by much.
The hardest grounds to visit
Using all 26 seasons, the spread between clubs s comes out at 0.074 points a game, around a league average of +0.33. That's small. Most clubs' true home advantage is probably within about 0.15 of the average, much closer together than the raw figures, which run from +0.08 to +0.61.
| Club | Raw (games) | Pooled |
|---|---|---|
| Hearts | +0.61 (451) | +0.45 |
| Livingston | +0.48 (221) | +0.38 |
| Dunfermline | +0.53 (152) | +0.37 |
| Hamilton | +0.12 (187) | +0.28 |
| St Johnstone | +0.08 (339) | +0.24 |
The number in brackets is the club's Premiership home games.
Hearts are the hardest side to visit, even after pooling: their home advantage is about 0.12 points a game above the league average. Dunfermline had a bigger raw figure than Livingston, but from fewer home games, so it's pulled back further and ends up just below. At the other end, St Johnstone's ground looks like no advantage at all in the raw table; pooled, it's still an advantage, just the smallest.
How much each club keeps
A club's k depends on how noisy its own figure is. Celtic keep the most of their own figure, 54%, not only because they have among the most home games but because their home results are steady: they win most of them, so their home points vary less. Gretna, with 18 Premiership home games, keep 5%: their raw +0.43 becomes +0.34, almost exactly the average. (Gretna played their one Premiership season, 2007/08, at Motherwell's Fir Park, and the club folded in 2008.)
Why it matters
- Raw league tables exaggerate differences. The biggest and smallest figures in any table are partly luck, and pooling pulls them back. That's why the best and worst on a list usually look less extreme next season.
- The data can choose its own prior. In Bayesian updating the prior was chosen by hand; here the league average and how strongly to lean on it are both estimated from the clubs themselves.
- It works anywhere there are many small groups. Players' penalty records, referees' card counts, a scout's ratings: all are noisier than they look, and all benefit from the same pull to the middle.
- Elo's one home advantage for everyone is nearly right. With a spread of 0.074 points a game between clubs, a single figure loses little, which is part of why tuned Elo does as well as it does.
Limitations
- Grounds change. 26 seasons cover rebuilt stands, new pitches and different managers. Each figure mixes all of those.
- Home advantage moves over time. It rose in recent seasons, and 2020/21 was played without crowds; both are averaged in here.
- A normal approximation. Points come in 0, 1 and 3, not a smooth scale. With hundreds of matches per club that's fine; for Gretna's 18 it's rough, but that's exactly the club the method trusts least.
- The Premiership only. Clubs' seasons in the Championship aren't included, so the clubs promoted and relegated most often have the least data.
Try it yourself
Take your club's home and away points over the last five seasons, and work out its home advantage: home points a game minus away points a game. Compare it with the league's. Unless your club's figure is far from the average over many seasons, the truth is probably closer to the average than it looks. A rough rule of thumb from this data, for a club with many Premiership seasons: keep about a third to a half of the difference.
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It runs in under a second:
Show the Python54 lines, ready to copy and run.
import csv
from collections import defaultdict
from math import log
from statistics import mean, pvariance
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
def home_advantage(seasons):
"""Each club's home advantage (home points a game minus away points a game), how noisy it is, and its home games."""
home, away = defaultdict(list), defaultdict(list)
for s in seasons:
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
for r in csv.DictReader(f):
if r.get("FTR") in POINTS:
h, a = POINTS[r["FTR"]]
home[r["HomeTeam"]].append(h)
away[r["AwayTeam"]].append(a)
return {t: (mean(home[t]) - mean(away[t]), # the raw figure
pvariance(home[t]) / len(home[t]) + pvariance(away[t]) / len(away[t]), # its noise (variance)
len(home[t]))
for t in home if len(home[t]) >= 10 and len(away[t]) >= 10}
def pool(clubs):
"""Partial pooling: find the spread between clubs that makes the raw figures most likely, then pull each club
towards the league average by an amount that depends on how noisy its own figure is."""
def fit(spread2):
weight = {t: 1 / (v + spread2) for t, (_, v, _) in clubs.items()}
average = sum(weight[t] * clubs[t][0] for t in clubs) / sum(weight.values())
likelihood = sum(-0.5 * log(v + spread2) - (y - average) ** 2 / (2 * (v + spread2)) for y, v, _ in clubs.values())
return likelihood, average
spread2 = max((x / 10000 for x in range(400)), key=lambda x: fit(x)[0])
average = fit(spread2)[1]
kept = {t: spread2 / (spread2 + v) for t, (_, v, _) in clubs.items()} # how much of its own figure each club keeps
return average, spread2 ** 0.5, {t: average + kept[t] * (y - average) for t, (y, _, _) in clubs.items()}, kept
# 1. the test: learn from one half of the seasons, predict each club's home advantage in the other half
halves = {"odd": names[1::2], "even": names[0::2]}
for learn, check in (("odd", "even"), ("even", "odd")):
a, b = home_advantage(halves[learn]), home_advantage(halves[check])
average, spread, pooled, _ = pool(a)
common = [t for t in a if t in b]
print(f"learn on {learn} seasons, check on {check}: {len(common)} clubs, spread between clubs {spread:.3f}")
for method, guess in (("no pooling", lambda t: a[t][0]), ("complete pooling", lambda t: average), ("partial pooling", lambda t: pooled[t])):
plain = mean((guess(t) - b[t][0]) ** 2 for t in common)
scaled = mean((guess(t) - b[t][0]) ** 2 / b[t][1] for t in common)
print(f" {method:17} mean squared error {plain:.4f}; scaled by each club's noise {scaled:.2f}")
# 2. every season together: the hardest and easiest grounds to visit
clubs = home_advantage(names)
average, spread, pooled, kept = pool(clubs)
print(f"all 26 seasons: league average {average:+.2f} points a game, spread between clubs {spread:.3f}")
for t in sorted(clubs, key=pooled.get, reverse=True):
print(f" {t:15} raw {clubs[t][0]:+.2f} from {clubs[t][2]:3} home games -> pooled {pooled[t]:+.2f} (keeps {kept[t]:.0%} of its own figure)")
Further reading
- Empirical Bayes method, Wikipedia. Letting the data choose the prior, the approach used here, and how it relates to hierarchical models.
- Stein's example, Wikipedia. The surprising result that, with three or more groups, pulling every estimate towards a common value beats using each group's own average.
- Shrinkage (statistics), Wikipedia. The general idea of pulling noisy estimates towards a common value, and where it's used.