1–0 up at half-time. Will they hold on? Bayes' theorem
A half-time lead is evidence, and what it means depends on who's leading. Bayes' theorem combines the pre-match odds with the half-time score, and 5,419 SPFL matches show it working.
Beginner Part 3 of Bayesian Thinking Through Football
Contents
The football question
It's half-time and a team is 1–0 up. Will they hold on?
It depends who they are. A heavy favourite leading at the break is almost home. An underdog leading at the break has done something unexpected, and the second half might put things back the way everyone thought. The half-time score is evidence; what it means depends on what you believed before kick-off.
Combining the two is exactly what Bayes' theorem does. Bayesian updating showed the idea with a penalty taker's record, and where priors come from showed where the starting belief comes from. This is the rule underneath both.
The concept
Three pieces:
- The prior: what you believed before the evidence. Here, a team's chance of winning before kick-off.
- The evidence: what you've seen since. Here, the half-time score.
- The posterior: what you believe now. Here, their chance of winning, knowing the half-time score.
Bayes' theorem gets you from one to the other:
$$P(\text{win} \mid \text{lead}) = \frac{P(\text{lead} \mid \text{win}) \times P(\text{win})}{P(\text{lead})}$$
In plain football
- P(win | lead) is what we want: the chance a team wins, given it's leading at half-time. The bar "|" reads "given".
- P(lead | win) is the other way round: of the teams that won, how many were leading at half-time?
- P(win) is the prior, how often a team wins at all; P(lead) is how often a team leads at half-time.
The flip that catches people out
Those two "givens" sound alike but aren't. In the Scottish Premiership from 2002/03 to 2025/26, 5,419 matches with bookmakers' odds:
- Of teams leading at half-time, 76.4% won.
- Of teams that won, 60.9% had been leading at half-time.
Four in ten winners came from level or behind at the break. Mixing up "the chance of A given B" with "the chance of B given A" is one of the commonest mistakes in probability, in football and far beyond it. Bayes' theorem is how you get from one to the other correctly. In these matches a team won 38.1% of the time and led at half-time 30.3% of the time, so:
$$\frac{0.609 \times 0.381}{0.303} = 0.764$$
In plain football
- 0.609: of winners, the share that were leading at half-time.
- × 0.381 ÷ 0.303: scale by how common winning is, compared with how common a half-time lead is.
- 0.764: of teams leading at half-time, the share that won. Exactly the count from the data.
A football example
Now use a better prior than "a team wins 38% of the time". Before every match, the bookmakers give each side a chance of winning. Turn their odds into probabilities, group the teams by that pre-match chance, and see what happened after each half-time score:
| Chance before | Behind | Level | Ahead |
|---|---|---|---|
| 15% (under 25%) | 2.4% | 15.3% | 52.3% |
| 35% (25–45%) | 9.0% | 30.0% | 70.6% |
| 53% (45–65%) | 16.0% | 43.3% | 83.1% |
| 75% (65%+) | 29.4% | 63.9% | 93.1% |
Same evidence, different conclusions. A half-time lead turns an underdog into a slight favourite, just over 50%, and a favourite into a near-certainty. Being level at half-time barely moves an underdog, but it knocks a favourite from 75% down to 64%: they were expected to be ahead by now. And a favourite trailing at half-time still wins 29% of the time, more often than an underdog wins at all.
Odds and a multiplier
There's a neater way to write Bayes' theorem, using odds instead of probabilities:
$$\begin{aligned} &\text{odds after} \\ &= \text{odds before} \times \text{likelihood ratio} \end{aligned}$$
In plain football
- Odds are chance of winning ÷ chance of not winning. A 15% underdog is at odds of 0.15 ÷ 0.85 ≈ 0.18.
- The likelihood ratio is how much more often winners were leading at half-time than non-winners. Across all these matches it's 5.3.
- So a half-time lead multiplies any team's odds by about 5. The underdog goes from 0.18 to about 0.93: a chance of about 48%.
The actual figure for underdogs was 52.3%, so one multiplier of about 5 gets close, but it isn't the whole story. The multiplier differs by group:
| Before kick-off | Half-time lead multiplies the odds by |
|---|---|
| Underdogs (under 25%) | 6.6 |
| 25% to 45% | 4.4 |
| 45% to 65% | 4.1 |
| Favourites (65%+) | 3.8 |
A lead is stronger evidence for an underdog. When a favourite leads at half-time, that's what everyone expected, so it tells you less. When an underdog leads, something unusual is going on, perhaps they're better on the day than the odds allowed, and the lead carries more information. That's Bayesian thinking in one line: the surprise in the evidence is what moves your belief.
For how the size of a half-time lead matters, see is 2–0 the most dangerous lead?
Why it matters
- Evidence means nothing without a prior. "They're 1–0 up" tells you one thing about Celtic away at a bottom side and something very different about the bottom side.
- Watch the flip. "Most teams who lead at half-time win" and "most winners led at half-time" are different claims with different numbers.
- In-play betting is Bayes' theorem. Live odds are the pre-match odds updated by the score, the clock and everything else, all match long.
- Surprise is informative. The less expected the evidence, the more it should change your mind.
Limitations
- Bookmaker odds aren't pure probabilities. They include a margin, removed here by scaling the three to add up to 1, which is approximate.
- Groups hide detail. Within "under 25%" there are 5% outsiders and 24% underdogs; finer groups need more matches.
- Half-time is one moment. The score at 44 minutes, red cards and injuries all change the picture; a real in-play model updates continuously.
- "Not winning" lumps together draws and defeats. A trailing favourite that scrapes a draw counts the same as one that loses.
Try it yourself
At half-time in the next match you watch, write down each side's pre-match chance (from the odds, or your own guess) and the score. Using the tables above, what's your updated chance for the team leading? At full time, see how it went. Over a weekend of matches, do the leading underdogs hold on about half the time?
Reproduce the analysis
The results and odds files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. Seasons without Bet365 odds (the first two) are skipped. Then:
import csv
# one row per team per match: the bookmaker's pre-match chance it wins, where it stood at half time, and whether it won
teams = []
for y in range(2000, 2026):
with open(f"SC0_{y % 100:02d}{(y + 1) % 100:02d}.csv", encoding="latin-1") as f:
for r in csv.DictReader(f):
try:
inverse = {k: 1 / float(r["B365" + k]) for k in "HDA"}
except (KeyError, ValueError):
continue # no odds this season
if r.get("FTR") not in inverse or r.get("HTR") not in inverse:
continue
for side in "HA":
half = "lead" if r["HTR"] == side else "level" if r["HTR"] == "D" else "trail"
teams.append((inverse[side] / sum(inverse.values()), half, r["FTR"] == side))
n = len(teams)
won = [t for t in teams if t[2]]
leading = [t for t in teams if t[1] == "lead"]
p_win, p_lead = len(won) / n, len(leading) / n
p_lead_given_win = sum(t[1] == "lead" for t in won) / len(won)
print(f"{n // 2} matches; P(win) {p_win:.1%}, P(lead at half time) {p_lead:.1%}")
print(f"P(win | lead) {sum(t[2] for t in leading) / len(leading):.1%}; P(lead | win) {p_lead_given_win:.1%}")
print(f"Bayes: {p_lead_given_win:.3f} x {p_win:.3f} / {p_lead:.3f} = {p_lead_given_win * p_win / p_lead:.1%}")
def ratio(group, half): # how much more often winners were in this half-time state than non-winners
w = [t for t in group if t[2]]
o = [t for t in group if not t[2]]
return (sum(t[1] == half for t in w) / len(w)) / (sum(t[1] == half for t in o) / len(o))
band = lambda p: "under 25%" if p < 0.25 else "25-45%" if p < 0.45 else "45-65%" if p < 0.65 else "65% and over"
lead_ratio = ratio(teams, "lead")
print(f"likelihood ratio: lead {lead_ratio:.2f}, level {ratio(teams, 'level'):.2f}, trail {ratio(teams, 'trail'):.2f}")
for b in ("under 25%", "25-45%", "45-65%", "65% and over"):
group = [t for t in teams if band(t[0]) == b]
prior = sum(t[0] for t in group) / len(group)
after = {h: [t[2] for t in group if t[1] == h] for h in ("lead", "level", "trail")}
odds = prior / (1 - prior) * lead_ratio
print(f"{b:13} {len(group):5} teams, prior {prior:.1%}: won when leading {sum(after['lead']) / len(after['lead']):.1%}, "
f"level {sum(after['level']) / len(after['level']):.1%}, trailing {sum(after['trail']) / len(after['trail']):.1%}; "
f"lead ratio {ratio(group, 'lead'):.2f}; Bayes with {lead_ratio:.2f} says {odds / (1 + odds):.1%}")
Further reading
- Bayes' theorem, 3Blue1Brown. The theorem as a picture, and why it's easy to get the "given" backwards.
- Bayes' theorem, Wikipedia. The statement, the odds form with the likelihood ratio, and worked examples.
- Confusion of the inverse, Wikipedia. The mistake of treating P(A | B) as if it were P(B | A).