# Sure of itself, or just better informed? Calibration in depth

Source: https://www.bryanmcguire.co.uk/learn/calibration-in-depth
Published: 2026-09-30

> A forecast can lose by being wrong about its own confidence, or by knowing less. Splitting the Brier score into calibration and resolution shows which. On five SPFL test seasons Elo is as well calibrated as the bookmakers, and loses only because they know more.

## The football question

[Tuned Elo](/learn/elo-tuned) scores 0.945 on the five test seasons; the bookmakers score 0.932. The gap could mean two quite different things. Maybe Elo is **wrong about its own confidence**: when it says 60%, it happens 50% of the time. Or maybe Elo is honest about its confidence, but **knows less**: it can't tell a 70% match from a 50% one as well as the bookmakers can.

**Which is it?** The first is easy to fix; the second isn't. And the same question applies to every forecast, from a model to a pundit.

## Two ways to be good

A forecaster can be good in two separate ways:

- **Calibration**: its chances mean what they say. Of all the matches it calls 30%, about 30% happen.
- **Resolution**: its chances vary, and in the right direction. It gives high chances to matches that go on to happen and low chances to ones that don't.

They're independent. A forecaster that always gives the league's usual rates for home, draw and away is almost perfectly calibrated and has no resolution at all: honest, and useless. A forecaster that says 90% whenever it fancies a home win might have plenty of resolution and be badly calibrated.

## Splitting the Brier score

The [Brier score](/learn/model-evaluation) splits exactly into three parts:

$$\begin{aligned} \text{Brier} = \;&\text{uncertainty} \\ &- \text{resolution} \\ &+ \text{calibration error} \end{aligned}$$

<div class="plain" markdown="1">
In plain football

- **Uncertainty** is how hard the matches are to call, whoever forecasts them: the Brier score you'd get knowing the true base rates of this set of matches and nothing else. It's the same for every forecaster on the same matches.
- **Resolution** is how much better than the base rates the forecaster's groups of matches turn out. It's subtracted: more resolution, lower Brier, better.
- **Calibration error** is how far, on average, what the forecaster said is from what happened. It's added: any miscalibration makes the score worse.
</div>

To measure the last two, group each forecaster's chances into tenths (every home win it called 30% to 40% together, and so on), and compare, in each group, the average chance it gave with how often it happened. Grouping like this is approximate, so the three parts add up to within 0.005 of the actual Brier score, not exactly.

## Base rates, Elo and the bookmakers

The same 990 test matches as every model in this series, 2021/22 to 2025/26:

<figure class="rank-chart">
<div role="img" aria-label="Two sets of bars. Resolution, higher is better: base rates 0.000, Elo 0.080, bookmaker 0.088. Calibration error, lower is better, on a scale ten times finer: base rates 0.0016, Elo 0.0055, bookmaker 0.0069.">

</div>
<figcaption>The two parts that differ between forecasters. Note the scales: the calibration errors are about a tenth the size of the differences in resolution.</figcaption>
</figure>

| | Resolution | Calibration error |
|---|---|---|
| Base rates | 0.000 | **0.0016** |
| Elo | 0.080 | 0.0055 |
| Bookmaker | **0.088** | 0.0069 |

Their log losses are 1.055, 0.945 and 0.932, and their Brier scores 0.637, 0.558 and 0.550. Uncertainty is 0.6355 for all three, because it depends only on the matches.

Three things stand out.

- **The base rates are the best calibrated of the three**, and the worst forecaster, because they have no resolution at all. Honesty isn't enough.
- **Elo is better calibrated than the bookmakers**: 0.0055 against 0.0069. When Elo says a chance, it happens about as often as it says.
- **The bookmakers win entirely on resolution**: 0.0880 against 0.0804. They tell matches apart better, because they know things Elo doesn't: team news, injuries, a whole market's opinion.

So the answer to the football question is the second one. Elo isn't overconfident or underconfident; it's less informed.

## The reliability diagram

The calibration part can be seen directly. For home wins, here's what each forecaster said against how often it happened:

<figure class="rank-chart">
<div role="img" aria-label="What Elo and the bookmaker said about a home win, grouped into tenths, against how often it happened. Both follow the diagonal closely. Both sit slightly above it in the middle: said 35 percent, happened 41 percent for both; said 54 or 55 percent, happened 62 percent for Elo and 66 percent for the bookmaker.">

</div>
<figcaption>A perfectly calibrated forecaster sits on the dashed diagonal. Dot size shows how many matches are in each group. Both forecasters sit a little above it in the middle: home wins happened more often than either expected.</figcaption>
</figure>

Both lines hug the diagonal, and both miss in the same place: **home wins in the middle range happened more than they said**. Said 35%, happened 41%, for both. Said about 55%, happened 62% for Elo and 66% for the bookmaker. That's the rise in home advantage from [tuning Elo](/learn/elo-tuned) again: the test seasons had more of it than anyone expected, bookmakers included.

### Draws

| Draw chance said | Elo: happened | Bookmaker: happened |
|---|---|---|
| About 10% | 9.3% | 5.7% |
| About 15% | 15.3% | 14.8% |
| About 20% | 18.3% | 17.1% |
| About 26% | 26.7% | 25.7% |

Elo's draw chances, from part 13's simple draw step, come out close. The bookmakers' lowest draw prices are the one clearly miscalibrated group: priced around 10.6%, those draws happened 5.7% of the time, in 53 matches. That's the longshot end of the [favourite–longshot bias](/research/bookmaker-odds): unlikely outcomes are priced a little too generously.

## Can Elo be recalibrated?

If calibration were Elo's problem, a quick fix would help: nudge its chances, making them sharper or softer and shifting home and away up or down, with the nudge fitted on the held-back seasons, 2016/17 to 2020/21.

| Recalibration | Held-back log loss | Test log loss |
|---|---|---|
| None | 0.9656 | **0.9449** |
| Fitted on held-back seasons | 0.9651 | 0.9460 |
| Fitted on the test itself (hindsight) | | 0.9400 |

The fitted nudge leaves the sharpness alone (1.0) and moves home and away chances down slightly, because the held-back seasons had less home advantage. On the test seasons that's the wrong way, and it makes Elo slightly **worse**. Fitted in hindsight on the test itself, the nudge goes the other way, home up, and would reach 0.940, but that's using the exam to mark itself.

Recalibration can only fix miscalibration that stays the same from one period to the next. Elo's small miscalibration here came from home advantage changing, which no fix fitted on the past could anticipate.

## Why it matters

- **Log loss and Brier mix two things.** Splitting them tells you what to work on: a badly calibrated model needs its confidence fixed; a well calibrated one with low resolution needs better information.
- **Better information beats better maths.** The bookmakers' edge is knowledge. It's the same conclusion as [gradient boosting](/learn/gradient-boosting) and [Dixon-Coles](/learn/dixon-coles-ratings) reached from the other direction: the limit is what goes in.
- **Being honest about uncertainty is a skill worth having.** A pundit or scout who's well calibrated, even if not very sharp, can be trusted: their 70% means 70%.
- **Check it yourself.** The [Score your own predictions](/models/score-your-predictions) tool shows the scoring rules on single matches; over a season, grouping your forecasts like this shows whether you're overconfident.

## Limitations

- **Grouping is approximate.** The parts come from forecasts grouped into tenths; finer or coarser groups move them slightly.
- **Five seasons.** 990 matches is enough to see the big picture, but a single group can hold as few as 40 matches, and its gap between said and happened can be luck.
- **Elo only, of the models.** Dixon-Coles and the machine learning models would split the same way; Elo is here because it's quick to run and scores level with the best of them.
- **Bookmaker chances with the margin removed evenly.** Removing the margin in proportion to each price is the simplest method, not the only one, and it can shift the longshots slightly.

## Try it yourself

Before each round of fixtures, write down your chance of a home win for every match. After a few weeks, group them: all your 30%s together, all your 60%s. If your 60%s win 60% of the time, you're well calibrated. If they win 45%, you're overconfident, and now you know by how much.

## Reproduce the analysis

The results files are published by [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php). Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as `SC0_2425.csv`; they aren't rehosted on this site. It runs in about ten seconds:

```python
import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import exp, log

names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

# every Premiership match in date order, with the bookmaker's chances (margin removed) and whether parts 8-13 tested it
matches, teams_in = [], {}
for s in names:
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    played = Counter()
    for r in games:
        h, a = r["HomeTeam"], r["AwayTeam"]
        inverse = {c: 1 / float(r[f"B365{c}"]) for c in "HDA"} if r.get("B365H") else None
        book = {c: v / sum(inverse.values()) for c, v in inverse.items()} if inverse else None
        matches.append({"season": s, "home": h, "away": a, "hg": int(r["FTHG"]), "ag": int(r["FTAG"]), "result": r["FTR"],
                        "book": book, "tested": played[h] >= 5 and played[a] >= 5})
        played[h] += 1
        played[a] += 1
    teams_in[s] = set(played)

# Elo exactly as tuned in part 13: K 20, home advantage 40, 90% carry-over, promoted sides at 1550, margin counted, draws 27%
rating, season = {}, None
for m in matches:
    if m["season"] != season:
        season = m["season"]
        rating = {t: 1500 + 0.9 * (rating[t] - 1500) if t in rating else 1550 for t in teams_in[season]}
    e = 1 / (1 + 10 ** (-(rating[m["home"]] + 40 - rating[m["away"]]) / 400))
    d = 0.27 * 4 * e * (1 - e)
    m["elo"] = {"H": e - d / 2, "D": d, "A": 1 - e - d / 2}
    step = 20 * ({"H": 1, "D": 0.5, "A": 0}[m["result"]] - e) * (log(abs(m["hg"] - m["ag"]) + 1) if m["hg"] != m["ag"] else 1)
    rating[m["home"]] += step
    rating[m["away"]] -= step

pick = lambda first, last: [m for m in matches if m["tested"] and first <= m["season"] <= last]
training, held_back, test = pick("0102", "1516"), pick("1617", "2021"), pick("2122", "2526")
base = {c: sum(m["result"] == c for m in training) / len(training) for c in "HDA"}  # knows nothing but how often each result happens
for m in matches:
    m["base"] = base

def log_loss(ms, who):
    return sum(-log(m[who][m["result"]]) for m in ms) / len(ms)

def brier_parts(ms, who):
    """Brier score split into calibration error, resolution and uncertainty, forecasts grouped into tenths, per result."""
    calibration = resolution = uncertainty = 0.0
    for c in "HDA":
        rate = sum(m["result"] == c for m in ms) / len(ms)
        uncertainty += rate * (1 - rate)
        groups = defaultdict(list)
        for m in ms:
            groups[min(int(m[who][c] * 10), 9)].append((m[who][c], m["result"] == c))
        for g in groups.values():
            said, happened = sum(p for p, _ in g) / len(g), sum(o for _, o in g) / len(g)
            calibration += len(g) / len(ms) * (said - happened) ** 2
            resolution += len(g) / len(ms) * (happened - rate) ** 2
    brier = sum(sum((m[who][c] - (m["result"] == c)) ** 2 for c in "HDA") for m in ms) / len(ms)
    return brier, calibration, resolution, uncertainty

print(f"{len(test)} test matches, 2021/22-2025/26")
for who, name in (("base", "Base rates"), ("elo", "Elo"), ("book", "Bookmaker")):
    brier, cal, res, unc = brier_parts(test, who)
    print(f"{name:10} log loss {log_loss(test, who):.3f}, Brier {brier:.4f}: calibration error {cal:.4f}, resolution {res:.4f}, "
          f"uncertainty {unc:.4f} (uncertainty - resolution + calibration = {unc - res + cal:.4f})")

# reliability: what each forecaster said about a home win, against how often it happened, in tenths
for who, name in (("elo", "Elo"), ("book", "Bookmaker")):
    groups = defaultdict(list)
    for m in test:
        groups[min(int(m[who]["H"] * 10), 9)].append((m[who]["H"], m["result"] == "H"))
    print(f"{name}, home win: " + "; ".join(f"said {sum(p for p, _ in g) / len(g):.0%} happened {sum(o for _, o in g) / len(g):.0%} ({len(g)})"
                                           for _, g in sorted(groups.items()) if len(g) >= 20))
    draws = defaultdict(list)
    for m in test:
        draws[round(m[who]["D"] * 20) / 20].append((m[who]["D"], m["result"] == "D"))
    print(f"{name}, draw: " + "; ".join(f"said {sum(p for p, _ in g) / len(g):.1%} happened {sum(o for _, o in g) / len(g):.1%} ({len(g)})"
                                       for _, g in sorted(draws.items()) if len(g) >= 20))

# recalibration: nudge Elo's chances (a sharpness T and home and away shifts), fitted on the held-back seasons
def nudge(p, t, h, a):
    z = {c: t * log(p[c]) + {"H": h, "D": 0, "A": a}[c] for c in "HDA"}
    top = max(z.values())
    return {c: exp(v - top) / sum(exp(w - top) for w in z.values()) for c, v in z.items()}

grid = [(t / 10, h / 50, a / 50) for t in range(8, 13) for h in range(-10, 11) for a in range(-10, 11)]
score = lambda ms, x: sum(-log(nudge(m["elo"], *x)[m["result"]]) for m in ms) / len(ms)
fitted = min(grid, key=lambda x: score(held_back, x))
hindsight = min(grid, key=lambda x: score(test, x))
print(f"recalibration fitted on held-back seasons: sharpness {fitted[0]}, home {fitted[1]:+.2f}, away {fitted[2]:+.2f}; "
      f"held-back {log_loss(held_back, 'elo'):.4f} -> {score(held_back, fitted):.4f}, test {log_loss(test, 'elo'):.4f} -> {score(test, fitted):.4f}")
print(f"in hindsight, fitted on the test itself: home {hindsight[1]:+.2f}, away {hindsight[2]:+.2f}; test {score(test, hindsight):.4f}")
```

## Further reading

- [Brier score](https://en.wikipedia.org/wiki/Brier_score), Wikipedia. The score, and its split into reliability, resolution and uncertainty, which is the calibration, resolution and uncertainty used here.
- [Probability calibration](https://scikit-learn.org/stable/modules/calibration.html), scikit-learn. Calibration curves (reliability diagrams) and the standard ways of recalibrating a model.
- [Scoring rule](https://en.wikipedia.org/wiki/Scoring_rule), Wikipedia. Log loss, Brier and the other ways of scoring probability forecasts, and why honest forecasts score best.
