# Tuning Elo honestly. Five dials and a moving target

Source: https://www.bryanmcguire.co.uk/learn/elo-tuned
Published: 2026-09-30

> Elo has dials to set, from how far one result moves a rating to how much home advantage is worth. Tuned on held-back seasons it draws level with Dixon-Coles on the test seasons. The dial that mattered most changed between eras, and no honest tuning could have seen it coming.

## The football question

Elo is the simplest team rating there is. Every side has a number; a win against a stronger side moves it up more than a win against a weaker one. The [Elo Ratings Simulator](/models/elo) shows how. But Elo comes with dials: how far one result should move a rating, how much home advantage is worth, what happens over the summer. **Set those dials carefully, and how good a forecaster is Elo?**

Choosing a model's dials is called **tuning**, and the dials are **hyperparameters**: settings the model doesn't learn by itself. Tuning honestly turns out to be harder than it looks, because the world the dials are set for doesn't stand still.

## The model

A match's **expectancy**, the share of the points the home side should take, counting a draw as half, comes from the rating gap:

$$E = \frac{1}{1 + 10^{-(R_{\text{home}} + H - R_{\text{away}})/400}}$$

<div class="plain" markdown="1">
In plain football

- **R** is a side's rating, and **H** is home advantage, counted as extra rating points for the home side.
- Equal ratings and no home advantage give **E = 50%**. A 400-point gap makes the stronger side ten times as likely to take the points.
</div>

After the match, both ratings move by the same amount, in opposite directions:

$$\begin{aligned} \text{change} &= K \times m \times (W - E) \\ m &= \ln(\text{winning margin} + 1) \end{aligned}$$

<div class="plain" markdown="1">
In plain football

- **W** is what happened for the home side: 1 for a win, ½ for a draw, 0 for a defeat. **W − E** is the surprise.
- **K** is how far one surprise moves a rating. Big K reacts fast; small K is steady.
- **m** makes a big win count for more than a narrow one: 0.69 for a one-goal win, 1.10 for two goals, 1.61 for four. It's 1 for a draw. Whether to use it at all is one of the dials.
</div>

### From expectancy to three chances

Expectancy alone can't be scored against home, draw and away: 60% could be 60% wins and no draws, or 50% wins and 20% draws. So one more step turns it into three chances, with draws likeliest when the sides are even:

$$\begin{aligned} P(\text{draw}) &= 4d \times E(1 - E) \\ P(\text{home}) &= E - P(\text{draw})/2 \end{aligned}$$

<div class="plain" markdown="1">
In plain football

- **d** is the draw chance between perfectly even sides. At E = 50%, 4 × E × (1 − E) = 1, so the draw chance is exactly d. It shrinks as one side becomes a clear favourite.
- **d** is learned from the training seasons, 2001/02 to 2015/16, and comes out at **27%**.
</div>

## Five dials

| Dial | What it does | Values tried |
|---|---|---|
| K | How far one result moves a rating | 15, 20, 25, 30 |
| Home advantage | Extra rating points at home | 20 to 70 |
| Carry-over | How much of a rating survives the summer | 0.8 to 1 |
| Promoted start | A promoted side's first rating | 1450, 1500, 1550 |
| Margin | Whether bigger wins count for more | Yes or no |

That's **720 combinations**. Elo runs through every Premiership match from 2000/01, and each combination is judged, as in [gradient boosting](/learn/gradient-boosting) and [Dixon-Coles](/learn/dixon-coles-ratings), on the **held-back seasons, 2016/17 to 2020/21**. The test seasons stay unseen until the end.

The winner: **K 20, home advantage 40, 90% carry-over, promoted sides start at 1550, margin counted**.

### A flat valley

The held-back seasons barely separate the settings. **63 of the 720** score within 0.001 of the best, and between them they include every promoted start tried, margin counted and not counted, and carry-over from 0.8 to 0.95. Only two dials are pinned down: K between 20 and 30, and home advantage between 30 and 50.

That's normal, and it's why the odd-looking start of 1550, above the league average, isn't worth worrying about: the held-back seasons would have been almost as happy with 1450.

## The test

| Model | Test accuracy | Test log loss |
|---|---|---|
| Logistic regression | 54.8% | 0.950 |
| **Elo, tuned on held-back seasons** | **54.9%** | **0.945** |
| Dixon-Coles | 55.3% | 0.945 |
| Bookmaker | 56.2% | 0.932 |

Tuned Elo draws level with [Dixon-Coles](/learn/dixon-coles-ratings) and gets past logistic regression. Like Dixon-Coles, it updates after every match, and it rates each side individually.

But there's a twist. The [Elo Ratings Simulator](/models/elo)'s own defaults, K 20 and home advantage 60, with ratings carried over in full and the margin ignored, score **worse** on the held-back seasons (0.9702) and slightly **better** on the test (0.944). The 20 best settings on the held-back seasons score anywhere from 0.945 to 0.951 on the test. Why doesn't careful tuning pick the best setting for the test?

## The dial that moved

Hold every other dial where tuning put it, and turn home advantage:

<figure class="rank-chart">
<div role="img" aria-label="Log loss as home advantage goes from 20 to 100 rating points. On the held-back seasons the best is 30 to 40, at 0.9656, rising to 0.9866 at 100. On the test seasons the best is 70, at 0.9400, with 60 and 80 almost as good; 40, the value chosen, scores 0.9449.">

</div>
<figcaption>Log loss, lower is better. The held-back seasons wanted 30 to 40 points of home advantage; the test seasons wanted about 70.</figcaption>
</figure>

| Home advantage | Held-back | Test |
|---|---|---|
| 40 (chosen) | **0.9656** | 0.9449 |
| 60 | 0.9686 | 0.9406 |
| 70 | 0.9716 | **0.9400** |

With 70 points of home advantage, Elo would have scored **0.940** on the test, better than Dixon-Coles. But 70 was the worst of these three on the held-back seasons. Nothing honest points to it; only the test does, and choosing a setting by looking at the test turns the exam into homework.

The reason is in the results themselves:

<figure class="rank-chart">
<div role="img" aria-label="Home goals minus away goals per match for every season from 2000/01 to 2025/26. It is mostly between plus 0.2 and plus 0.4. The held-back seasons include some of the lowest values, including 2020/21, played behind closed doors, at plus 0.11. The test seasons, 2021/22 to 2025/26, include the highest, plus 0.46 in 2022/23.">

</div>
<figcaption>Home advantage, season by season. The held-back seasons, in white, include some of the lowest, and 2020/21 behind closed doors; the test seasons, in gold, some of the highest.</figcaption>
</figure>

| Seasons | Home wins | Home goal edge |
|---|---|---|
| 2001/02 to 2015/16, training | 43.5% | +0.28 a match |
| 2016/17 to 2020/21, held back | 42.0% | +0.25 a match |
| 2021/22 to 2025/26, test | 46.0% | +0.38 a match |

**Home advantage moved.** The test seasons had more of it than any earlier stretch: home sides won 46% of matches and outscored visitors by 0.38 goals a game, against 42% and 0.25 in the held-back seasons. Tuning did its job; it found the best home advantage for the recent past. The recent past just wasn't like the future.

## Why it matters

- **Tuning is choosing on data you haven't trained on, and not on the test.** The held-back seasons are there so the test stays a fair exam. Picking 70 here would make the score look better and mean less.
- **Flat valleys are normal.** Many settings do almost equally well; the exact winner matters less than getting the important dials roughly right.
- **The world drifts.** Home advantage, styles of play and the gap between clubs all change. A model tuned on one era is a snapshot, which is why [walking forward through the seasons](/learn/cross-validation) and re-tuning regularly beats tuning once.
- **Simple can be enough.** A rating system from chess, with five dials, matches a model fitted by maximum likelihood.

## Limitations

- **One home advantage for everyone.** Some grounds are harder to visit than others; Elo here gives every side the same boost. [Partial pooling](/learn/partial-pooling) shows the clubs differ by less than you'd think.
- **The draw step is simple.** One draw rate for even matches, shrinking with the gap; it doesn't know that some sides draw more than others.
- **One held-back block.** Tuning on several blocks, as in [cross-validation](/learn/cross-validation), would be steadier, though it couldn't have foreseen the rise in home advantage either.
- **Results only.** No injuries, transfers or managers, like every model in this series.

## Try it yourself

Open the [Elo Ratings Simulator](/models/elo) and set two equal ratings. Move home advantage from 40 to 70 and watch the home expectancy: 55.7% to 59.9%. That four-point shift, on every home match, is the difference between this part's two scores.

## Reproduce the analysis

The results files are published by [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php). Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as `SC0_2425.csv`; they aren't rehosted on this site. It tries all 720 settings, which takes about ten seconds:

```python
import csv
from collections import Counter
from datetime import datetime
from itertools import product
from math import log

names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
label = lambda s: f"20{s[:2]}/{s[2:]}"

# every Premiership match in date order; `tested` marks the ones parts 8-12 were scored on (both sides 5+ games in)
matches, teams_in = [], {}
for s in names:
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in ("H", "D", "A")]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    played = Counter()
    for r in games:
        h, a = r["HomeTeam"], r["AwayTeam"]
        matches.append((s, h, a, int(r["FTHG"]), int(r["FTAG"]), r["FTR"], played[h] >= 5 and played[a] >= 5))
        played[h] += 1
        played[a] += 1
    teams_in[s] = set(played)

def elo(k, home, carry, start, margin):
    """Run Elo through every match; return each match's home expectancy, made before the match."""
    rating, season, out = {}, None, []
    for s, h, a, hg, ag, result, tested in matches:
        if s != season:  # summer: ratings drift back towards 1500; promoted sides start at `start`
            season = s
            rating = {t: 1500 + carry * (rating[t] - 1500) if t in rating else start for t in teams_in[s]}
        e = 1 / (1 + 10 ** (-(rating[h] + home - rating[a]) / 400))
        out.append((s, e, result, tested))
        w = {"H": 1, "D": 0.5, "A": 0}[result]
        step = k * (w - e) * (log(abs(hg - ag) + 1) if margin and hg != ag else 1)  # bigger wins move ratings further
        rating[h] += step
        rating[a] -= step
    return out

def chances(e, draw):  # expectancy to home, draw and away: draws likeliest when the sides are even
    d = draw * 4 * e * (1 - e)
    return {"H": e - d / 2, "D": d, "A": 1 - e - d / 2}

def log_loss(out, draw, first, last):
    picked = [(e, r) for s, e, r, tested in out if tested and first <= s <= last]
    return sum(-log(chances(e, draw)[r]) for e, r in picked) / len(picked)

def accuracy(out, draw, first, last):
    picked = [(e, r) for s, e, r, tested in out if tested and first <= s <= last]
    return sum(max("HDA", key=chances(e, draw).get) == r for e, r in picked) / len(picked)

def fit(settings):  # the draw rate is learned on the training seasons, 2001/02-2015/16
    out = elo(*settings)
    draw = min((d / 100 for d in range(22, 32)), key=lambda d: log_loss(out, d, "0102", "1516"))
    return out, draw

# 1. every combination of the five dials, judged on the held-back seasons 2016/17-2020/21
grid = []
for settings in product((15, 20, 25, 30), (20, 30, 40, 50, 60, 70), (0.8, 0.85, 0.9, 0.95, 1.0), (1450, 1500, 1550), (False, True)):
    out, draw = fit(settings)
    grid.append((log_loss(out, draw, "1617", "2021"), log_loss(out, draw, "2122", "2526"), settings, draw))
grid.sort()
held, test, best, draw = grid[0]
print(f"{len(grid)} settings. Best on the held-back seasons: K {best[0]}, home advantage {best[1]}, carry-over {best[2]}, "
      f"promoted sides start at {best[3]}, margin {'counted' if best[4] else 'ignored'}, draw rate {draw:.2f}")
print(f"  held-back log loss {held:.4f}; test log loss {test:.3f}, accuracy {accuracy(fit(best)[0], draw, '2122', '2526'):.1%}")
print(f"  settings within 0.001 of the best on the held-back seasons: {sum(g[0] <= held + 0.001 for g in grid)}; "
      f"the 20 best on held-back score {min(g[1] for g in grid[:20]):.3f} to {max(g[1] for g in grid[:20]):.3f} on test")
near = [g[2] for g in grid if g[0] <= held + 0.001]
for i, dial in enumerate(("K", "home advantage", "carry-over", "promoted start", "margin counted")):
    print(f"  {dial} among those {len(near)}: {sorted(set(x[i] for x in near))}")
page = fit((20, 60, 1.0, 1500, False))
print(f"  the Elo page's defaults (K 20, home advantage 60, ratings carried over in full, margin ignored): held-back {log_loss(page[0], page[1], '1617', '2021'):.4f}, "
      f"test {log_loss(page[0], page[1], '2122', '2526'):.3f}")

# 2. home advantage, the dial that matters: everything else as chosen
print("home advantage: held-back log loss / test log loss")
for home in range(20, 101, 10):
    out, d = fit((best[0], home, best[2], best[3], best[4]))
    print(f"  {home:3}: {log_loss(out, d, '1617', '2021'):.4f} / {log_loss(out, d, '2122', '2526'):.4f}")

# 3. why: home advantage moved. Home win share and home goal edge by season, and by period
per = {}
for s in names:
    ms = [m for m in matches if m[0] == s]
    per[s] = (len(ms), sum(m[5] == "H" for m in ms), sum(m[3] - m[4] for m in ms))
for name, first, last in (("2001/02-2015/16, training", "0102", "1516"), ("2016/17-2020/21, held back", "1617", "2021"),
                          ("2021/22-2025/26, test", "2122", "2526")):
    ss = [s for s in names if first <= s <= last]
    n = sum(per[s][0] for s in ss)
    print(f"{name}: home wins {sum(per[s][1] for s in ss) / n:.1%}, home goal edge {sum(per[s][2] for s in ss) / n:+.2f} a match")
print("by season (home goal edge):", {label(s): round(per[s][2] / per[s][0], 2) for s in names})
```

## Further reading

- [Elo rating system](https://en.wikipedia.org/wiki/Elo_rating_system), Wikipedia. The system's origins in chess, the formula, and how the K-factor is chosen.
- [World Football Elo Ratings](https://en.wikipedia.org/wiki/World_Football_Elo_Ratings), Wikipedia. Elo for national teams: how match importance and the number of goals change the size of each update, with worked examples.
- [Home advantage](https://en.wikipedia.org/wiki/Home_advantage), Wikipedia. What causes it, how it's measured, and what happened to it in matches played without crowds.
- [Hyperparameter optimization](https://en.wikipedia.org/wiki/Hyperparameter_optimization), Wikipedia. Grid search, the method used here, and the alternatives.
