How our Premier League predictions work
No hunches and no tipsters: these are probabilities from a model trained on every Premier League match since 1993, refreshed automatically several times a day.
1. The data
- Results since 1993-94, more than 12,700 matches, with shots, shots on target, corners and cards.
- Fixtures and live results, updated every couple of hours.
- Expected goals (xG) where it's available.
- Bookmaker closing odds, used only to grade our predictions afterwards, never as an input. A model that copies the bookmakers can't tell you anything they don't.
2. Team strength: Elo ratings
Every club carries a rating that goes up when it beats expectations and down when it doesn't. Bigger wins move it more. Ratings drift a third of the way back towards average each summer, and promoted sides start at the level of the teams they replaced.
3. The score model: Dixon-Coles
Each team gets an attacking and a defensive strength, plus a home advantage. Together they give the expected goals for each side, and a Poisson distribution turns those into a probability for every possible scoreline. The Dixon-Coles correction fixes the plain Poisson model's known weakness with 0-0, 1-0, 0-1 and 1-1 scores. Recent games count for more (a match a year ago carries about a third of the weight of last weekend's). Teams with little recent top-flight data are pulled towards the average rather than trusted blindly.
The strengths are fitted mostly on expected goals (xG) rather than actual goals. xG measures the quality of chances created and conceded, so it is less noisy than the scoreline: a lucky 1-0 counts for less than a dominant 1-0. In our backtest a 70/30 blend of xG and real goals beat goals alone.
On top of team strength, the model uses each side's Elo gap and its recent shots-on-target dominance. In testing, shots on target predicted future results better than goals alone. We tried other inputs too, such as days of rest; anything that didn't improve the backtest was left out.
What about transfer spending? We tested summer net spend and squad market values. Spending added nothing. Last season's squad value didn't either: money matters, but it already shows up in results, xG and ratings. This season's market values did seem to help, but those figures are updated during the season based on how teams perform, so the gain was mostly hindsight. We left money out of the model rather than claim an edge we can't test honestly. Team pages still show squad values and spending for context.
4. From scores to everything else
The full scoreline grid gives the win/draw/loss chances, the most likely score, over/under 2.5 goals and both-teams-to-score. Simulating the rest of the season 10,000 times with the same grid produces the predicted table.
5. An AI second opinion: Jev
For every match we also ask Jev, a model from TypeSafe AI built to give probabilities rather than chat. It gets a written briefing like a pundit's notes: the league table, each side's last 12 results with expected goals, home and away records, the last ten meetings, rest days and transfer spending. It never sees our model's forecast or the betting odds.
How you ask matters enormously. In testing on 200 matches it had never seen, asking Jev to pick the result made it wildly overconfident: when it said 90%, it was right about half the time. Asking it three separate yes/no questions ("Will the home side win?", "Will it be a draw?", "Will the away side win?") made it as well calibrated as our statistical model, and on those 200 games it matched our model and edged the bookmakers.
The two views agree most of the time, but not always, so we also track a consensus (the simple average of the two). All three are scored on the accuracy page. If the consensus keeps winning once a full set of live predictions has built up, it will become our headline prediction.
6. Honesty rules
- Predictions lock at kick-off and are never changed after the event.
- Games played before launch are marked as pre-launch: generated afterwards, but using only data from before each match.
- The model's settings were chosen on 2005–2017 data and tested unchanged on 2018 onwards.
Over 8,030 matches in that backtest, the model picked the right result 54% of the time. The bookmakers' closing favourite was right 55% of the time. Closing odds are the sharpest public forecast there is, so getting close to them is the realistic goal.
These are probabilities, not certainties. A 60% favourite fails to win four times in ten. Nothing here is betting advice.
Questions
How are the Premier League predictions made?
A Dixon-Coles model, fitted mostly on expected goals (xG) with Elo ratings and shots-on-target form, turns each team's attack and defence into a probability for every scoreline. It is refitted several times a day on every Premier League result since 1993.
What is Jev?
Jev is an AI model from TypeSafe AI that gives calibrated probabilities. We give it a written briefing on each match (never the odds or our own forecast) and ask three yes/no questions, as an independent second opinion.
Do the predictions use betting odds?
No. Bookmaker closing odds are only used afterwards, to grade the predictions.