What Football Prediction Accuracy Actually Means (And Why 90% Claims Are Fake)
Hit rate is the most misleading number in football prediction. Here is how forecast quality is really measured — calibration, Brier score, log loss, closing line value — and how to audit any site's claim in five minutes.
A prediction site advertising 90% accuracy is almost always measuring something meaningless. Accuracy on its own is trivially gamed: predict "under 5.5 goals" on every match and you will be right nearly every time, while losing money steadily. This article covers the metrics that cannot be gamed, and how to check a claim yourself.
The three ways accuracy claims get inflated
1. Choosing easy markets
Not all predictions are equally hard. "Over 0.5 goals" lands the overwhelming majority of the time. Double chance on a strong home side lands most of the time too. A blended hit rate across markets like these reports how easy the markets were, not how good the model is.
2. Survivorship in the record
If picks are published but the archive is compiled afterwards, losers quietly disappear. The test is whether each pick was timestamped and visible before kickoff, and whether the archive contains everything since day one.
3. Counting non-losses as wins
Postponements, Asian handicap pushes and cashed-out positions all get recorded as "not a loss". Across a season that alone can move a headline number several points.
What forecast quality actually measures
Calibration
Take every match the model called 70%. Did roughly 70% of them happen? Repeat for the 30% bucket, the 50% bucket, and so on. A calibrated model's confidence means what it says. It is the single most useful check, and almost no prediction site publishes it.
Calibration also catches the opposite failure: a model that is often right but always says 95% is badly calibrated and dangerous to stake against, because you cannot size a bet using a number that lies.
Brier score
The mean squared error of the probability forecast. Say 0.7, the event happens, you are charged (1 − 0.7)² = 0.09. Lower is better, zero is perfect. It rewards being confident and correct, and punishes confident errors much harder than cautious ones.
Log loss
The same idea with a harsher penalty for confident mistakes — saying 99% and being wrong is close to catastrophic. Log loss is what most models are trained against, precisely because it forces honest probabilities.
| Metric | Answers | Gameable? |
|---|---|---|
| Hit rate | How often the top pick was right | Yes — pick easy markets |
| Calibration | Whether the confidence numbers mean anything | No |
| Brier score | How far off the probabilities were | No |
| Log loss | The same, punishing confident errors | No |
| Closing line value | Whether it found real edge | No |
The benchmark that settles arguments
The closing odds — the price at kickoff, after all the money is in — are the most accurate public forecast of a football match that exists. They already contain team news, weather, sharp money and everything else.
So the honest test of any model is whether it beats the closing line. If it consistently backs selections at prices better than where those prices close, it is finding real information. If it does not, it can still post a pretty hit rate while quietly losing to the margin.
This is a demanding benchmark, which is why any serious claim of edge should be modest. A model beating the close by a few percent is doing genuinely well.
How to audit a prediction site in five minutes
- Find the full archive. No public record of every pick, no conversation. Ours is at picks history.
- Check the sample size. Under a few hundred settled picks the record is noise. A 60% hit rate over 40 picks is something a coin does regularly.
- Check the markets. Is the record built on 1X2 calls, or on over 0.5 goals?
- Check for odds. A record without the price taken cannot be turned into profit or loss, which is the only number that matters.
- Check the confidence buckets. If high-confidence picks do not outperform low-confidence ones, the confidence score is decoration.
What a believable number looks like
On 1X2 across competitive leagues, a strong model lands somewhere in the high 50s, and higher on the subset where it is most confident. Anything advertised far above that, on hard markets, over a long sample, is a claim to verify rather than a fact.
The uncomfortable truth is that a good model is not one that is right most of the time. It is one whose 55% actually means 55%. That is harder to build and much less exciting to advertise.
Related: what AI can and cannot do and how to evaluate a prediction site.
Équipe Analytique Prodict
Ingénieurs IA Data & Prédiction
Cette analyse est produite par le modèle d'intelligence artificielle principal de Prodict. En traitant des millions de données football historiques et en temps réel, le modèle détecte les paris valeur et les avantages algorithmiques indépendamment du biais humain.