Chapter 3 · 7 min read
Chapter 3 — Learning from past results without fooling yourself
When the match ends, the most important and least-practiced part begins: learning. Most bettors judge their picks by outcome only (won = good, lost = bad). That's wrong. A good decision on a random event can lose; a bad one can win. What matters is the PROCESS. Here's how to evaluate it properly.
1. Was it deserved? — final xG as referee
When a match ends, the first question isn't 'who won' but 'who should have'. Home xG 2.1 vs away xG 0.6 with score 0-1 tells a clear story: the away side stole it. Expecting the same to repeat over the next 3 matches is betting against mean reversion. PronoStats stores final xG for every match and recap pages show the delta between actual and expected score. A team that lost three in a row with positive xG is a value buy in the next market.
2. Brier score — the right way to grade a prediction
If the model says '70% home win' and the away wins, did it fail? Not necessarily. That 0-1 was predicted at 15% — it happens. The right question isn't 'was it right?' but 'were the probabilities calibrated?'. The tool is Brier score: how far predicted probabilities are from real outcomes. PronoStats computes Brier for both the ML model and the Groq correction on every match. The month's Brier tells the truth about model reliability: close to 0 = great, above 0.30 = room to improve.
3. Sample size — don't trust 5 matches
The most common mistake: 'The AI got 3 in a row wrong, it sucks'. Three matches aren't a sample. To judge 1X2 accuracy you need at least 50 matches. For Over/Under 30. For xG-based markets 80. Below that, you're a victim of variance. PronoStats keeps 1,200+ matches with logged feedback — those numbers speak. The last 5 are curiosities, not statistics.
4. Accuracy trend — is the model getting better?
Static accuracy is a snapshot. Trend is the movie. PronoStats tracks daily accuracy for every market (1X2, O2.5, BTTS) and charts it on /accuracy. A model getting worse week over week likely has data drift or a broken feature. One getting better is learning from feedback. Week-on-week is too noisy — read 30-day rolling windows.
5. Per-league accuracy — where is the AI strongest?
Not all leagues are equal. The model performs best where it has the most data and the most tactical stability (Bundesliga Over 2.5: ~70%, Serie A Over 2.5: ~62%). It performs worst with tactical fragmentation or short history (MLS, lower divisions). The /accuracy page has a 'By league' section that tells you exactly where to back the AI and where to skip. Operational rule: bet where the AI has a solid track record, not because you like the match.
6. Calibration — the model's ultimate test
A calibrated model is one where, when it says 70%, that event actually happens 70% of the time. Sounds obvious but it's rare. Many models are over-confident (say 80%, reality 65%) or over-cautious (say 55%, reality 70%). PronoStats blends ML probabilities with bookmaker odds 60-40 specifically to improve calibration. When you see a flagged edge, the first question isn't 'how big' but 'how calibrated is the model on that market in that league'.
Chapter quiz
Check if the concepts stuck — nothing tracked, just for you.
1.The AI got 3 in a row wrong. Is it clear the model is bad?
2.A Brier score of 0.20 is...
3.In which league does PronoStats currently have the best accuracy?
Frequently asked questions
What is a 'good' Brier score?
Below 0.25 is great for 1X2 football (3 outcomes). Below 0.20 is excellent. The PronoStats ML model sits around 0.79 raw, ~0.64 after Groq correction.
Why is the ML Brier worse than Groq's?
Because Groq sits ON TOP of ML: it reads the ML probabilities, adds context (news, last-hour injuries, motivation) and corrects. That's the value-add.
Where can I see historical accuracy for a single market?
On the /accuracy page with filters for market (1X2, O2.5, BTTS) and period (today, week, month, all-time).