Open, transparent research implementation of state-of-the-art statistical football match prediction methodology. NOT a betting tool — built for methodological transparency and learning.
This repository implements and benchmarks the academically-proven stack: pi-ratings + Dixon-Coles Poisson + gradient-boosted trees, combined via a calibrated stacked ensemble, with honest chronological backtesting.
| Component | Source |
|---|---|
| Pi-ratings | Constantinou & Fenton (2013) — Section 2–3; cross-checked against R reference |
| Dixon-Coles | Michels et al. (2023) arXiv:2307.02139 Section 2.1; regista reference |
| GBT + stacking | Survey: arXiv:2403.07669; Evaluation: arXiv:2309.14807 |
| Benchmark structure | jdgoated1/football-predictor (organization reference, not code copy) |
| Data | martj42/international_results |
- Four venue-specific ratings per team:
home_attack,home_defense,away_attack,away_defense - Expected goal difference uses the paper's exponential transform (Section 2.3)
- Error dampening: ψ(e) = 3 × log₁₀(1 + |e|) (Section 2.2)
- Learning rates λ (performance) and γ (cross-context) — tunable; paper EPL optimum ≈ λ=0.035, γ=0.7
- Attack/defense parameters with time decay w(t) = exp(−ξ × days)
- Low-score τ correction for (0-0), (0-1), (1-0), (1-1) per arXiv:2307.02139
- Base models: CatBoost + LightGBM + XGBoost (with sklearn fallbacks when OpenMP libs unavailable)
- Meta-learner: logistic regression + isotonic calibration on validation-fold predictions
python3.11 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
# macOS: if LightGBM/XGBoost fail to load, install OpenMP:
# brew install libomp
python train.py # ~3–4 min on laptop CPU (trains weighted + baseline GBT)
python predict.py "Brazil" "Argentina" --tournament "Copa America" --neutral
python predict.py "Brazil" "Argentina" --decision-rule compare-allUsing MatchBench in Claude Desktop or Cursor? See MatchBench agent setup & usage below.
Split: train pre-2018 · val 2018–2021 · test 2022–2024
Comparable difficulty to academic benchmarks. Do not use aggregate test accuracy — ~48% of test fixtures are heavy Elo mismatches.
| Model | Accuracy | RPS |
|---|---|---|
| Dixon-Coles alone | 0.387 | 0.245 |
| GBT ensemble | 0.411 | 0.225 |
| Stacked ensemble (baseline argmax) | 0.428 | 0.225 |
Default production config: stacked ensemble trained with draw_weight=1.5 and post-hoc decision rule draw_margin=0.075 (validation-tuned on close matches). See draw-recall section below.
Tuning methodology: scripts/tune_draw_decision.py — validation-only search, single test evaluation, RPS-first selection.
| Config | Acc | RPS | Draw recall | Home recall | Away recall |
|---|---|---|---|---|---|
| Baseline argmax (unweighted) | 0.428 | 0.225 | 3.4% | 60.5% | 54.9% |
Margin only (margin=0.05, no retraining) |
0.426 | 0.225 | 9.3% | 56.6% | 53.8% |
Default (draw_weight=1.5 + margin=0.075) |
0.428 | 0.224 | 18.3% | 51.6% | 52.7% |
Draw recall improvements are achieved by shifting decision boundaries and modest retraining reweighting, not by improving underlying probability calibration — P(draw) was already well-calibrated at ~28% before any tuning.
Rejected: per-class threshold rule (th_H=0.36, th_D=0.18, th_A=0.44) — artificially inflates draw recall (~40%) while destroying away-win recall (~20% on test) by over-predicting draws (~35% vs ~28% actual). Not deployed.
Use python predict.py "Brazil" "Argentina" --decision-rule compare-all to see all three configs side by side for a single fixture.
| Model | Accuracy | RPS |
|---|---|---|
| Dixon-Coles alone | 0.558 | 0.192 |
| GBT ensemble | 0.598 | 0.170 |
| Stacked ensemble | 0.603 | 0.171 |
World Cup finals only (n=96): stacked 0.542 acc / 0.211 RPS — in line with published benchmarks.
| Benchmark | Accuracy | RPS |
|---|---|---|
| CatBoost+pi-ratings, 2017 Soccer Prediction Challenge (arXiv:2309.14807) | 0.558 | 0.193 |
| Reference repo — Dixon-Coles alone (football-predictor) | 0.490 | 0.212 |
| Reference repo — CatBoost alone | 0.521 | 0.201 |
| Reference repo — Stacked ensemble | 0.531 | 0.198 |
| Reference repo — Bookmaker baseline | 0.536 | 0.195 |
Honest caveat: No temporal feature leakage was found (see audit below), but aggregate test accuracy is misleading because the international test set contains many minnow-vs-top-team fixtures. On close matches, stacked accuracy (~0.43) is below the published benchmark (~0.56). Bookmaker odds remain the realistic ceiling (~53.6% on league data).
This repo does not include a live fixture API. There is no "what's playing this weekend" lookup — you supply the teams, tournament, and date yourself. For a separate live schedule feed you would need an external API (not implemented here).
# Predict a future friendly (defaults to today's date, auto-refreshes data first)
python predict.py "Brazil" "Argentina" --date 2026-06-25 --tournament "Friendly"
# Compare decision rules side by side (logs default config only)
python predict.py "Brazil" "Argentina" --decision-rule compare-all
# After matches finish, resolve open predictions against martj42 results
python reconcile.pyHow it works
predict.pycallsrefresh_latest()before every run — re-downloads martj42/international_results and incrementally updates cached pi-ratings/Elo for new rows only (seconds, not a full refit).- Features are built with
build_fixture_features()using each team's state strictly before--date(default: today). - A prominent staleness warning prints if the dataset's latest match is before your requested date — ratings won't reflect intervening fixtures until the source CSV updates.
- Every run appends one row to
predictions_log.csv(gitignored, local audit trail).
predictions_log.csv columns: timestamp_predicted, teams, match_date, tournament, p_home/p_draw/p_away, decision_rule_used, predicted_outcome, model_version, then empty actual_* fields until reconcile.py fills them and sets resolved=True.
reconcile.py refreshes data, finds unresolved rows where match_date < today, matches exact home/away/date in the dataset, fills actual scores/outcome, and prints rolling accuracy over the last 50 resolved predictions.
MatchBench is an MCP server that exposes the statistical prediction pipeline to AI agents (Claude Desktop, Cursor, etc.).
What the agent does vs what MatchBench does
| Agent (Claude, Cursor, …) | MatchBench (this MCP server) |
|---|---|
| Finds fixtures, dates, tournaments from web/search | No fixture list — you must pass home_team, away_team, date |
| Researches injuries, news, squad changes | No news — historical stats only |
| Synthesizes a final answer | Returns H/D/A probabilities, ratings, staleness metadata |
The intended workflow: agent researches the match → calls MatchBench for probabilities → combines both in its reply.
From the repo root:
git clone <your-repo-url> football-prediction-research
cd football-prediction-research
python3.11 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
# macOS only — if LightGBM/XGBoost fail to load:
# brew install libomp
python train.py # required once — creates models/predictor.joblib (~3–4 min)
chmod +x mcp_server/run_mcp.sh # launcher for Claude Desktop / CursorVerify the server starts (Ctrl+C to stop):
source venv/bin/activate
mcp dev mcp_server/server.py # opens MCP Inspector in browser — try refresh_data, predict_match- Do not use
mcp installalone — it spawns an isolated Python env withoutpandasor your trained models. - Edit
~/Library/Application Support/Claude/claude_desktop_config.json(create the file if missing):
{
"mcpServers": {
"MatchBench": {
"command": "/Users/YOU/projects/football-prediction-research/mcp_server/run_mcp.sh",
"args": []
}
}
}Replace /Users/YOU/projects/football-prediction-research with your clone path. If you already have other MCP servers, add the "MatchBench" block inside the existing "mcpServers" object — do not replace the whole file.
- Fully quit and reopen Claude Desktop (not just close the window).
- Start a new chat. You should see MatchBench connected with four tools:
predict_match,get_team_rating,refresh_data,get_prediction_history.
Alternative (same venv, no shell script):
"MatchBench": {
"command": "/Users/YOU/projects/football-prediction-research/venv/bin/python",
"args": ["/Users/YOU/projects/football-prediction-research/mcp_server/server.py"]
}Troubleshooting: If MatchBench shows "failed", open Claude → Settings → Developer → View logs. Common fixes: wrong path, forgot python train.py, or venv not activated when paths were copied.
- Open Cursor Settings → MCP (or edit
.cursor/mcp.jsonin your project / user config). - Add:
{
"mcpServers": {
"MatchBench": {
"command": "/Users/YOU/projects/football-prediction-research/mcp_server/run_mcp.sh",
"args": []
}
}
}- Restart Cursor or reload MCP servers from settings.
- In Agent or Composer mode, the model can call MatchBench tools when relevant.
Example: "What does the model think about Brazil vs Argentina on 2026-06-25, Copa America, neutral?"
- You ask in Claude Desktop or Cursor agent chat.
- Agent researches (optional but recommended): confirms fixture date, venue, any injury news from the web.
- Agent calls MatchBench, e.g.:
refresh_data— optional;predict_matchauto-refreshes anywaypredict_match(home_team="Brazil", away_team="Argentina", date="2026-06-25", tournament="Copa America", neutral=true)get_team_rating("Brazil")/get_team_rating("Argentina")— if comparing strength
- MatchBench returns JSON, e.g.:
{
"probabilities": { "home_win": 0.32, "draw": 0.27, "away_win": 0.42 },
"expected_score": "2-1",
"model_agreement": "low (models disagree)",
"pi_rating_diff": 0.1,
"elo_diff": -118.0,
"data_current_as_of": "2026-06-19",
"staleness_warning": "...",
"decision_rule_used": "default"
}- Agent interprets: cites probabilities, explains staleness if
requested date > data_current_as_of, weighs injury news against the model, gives a plain-language summary.
After the match: run python reconcile.py locally (CLI) to backfill results in predictions_log.csv. Then ask the agent: "Check my prediction history for Brazil — was the Argentina call right?" → it calls get_prediction_history(team_name="Brazil").
| Tool | When the agent should call it |
|---|---|
predict_match |
User asks for win/draw/loss probabilities for a specific fixture |
get_team_rating |
Compare team strength, pi-ratings, or Elo as of a date |
refresh_data |
User asks how fresh the dataset is; optional before other calls |
get_prediction_history |
Audit past predictions, resolved results, rolling accuracy |
All tools return structured JSON (not printed text). Team names must match martj42/international_results spelling (e.g. "Brazil", "Côte d'Ivoire").
This MCP server (MatchBench) is designed to be used by an AI agent (Claude Desktop, Cursor, or any MCP-compatible client) that combines its own research with MatchBench's statistical output. MatchBench only knows historical results — it has no knowledge of injuries, fixtures, or news. The agent is expected to fill that gap.
- "Predict the outcome of Brazil vs Argentina on [date], neutral venue, [tournament]"
- "Use MatchBench to forecast [Team A] vs [Team B] next Saturday"
- "What does the model give Germany vs Ivory Coast for tonight's match?"
- "Run a prediction for [Team A] vs [Team B] at [Team A]'s home ground"
- "Predict [Team A] vs [Team B] for [date], and check if there's any injury news that might change the call"
- "What does the model say about tomorrow's [Team A] vs [Team B] match, and does current squad news support or contradict it?"
- "Give me a prediction for [match], then tell me how confident I should actually be once you factor in recent form and injuries"
- "Predict this match, and flag anything in the news that could make the model wrong"
- "What's [Team A]'s current pi-rating and Elo? How does it compare to [Team B]?"
- "Which team has the stronger home record right now: [Team A] or [Team B]?"
- "Show me the rating gap between [Team A] and [Team B] and explain what's driving the model's lean"
- "Predict [match] and tell me if the underlying models actually agree with each other, or if this is a low-confidence call"
- "Is this prediction based on fresh data, or is something stale? Check before you answer"
- "How recent is the data behind this prediction? Tell me if anything important could have happened since the last refresh"
- "Look up my prediction history for [Team A] — how has the model done on their matches so far?"
- "Pull my recent prediction log and tell me the rolling accuracy"
- "Has the prediction for [match I asked about earlier] been resolved yet? What was the actual result?"
- "Predict [match] using the default decision rule, then show me how it changes with plain argmax instead — is the model assuming a draw is more or less likely with each approach?"
- "Run [match] through all available decision rules and compare them"
For decision-rule comparisons via CLI (outside MCP): python predict.py "Brazil" "Argentina" --decision-rule compare-all
Same pipeline, no MCP — useful for scripts and debugging:
python predict.py "Brazil" "Argentina" --date 2026-06-25 --tournament "Copa America" --neutral
python reconcile.pyClaude.ai custom connectors in the browser expect remote HTTP/SSE. MatchBench uses local stdio — use Claude Desktop or Cursor as above. Remote deployment would require a separate MCP HTTP gateway (not included in this repo).
- No bookmaker odds in v1 (reference repo shows ~53.6% acc ceiling with odds)
- International friendlies dominate some eras; tournament-specific performance varies
- Dixon-Coles MLE uses post-1990 matches for computational practicality
- Predictions are probabilistic estimates with substantial irreducible variance in football
International results: github.com/martj42/international_results (see upstream README for license).
- TimesFM — research shows no proven accuracy edge over CatBoost+pi-ratings for this data type; noted as an unexplored experimental direction, not implemented in v1
football-prediction-research/
├── docs/
│ └── matchbench-banner.png
├── train.py # End-to-end training
├── predict.py # CLI predictions (--date, --decision-rule, auto-refresh)
├── reconcile.py # Resolve predictions_log.csv against new results
├── mcp_server/
│ ├── server.py # MatchBench MCP server (stdio)
│ └── run_mcp.sh # Launcher for Claude Desktop / Cursor
├── scripts/
│ └── tune_draw_decision.py # Validation-only draw threshold / class-weight tuning
├── data/fetch.py # Download & cache CSV data
├── src/ratings/ # Pi-ratings, Elo
├── src/models/ # Dixon-Coles, GBT
├── src/features.py # 28 chronological features
├── src/draw_decision.py # Post-hoc H/D/A decision rules
├── src/ensemble.py # Stacked meta-learner
├── src/evaluate.py # RPS, calibration, metrics
├── tests/ # Including temporal leakage tests
└── notebooks/ # Backtest analysis
MIT — research and educational use.
