Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

MatchBench — open-source football match prediction research tool

MatchBench

Open, transparent research implementation of state-of-the-art statistical football match prediction methodology. NOT a betting tool — built for methodological transparency and learning.

This repository implements and benchmarks the academically-proven stack: pi-ratings + Dixon-Coles Poisson + gradient-boosted trees, combined via a calibrated stacked ensemble, with honest chronological backtesting.

Methodology

Component Source
Pi-ratings Constantinou & Fenton (2013) — Section 2–3; cross-checked against R reference
Dixon-Coles Michels et al. (2023) arXiv:2307.02139 Section 2.1; regista reference
GBT + stacking Survey: arXiv:2403.07669; Evaluation: arXiv:2309.14807
Benchmark structure jdgoated1/football-predictor (organization reference, not code copy)
Data martj42/international_results

Pi-ratings (Section 2)

  • Four venue-specific ratings per team: home_attack, home_defense, away_attack, away_defense
  • Expected goal difference uses the paper's exponential transform (Section 2.3)
  • Error dampening: ψ(e) = 3 × log₁₀(1 + |e|) (Section 2.2)
  • Learning rates λ (performance) and γ (cross-context) — tunable; paper EPL optimum ≈ λ=0.035, γ=0.7

Dixon-Coles

  • Attack/defense parameters with time decay w(t) = exp(−ξ × days)
  • Low-score τ correction for (0-0), (0-1), (1-0), (1-1) per arXiv:2307.02139

Stacked ensemble

  • Base models: CatBoost + LightGBM + XGBoost (with sklearn fallbacks when OpenMP libs unavailable)
  • Meta-learner: logistic regression + isotonic calibration on validation-fold predictions

Setup

python3.11 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

# macOS: if LightGBM/XGBoost fail to load, install OpenMP:
# brew install libomp

python train.py          # ~3–4 min on laptop CPU (trains weighted + baseline GBT)
python predict.py "Brazil" "Argentina" --tournament "Copa America" --neutral
python predict.py "Brazil" "Argentina" --decision-rule compare-all

Using MatchBench in Claude Desktop or Cursor? See MatchBench agent setup & usage below.

Results (international matches, chronological split)

Split: train pre-2018 · val 2018–2021 · test 2022–2024

Primary — close matches (|Elo diff| < 100, n=1,290)

Comparable difficulty to academic benchmarks. Do not use aggregate test accuracy — ~48% of test fixtures are heavy Elo mismatches.

Model Accuracy RPS
Dixon-Coles alone 0.387 0.245
GBT ensemble 0.411 0.225
Stacked ensemble (baseline argmax) 0.428 0.225

Default production config: stacked ensemble trained with draw_weight=1.5 and post-hoc decision rule draw_margin=0.075 (validation-tuned on close matches). See draw-recall section below.

Draw recall on close matches (test set, one-shot evaluation)

Tuning methodology: scripts/tune_draw_decision.py — validation-only search, single test evaluation, RPS-first selection.

Config Acc RPS Draw recall Home recall Away recall
Baseline argmax (unweighted) 0.428 0.225 3.4% 60.5% 54.9%
Margin only (margin=0.05, no retraining) 0.426 0.225 9.3% 56.6% 53.8%
Default (draw_weight=1.5 + margin=0.075) 0.428 0.224 18.3% 51.6% 52.7%

Draw recall improvements are achieved by shifting decision boundaries and modest retraining reweighting, not by improving underlying probability calibration — P(draw) was already well-calibrated at ~28% before any tuning.

Rejected: per-class threshold rule (th_H=0.36, th_D=0.18, th_A=0.44) — artificially inflates draw recall (~40%) while destroying away-win recall (~20% on test) by over-predicting draws (~35% vs ~28% actual). Not deployed.

Use python predict.py "Brazil" "Argentina" --decision-rule compare-all to see all three configs side by side for a single fixture.

All test matches (n=4,600 — inflated by easy mismatches)

Model Accuracy RPS
Dixon-Coles alone 0.558 0.192
GBT ensemble 0.598 0.170
Stacked ensemble 0.603 0.171

World Cup finals only (n=96): stacked 0.542 acc / 0.211 RPS — in line with published benchmarks.

External benchmarks (different datasets — sanity-check only)

Benchmark Accuracy RPS
CatBoost+pi-ratings, 2017 Soccer Prediction Challenge (arXiv:2309.14807) 0.558 0.193
Reference repo — Dixon-Coles alone (football-predictor) 0.490 0.212
Reference repo — CatBoost alone 0.521 0.201
Reference repo — Stacked ensemble 0.531 0.198
Reference repo — Bookmaker baseline 0.536 0.195

Honest caveat: No temporal feature leakage was found (see audit below), but aggregate test accuracy is misleading because the international test set contains many minnow-vs-top-team fixtures. On close matches, stacked accuracy (~0.43) is below the published benchmark (~0.56). Bookmaker odds remain the realistic ceiling (~53.6% on league data).

Live prediction & verification

This repo does not include a live fixture API. There is no "what's playing this weekend" lookup — you supply the teams, tournament, and date yourself. For a separate live schedule feed you would need an external API (not implemented here).

# Predict a future friendly (defaults to today's date, auto-refreshes data first)
python predict.py "Brazil" "Argentina" --date 2026-06-25 --tournament "Friendly"

# Compare decision rules side by side (logs default config only)
python predict.py "Brazil" "Argentina" --decision-rule compare-all

# After matches finish, resolve open predictions against martj42 results
python reconcile.py

How it works

  1. predict.py calls refresh_latest() before every run — re-downloads martj42/international_results and incrementally updates cached pi-ratings/Elo for new rows only (seconds, not a full refit).
  2. Features are built with build_fixture_features() using each team's state strictly before --date (default: today).
  3. A prominent staleness warning prints if the dataset's latest match is before your requested date — ratings won't reflect intervening fixtures until the source CSV updates.
  4. Every run appends one row to predictions_log.csv (gitignored, local audit trail).

predictions_log.csv columns: timestamp_predicted, teams, match_date, tournament, p_home/p_draw/p_away, decision_rule_used, predicted_outcome, model_version, then empty actual_* fields until reconcile.py fills them and sets resolved=True.

reconcile.py refreshes data, finds unresolved rows where match_date < today, matches exact home/away/date in the dataset, fills actual scores/outcome, and prints rolling accuracy over the last 50 resolved predictions.

MatchBench agent setup & usage

MatchBench is an MCP server that exposes the statistical prediction pipeline to AI agents (Claude Desktop, Cursor, etc.).

What the agent does vs what MatchBench does

Agent (Claude, Cursor, …) MatchBench (this MCP server)
Finds fixtures, dates, tournaments from web/search No fixture list — you must pass home_team, away_team, date
Researches injuries, news, squad changes No news — historical stats only
Synthesizes a final answer Returns H/D/A probabilities, ratings, staleness metadata

The intended workflow: agent researches the match → calls MatchBench for probabilities → combines both in its reply.


One-time setup

From the repo root:

git clone <your-repo-url> football-prediction-research
cd football-prediction-research

python3.11 -m venv venv
source venv/bin/activate          # Windows: venv\Scripts\activate
pip install -r requirements.txt

# macOS only — if LightGBM/XGBoost fail to load:
# brew install libomp

python train.py                   # required once — creates models/predictor.joblib (~3–4 min)
chmod +x mcp_server/run_mcp.sh    # launcher for Claude Desktop / Cursor

Verify the server starts (Ctrl+C to stop):

source venv/bin/activate
mcp dev mcp_server/server.py      # opens MCP Inspector in browser — try refresh_data, predict_match

Connect to Claude Desktop

  1. Do not use mcp install alone — it spawns an isolated Python env without pandas or your trained models.
  2. Edit ~/Library/Application Support/Claude/claude_desktop_config.json (create the file if missing):
{
  "mcpServers": {
    "MatchBench": {
      "command": "/Users/YOU/projects/football-prediction-research/mcp_server/run_mcp.sh",
      "args": []
    }
  }
}

Replace /Users/YOU/projects/football-prediction-research with your clone path. If you already have other MCP servers, add the "MatchBench" block inside the existing "mcpServers" object — do not replace the whole file.

  1. Fully quit and reopen Claude Desktop (not just close the window).
  2. Start a new chat. You should see MatchBench connected with four tools: predict_match, get_team_rating, refresh_data, get_prediction_history.

Alternative (same venv, no shell script):

"MatchBench": {
  "command": "/Users/YOU/projects/football-prediction-research/venv/bin/python",
  "args": ["/Users/YOU/projects/football-prediction-research/mcp_server/server.py"]
}

Troubleshooting: If MatchBench shows "failed", open Claude → Settings → Developer → View logs. Common fixes: wrong path, forgot python train.py, or venv not activated when paths were copied.


Connect to Cursor

  1. Open Cursor Settings → MCP (or edit .cursor/mcp.json in your project / user config).
  2. Add:
{
  "mcpServers": {
    "MatchBench": {
      "command": "/Users/YOU/projects/football-prediction-research/mcp_server/run_mcp.sh",
      "args": []
    }
  }
}
  1. Restart Cursor or reload MCP servers from settings.
  2. In Agent or Composer mode, the model can call MatchBench tools when relevant.

How to use it in an agent (step by step)

Example: "What does the model think about Brazil vs Argentina on 2026-06-25, Copa America, neutral?"

  1. You ask in Claude Desktop or Cursor agent chat.
  2. Agent researches (optional but recommended): confirms fixture date, venue, any injury news from the web.
  3. Agent calls MatchBench, e.g.:
    • refresh_data — optional; predict_match auto-refreshes anyway
    • predict_match(home_team="Brazil", away_team="Argentina", date="2026-06-25", tournament="Copa America", neutral=true)
    • get_team_rating("Brazil") / get_team_rating("Argentina") — if comparing strength
  4. MatchBench returns JSON, e.g.:
{
  "probabilities": { "home_win": 0.32, "draw": 0.27, "away_win": 0.42 },
  "expected_score": "2-1",
  "model_agreement": "low (models disagree)",
  "pi_rating_diff": 0.1,
  "elo_diff": -118.0,
  "data_current_as_of": "2026-06-19",
  "staleness_warning": "...",
  "decision_rule_used": "default"
}
  1. Agent interprets: cites probabilities, explains staleness if requested date > data_current_as_of, weighs injury news against the model, gives a plain-language summary.

After the match: run python reconcile.py locally (CLI) to backfill results in predictions_log.csv. Then ask the agent: "Check my prediction history for Brazil — was the Argentina call right?" → it calls get_prediction_history(team_name="Brazil").


MCP tools reference

Tool When the agent should call it
predict_match User asks for win/draw/loss probabilities for a specific fixture
get_team_rating Compare team strength, pi-ratings, or Elo as of a date
refresh_data User asks how fresh the dataset is; optional before other calls
get_prediction_history Audit past predictions, resolved results, rolling accuracy

All tools return structured JSON (not printed text). Team names must match martj42/international_results spelling (e.g. "Brazil", "Côte d'Ivoire").


Example agent prompts

This MCP server (MatchBench) is designed to be used by an AI agent (Claude Desktop, Cursor, or any MCP-compatible client) that combines its own research with MatchBench's statistical output. MatchBench only knows historical results — it has no knowledge of injuries, fixtures, or news. The agent is expected to fill that gap.

Basic match prediction

  • "Predict the outcome of Brazil vs Argentina on [date], neutral venue, [tournament]"
  • "Use MatchBench to forecast [Team A] vs [Team B] next Saturday"
  • "What does the model give Germany vs Ivory Coast for tonight's match?"
  • "Run a prediction for [Team A] vs [Team B] at [Team A]'s home ground"

Prediction + live research combined (the intended primary use case)

  • "Predict [Team A] vs [Team B] for [date], and check if there's any injury news that might change the call"
  • "What does the model say about tomorrow's [Team A] vs [Team B] match, and does current squad news support or contradict it?"
  • "Give me a prediction for [match], then tell me how confident I should actually be once you factor in recent form and injuries"
  • "Predict this match, and flag anything in the news that could make the model wrong"

Comparing teams / ratings

  • "What's [Team A]'s current pi-rating and Elo? How does it compare to [Team B]?"
  • "Which team has the stronger home record right now: [Team A] or [Team B]?"
  • "Show me the rating gap between [Team A] and [Team B] and explain what's driving the model's lean"

Checking the model's own confidence

  • "Predict [match] and tell me if the underlying models actually agree with each other, or if this is a low-confidence call"
  • "Is this prediction based on fresh data, or is something stale? Check before you answer"
  • "How recent is the data behind this prediction? Tell me if anything important could have happened since the last refresh"

Tracking accuracy over time

  • "Look up my prediction history for [Team A] — how has the model done on their matches so far?"
  • "Pull my recent prediction log and tell me the rolling accuracy"
  • "Has the prediction for [match I asked about earlier] been resolved yet? What was the actual result?"

Decision rule comparisons (for more advanced users)

  • "Predict [match] using the default decision rule, then show me how it changes with plain argmax instead — is the model assuming a draw is more or less likely with each approach?"
  • "Run [match] through all available decision rules and compare them"

For decision-rule comparisons via CLI (outside MCP): python predict.py "Brazil" "Argentina" --decision-rule compare-all

CLI without an agent

Same pipeline, no MCP — useful for scripts and debugging:

python predict.py "Brazil" "Argentina" --date 2026-06-25 --tournament "Copa America" --neutral
python reconcile.py

Claude.ai web note

Claude.ai custom connectors in the browser expect remote HTTP/SSE. MatchBench uses local stdio — use Claude Desktop or Cursor as above. Remote deployment would require a separate MCP HTTP gateway (not included in this repo).

Limitations

  • No bookmaker odds in v1 (reference repo shows ~53.6% acc ceiling with odds)
  • International friendlies dominate some eras; tournament-specific performance varies
  • Dixon-Coles MLE uses post-1990 matches for computational practicality
  • Predictions are probabilistic estimates with substantial irreducible variance in football

Data credit

International results: github.com/martj42/international_results (see upstream README for license).

Future work

  • TimesFM — research shows no proven accuracy edge over CatBoost+pi-ratings for this data type; noted as an unexplored experimental direction, not implemented in v1

Project layout

football-prediction-research/
├── docs/
│   └── matchbench-banner.png
├── train.py              # End-to-end training
├── predict.py            # CLI predictions (--date, --decision-rule, auto-refresh)
├── reconcile.py          # Resolve predictions_log.csv against new results
├── mcp_server/
│   ├── server.py         # MatchBench MCP server (stdio)
│   └── run_mcp.sh        # Launcher for Claude Desktop / Cursor
├── scripts/
│   └── tune_draw_decision.py  # Validation-only draw threshold / class-weight tuning
├── data/fetch.py         # Download & cache CSV data
├── src/ratings/          # Pi-ratings, Elo
├── src/models/           # Dixon-Coles, GBT
├── src/features.py       # 28 chronological features
├── src/draw_decision.py  # Post-hoc H/D/A decision rules
├── src/ensemble.py       # Stacked meta-learner
├── src/evaluate.py       # RPS, calibration, metrics
├── tests/                # Including temporal leakage tests
└── notebooks/            # Backtest analysis

License

MIT — research and educational use.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages