Simulating all 104 matches of the first 48-team World Cup 10,000 times — with a Monte Carlo engine built on real qualification data from all six confederations, FIFA Elo ratings, and a composite strength model that weights both pedigree and recent form.
Every number in this README is an output of the model, not a hand-picked prediction. Run the code and you get these results.
⚙️ All data ingestion and aggregation run on PardoX, a Rust-backed DataFrame engine I built — not pandas.
📝 The story behind this project, in three parts:
- I Built a World Cup 2026 Predictor From Scratch — Here's What the Data Actually Says — the methodology
- I Tried to Predict the World Cup — What I Learned Had Nothing to Do With Football — the engineering lessons
- Thank You, Football — What the 2026 World Cup Taught Me, and Why I'm Already Working on 2030 — the reflection
Every World Cup cycle produces the same prediction content, and most of it looks backward — treating 2026 as an extension of Qatar 2022, ranking teams by what they did four years ago.
I wanted to start from a more defensible premise: the process that gets a team into the World Cup matters just as much as what happened in the previous one. Team strength should be modeled from a combination of long-term quality, recent qualification form, host advantage, and full tournament simulation — not nostalgia.
There's also a genuine engineering challenge underneath. The 2026 World Cup is the biggest in history: 48 teams, 12 groups, 104 matches, three host nations. That scale makes prediction harder, not easier — more matches, more paths, more variance, plus FIFA's real group-eligibility rules for third-place qualifiers, a penalty-shootout model, and a host-nation correction. The football is the fun part; the interesting engineering is underneath.
So I built a transparent, reproducible model where every assumption is written down, every probability traces back to a formula, and anyone can change an input and watch the tournament re-simulate.
Three Python scripts form the complete pipeline:
| Script | Purpose |
|---|---|
tournament_simulator_2026.py |
Monte Carlo engine — 10,000 full-tournament simulations, championship probability map |
match_predictor_2026.py |
Deterministic predictor — most-likely scoreline and winner for all 104 matches, full knockout bracket |
mexico_analysis_2026.py |
Mexico deep-dive — corrects host-nation form with official A-match data, re-runs Monte Carlo, and derives Mexico's live bracket path |
Every prediction starts from one number per team:
composite_strength = FIFA_Elo + 0.35 × (form_PPG − 1.5) × 300
- FIFA Elo (65%) — official ratings as of January 2026. Captures long-term pedigree and consistency.
- Qualifier form (35%) — recency-weighted points-per-game over each team's last 10 qualification matches. The most recent match carries weight 10; the oldest, weight 1.
- Host bonus — USA, Mexico, and Canada receive +80 Elo in every match they play.
A team in perfect form (3.0 PPG) gains roughly +157 Elo-equivalent points over a team at neutral form (1.5 PPG).
diff = strength_t1 − strength_t2 (+80 if host nation)
E1 = 1 / (1 + 10^(−diff / 800))
p_draw = 0.27 × exp(−(diff / 500)²) [capped 5%–33%]
p_win_t1 = E1 × (1 − p_draw)
p_win_t2 = (1 − E1) × (1 − p_draw)
The divisor of 800 — rather than the more aggressive 400 common in Elo systems — deliberately reflects football's high inherent variance. A larger divisor flattens the favorite's edge, which matches reality: upsets happen far more often in football than in chess or tennis. The three outcome probabilities always sum to exactly 1.0.
- Deterministic predictor — draw probability is redistributed proportionally between the two teams' win shares, producing a guaranteed winner and a complete projected bracket.
- Monte Carlo simulator — draws are simulated explicitly. If a knockout match ends level, a separate penalty-shootout model resolves it, with the stronger side holding at most a 7.5% edge above the 50/50 baseline (penalties are close to a coin flip, and the model respects that).
In the deterministic layer, standings come from expected values across the full outcome distribution:
expected_points(team) = 3 × P(win) + 1 × P(draw)
expected_GD(team) = Σ [xG_for − xG_against] across all group matches
Ranking order: (1) expected points → (2) expected goal difference → (3) expected goals scored → (4) raw Elo.
The 8 best third-place finishers are selected by the same criteria across all 12 groups. Assigning them to the correct Round-of-32 slots is genuinely hard — FIFA's group-eligibility rules constrain which third-place team can go where — so a backtracking algorithm solves the assignment.
For the deterministic predictor, expected goals are derived from the strength differential, and a Poisson model picks the most likely scoreline:
predicted_score = argmax P(X = g1) × P(Y = g2) over g1, g2 ∈ {0..5}
The Monte Carlo simulator skips per-match Poisson scorelines — it resolves outcomes through probability-weighted random draws, which is faster and sufficient for aggregated tournament probabilities.
These figures are generated by running the pipeline against the compiled CSV data. They emerge from the model — nothing here is hardcoded.
France vs England — MetLife Stadium, East Rutherford — July 19, 2026
England 50.3% | France 49.7%
| # | Team | P(Champion) | P(Final) |
|---|---|---|---|
| 1 | England | 10.05% | 16.18% |
| 2 | France | 9.45% | 15.89% |
| 3 | Spain | 9.19% | 15.58% |
| 4 | Argentina | 7.08% | 12.27% |
| 5 | Morocco | 6.01% | 11.13% |
| 6 | Germany | 5.62% | 10.27% |
England's composite strength tops the field — elite base Elo plus perfect recent qualifying form. Note how flat the top of the distribution is: even the favorite wins only ~1 in 10 simulations, which is exactly what you'd expect from a 48-team tournament with football's variance. Any model claiming a 30%+ favorite would be overfitting.
Mexico qualified automatically as co-host, so the base simulator has no qualifier data for them. By default, any team without qualifier history gets a fallback form score of 1.20 PPG — a reasonable placeholder, but one that clearly understates a host nation that's been playing competitive football.
mexico_analysis_2026.py loads Mexico's 54 official A-matches from 2023–2026 and applies the same recency-weighted form formula used for every other team.
| Parameter | Base simulator | With real match data |
|---|---|---|
| Form PPG (last 10) | 1.200 | 1.636 |
| Form Elo adjustment | −31.5 | +14.3 |
| Composite strength | 1644.2 | 1690.1 |
| Strength gain | — | +45.8 Elo pts |
After the correction, the full 10,000-run Monte Carlo is re-executed, lifting Mexico's championship probability to 3.1%.
| Round | Opponent | Venue | P(Mexico) | Cumulative |
|---|---|---|---|---|
| Round of 32 | Scotland | Estadio Azteca, Mexico City | 63.4% | 63.4% |
| Round of 16 | England | Estadio Azteca, Mexico City | 34.6% | 21.9% |
The model sees Mexico as a legitimate Round-of-16 side on home soil: a solid favorite to clear the first knockout, then a much steeper climb against England.
The 2026 format quietly changes the meaning of the most repeated phrase in Mexican football.
Under the old 32-team bracket, the famous quinto partido (the "fifth match") effectively meant reaching the quarterfinals — the barrier Mexico has failed to cross time and again. The new 48-team format adds a round:
3 group matches → Round of 32 → Round of 16 → Quarterfinals → Semifinals → Final
So there's now a subtle but real distinction:
- Literal fifth match = the Round of 16 (win the group, beat Scotland, and the England match is game five)
- Historical equivalent of the old breakthrough = the quarterfinals, now the sixth match
The model gives Mexico a credible path to a literal fifth match. But if Mexico clears the England test, it's no longer just reaching a fifth game — it's doing the harder thing Mexican football has actually been chasing all along: reaching the quarterfinals, a sixth match. And that is exactly why we watch.
# Full Monte Carlo simulation (10,000 runs) + championship probability map
python tournament_simulator_2026.py
# Deterministic predictor — all 104 matches, full projected bracket
python match_predictor_2026.py
# Mexico deep-dive with corrected form and live bracket path
python mexico_analysis_2026.pyRequirements: Python 3.9+, PardoX (Rust-backed DataFrame engine), NumPy.
Reproducibility: fixed random seed — run it twice, get the same numbers. Every probability, standing, and bracket outcome shown here is a pipeline output, not a manually inserted assumption.
| File | Description |
|---|---|
csv/fifa_ranking_jan2026.csv |
FIFA Elo rankings — January 2026 |
csv/qualifiers_{afc,caf,concacaf,conmebol,ofc,uefa}_2026.csv |
Qualification results, all six confederations |
csv/intercontinental_playoff_2026.csv |
Intercontinental playoff results |
csv/mexico_official_matches_2023_2026.csv |
Mexico official A-matches, 2023–2026 |
csv/worldcup_2026_schedule_full.csv |
Full 104-match schedule |
csv/worldcup_2026_knockout_bracket.csv |
Knockout bracket with slot assignments |
| File | Description |
|---|---|
group_stage_predictions.csv |
All 72 group matches — probabilities and predicted scores |
group_standings_predicted.csv |
Predicted final standings, all 12 groups |
knockout_bracket_predictions.csv |
Full projected knockout bracket |
all_matches_predictions.csv |
Unified 104-match output |
championship_probabilities.csv |
Championship and finalist probabilities, all 48 teams |
mexico_path.csv |
Mexico's projected live bracket path |
Detailed markdown reports are also generated: analisis.md (tournament-wide summary), predictions_detailed.md (all 104 matches), and analisis_mexico.md (Mexico deep-dive).
- 48 teams, 12 groups (A–L), 4 teams each
- 104 matches — 72 group stage, 32 knockout
- Co-hosts: USA, Mexico, Canada
- Dates: June 11 – July 19, 2026
- Monte Carlo uses a fixed random seed for reproducibility.
- Composite strength weights FIFA Elo at 65%, qualifier form at 35%.
- Penalty shootouts carry only a small strength-based edge over a 50/50 baseline.
- Third-place slot assignment uses backtracking to satisfy FIFA's eligibility rules.
- All data ingestion and aggregation run on PardoX, a Rust-backed DataFrame engine, rather than pandas.
This started as a football question and turned into something else — a lesson in how much modeling choices, data quality, and honest assumptions shape a result. I wrote about what the project actually taught me (spoiler: it had little to do with football), and why I'm already working on the 2030 edition.
You're welcome to use, reference, share, or build on this study for educational, analytical, or editorial purposes. If you reuse the findings, tables, or methodology, a credit and a link back to this repository are greatly appreciated.
This is a technical and methodological exercise from a data-engineering perspective — not a definitive prediction. The results rest on statistical assumptions, probabilistic models, historical data, and simulation criteria that, by design, simplify a far more complex reality. Football remains subject to variables no model fully captures. Read this as a serious analytical approximation, not a forecast set in stone.