Publishing predictions has one job: making selective memory impossible.
Football prediction models are often remembered through a handful of screenshots. The winning calls travel. The losing probabilities disappear into old feeds and overwritten files.
A serious forecast needs a public record that preserves the probability available before each match and every later update.
Four of the first five teams on the signed title board reached the semifinals. Across 104 matches, the highest-probability 90-minute outcome landed 62 times. The same model backed the wrong Golden Boot winner, underestimated the United States and moved the eventual champion down its title ranking.
Every forecast was published before the relevant cutoff and preserved in a signed GitHub history. The record shows which forecast types travelled well and where updating failed.
What a public record prevents
A permanent forecast history closes three easy escape routes:
- Failed calls disappearing after the tournament.
- Published probabilities changing once the result is known.
- A later accurate update replacing an earlier miss in the final account.
Accuracy still has to be earned. The record lets readers verify it for themselves.
How much does four of five prove?
France headed the first signed board at 12.02%. Spain and Argentina were tied on 9.78%, then came Germany at 7.61% and England at 7.52%. Those percentages were modest for a reason. Forty-eight teams and an extra knockout round made certainty scarce.
View chart data as a table
| Rank | Team | Title chance | Actual finish |
|---|---|---|---|
| 1 | France | 12.02% | Fourth place |
| 2 | Spain | 9.78% | Champion |
| 3 | Argentina | 9.78% | Runner-up |
| 4 | Germany | 7.61% | Round of 32 |
| 5 | England | 7.52% | Third place |
Their five marginal semifinal probabilities added up to 1.30. Four arrived. The gap between expectation and outcome is the strongest statistical measure the available data can support.
This success also exposes a flaw in the forecasting process. The path-level simulation draws were not retained, so the archive cannot answer how often all four eventual semifinalists appeared together. Re-running the tournament with a later model would contaminate the test. An output table can keep the ranking while losing the evidence needed to evaluate a joint event.
Spain won the final 1-0 after extra time, with Argentina second, England third and France fourth. Germany were the exception, eliminated by Paraguay on penalties in the Round of 32. The ranking held up deep into the bracket, even though France were first and the eventual champion began second in this signed version.
The snapshot distinction matters. The headline uses the signed 5 June repository snapshot. The 30 May editorial preview used an earlier run: Spain led France, and the percentages were higher. The five names stayed the same while their order and probabilities changed. The USA route article later displayed a separate pre-tournament baseline rounded to 20%; the earlier preview figure was 15.54%. Neither number is substituted into the 104-match audit.
The same files carry a harder case. Spain won the tournament while six weeks of public updates moved their title chance lower, and the match-level record below is what stops that miss from being smoothed over later.
How can a model miss 42 matches and still be good?
One tournament is a limited test of model quality. It can still show whether confidence rises with accuracy and whether the published probabilities punish confident errors.
The conventional test is 90-minute 1X2. Under that rule, the highest-probability outcome was correct in 62 of 104 matches, or 59.6%. The group stage produced 43/72. The knockouts produced 19/32 because five eventual winners needed extra time after a draw at 90.
Many readers remember those games by the team that advanced. On that looser reading, carrying the score through extra time while leaving shootouts as draws, the total becomes 67/104 and the knockout record becomes 24/32. For probabilities labelled home win, draw and away win, the formal grade remains the 90-minute figure.
View chart data as a table
| Phase | 90-minute hits | 90-minute rate | Through ET hits | Through ET rate |
|---|---|---|---|---|
| All matches | 62/104 | 59.6% | 67/104 | 64.4% |
| Group stage | 43/72 | 59.7% | 43/72 | 59.7% |
| Knockout | 19/32 | 59.4% | 24/32 | 75.0% |
Accuracy alone rewards caution and ignores how much probability sat behind the wrong alternatives. The strict scorecard's multiclass Brier score was 0.533, compared with 0.667 for assigning one third to every outcome. Log loss was 0.903. Lower is better for both measures.
| Forecast | Top-outcome accuracy | Brier | Log loss |
|---|---|---|---|
| Equal thirds | No unique top outcome | 0.667 | 1.099 |
| uanalyse | 59.6% | 0.533 | 0.903 |
Equal thirds is a basic sanity check and a weak football benchmark. No rating-only baseline was frozen before kickoff. Creating one after seeing the results would make the comparison retrospective.
Confidence stayed restrained, with no forecast reaching 80%. Every band is shown here: below 45% landed 11/28, 45 to 55% landed 16/31, 55 to 65% landed 20/24, and 65 to 80% landed 15/21. Their realised rates were 39.3%, 51.6%, 83.3% and 71.4%.
View chart data as a table
| Band | Matches | Average confidence | Correct | Realised rate |
|---|---|---|---|---|
| Below 45% | 28 | 40.8% | 11 | 39.3% |
| 45 to 55% | 31 | 50.6% | 16 | 51.6% |
| 55 to 65% | 24 | 60.1% | 20 | 83.3% |
| 65 to 80% | 21 | 69.5% | 15 | 71.4% |
Search the row-level record below or download the frozen files. Home, draw and away probabilities appear in fixture order. The observed probability is the probability assigned to the actual 90-minute result.
| Match | Snapshot | Home | Draw | Away | Highest | 90-minute result | Correct? | Final | Observed p |
|---|---|---|---|---|---|---|---|---|---|
| South Korea vs Czechia2026-06-12 · Group stage | 2026-06-11 | 42.7% | 26.2% | 31.2% | South Korea (42.7%) | 2-1 | yes | 2-1 | 42.7% |
| Mexico vs South Africa2026-06-11 · Group stage | 2026-06-10 | 65.2% | 20.6% | 14.2% | Mexico (65.2%) | 2-0 | yes | 2-0 | 65.2% |
| Canada vs Bosnia-Herzegovina2026-06-12 · Group stage | 2026-06-11 | 63.6% | 21.7% | 14.8% | Canada (63.6%) | 1-1 | no | 1-1 | 21.7% |
| United States vs Paraguay2026-06-13 · Group stage | 2026-06-12 | 41.4% | 27.7% | 30.9% | United States (41.4%) | 4-1 | yes | 4-1 | 41.4% |
| Haiti vs Scotland2026-06-14 · Group stage | 2026-06-13 | 29.7% | 25.5% | 44.9% | Scotland (44.9%) | 0-1 | yes | 0-1 | 44.9% |
| Brazil vs Morocco2026-06-13 · Group stage | 2026-06-12 | 47.1% | 26.8% | 26.1% | Brazil (47.1%) | 1-1 | no | 1-1 | 26.8% |
| Qatar vs Switzerland2026-06-13 · Group stage | 2026-06-12 | 10.6% | 19.2% | 70.2% | Switzerland (70.2%) | 1-1 | no | 1-1 | 19.2% |
| Australia vs Türkiye2026-06-14 · Group stage | 2026-06-13 | 23.8% | 27.2% | 49.0% | Türkiye (49.0%) | 2-0 | no | 2-0 | 23.8% |
Lessons from forecasts that held up
The grade needed its probability
The ledger marks Mexico against Ecuador, England against Norway and Spain in the final as successful calls. Their probabilities tell three different stories. Mexico's 51.5% advancement chance and Spain's 50.56% trophy chance were almost coin flips. England's 65.7% chance was a firmer position. A result-only list makes those forecasts look equal.
The same caution applies to the correct pre-tournament calls. Lionel Messi had a 53.4% chance to pass Miroslav Klose's career scoring record, while Mexico were 58.2% favourites against a Czechia side that needed a result. Both happened. Each outcome supports the process only as strongly as the probability committed beforehand.
Conditional forecasts handled tournament structure well
Algeria's draw scenario said one point against Austria would lift qualification to 99.9%. The match finished 3-3 and Algeria advanced. Norway's rotation left them with Ivory Coast instead of Sweden, a branch the simulation considered only marginally worse. Norway beat Ivory Coast and then Brazil. Those forecasts mapped the consequences of a decision before either route was known, which made them useful even without a single unconditional call.
What failed, and why
Spain exposed a failure of model observability
Across six weeks of public updates, the eventual champion moved down the title board. The 30 May editorial run put Spain at 13.65%; the first signed repository board had cut that to 9.78%, and the post-group update stood at 9.63%. Spain then conceded once in eight matches and held Argentina scoreless for 120 minutes in the final, as FIFA's match report records.
That is uncomfortable to publish. Hiding it would make the audit meaningless. The public files show the movement and contain no input-level decomposition, so they cannot reveal whether route changes, team-strength inputs or another component caused the downgrade. The files documented the miss and omitted enough of its causes to prevent a diagnosis.
Player awards need a separate standard
The Golden Boot board missed twice on the same winner. The preview made Messi the favourite at 16.3% and assigned Kylian Mbappé 12.53%. With two games left, the update widened the gap to 61.1% against 36.4%. Mbappé scored a hat trick in the third-place match and finished with 10 goals, as FIFA's award list records.
A 36.4% outcome was always live. The miss mattered because the board ranked Messi ahead of Mbappé in two separate forecasts. Player awards depend on minutes, penalties, teammates and a small collection of shots. Their calibration deserves its own sample and should not borrow authority from a team-ranking forecast that happened to perform well.
A later correction cannot erase an earlier miss
The United States began as the least likely Group D winner in the 30 May preview, at 15.54%. They won the group. The later route article recovered: Bosnia was the likeliest first knockout opponent, and the USA beat them 2-0 before Belgium ended the run 4-1. Continuous updating is useful only when every version survives. Otherwise the accurate route forecast can quietly replace the failed group forecast in the story.
Extra time exposed the limit of 1X2
Five eventual knockout winners were level after 90 minutes: Belgium against Senegal, Argentina against Cape Verde, England against Norway, Argentina against Switzerland, and Spain in the final. That is why the strict knockout record is 19/32 while the through-extra-time reading reaches 24/32. The tournament recap remembers the advancing teams; a 1X2 audit has to remember the draws. Regulation and advancement probabilities answer different questions and should be published together.
The binary grade stays visible, though the probability assigned to what happened carries more information. Conditional routes stay conditional. The tournament shot map and the 1-1 trap were retrospective, so neither enters the score. The first six calls appear below; the full ledger can be expanded.
| Article | Grade | Published call | What happened | Observed p | Log score |
|---|---|---|---|---|---|
| Before kickoff2026-05-30Direct forecast | landed | Spain, France, Argentina, Germany and England formed the 30 May title top five. | Four reached the semifinals; Germany went out in the Round of 32. | Not published for this joint event | Not scored |
| Golden Boot preview2026-06-05Direct forecast | missed | Messi led at 16.3%; Mbappé was second at 12.5%. | Mbappé won the Golden Boot with 10 goals. | 12.5% | 2.077 |
| Messi record call2026-06-05Direct forecast | landed | Messi had a 53.4% chance to pass Klose's career World Cup record. | He passed it and finished on 21 career World Cup goals. | 53.4% | 0.627 |
| Czechia vs Mexico2026-06-19Direct forecast | landed | Mexico were 58.2% to win; Czechia had a 20.3% qualification chance. | Mexico won 3-0 and Czechia were eliminated. | 58.2% | 0.541 |
| USA route2026-06-21Direct forecast | landed | The USA were 76.5% to beat Bosnia after Belgium became their likeliest last-16 opponent. | The USA beat Bosnia 2-0, then lost 4-1 to Belgium in the next round. | 76.5% | 0.268 |
| Messi versus Fontaine2026-06-23Direct forecast | landed | Messi had only a 4.4% chance to reach Fontaine's 13-goal tournament record. | He finished the tournament with eight goals. | 95.6% | 0.045 |
Spain, France, Argentina, Germany and England formed the 30 May title top five.
Four reached the semifinals; Germany went out in the Round of 32.
Messi led at 16.3%; Mbappé was second at 12.5%.
Mbappé won the Golden Boot with 10 goals.
Messi had a 53.4% chance to pass Klose's career World Cup record.
He passed it and finished on 21 career World Cup goals.
Mexico were 58.2% to win; Czechia had a 20.3% qualification chance.
Mexico won 3-0 and Czechia were eliminated.
The USA were 76.5% to beat Bosnia after Belgium became their likeliest last-16 opponent.
The USA beat Bosnia 2-0, then lost 4-1 to Belgium in the next round.
Messi had only a 4.4% chance to reach Fontaine's 13-goal tournament record.
He finished the tournament with eight goals.
Each match uses the newest signed snapshot dated no later than the previous UTC calendar day. The strict score covers the first and second half; extra-time and shootout matches are draws at 90. The undated third-place and final rows are matched from the 15 July snapshot by stage and teams. The record covers 48 initial teams, 104 matches and 33 forecast dates.
One tournament is a small test. Top-outcome accuracy discards the distance between a 39% lean and a 70% favourite, while the through-extra-time view applies 90-minute probabilities outside their formal target. The original simulation paths and a pre-registered rating-only baseline were not preserved. Conditional scenarios remain separate from direct calls, and retrospective pieces remain unscored.
What changes before the next tournament
Three requirements follow from this record. Each one addresses evidence that was missing when a result needed explaining.
Publish 1X2 and advancement together
Every knockout forecast should carry two clearly labelled views: the result after 90 minutes and the team that advances. Five extra-time winners showed how easily readers can remember a correct route while the formal 1X2 call was a draw.
Freeze a rating-only baseline before kickoff
Equal thirds tests whether a probability model clears the lowest possible bar. A signed rating-only forecast using the same fixtures and cutoffs would show whether the full process added value beyond team strength and venue.
Show calibration while the tournament is running
Confidence bands should update after each match with their exact counts visible. A live chart would expose drift during the tournament, when forecasts can still be challenged.
The minimum standard
Every forecast in this article can be traced to a public snapshot published before kickoff. Every miss remains visible, and the sequence of updates survives.
Football forecasting should meet that minimum standard. Readers should treat unverifiable records with caution.