Appunti

Forecasting

VARmageddon

I did not want to predict 104 World Cup scorelines by hand. Is it Brazil 2–0 or 2–1, does Japan nick a draw — it’s the same joyless little exercise, 104 times over. So, naturally, I spent far longer building a machine to do it for me.

Mon Petit Prono is not a bookmaker — you stake nothing and win nothing but pride and a place on a leaderboard — but it keeps score with a bookmaker’s instinct. Each match comes with three cotes: point values for a home win, a draw, and an away win. They are inverse to probability, so a Germany win pays little and a Curaçao win pays a fortune. Call the result right and you bank that result’s cote. Call the exact score right and you get a rarity bonus on top — larger the fewer people who also nailed it. You get one “double points” joker to spend, once, on a single match all tournament.

One rule makes the whole thing interesting: the cotes are frozen the day the market opens, and they never move again. The group-stage book is printed well before a ball is kicked, and that is the book you play against for a month.

Two markets, one match

A frozen book is a standing offer. Meanwhile the world keeps learning things — an injury in training, a manager’s press conference, a back line that suddenly can’t defend a corner. Real prediction markets price all of it in continuously, and because people put money down, they price it honestly. Polymarket and Kalshi will both sell you a “Germany to win” contract whose price is the crowd’s live probability, updated to the minute.

So for every match I have two numbers that disagree: a frozen payout (the cote, printed a fortnight ago) and a live probability (the market, as of five minutes ago). The entire game is the distance between them. Betting is just multiplying one by the other.

Polymarket Kalshi live probability goal model full scoreline grid frozen cote book · never moves E[points] P·cote + P·bonus → pick the score MPP
The market sets how likely each score is; the frozen book sets how much it's worth. The engine multiplies the two, keeps the best score, and posts that pick to Mon Petit Prono.

The engine reads the live win/draw/loss prices from both markets — a “Germany to win” contract at 0.57 simply is a 57% probability — and fits a goal model to them: from those three numbers it infers a scoring rate for each side, then turns those rates into a probability for every scoreline, 1–0, 2–1, 0–0 and on up, a whole grid (more on how, further down). Then, for each candidate score, it computes the expected Mon Petit Prono points — the chance of getting the result right times that result’s cote, plus the chance of the exact score times its rarity bonus — and picks the score that maximises it.

And none of this is done by hand. It runs on a timer all tournament: it re-reads Polymarket every five minutes and Kalshi every thirty, re-picks as the prices move, and posts the current best score to Mon Petit Prono itself, no click from me. Given why I built it, automating the last click was not optional.

The one thing it never does is pick the most likely score; it picks the most valuable one.

What value looks like

Ecuador played Germany in the group stage. The market was in no doubt: Germany 57% to win, Ecuador 22%. If you were predicting who’d win, you’d say Germany and move on.

But look at the frozen book. Ecuador’s win was worth 145 points; Germany’s, 42. Multiply each result by its live probability and the ranking flips: a 22% shot at 145 is 32 expected points, a 57% shot at 42 is 24.

HOW LIKELY — market Ecuador draw Germany 22% 21% 57% WHAT IT PAYS — E[points] Ecuador draw Germany 32 27 24 ↑ Ecuador — the model's pick
Same match, two questions. Germany is the longest bar on the left — the likeliest result. Ecuador is the longest on the right — the best-value bet. The engine plays the right-hand chart.

So the engine backed Ecuador. Ecuador won, 2–1. 145 points, on a match the whole room had handed to Germany.

This is value betting, not arbitrage — positive expected value with real variance. Ecuador winning was still the less likely outcome; it simply paid enough to be worth the risk, and bets like it only come good on average, over a whole tournament. Some of them lose. You keep taking the longest right-hand bar and let the month average it out.

The goal model

Turning three result-probabilities into a grid over scorelines needs a goal model — a bivariate Poisson. The plain version treats each team’s goals as an independent Poisson draw around a scoring rate; real matches aren’t independent, though — open games run end to end and both teams score, cagey ones die 0–0 together. The bivariate version fixes that with a shared term: each side’s goals are its own Poisson draw plus a common one both teams share.

home goals = X + Z away goals = Y + Z
X ∼ Poisson(λhome) Y ∼ Poisson(λaway) Z ∼ Poisson(λshared)

That shared Z — the goals the match produces rather than either side — couples the two totals and puts the right weight on the 1–1s and 2–2s independent Poisson under-counts. Its rate, λshared ≈ 0.15, was fitted by maximum likelihood on the tournament’s own finished matches midway through.

The load-bearing part is the calibration. The two team rates are not guessed: they are solved numerically, match by match, so that summing the finished grid back up reproduces the market’s win/draw/loss split to the decimal. The model is never allowed to disagree with the market about who wins — it only fills in the scorelines the market doesn’t quote a price for. And where Polymarket runs liquid exact-score and over/under markets, the engine reads the low scores and the total-goals level straight off those too, so even the shape of the grid is pinned to live prices rather than the model’s own tail (which, left to itself, invents too many blowouts). The goal model is less a forecaster than a way of spreading the market’s own numbers across every scoreline consistently.

Knockouts add a wrinkle: the market prices 90 minutes, but Mon Petit Prono scores the 120-minute result. So for a knockout the grid is projected forward. Every non-draw score is already settled at 90’ and passes straight through; each draw is handed its extra half-hour. Only about one in five ties still level at 90’ gets decided in extra time — the rest go to penalties, which MPP records as a draw — so the model breaks exactly that fifth of the draw mass into a win, shaped so the stronger side is likelier to nick it. The effect is small but one-directional: knockout draws get discounted, and the pick leans toward whoever is better placed to win the extra period.

The gap, measured

Why should this gap exist at all? Because the frozen book and the live market are answering the same question weeks apart, and the market has since moved. Before a ball was kicked, I measured it across all 72 group-stage matches: the typical value pick’s result paid about 16% more points than a fair, up-to-the-minute cote would have. In 61 of 72 matches the market trusted the favourite more than the frozen book did; the book systematically under-prices favourites and over-prices draws.

The tournament then confirmed it from both ends. Scored as forecasters over the 102 finished matches, the live market beats the frozen book on ranked probability score, 0.141 to 0.162 (lower is better). And the further the market had drifted off a cote, the more the pick banked: the 19 matches where the live probability sat more than 30% above the cote-implied one averaged 59 points; the 45 where the two books still roughly agreed averaged 33.

The gap can’t close by convergence the way a mispriced stock does, because one side of it isn’t allowed to move. The market just drifts a little further away each day.

The pick moves with the market

Because the market never stops, neither does the pick. Congo against Uzbekistan opened close to even — Congo around 40% — and at that split it was Uzbekistan’s higher cote that carried the better expected points, so the board spent nearly two weeks on an Uzbekistan win: mostly a 0–2, swapping with the 0–1 whenever the two ran near-tied. Then the market made up its mind: Congo firmed into the high 50s, the expected points overtook, and three days before kickoff the pick crossed to a Congo win, 2–1. Congo won 3–1, banked for 63 points. Nobody told the model to switch sides; it just kept multiplying the latest prices by the frozen payouts and taking the largest.

pick flips → Congo 40% Uzbekistan 31% 58% 18% the pick Uzbekistan Congo 2–1 book opens kickoff → won 3–1 · 63 pts
Seventeen days of one match — the market above, the board's pick below.

The same slow tilt decided my league, in the two matches that mattered most. Knockout books are short-lived (the semi-final cotes were printed once the quarter-finals settled the bracket, four days before kickoff) and both semis were coin-flips the frozen book paid well for. In both, the pick started on the home side and ended on the away side, not out of any taste for upsets, but because as each match tightened the higher-priced result overtook on expected points. England led Argentina in the market the entire run-up, the draw wedged between them by the end — yet the pick crossed to Argentina for good a day out, and settled on the scoreline, 1–2, only minutes before kickoff. Argentina won 1–2 — the exact score, 129 points. France led Spain more comfortably still; the pick tipped to Spain barely two hours before the whistle, once the match was close enough that Spain’s richer cote paid more in expected points. Spain won 2–0, and that cote returned 115. Two tight knockouts, two late crossings on value, both right: 244 points across one weekend, and, it turned out, the cushion that won me my league.

England–Argentina England 38% Argentina 30% 35% 31% the pick England Argentina book opens kickoff → 1–2 exact · 129 pts France–Spain France 41% Spain 29% 37% 31% the pick France Spain → book opens kickoff → 0–2 · 115 pts
The two semi-finals, run-up to kickoff: market above, pick below. Argentina crossed a day out (after one overnight wobble); Spain, two hours before the whistle.
A recycling bin lying on its side across the pavement on Tottenham Court Road at night, cigarette ends and broken glass scattered around it; a red double-decker bus, the Tube station and passers-by in the background.
Tottenham Court Road, London — July 15, around 10:45pm. “it’s coming home”.

Right direction, wrong crowd

Notice what the model did not do in that game: it didn’t pick the single most likely score. It picked a Congo win — the right direction — but 2–1, a scoreline much of the crowd skips for a tidier 1–0. That’s deliberate. The rarity bonus rewards being right where few others are, so among the scorelines in the direction the market favours, the engine leans to the one the majority underweights.

How far to lean is the one knob I turn, and it really is one number. A pick’s points are a three-outcome lottery: nothing if the result is wrong, the cote if the result lands but the exact score doesn’t, cote plus rarity bonus if it does. Two numbers summarise any lottery: its mean E (what the ticket pays on average) and its standard deviation σ (how far from that average a single run of it tends to land). A ticket paying 40 every time and one paying 0 or 80 on a coin-flip have the same E; only σ tells them apart, and it’s σ that says how much of a gamble the pick is. The pick maximises E + β·σ over every cell of the scoreline grid, with β the dial — a risk-appetite knob, no relation to the goal rates λ from earlier. Cautious (β = −0.5) penalises spread and settles on the likeliest realistic scoreline — realistic meaning at least a 2% chance, so it can’t hide in some 7–6 that merely banks the result — which is also the score most people share, so the bonus is small. Bold (β = +0.5) rewards spread and reaches for the widest lottery on the board: a rarer score, sometimes the other team entirely. Neutral (β = 0) ignores spread and takes the mean. So cautious lands more exact scores for fewer points; bold lands fewer for more, and pays for the reach by getting the plain result wrong more often.

Here is what that means on a real grid — Congo–Uzbekistan again, exactly as the engine held it at kickoff:

Congo–Uzbekistan · the grid at kickoff (58 / 24.5 / 17.5) Uzbekistan goals → Congo goals ↓ 0 0 1 1 2 2 3 3 4 4 8.5 6.6 2.7 0.8 0.1 14.7 11.5 4.6 0.9 0.3 11.9 10.0 4.0 0.9 0.3 5.8 5.2 2.4 0.5 0.2 1.9 1.9 0.9 0.3 0.1 C N B × C cautious · β = −0.5 → 1–0 E 41.0 · σ 36 — the likeliest realistic score; 29% of the crowd shares it, so the bonus is only +30 N neutral · β = 0 → 2–1 E 41.5 · σ 38 — half a point more mean, a rarer score, bonus +50. This is the pick that got posted. B bold · β = +0.5 → 0–2 E 29.0 · σ 63 — crosses to the longshot: the 158 cote is the widest lottery, and E + β·σ wins by 0.1 cells: P(scoreline) % · × = the final score (3–1)
Each cell is a scoreline's probability; E and σ follow from the cell, the frozen cotes (63 / 107 / 158) and the crowd's vote. Cautious gives up half a point of mean for the narrowest realistic lottery, 1–0. Neutral takes the grid's best mean, the rarer 2–1 — the pick that was posted. Bold walks across the diagonal to 0–2, trading twelve points of mean for twenty-five of spread. Congo won 3–1 (×), so the 2–1 banked the cote.

I learned how much that dial matters by accident. I’d flipped it to bold to show a friend how the thing worked, then forgot to flip it back — and one of the picks it made while stuck there was Egypt–Iran, 1–1, which came in exactly and paid 134. Pure luck that it landed, but a bold pick I never meant to leave on banked one of my best group-stage hauls.

The backtest is blunt on one point: the dull strategy — same calibrated grid, but always submitting its single most-likely score — is very good. Re-run over the 102 finished matches it banks 48.1 points a match against my 48.5, and lands more exact scores doing it, nineteen to my fifteen (it’s the safe-score strategy; of course it does). All the value-hunting bought me less than half a point a match over just taking the chalk. The machine’s real margin is over anyone without a market-calibrated grid at all: backing the live-market favourite with a stock scoreline averages 43.5 a match, and backing the frozen book’s own favourite about 40.5. Most of the edge is in knowing the true probabilities; the value layer only decides where the points arrive — in fewer, rarer, larger lumps. It so happens that leagues get decided in exactly the matches where that difference shows.

Same caller, different shape

Here is that difference, drawn. Three ways to pick — the model’s own pick, the market favourite backed with a stock scoreline, the frozen cote favourite backed the same way — scored across all 102 matches on two questions: how often does each call the actual result, and how do the points it banks pile up?

POINTS BANKED PER MATCH — 102 matches calls result 0 50 100 150 model 69% market fav. 69% cote fav. 66% median mean box = middle half · whisker to the single best match
Points banked per match, over 102 matches — a box plot for each way of picking. Every box starts at 0 because all three blank about a third of matches, and all three share the same median (45 for the model and the market favourite, 42 for the cote book), so the typical match is a wash. They part in the upper half: the model's box runs to 78 against the favourites' 70 and 68, and its whisker reaches 145 where both stop at 122. Its mean (the diamond) sits right of its median; the favourites' sit left. Same middle, heavier tail. The right-hand figures are how often each called the actual result — against 33% for a random one-in-three guess — with the model and the market favourite level to the match and the frozen book a touch behind, none of them out-forecasting a market the model is built from.

So the value engine buys nothing in the typical match — the medians line up — and everything in the tail: a longer box, a whisker reaching further, a mean pulled right of its median. It’s the boldness dial’s mean-for-variance trade, made once a match and compounded over a tournament, and it’s the whole margin. The accuracy figures on the right are the control: calibrated to the market, the model can’t out-forecast it and doesn’t try.

Did it work

Through the semi-finals — 102 of the 104 matches, with only the third-place play-off and the final still to come — here is the scorecard.

Matches scored 102 of 104
Total points 4,945
Average per match 48.5
Got the result (or better) 70 of 102 · 69%
Nailed the exact score 15 of 102 · 15%
Best single match 182 — the double-points joker
Rarest exact hit Uruguay–Spain · top 0.5–5% of players
Global rank 15,805th of 3,792,423 · top 0.42%
50 100 150 joker ×2 · 182 groups · 45/match knockouts · 56/match exact-score hit (15) — cote + rarity bonus result right (55) — cote only · no bar = blanked (32) match 1 semi-finals
Match by match, the shape of the income: a third of the matches bank nothing, most of the rest bank their cote, and the accent bars — the fifteen exact hits — are where the rarity bonus stacks on top: four of the five best matches of the run. The knockouts run taller because the cotes do.

Sixty-nine percent is a lot of correct results, and the tail did its part: the fifteen matches where the exact score came in added 500 points of rarity bonus between them, a tenth of the total, on top of the cotes they already paid.

Where that put me

That last line of the scorecard — 15,805th of 3.79 million — is the global picture. But you don’t feel the globe; you feel the private leagues — VARmageddon, people who take this far too seriously.

For long stretches I was losing it. To read that honestly, the chart below plots everyone’s running total minus mine: my line is the flat zero, and every other line is how far ahead of me (up) or behind me (down) they were, match-day by match-day. It’s the delta that’s hard to see on a normal cumulative plot, where six near-identical lines just pile up. The others’ lines carry a dot at each reading because MPP only reveals their totals by game-week — seven readings, connected.

+500 −500 −1000 ↑ ahead of me · behind me ↓ groups knockouts me Thomas −204 Guillaume −752 Jimmy −829 Sebastiano −952 the others — dots mark each game-week model (bold, neutral)
Everyone's running total minus mine — my line is flat zero. Jimmy peaked at +364 after 48 matches; the others simply fall away. Dashed: the model re-run at bold and at neutral.

The dashed lines are the machine’s counterfactual selves — the same engine re-run at bold and at neutral, no joker, no accidental toggles. Bold is the one worth staring at: fifty points a match through the groups, 460 up on me at the readings — it would have led VARmageddon outright — and then it bottled it, 44 a match through the knockouts against my 56, crossing the line dead level. That is what a high-variance strategy looks like once the sample shrinks: seventy-two group matches average the misses out; a thirty-match knockout run doesn’t have to. Neutral, meanwhile, shadowed me a hundred-odd points back the whole way.

That’s the shape of it. For half the tournament I was chasing — Jimmy ran away early, and Thomas sat glued to me the whole way and was ahead going into the final weekend. Then the two semis: Thomas scored zero, I had flipped to Argentina and Spain and banked 244, and the gap opened for good with two matches to spare. This is not a machine that dominates. It stays in touch all month. Thomas landed one more exact score than me over the whole run, sixteen to fifteen; my margin wasn’t the rare scorelines everyone remembers but the boring results, banked by a thing with no nerve to lose that switched sides the moment the numbers turned.

The jackpots

The bonus is where a good exact-score call compounds. Get a scoreline that most of the room also guessed and it adds a token 20; get one almost nobody saw and it can more than double your haul. Fifteen exact scores landed.

Match Score Points of which rarity bonus
Paraguay–Australia 0–0 154 50
Egypt–Iran 1–1 134 20
England–Argentina (SF) 1–2 129 20
Uruguay–Spain 0–1 127 70
Portugal–Spain 0–1 122 50
Scotland–Brazil 0–3 98 50
Jordan–Algeria 1–2 97 30
Haiti–Scotland 0–1 94 50
Switzerland–Algeria 2–0 91 30
Brazil–Japan 2–1 85 20
Spain–Belgium 2–1 80 20
Curaçao–Ivory Coast 0–2 73 20
Mexico–South Africa 2–0 69 20
Panama–England 0–2 62 30
Brazil–Haiti 3–0 41 20

Three of them deserve the detail.

Uruguay–Spain, 0–1. Spain was the value pick and the favourite both — sometimes the longest right-hand bar really is the favourite. But the model didn’t just call Spain; it called the tight 0–1, and almost nobody else did. Fewer than one correct-result player in twenty had the score — the rarest tier any of my picks reached — worth a +70 bonus and 127 points in all.

Paraguay–Australia, 0–0. A near-even match where the model bet the draw — expected points of 41 against 28 for the home win — and then picked the 0–0 that nobody ever wants to write down. 154 points, my best of the run bar the joker. The unloved result paid twice: once for being right, once for being lonely.

Tunisia–Japan, 0–4. The one double-points joker isn’t placed by hand — it automatically rides whichever unplayed match tops the board on expected points, and when Tunisia–Japan came around it was that match. So the joker locked in there. And it wasn’t even an exact hit: I’d read Japan as the stronger side and picked 1–2, but they won 0–4, nothing like it. No matter — I still had the result, Japan’s win cote was a fat 91, and the doubling took it to 182, the biggest single haul of the month, on a score I got wrong.

Extra time

So I built all this to dodge an afternoon of guessing scorelines, and it cost me a hundred afternoons instead. On pure forecasting the machine never beat the market it feeds on — it was never asked to. It holds two numbers side by side, one frozen a fortnight ago, one alive to the minute, and bets the gap between them a few thousand times, without once getting bored or sentimental. That was worth a top-0.4% finish out of nearly four million, and a league.

The final hasn’t been played — and I am still not doing the guessing myself.