Explainer·September 2, 2026

What Model Calibration Means, and Where Ours Is Wrong

Our reliability curve across 51,494 settled positions: forecast 42.5%, won 36.9%. The bets where our probabilities were worst returned +37%, and why both are true at once.

A model that says 30% should win about 30 of every 100 times it says it. That is calibration, and it is a different question from whether the model is any good.

A calibrated model can be nearly useless: forecast the base rate every time and you will be perfectly calibrated and completely uninformative. And a genuinely sharp model can be badly calibrated, right about which bets are better while being wrong about how often they land. You want both, and almost nobody publishes the second.

Here is ours, across 51,494 settled positions in 12 events.

The reliability curve

Every position we published, grouped by what the model forecast, against what actually happened.

We saidPositionsAverage forecastActually wonGap
0% to 5%3734.0%5.4%+1.3 pts
5% to 10%2,2437.7%6.6%-1.2 pts
10% to 15%3,78812.7%9.3%-3.3 pts
15% to 20%4,63517.5%13.5%-4.0 pts
20% to 30%8,33724.7%20.2%-4.6 pts
30% to 40%6,74334.8%28.4%-6.5 pts
40% to 60%10,10848.8%41.4%-7.4 pts
60%+15,26772.5%66.1%-6.5 pts

Overall the model forecast 42.5% and won 36.9%, so it runs 5.6 points overconfident. The gap is small at low probabilities and widens through the middle, worst in the 40% to 60% band at -7.4 points.

That is the honest headline and we are not going to dress it up. But it is also not the whole story, and the rest is more interesting than the number.

These are the bets we chose, not every forecast we made

The table above covers positions we published, which means positions where our probability was high enough above the market's price to clear an expected-value floor. That is a filtered sample, and filtering it that way makes it look overconfident even when the underlying model is well calibrated.

Think about what has to happen for a bet to reach that table. The market prices a hole at 20%. We publish only when our number is meaningfully higher. Our estimate carries error, like every estimate. Cases where the error runs high get selected in. Cases where it runs low get filtered out. We are not sampling our forecasts, we are sampling our forecasts on the occasions they happened to be highest, and the average of a maximum is biased upward. Statisticians call it the winner's curse. Every bettor experiences it and most blame the model.

The pattern selection would leave behind

If selection is doing the work, the effect should scale: the further above the market we sat, the more of that distance should be our error, and the worse the gap should get.

How far we disagreed with the bookPositionsWe saidActually wonGap
0% to 3% higher11,54521.2%19.0%-2.2 pts
3% to 6% higher19,40645.1%41.3%-3.8 pts
6% to 10% higher12,49750.4%43.9%-6.5 pts
10% to 20% higher7,00953.5%41.9%-11.6 pts
more than 20% higher1,03762.5%37.3%-25.2 pts

It is monotone. Every step further from the market's price makes the forecast less accurate, from -2.2 points where we barely disagreed to -25.2 points where we disagreed most.

That is the shape selection produces. Near the market's price our forecasts are close to honest. Far from it, most of the distance is us being wrong.

The obvious objection, tested

Bets where we disagreed most also tend to be different bets, at different probabilities, so that table could be the reliability curve again in disguise rather than anything to do with disagreement.

It is not. Restricted to positions we forecast between 45% and 70%, so every row below is a like-for-like comparison of similar bets:

How far we disagreedPositionsWe saidActually wonGap
0% to 3% higher74550.1%52.2%+2.2 pts
3% to 6% higher5,79059.7%56.5%-3.2 pts
6% to 10% higher4,05159.7%54.6%-5.1 pts
10% to 20% higher3,07157.9%46.9%-11.0 pts
more than 20% higher50257.8%33.3%-24.5 pts

Same pattern, with the forecast level held fixed. The confound does not explain it, and the effect survives inside every probability band we checked, not only this one.

The obvious next question: did those bets lose money?

They did not, and this is the part that makes calibration genuinely counterintuitive.

How far we disagreedPositionsFlat return95% interval
0% to 3% higher11,545+0.8%-4.6% to +6.3%
3% to 6% higher19,406+2.8%+0.1% to +5.4%
6% to 10% higher12,497+3.7%+0.8% to +6.6%
10% to 20% higher7,009+7.6%+3.0% to +12.3%
more than 20% higher1,037+37.1%+20.8% to +53.5%

Returns move in the opposite direction to calibration. The bets where our probabilities were worst are the bets that made the most money, at +37.1% for the widest disagreements. That is not a small sample fluke: the interval excludes zero, and the bucket was profitable in 10 of the 12 events in the record.

Both things are true at once, and the reason is that these are plus-money props. Suppose we say 62%, the truth is 37%, and the book is pricing 30%. We are 25 points overconfident and still correct that the price is too long. Being wrong about the magnitude and right about the direction are different failures, and only the second one costs you money.

It also puts a limit on the explanation offered above. If selection were the whole story, the apparent edge should have disappeared along with the calibration, because the winner's curse eats the edge and the accuracy together. The edge survived. So selection is inflating our probabilities, and underneath that there is a real signal it has not accounted for. We would rather say that than claim a tidier mechanism than the data supports.

What that means if you are betting

The lesson is not that big edges are fake. It is more specific and more useful than that.

  • Do not read our probability as a probability, especially when it sits far from the market price. Treat it as a ranking, not a forecast. That is the same conclusion our page on expected value reached from the other direction.
  • The direction still carries information, including in the region where the calibration is worst. The model is finding genuinely mispriced positions even while overstating how often they land.
  • Size accordingly. A forecast you know is inflated should not get a large stake, however big the apparent edge. That is exactly what fractional Kelly and the confidence filter are for.

The trap is seeing a calibration gap and concluding the model is broken. A model can be miscalibrated and profitable at the same time. Those are separate questions and they need separate evidence, which is why both tables are on this page.

What we do about it

Two things, neither a full fix.

A calibration layer fitted per course. Raw simulation output is corrected toward that venue's observed scoring before anything is published, so the numbers on the board are already adjusted rather than straight out of the simulation. It needs a couple of seasons of tracked shots at a venue to fit, so at a course we have never captured it cannot run and the raw output is what you get.

A confidence rating on every signal. Each one is scored on how far it reaches past the model's own predictions for other players on the same hole, and shown as Consensus, Lean or Outlier. In backtest, signals where the model agreed with its own field returned +11.5% on quarter-Kelly staking while the most extreme ones returned -11.5%, so the board can be filtered by it.

Worth being precise about, because it sounds like it contradicts the table above and does not. Two different kinds of reaching:

  • Reaching past the market (the tables above) was rewarded. The model found genuinely mispriced positions.
  • Reaching past its own field, calling one player wildly different from the others on the same hole, was punished. That pattern is usually a quirk in one player's inputs rather than a real read.

Disagreeing with the bookmaker and disagreeing with yourself are not the same act, and our data says to treat them differently.

Neither removes the effect. Selection bias is a property of choosing which bets to make at all, so the only complete escape is not choosing, which defeats the purpose of having a model.

The caveats worth stating

12 consecutive events, the back half of the 2026 PGA Tour season, every one of them settled position by position. Enough to establish the shape. Not enough to pin the magnitude.

Worth separating from the rest of the data, because they measure different things. The betting record starts when the current pipeline went live. The shot corpus behind the model is far older and larger: 3,508,763 tracked shots across 29 events and 7 seasons. Calibration can only be measured on bets that settled, so this page is the smaller number by necessity.

We cannot fully separate two explanations. The widening gap is consistent with our estimates being noisy, and equally consistent with the market being sharper than us exactly where we disagree with it most. Both probably contribute. Usefully, the practical advice is identical either way.

This is the curated set, the same positions the live board surfaced, not every number the model produced. Every one of them is on our public results page, wins and losses, so the curve above can be checked rather than taken on trust.


PropGolf publishes every model probability with its settled outcome, so the calibration on this page stays checkable. See the live board.

See it live on PropGolf

Track every hole-by-hole signal in real time, with multi-book odds and per-group ETAs.