How Accurate Are Betting Models?
Per-game accuracy is mostly variance, so it is the wrong question. What separates a useful betting model from a tout is calibration over many games, graded against results, and published where you can check it.
Updated Sep 2026 · Part of the getting started series
Statistical estimates, not betting advice. Past results do not predict future results. 21+. Call or text 1-800-MY-RESET (1-800-697-3738).
Why is per-game accuracy the wrong question?
Because chance decides most of any single game, so the answer reads the same for a good model and a lucky one. No sample a bettor will ever see is large enough to tell the two apart on per-game hit rate.
Ask how accurate a betting model is, and most people mean how often it calls the game correctly. Per-game results are dominated by variance, the swings in outcome driven by chance rather than skill, covered in Variance in betting. That per-game question cannot separate a good model from a lucky one on any sample size a bettor will ever actually see.
What is calibration, and why does it replace accuracy?
Calibration asks whether a stated probability comes true at the stated rate. If the model gives a side a 60% chance, that side should win about 60 times in 100 similar calls. That check accumulates over hundreds of games, which per-game correctness never does.
A checkable question replaces the per-game one. When the model gives a side a 60% chance to win, does that side actually win about 60% of the time across hundreds of similar calls? That check is calibration. Calibration accumulates over many games. A single game’s outcome does not. Per-game correctness stays noise no matter how many times you check it.
Why compare a model against the closing line?
Because the closing line is the best public forecast of a game that exists. By kickoff it has absorbed injury news, confirmed lineups, weather, and the money of bettors who have beaten the market before. A model claim has to be measured against that, not against a coin flip.
The closing line at a sharp book gives calibration a benchmark to compare against. By kickoff, the price has absorbed injury news, confirmed lineups, weather, and the money of bettors who have historically beaten the market. That combination makes it the best publicly available forecast for a game’s outcome. That is the idea Closing line value covers in full. Academic comparisons have tested well-known public models against that closing line across tens of thousands of games, and the market price won.
Can a model beat the market?
A public-data model that ties the close over a large sample is performing at the practical ceiling for what public information can produce. Any claim to broadly beat it deserves skepticism, Fairline included. Fairline claims no such win.
One MLB pocket is tracked forward rather than claimed. Its status in the claims registry is a tracked hypothesis, which means it is under test and has not passed. The MLB underdog edge sets out how a bet enters that pocket and where the record stops.
Which three numbers actually measure a model?
A calibration curve, a Brier score, and the closing line value of the model’s own picks. The first shows whether stated probabilities come true at their stated rates. The second grades those probabilities against simply predicting the historical average every time. The third shows whether the picks beat the market before any result lands.
Three numbers measure any model, checkable without taking the builder’s word for anything. A calibration curve plots predicted probability against realized outcome by bucket, every pick made in the 60% range grouped together and checked against how often that bucket actually won, repeated bucket by bucket from long shots to near locks. A Brier score grades those same predictions against the base rate, the accuracy of just predicting the historical average outcome every time, and shows whether the model adds anything beyond that plain baseline. The CLV of the model’s picks shows whether its specific selections beat the market before a single result gets decided by variance.
A model seller who publishes none of these three is asking you to trust an adjective.