Research note · August 2026

The prop model that didn't work.

We built a strikeout model, graded it against the props books actually post, and found it has no edge there. This is the whole result — including the three explanations we got wrong on the way, and the one this page itself got wrong and has since corrected.

Our model disagrees loudly with the prices books actually post, and it is wrong about as often as a coin. Across 296 real prop lines it disagreed with the market by 11.1 points on average and its side won 44.7% of the time. We are not selling it, and this page exists because the measurement is worth more than the model was.

WHAT WE TESTED

A prop bet is a wager on one player's line — how many batters a pitcher strikes out, how many bases a hitter records. Props are widely believed to be the softest numbers a book hangs, because there are thousands of them and less money shaping each one. That belief is why we built a model for them.

The model projects the full distribution of a pitcher's strikeouts rather than a single guess, and publishes the probability of going over the line next to the market's own de-vigged probability. The gap between those two numbers is the entire claim: if ours is better, the gap is an edge.

We graded it the only way that means anything — walk-forward, where every projection is built solely from starts that happened before the game being predicted. No result on this page uses information the model would not have had at the time.

THE RESULT

We reconstructed every strikeout prop our capture actually saw between 8 August 2026 and 1 September 2026 — real lines, real prices from three or more books on both sides, outcomes known. That is 296 props with a gradeable result.

On the real board propsModelMarket
Predicted the over46.8%49.1%
Actually went over44.3%44.3%
Average disagreement with the price11.1 pts
Props where we disagreed by 3+ points237 of 296
Our side won those44.7%

A coin flip is 50%. Breaking even at standard prop pricing needs more than that.

The size of the disagreement is what kills it. The model was not quietly hedging near the market and losing narrowly. It was claiming an eleven-point edge, on 80% of the props it saw, and then winning fewer than half of them. An edge that large would be unmissable if it were real.

Stated honestly, because it matters: with 237 decisions the margin of error is about 3.2 points, so 44.7% is close enough to a coin flip that we cannot claim the model is actively bad. What this sample does rule out is the large edge the model was advertising.

WHY — AND THIS IS THE PART THAT GENERALISES

The model is not useless in the abstract. Graded against one league-average line applied to every pitcher alike, it carries real predictive information. Graded against the line a book actually posted, that information is gone.

One statistic captures it. Fit the model's own output against what happened, and ask how much weight the answer deserves. A value of 1.0 means the model's probabilities can be taken at face value. A value of 0 means they carry no information at all.

What the model was graded againstSampleInformation in the model
Every start, one league-median line for every pitcher2,647+0.51
Every start, the line that pitcher usually gets2,631+0.22
Only games books posted a prop on, at the price they hung296-0.10

95% interval on the last figure is -0.44 to +0.24. Wide, because 296 is not many, but it excludes the values the model would need to be worth publishing.

The drop that is real is the first one, and it is not about which games books choose to hang. It is about what a posted line already contains. Most of what looked like the model's information was knowing that different pitchers strike out at different rates, and the number a book hangs is built around exactly that. Withhold that from the line and hand it only to the model, and the model looks informed. Give it to both and 57% of the effect disappears. That single step is worth 4.6 standard errors.

The rest of the fall is not established. Holding the line rule fixed and comparing the games books posted a prop on (+0.07 on 294 of them) against every other start moves the figure by 0.8 standard errors. Swapping that pitcher's usual line for the exact price the book actually hung moves it by another 0.7. At 296 props neither of those is a finding, and we are not going to present them as one.

Our edge existed where no market existed. What we got wrong for eight days was what "no market" meant. It does not mean a game the books ignored. It means a line that does not know who is pitching — and no real line is ever that.

The generalisable warning survives the correction, and it is not the one we first wrote down: a model validated against a benchmark weaker than the one it will face is measuring the benchmark, not itself. Every figure we quoted before rebuilding this test around real board prices was measured on an easier sample than the real one.

WE TRIED TO SALVAGE IT

A model can fail to beat the price and still be worth showing — an honest probability next to the market's is useful information even when it is not a better bet. So we tried to publish that instead.

Recalibrating the model against real board props does fix its accuracy. It fixes it by throwing the model away.

Calibration errorBrier score
The model as built15.1 pts0.2742
After recalibration6.4 pts0.2476
The market itself7.7 pts0.2454

Lower is better on both. Scored out of sample, in 5 folds, so no prop is graded by a correction fitted on itself. The recalibrated model looks excellent.

The correction the fit chooses is to stop listening. Across the folds the weight it puts on the model's own opinion lands between -0.17 and -0.03: nought, to within the noise. What comes out the other side sits between 38% and 54% on every prop, averaging 44.3%, which is the base rate. It scores well because predicting the base rate always scores well. It is a constant wearing a model's clothes, and publishing it would be publishing the base rate with extra steps.

This is worth stating plainly because it is the trap most people never check: calibration and information are different properties. The Brier score improved, and that is usually where the checking stops.

THREE EXPLANATIONS WE GOT WRONG

Each of these was stated confidently before it was tested, and each was abandoned because a measurement contradicted it. They are here because the revisions are the honest part of the record. Unlike everything above, these three are summaries of one-off diagnostics from August 2026 rather than figures this page regenerates, so they are described and not quantified.

1. "Per-player error swamps the signal."

Measured in-sample and over-read. Separating out the noise you would expect from small samples, per-player deviation in both models is entirely consistent with chance. At 8–28 starts per pitcher this simply cannot be measured. Withdrawn.

2. "The probabilities are too extreme because we ignore uncertainty."

The fix would have been a wider distribution. We fitted one, and there was almost nothing to widen: strikeouts per start are very close to the simple distribution we already assumed, and the calibration error barely moved. The story was wrong.

3. "Shrink each pitcher's rate toward league average by sample size."

Made everything worse. More useful than the failure: pushing that correction to the limit of what the data allows still could not fix the defect, which rules out the entire approach rather than one setting of it.

What survived

The model's projections move further than reality does, and the noise enters through how it estimates a pitcher's workload from his recent starts. Weighting recent form — the thing that feels most like insight — measurably made the model worse than simply using the season to date.

WHAT THIS MEANS FOR THE SITE

We do publish a daily prop post, and this result decides what it can say. It is selected on price — the best available number across the books pricing each prop, measured against their own de-vigged consensus. It is not selected on model edge, because the model does not have any. Those are different products and this page is the reason we ship the first one.

So the daily post will tell you where the best number is, and what a player has actually done in his last five, last ten and season starts. It will not tell you a prop is mispriced, because we measured our ability to know that and it came back at 44.7%. Where our model probability appears it is labelled as context, not as a forecast.

Worth being blunt about what a price edge is and is not: the vig means the best available price is usually still worse than the consensus fair price. Picking the best number reduces what you give up. It does not turn a losing proposition into a winning one, and we are not going to imply otherwise.

This is the second time we have published our own model failing. Our NFL prediction model hit 49.8% against the spread across 2,608 decided walk-forward games, where breakeven is 52.4%. That measurement is why this site sells price transparency rather than picks, and it is why the rest of the numbers here are worth something. Results that flatter us and results that do not get published the same way.

If you want the bar for reopening this: a model would have to show real information against the prices books actually post, on a sample larger than 296. That is the current sample, it is written down, and it moves only when this page is regenerated from the capture, so nobody — including us — gets to move it quietly.

TWO THINGS WE DID NOT FIND

Recorded so they are not later rediscovered as news.

The market itself predicted 49.1% against a 44.3% actual on this sample. That looks like a bias toward the under. At 296 props it is about 1.7 standard deviations: noise, not a finding, and we are not trading on it.

We built a total-bases model too, and it points the same way, but it was never tested on a board population. We captured 160 total-bases rows against 4,125 for strikeouts. Its exclusion rests on weaker evidence than the strikeout case and should not inherit the strikeout case's certainty.

WHERE THESE NUMBERS COME FROM, AND WHAT CHANGED

Until 28 August 2026 every figure above was typed into this page by hand, from an analysis that existed once and was then thrown away. On a site whose whole position is that its numbers are checkable, that made this the least checkable page we had — and it is the page most worth checking, because it is the one arguing we are honest about a failure.

It is now regenerated by scripts/props_model_note.py from two files committed in this repository: our own append-only prop capture, and a cached copy of the free MLB game logs. Anyone who clones it reproduces every number offline. The script also writes the figures into this page, so no digit here is authored in the markup, and a test fails the build if page and payload ever disagree. The payload itself is at props_model_note.json.

Regenerating it moved things, and the moves are the reason this section exists rather than a quiet edit.

The sample grew and the result got slightly worse. The original run graded 194 props and reported our side winning 48.1% of its disagreements. This one grades 296 and reports 44.7%. Part of the difference is five more days of capture; part is that the original could not resolve about a dozen pitchers whose names are shared with another player, so working starters were silently dropped. The conclusion is unchanged.

One claim did not survive at all. This page used to report that the model carries information across all pitcher-starts (0.48) and none on the games books post a prop on (−0.07), and it blamed books for choosing which games to hang. Regenerating showed that figure was measured against a single league-median line applied to every pitcher alike. Let the line know which pitcher is on the mound — which every real line does — and most of that information is gone before books select anything. The selection effect we described is not measurable at this sample size. The section above now says so.

One statistic was replaced. The salvage table used to quote the worst of ten calibration buckets. On a few hundred props the worst bucket often holds two of them, and the number swings tens of points between runs. It now quotes calibration error across five equal-sized buckets, which is stable, and is scored out of sample.

Nothing here was reconciled by hand. Where the new run disagreed with what was published, the new run is what is published, and the disagreement is written above.