How a losing bet hides inside a winning average
My model grades its own picks. On player props that turned up one prop it's genuinely good at — and one it's confidently, expensively bad at. The average hid both.
Last time I wrote about closing-line value — whether the model beats the market. This post is the mirror image: a place the model beats itself. And I only found it because the model reconciles every pick it makes against what actually happened.
Here's how that works. On player props, the model predicts something concrete — this hitter gets a hit, this pitcher goes over X strikeouts. Then, once the box score is in, it grades that prediction: right or wrong. Over a month that's thousands of graded picks — a running report card the model keeps on itself. That report card is where you find your own leaks, if you're willing to read it.
The trap: an average that meant nothing
I pulled the model's highest-conviction prop bets — the ones where it saw a 10-point-plus edge over the market. 517 graded legs across six weeks of slates. As a group, they hit about 53% of the time. The market had priced them at about 52%.
Edgeless. If I'd stopped there, I'd have concluded the whole prop model was worthless and shut it off.
But an average is a great place to hide. Split those exact same bets by what kind of prop they were, and two opposite stories fall out:
| Prop | bets | model said | market priced | actually hit |
|---|---|---|---|---|
| Batter — hits | 265 | 66% | 53% | 58% ✓ |
| Batter — strikeouts | 82 | 65% | 51% | 56% ✓ |
| Pitcher — strikeouts | 149 | 67% | 51% | 46% ✗ |
Read across the batter rows first. The model won ~57% where the market priced ~52% — a real, roughly 5-point edge over the market, across ~350 bets. That's the good half, and it's exactly the kind of edge the blended average told me didn't exist.
Now the pitcher line. The model said 67%. The market said 51%. Reality came in at 46%. The model wasn't a little off — it was 21 points overconfident, and it lost badly. (Last post I mentioned the model runs "about a dozen points" overconfident on pitcher strikeouts overall; on the bets where it's most sure, the gap nearly doubles.)
That's the whole lesson in one table. The "edgeless" blend wasn't edgeless — it was a genuine batter-prop edge and a genuine pitcher-strikeout hole, averaged into a number that described neither. Trust the blend and you make the wrong move twice: you abandon a real edge and you keep funding a loser.
Why pitcher strikeouts?
My best guess is bullpen usage. A starter who could pile up strikeouts increasingly doesn't get the chance — pitch counts, the third-time-through-the-order penalty, openers and bullpen games. The market has fully absorbed that the strikeout over is a trap; the model was still betting the talent, not the leash it's kept on.
That's a hypothesis, not a finished fix — which is the honest state of it. What I've done in the meantime is boring and correct: down-weight pitcher strikeouts, keep the batter props. And I've pre-committed the rule that decides its fate, in writing, before seeing the data that will judge it: when the sample reaches 150 of these high-conviction bets, if it's still underwater against the market, it stops getting real weight — no re-litigating, no "but it feels due." As I write this the count sits at 149. The verdict lands within days.
The part I'm actually proud of
It isn't the fix. It's that the model will tell you this. Ask good-sport "is the model overconfident on pitcher strikeouts?" and it answers yes — by how much, and that it's the one prop type running a loss, and that you should down-weight it. It confesses its worst prop unprompted.
That's the line between a tout and a tool. A tout shows you the batter-hits edge and never mentions the pitcher-strikeout hole. A tool hands you the whole report card — and trusts you to read the second column.
The usual asterisk: this is six weeks of graded bets. Solid for batter hits, thinner for the pitcher cut. I'll keep re-grading and I'll say so here if the story changes. (The same cut-by-prop-type habit is already flagging my tennis props as the next suspect — the sample there is still tiny, so that's a future post if it holds.)