My model's worst bet died on schedule
The first audit under pre-committed rules: nothing fired on audit day, a kill landed three days later at exactly n=150, and a bug ate two weeks of history. This is the report card.
Two weeks ago I wrote that my model's pitcher-strikeout picks were sitting at 149 of a 150-bet kill gate — one graded bet away from a verdict I'd pre-committed to in writing on July 1, before seeing the data that would decide it. That post ended with "the verdict lands within days."
It landed. This is the first of the monthly report cards: the full audit table, what the rules did about it, and the two things that went wrong on the way.
The setup, briefly
Every prop family the model bets gets graded against box scores. The high-conviction subset — every graded bet where the model claimed a 10-point-plus edge over the market — is the honest test of whether a family actually beats the market. And each family has a written rule: a sample-size gate, a threshold, and a pre-committed consequence. The rules were locked before the deciding windows closed. No re-litigating, no "but it feels different now."
Audit day: nothing happened
The scheduled audit ran July 15. Result: no gates fired. Not because everything was fine — because both decidable gates were short of their sample-size floors.
Pitcher strikeouts sat at 149 of 150. The tennis families, pooled, sat at 21 of the 30 the tripwire requires — at −27 points versus the market. Twenty-one bets at minus twenty-seven is the kind of number that makes you want to act right now. The rule says n=30. So nothing happened.
That's not a bug in the process. That is the process. If I'll bend the sample-size floor when the number looks awful, I'll bend it when the number looks great, and then the rules are decoration.
The confession
Part of why tennis was still at 21: a pipeline bug of mine silently deleted graded history for two weeks. A scheduling change I made on June 30 caused a nightly job to erase each day's graded prop rows hours after they were written. It ran every night from June 30 to July 12 before I caught it. The batter samples were frozen for two weeks, and the tennis bets that would have pushed the tripwire past 30 — Wimbledon's second week — are gone for good. The bug is fixed and that failure mode is now structurally impossible (deletes can no longer touch a graded row), but the sample math in the table below carries the scar, and you should know that when you read it.
Three days later, the axe
On July 18 the 150th high-conviction pitcher-strikeout bet graded, and the gate fired.
Final line: 150 bets, 46.0% hit rate against a 51.1% market price — 5 points worse than the market — while the model claimed 67%. Here's the detail I like most: by the time bet #150 settled, the verdict was mathematically locked either way. Even a win left the family below water. The rule didn't need the last result, and it didn't get a vote from me.
| Family | bets | model said | market priced | actually hit | verdict |
|---|---|---|---|---|---|
| Batter — hits | 324 | 66% | 52% | 58% ✓ | active, audit 9/1 |
| Batter — strikeouts | 101 | 66% | 51% | 56% ✓ | active, audit 9/1 |
| Pitcher — strikeouts | 150 | 67% | 51% | 46% ✗ | killed 7/18 |
| Tennis props (pooled) | 22 | 74% | 51% | 23% ⏳ | tripwire at n=30 |
(Tennis "model said" is the pooled average of two families; its verdict waits on the n=30 tripwire. Batter families face their own keep/kill bars on September 1.)
What "killed" actually means
The family isn't deleted — it's demoted to report-only, and the demotion is enforced in code, not in my intentions: a state switch the bet optimizer checks on every run. The model keeps predicting pitcher strikeouts and keeps grading itself, because the re-entry rule — also pre-committed — needs the sample: 100 post-kill high-conviction bets at 2+ points over the market earns back half stakes; 100 more at 3+ earns full. What it can't do anymore is put money on them, or recommend that you do.
The honest asterisks
Six-ish weeks of slates per family. One book's pick'em pricing as the market reference. The batter edges (+5 to +6 points over market on ~425 bets) are real so far but face their own pre-committed bars in September — the same axe hangs over the winners. And tennis at −28 on 22 bets is almost certainly going to fire its tripwire; the report card's job is to say "almost certainly" is not the standard, 30 is.
Next report card ~August. If you want to interrogate the model about its own killed family — it will tell you, caveat included — that's good-sport. Research tool, not picks. 21+.