The first bet my rules refused to kill
The August audit: batter hits becomes the first family to pass a pre-committed keep gate, the family I killed last month is quietly beating the market from beyond the grave, and tennis is still five bets from its tripwire. This is the report card.
Last month's report card was about an execution: my pitcher-strikeout picks hit a written kill gate at exactly 150 bets and were demoted to report-only, no appeal. A fair question after that post: do the rules only ever swing the axe? If every gate is a kill gate, "pre-committed rules" is just a slow way of shutting the whole model off.
This month one answered. The August 1 audit made the batter-hits gate decidable — and the same machinery that killed pitcher-K voted keep, at full weight.
The table
Same canonical cut as every audit: every graded bet where the model claimed a 10-point-plus edge over one book's pick'em market, by family, cumulative. No blending — the blend is how a losing family hides inside a winning average.
| Family | bets | model said | market priced | actually hit | verdict |
|---|---|---|---|---|---|
| Batter — hits | 413 | 65% | 52% | 56% ✓ | kept 8/1 |
| Batter — strikeouts | 134 | 67% | 51% | 56% ✓ | gate at n=200 |
| Pitcher — strikeouts | 224 | 67% | 51% | 49% ✗ | killed 7/18 |
| Tennis props (pooled) | 25 | 74% | 51% | 28% ⏳ | tripwire at n=30 |
What passing actually took
The bar, written on July 1: at 400 graded high-conviction bets, a family keeps full weight only if its realized edge is 3+ points over the market. Batter hits crossed 400 this week and the deciding window read +4.1 points on 296 bets (cumulative: +4.5 on all 413). Above the bar. Kept.
Here's the part a tout would not tell you: the family we kept is still overconfident. The model claimed 65% on these bets; they hit 56%. The keep gate measures edge against the market's price, not against the model's self-regard — batter hits beats the market while exaggerating its own ability, which is why its probabilities get shrunk before anything is staked. Passing the gate means the market is beatable here. It does not mean the model is honest yet.
The dead family that won't stay dead
The awkward result of the month belongs to pitcher strikeouts — the family the rules executed in July. "Killed" means report-only: the model keeps predicting and grading them, it just can't stake them, because the re-entry rule (also written July 1) needs a post-kill sample: 100 high-conviction bets at 2+ points over market earns back half stakes.
Since the kill, those unstaked picks are 70 bets at +4.0 points over the market. Killed on a −5.1 verdict; beating the market ever since. Maybe it's the fixed grading window. Maybe it's variance — 70 bets is a coin-flip's throw of noise. The ladder exists precisely so I don't have to guess: 30 more bets and the rule decides, not the streak. If it re-enters, it re-enters at half stakes with its record public. Watching your own kill decision get argued with by the corpse is exactly what this system is for.
Still no tennis verdict
Third card in a row: the pooled tennis families sit at 25 of the 30-bet tripwire, at −23.2 points versus the market. Twenty-five bets at minus twenty-three. Every instinct says call it. The rule says n=30, and the whole point of writing the rule in June was to bind the version of me reading the number in August. Sparse grading — many tennis matches never produce the serve stats a bet needs — means these last five bets have taken six weeks to accrue. The next audit (August 15) almost certainly resolves it, and "almost certainly" is still not the standard.
Changes since last card
Two families were benched in late July that do not appear as fired gates above, and the distinction matters for honest bookkeeping: batter home runs and batter RBIs went report-only by an owner-approved calibration rail, not a tripped threshold — the full graded board showed HR losing 7.6% a unit with RBI claiming negative average edge, and I chose not to wait for a formal gate. The rules constrain the model; they don't forbid me from being more conservative than they require. Also shipped: stakes were cut to quarter-Kelly across the board and sub-coin-flip game picks are auto-skipped, both pending a calibration overhaul that's mid-flight.
The honest asterisks
Seven-ish weeks of slates. One book's pick'em pricing as the market. The pitcher-K resurrection is 70 bets — the exact sample size at which last year's me would have written a victory post. Batter strikeouts at +5.2 looks like a second keep, but its gate is n=200 and it's at 134; the axe hangs there until it doesn't. And every "model said" column in that table is still inflated, which is the problem the next few months of engineering are actually about.
Next report card ~September 1, when batter strikeouts' gate and the tennis verdict should both be in. If you want to interrogate the model about any of this — including the families it's not allowed to bet — that's good-sport. Research tool, not picks. 21+.