Before anyone trusts a lean, we replay the summer: models rebuilt each morning with only what was knowable, graded against 850 real games and the actual DraftKings prices captured before first pitch, including the purchased strikeout prop price history. Two model generations so far. This page is what the iteration loop has found, including the parts that hurt.
| Bet type | Model Brier v1 / v2 / v3 / v4 | Market | Sim ROI v1 / v2 / v3 / v4 | Call |
|---|---|---|---|---|
| Totals (809 games) | .2582 / .2512 / .2503 / .2501 | .2500 | -3.3% / -0.8% / -0.45% / +3.05% | First positive ledger; forward CLV decides if it is real |
| Moneyline (838 games) | .2487 / .2478 / .2474 / .2474 | .2436 | -1.5% / -5.1% / -3.6% / -3.6% | Parked |
| Strikeout props (test half) | — / .2498 / .2446 / .2446 | .2451 | no volume at the chosen shrink | Information only, deepening continues |
The totals crossing comes with its own caveat, stated the way we state everything: plus 3.05 percent on 571 simulated bets is inside one standard error of zero, and the model still trails the market by a hair on pure forecasting at the market line. We out-earned the close without out-forecasting it, which is exactly the pattern that forward closing-line-value tracking, now running on every posted lean, exists to arbitrate. No real-lean designation until CLV agrees.
Umpire strike zones were the winner: the home plate umpire's shadow-zone called-strike tendency, built from our own pitch data and validated on seasons the fit never saw, moves totals by enough to matter and is announced before first pitch. Bullpen fatigue and catcher framing both failed validation honestly and were not used; a train-half feature idea for strikeout props failed its frozen test half and is reported as noise.
Generation two prices every plate appearance from the pitcher's arsenal against each hitter's swing profile by pitch class, with platoon, times-through-order, and the actual posted lineups. It beat generation one on every bet type, halved the number of far-from-market disagreements, and removed the totals model's structural over-bias. It is the best strikeout forecaster this project has produced.
The de-vigged closing-quality price is still the best forecaster on every market, and no simulated policy is profitable yet. The strikeout model's neutral-line skill did not survive contact with real prices: where we disagree most with the line, the line has been right, which means the market is pricing information we do not yet hold, like rest plans, pitch-count intentions, and late scratches. Totals sit within one Brier point of the market and recover most of the vig; that is the current front of the iteration.
Generation one priced games from team strength alone. Generation two priced every plate appearance from arsenal against swing profile with real lineups, and beat one everywhere. Generation three added game-time weather to totals (temperature, wind toward center, air density; wind direction beat plain wind speed in validation), refit the moneyline blend, and proved the strikeout residual idea dead with a frozen out-of-sample test while a small shrink-to-market sliver held. Every graded lean now also records closing line value, the leading indicator this loop watches. The K props Brier marked with a star is the out-of-sample test half only.
Generation five moved the winning configuration into the live daily board (weather and umpire terms price every total each morning; crews not yet announced are flagged and re-priced when they post), upgraded the strikeout reference model with the umpire zone inside the per-PA machinery, and rejected every new totals candidate when each lost its one pre-registered frozen test. It also caught a fragility worth publishing: swapping forecast hours for archive hours in the weather feed collapsed the totals profit to zero until the merge was fixed, which is why closing line value, accumulating on every posted lean from August 23, is the arbiter before any lean is called real.
Replaying the frozen prop models over 2021 through 2025: across 19,448 starts, the strikeout model beats the naive baseline by the same two Brier points every single season (pooled .2360 against .2577, model accuracy flat across years), and across 145,487 batter games the hit and home run models beat their baselines with the two strongest seasons being the most recent two. The skill shown in the 2026 window is persistent, not luck. What that skill has not yet done is beat market prices, which price information beyond performance; that remains the gap the iteration loop is working.
We then bought the entire 2025 season of pre-start DraftKings prop prices and graded the frozen models against them: 2,515 strikeout prop bets and 16,814 hits prop bets under the standing policy. Both lost, 6.2 percent and 5.3 percent, with no relationship between claimed edge and realized win rate. One finding matters more than the losses: an early read of the hits props showed a large paper profit that turned out to be leakage, the evaluation frame quietly knowing which hitters ended up with three or more plate appearances, information no bettor has at post time. The corrected pregame-legal version flipped the sign. We publish that catch because it is exactly the kind of self-deception this page exists to prevent.
The comparison also answered a design question: the hitting projection stack, which has been through its full research-panel refinement, sits three times closer to the market than the pitching stack, which has not. Refinement measurably closes the gap; it has not yet crossed it. What crosses it is information the market holds and we do not: playing-time intentions, rest plans, and scratches.
Nothing on this site is called an edge until this table shows a bet type winning at real prices. Leans on the boards are model tracking that feeds the public ledger. The loop continues: deeper models, another replay, whatever the table says, posted here.
Prices are the latest pre-start DraftKings snapshot, one book. Closing-line value tracking on forward leans began August 22 from the twice-daily odds captures. All fits use 2015-2023 only; the evaluation window was never fit on. Everything regenerates from cached snapshots; nothing is retroactive.
21 and over. Gambling problem? Call 1-800-GAMBLER. Everything here is analysis and entertainment, not individualized betting advice. Project Hardball takes no bets, sells no picks, and has no sportsbook relationship.