The public, graded record.
Player-level probabilities for every MLB slate: pitcher strikeouts, batter hits, batter home runs. Published before first pitch, graded against Statcast that night, never edited. The goal is calibration, not beating the market. When this engine says 30%, it should happen about 30% of the time, and this page is the test. Probabilities, not picks. Not betting advice.
Void rules, stated in advance: published late voids. Postponed voids. A pitcher who does not start, or a batter outside the starting nine, voids. Voids stay on the record.
Calibration, the whole point
When the engine says 30%, does it happen 30% of the time? Points on the dashed diagonal are perfectly calibrated. Above it the model was too low, below it too high. Bins thinner than 25 samples live in the tables, not the curves.
Strikeouts
Live record
9.5% calibration error, 0.207 Brier, 252 grades
Reliability table
| Bin | Predicted | Observed | n |
|---|---|---|---|
| 0 to 10% | 9% | 0% | 3 |
| 10 to 20% | 16% | 22% | 37 |
| 20 to 30% | 25% | 18% | 38 |
| 30 to 40% | 35% | 44% | 55 |
| 40 to 50% | 46% | 52% | 48 |
| 50 to 60% | 55% | 74% | 34 |
| 60 to 70% | 65% | 77% | 26 |
| 70 to 80% | 74% | 91% | 11 |
Engine change, 2026-07-18 → 2026-07-19. For one day the strikeout model carried an extra feature — the trailing strikeout rate of the opposing starting nine. An adversarial review found it was trained on the lineup each pitcher actually faced but served a posted-or-projected one, and that on the projected path it made the model worse than having no such feature at all; its apparent gain was never distinguishable from zero. It was removed on 2026-07-19 and the model retrained without it. 81 graded line-results (27 starts, all from 7/18) plus the 7/19 board were produced by that version. Those predictions stand exactly as published and stay in the running totals — they were made in good faith before first pitch, and deleting inconvenient history is the one thing this record will not do. Method: docs/K-FEATURE-VERDICT.md.
Backtest, not the live record
2.2% pooled calibration error, per-line 2.2% / 3.9% / 2.6%, 0.205 Brier, 1090 holdout starts
Trained on real starts before 2026-06-01, openers excluded. The three lines per start are nested views of one outcome, so the pooled number can flatter. It earns nothing until the live record proves it.
Reliability table
| Bin | Predicted | Observed | n |
|---|---|---|---|
| 0 to 10% | 10% | 0% | 1 |
| 10 to 20% | 15% | 14% | 493 |
| 20 to 30% | 25% | 23% | 657 |
| 30 to 40% | 35% | 37% | 606 |
| 40 to 50% | 45% | 47% | 614 |
| 50 to 60% | 55% | 56% | 462 |
| 60 to 70% | 65% | 70% | 315 |
| 70 to 80% | 74% | 75% | 114 |
| 80 to 90% | 81% | 88% | 8 |
| Model | Log-loss | Brier >4.5 | Brier >5.5 | Brier >6.5 |
|---|---|---|---|---|
| model | 2.252 | 0.2279 | 0.2137 | 0.1737 |
| marginal | 2.288 | 0.2472 | 0.2367 | 0.1921 |
| persistence | 2.237 | 0.2341 | 0.2185 | 0.1743 |
Read it straight: the persistence baseline is ahead of the model on full-distribution log-loss. Stated precisely, because the earlier wording overstated it — the gap is +0.0169 with a 95% paired bootstrap interval of [-0.0073, +0.0410], so persistence leads on the point estimate but the difference is not distinguishable from zero at this sample size. The model does win the P(over) Briers at every published line. The live record is the referee.
Hits
Live record
2.1% calibration error, 0.198 Brier, 1282 grades
Reliability table
| Bin | Predicted | Observed | n |
|---|---|---|---|
| 10 to 20% | 17% | 11% | 161 |
| 20 to 30% | 24% | 23% | 478 |
| 30 to 40% | 31% | 0% | 2 |
| 40 to 50% | 49% | 62% | 8 |
| 50 to 60% | 56% | 52% | 180 |
| 60 to 70% | 65% | 65% | 440 |
| 70 to 80% | 71% | 38% | 13 |
Backtest, not the live record
0.5% pooled calibration error, per-line 0.6% / 0.4%, AUC 0.56 / 0.57, 9946 holdout batter-games
Calibrated but only slightly sharper than the base rate. Projected lineups miss the actual nine about 24% of the time; that scratch rate is declared here in advance. Judge the live record against the live-gradeable subset: log-loss 1.2159, ECE 0.6%. Night-one predictions of 7/17 predate that evening's engine fixes; immutable either way.
Reliability table
| Bin | Predicted | Observed | n |
|---|---|---|---|
| 10 to 20% | 17% | 17% | 3065 |
| 20 to 30% | 24% | 24% | 6797 |
| 30 to 40% | 31% | 35% | 84 |
| 40 to 50% | 49% | 59% | 139 |
| 50 to 60% | 56% | 56% | 3595 |
| 60 to 70% | 64% | 65% | 6083 |
| 70 to 80% | 71% | 73% | 129 |
| Model | Log-loss | Brier >0.5 | Brier >1.5 |
|---|---|---|---|
| model | 1.1992 | 0.2339 | 0.1706 |
| marginal | 1.2052 | 0.2359 | 0.1724 |
| persistence | 1.2144 | 0.2380 | 0.1732 |
On this market the model beats both cheap baselines on every metric. On the K market it does not, and both statements are published with the same straight face.
Home runs
Live record
2.2% calibration error, 0.101 Brier, 641 grades
Reliability table
| Bin | Predicted | Observed | n |
|---|---|---|---|
| 0 to 10% | 8% | 2% | 162 |
| 10 to 20% | 14% | 15% | 452 |
| 20 to 30% | 22% | 15% | 26 |
| 30 to 40% | 32% | 0% | 1 |
Backtest, not the live record
0.6% calibration error, 0.590 AUC, 10462 holdout batter-games
Aggregate rate bias -0.3%: predicted 12.7% vs observed 13.0%. The mid-summer power surge outran the training window, reported straight. Live-gradeable subset (76%): log-loss 0.4268, ECE 0.4%.
Reliability table
| Bin | Predicted | Observed | n |
|---|---|---|---|
| 0 to 10% | 8% | 9% | 3548 |
| 10 to 20% | 14% | 14% | 6069 |
| 20 to 30% | 22% | 21% | 835 |
| 30 to 40% | 32% | 20% | 10 |
| Model | Log-loss | Brier | AUC |
|---|---|---|---|
| model | 0.4141 | 0.1120 | 0.590 |
| marginal | 0.4202 | 0.1134 | 0.500 |
| persistence | 0.4193 | 0.1130 | 0.572 |
What the weather features bought (goal HR-2, published 2026-07-19). Split result, reported straight: weather did not measurably improve discrimination — log-loss moved -0.0006 and Brier -0.0001, both smaller than this engine's own run-to-run variation, so they are not evidence of added skill. What it fixed was the level: the summer under-prediction shrank from -0.85pt to -0.31pt, absorbing about 64% of the seasonal bias, and calibration error fell 1.08% → 0.61%. Mostly air density — temperature, humidity and elevation combined into the quantity that actually governs carry. Caveat: the backtest trains on reanalysis (what the weather was) while live serving uses forecasts (what it was predicted to be), and day-ahead wind direction is wrong often enough to flip blowing-out/blowing-in in roughly one game in five. The bias fix rides on density, which forecasts well.
| Metric | Without wx | With wx | Delta |
|---|---|---|---|
| Log-loss | 0.4147 | 0.4141 | -0.0006 |
| Brier | 0.1121 | 0.1120 | -0.0001 |
| AUC | 0.5883 | 0.5898 | 0.0015 |
| ECE | 0.0108 | 0.0061 | -0.0047 |
| Rate bias | -0.0085 | -0.0031 | +0.0054 |
Totals (over/under 8.5)
Live record
Accumulating: 0 of 60 grades before the live curve draws.
Scored by CRPS over the whole distribution; the curve here is the 8.5 line only. What clearing the launch gate does and does not mean is stated on the Totals board.
Backtest, not the live record
| Mean CRPS | Score |
|---|---|
| Model | 2.5466 |
| Climatology (the gate) | 2.5678 |
| Null: is_home only | 2.5604 |
| Persistence | 2.6783 |
-0.0212 vs climatology, 95% CI [-0.0413, -0.0007]
Clears the pre-registered floor on 100% of 100 refits. Seed stability is refit reproducibility, not a confidence statement — the interval above is the confidence statement, and it crosses zero.
Mean total 9.06 predicted vs 9.35 actual (-0.29 runs/game, -1.5 sigma): the model tracks the training run environment and the holdout ran hotter. Disclosed because a calibration product that hides a level miss is not a calibration product.
Backtest weather is reanalysis (what actually happened); live serving uses the forecast at publish time, so live lift will be smaller and that gap is not yet measured.
One day, worked
93 graded lines on 2026-07-19Every probability below was published before first pitch and graded that night against Statcast. Nothing here was edited after the fact; the full log of all 93 lines for this day, and every other day, is in the raw artifacts. Sorted by Brier — lowest is best — so the top of this table is the engine at its most accurate and the bottom is it being wrong in public.
| Pitcher | Team | Opponent | Line | P(over) | Actual K | Result | Brier |
|---|---|---|---|---|---|---|---|
| Jacob Lopez | ATH | WSH | 6.5 | 13% | 6 | under | 0.017 |
| Germán Márquez | SD | KC | 6.5 | 13% | 2 | under | 0.018 |
| Andre Pallante | STL | AZ | 6.5 | 14% | 3 | under | 0.020 |
| Eduardo Rodriguez | AZ | STL | 6.5 | 17% | 3 | under | 0.030 |
| Grant Holmes | ATL | TEX | 6.5 | 17% | 2 | under | 0.030 |
| Ryan Yarbrough | NYY | LAD | 6.5 | 18% | 1 | under | 0.034 |
| Ryan Johnson | LAA | DET | 6.5 | 19% | 5 | under | 0.035 |
| Shane McClanahan | TB | BOS | 6.5 | 19% | 3 | under | 0.036 |
| Zebby Matthews | MIN | CHC | 6.5 | 19% | 4 | under | 0.037 |
| Logan Gilbert | SEA | SF | 4.5 | 79% | 10 | over | 0.045 |
| Noah Cameron | KC | SD | 6.5 | 22% | 1 | under | 0.047 |
| Ryan Feltner | COL | CIN | 6.5 | 22% | 3 | under | 0.047 |
A Brier score is the squared error of a probability: 0 is perfect, 0.25 is a coin flip, 1 is confidently wrong.
Pre-registered goals
dated, falsifiable, failures shownEvery goal below was written down with its threshold and
its deadline before the result was known, in
docs/BENCHMARKS.md. A failed goal is marked FAILED and stays
visible; it is never deleted or quietly re-scoped. That rule is the whole
point — a record that only kept its wins would prove nothing.
| Goal | Standing |
|---|---|
| T-1 Totals launch gate: backtest CRPS < 2.5678 | MET — 2.5466, clearing on 100 of 100 refits. But the gate is a FLOOR: a model knowing only which team is home scores 2.5604 against it, so clearing it means "not worse than knowing nothing". |
| W-1 Winners launch gate: backtest Brier < 0.2488 | FAILED — not published. Clears on only 73 of 100 refits, and AUC 0.539 against ~0.60 for elite public models. Root cause is structural: deriving win probability from two runs-scored marginals fights the ninth-inning truncation. Held until a v1.1 fix clears it robustly. |
| K-1 Live K log-loss below the persistence baseline | NOT MET in backtest, and the gap is not distinguishable from zero: 2.2587 vs 2.2417, 95% CI [−0.0103, +0.0371]. The honest statement is that we cannot separate the model from persistence, not that persistence wins. |
| HR-2 Disclose what the weather features bought | MET — and the answer was mostly nothing: discrimination moved inside the noise band. What weather bought was calibration, absorbing 64% of the summer under-prediction. |
Full table with sample sizes, deadlines and method notes:
docs/BENCHMARKS.md in the repository.
Game prep
The scouting packet as a page, one per game: both starters with percentile sliders, arsenal tables, movement and location charts, lineup grids fused with the engine's published board numbers, and bullpen availability. Same data modules as everything above.
- 2026 07 20 CWS at TEX
- 2026 07 19 CWS at TOR
- 2026 07 18 CWS at TOR
- 2026 07 17 CWS at TOR
- CWS at TOR 2026 07 17 PDF, series format, archived
Graded log, strikeouts
Every K prediction ever published: probability, outcome, Brier. Batter-market logs live in the data archive as full CSVs.
| Date | Pitcher | Matchup | K | P(K>4.5) | P(K>5.5) | P(K>6.5) | Brier |
|---|---|---|---|---|---|---|---|
| 2026-07-19 | Alan Rangel | PHI vs NYM | 4 | 57% U | 39% U | 28% U | 0.183 |
| 2026-07-19 | Andre Pallante | STL vs AZ | 3 | 40% U | 24% U | 14% U | 0.080 |
| 2026-07-19 | Brandon Young | BAL vs HOU | 7 | 51% O | 31% O | 16% O | 0.475 |
| 2026-07-19 | Cam Schlittler | NYY vs LAD | 8 | 73% O | 61% O | 50% O | 0.159 |
| 2026-07-19 | Casey Mize | DET vs LAA | 5 | 70% O | 49% U | 34% U | 0.149 |
| 2026-07-19 | Eduardo Rodriguez | AZ vs STL | 3 | 47% U | 33% U | 17% U | 0.120 |
| 2026-07-19 | Eury Pérez | MIA vs MIL | 9 | 65% O | 50% O | 33% O | 0.274 |
| 2026-07-19 | Foster Griffin | WSH vs ATH | 2 | 67% U | 50% U | 32% U | 0.266 |
| 2026-07-19 | Germán Márquez | SD vs KC | 2 | 41% U | 24% U | 13% U | 0.081 |
| 2026-07-19 | Grant Holmes | ATL vs TEX | 2 | 47% U | 29% U | 17% U | 0.111 |
| 2026-07-19 | Hunter Brown | HOU vs BAL | 4 | 65% U | 47% U | 27% U | 0.238 |
| 2026-07-19 | Hunter Greene | CIN vs COL | 7 | 66% O | 54% O | 36% O | 0.248 |
| 2026-07-19 | Jacob Lopez | ATH vs WSH | 6 | 40% O | 21% O | 13% U | 0.333 |
| 2026-07-19 | Joey Cantillo | CLE vs PIT | 7 | 66% O | 55% O | 40% O | 0.227 |
| 2026-07-19 | Logan Gilbert | SEA vs SF | 10 | 79% O | 67% O | 51% O | 0.131 |
| 2026-07-19 | Nathan Eovaldi | TEX vs ATL | 2 | 74% U | 62% U | 44% U | 0.375 |
| 2026-07-19 | Noah Cameron | KC vs SD | 1 | 62% U | 44% U | 22% U | 0.206 |
| 2026-07-19 | Nolan McLean | NYM vs PHI | 10 | 69% O | 50% O | 36% O | 0.255 |
| 2026-07-19 | Paul Skenes | PIT vs CLE | 8 | 70% O | 59% O | 43% O | 0.194 |
| 2026-07-19 | Robbie Ray | SF vs SEA | 6 | 63% O | 49% O | 29% U | 0.160 |
| 2026-07-19 | Robert Gasser | MIL vs MIA | 5 | 56% O | 39% U | 23% U | 0.133 |
| 2026-07-19 | Ryan Feltner | COL vs CIN | 3 | 56% U | 35% U | 22% U | 0.162 |
| 2026-07-19 | Ryan Johnson | LAA vs DET | 5 | 57% O | 39% U | 19% U | 0.124 |
| 2026-07-19 | Ryan Yarbrough | NYY vs LAD | 1 | 40% U | 25% U | 18% U | 0.085 |
| 2026-07-19 | Sean Burke | CWS vs TOR | 5 | 66% O | 47% U | 29% U | 0.142 |
| 2026-07-19 | Shane McClanahan | TB vs BOS | 3 | 54% U | 37% U | 19% U | 0.157 |
| 2026-07-19 | Shota Imanaga | CHC vs MIN | 4 | 67% U | 49% U | 35% U | 0.269 |
| 2026-07-19 | Sonny Gray | BOS vs TB | 5 | 54% O | 42% U | 23% U | 0.147 |
| 2026-07-19 | Trey Yesavage | TOR vs CWS | 9 | 56% O | 39% O | 20% O | 0.400 |
| 2026-07-19 | Yoshinobu Yamamoto | LAD vs NYY | 7 | 73% O | 56% O | 35% O | 0.229 |
| 2026-07-19 | Zebby Matthews | MIN vs CHC | 4 | 59% U | 39% U | 19% U | 0.181 |
| 2026-07-18 | Brandon Pfaadt | AZ vs STL | 3 | 35% U | 23% U | 13% U | 0.065 |
| 2026-07-18 | Bryan Woo | SEA vs SF | 7 | 60% O | 45% O | 30% O | 0.319 |
| 2026-07-18 | Davis Martin | CWS vs TOR | 5 | 46% O | 29% U | 16% U | 0.132 |
| 2026-07-18 | Dustin May | STL vs AZ | 6 | 48% O | 33% O | 19% U | 0.250 |
| 2026-07-18 | Emmet Sheehan | LAD vs NYY | voided: postponed | ||||
| 2026-07-18 | Grayson Rodriguez | LAA vs DET | 3 | 47% U | 29% U | 15% U | 0.110 |
| 2026-07-18 | Griffin Canning | SD vs KC | 4 | 53% U | 37% U | 19% U | 0.151 |
| 2026-07-18 | Ian Seymour | TB vs BOS | 4 | 43% U | 33% U | 22% U | 0.114 |
| 2026-07-18 | J.T. Ginn | ATH vs WSH | 7 | 58% O | 42% O | 29% O | 0.341 |
| 2026-07-18 | Jesús Luzardo | PHI vs NYM | 7 | 77% O | 65% O | 49% O | 0.147 |
| 2026-07-18 | Logan Allen | CLE vs PIT | 3 | 51% U | 35% U | 26% U | 0.152 |
| 2026-07-18 | Logan Webb | SF vs SEA | 5 | 49% O | 31% U | 21% U | 0.134 |
| 2026-07-18 | MacKenzie Gore | TEX vs ATL | 7 | 64% O | 48% O | 35% O | 0.273 |
| 2026-07-18 | Matthew Boyd | CHC vs MIN | 4 | 55% U | 36% U | 21% U | 0.159 |
| 2026-07-18 | Max Meyer | MIA vs MIL | 5 | 64% O | 46% U | 32% U | 0.151 |
| 2026-07-18 | Patrick Sandoval | BOS vs TB | 5 | 49% O | 30% U | 15% U | 0.124 |
| 2026-07-18 | Rhett Lowder | CIN vs COL | 2 | 45% U | 29% U | 14% U | 0.101 |
| 2026-07-18 | Ryan Weathers | NYY vs LAD | voided: postponed | ||||
| 2026-07-18 | Sean Manaea | NYM vs PHI | 7 | 58% O | 39% O | 18% O | 0.412 |
| 2026-07-18 | Shane Bieber | TOR vs CWS | 6 | 45% O | 30% O | 17% U | 0.270 |
| 2026-07-18 | Shane Drohan | MIL vs MIA | 9 | 49% O | 33% O | 19% O | 0.452 |
| 2026-07-18 | Spencer Arrighetti | HOU vs BAL | 6 | 68% O | 51% O | 30% U | 0.144 |
| 2026-07-18 | Taj Bradley | MIN vs CHC | 6 | 74% O | 51% O | 39% U | 0.152 |
| 2026-07-18 | Tarik Skubal | DET vs LAA | 9 | 71% O | 58% O | 45% O | 0.191 |
| 2026-07-18 | Tomoyuki Sugano | COL vs CIN | 3 | 54% U | 32% U | 20% U | 0.145 |
| 2026-07-18 | Trevor Rogers | BAL vs HOU | 8 | 48% O | 30% O | 16% O | 0.493 |
| 2026-07-18 | Zack Littell | WSH vs ATH | 4 | 38% U | 22% U | 12% U | 0.070 |
| 2026-07-17 | Anthony Kay | CWS vs TOR | 5 | 38% O | 23% U | 12% U | 0.152 |
| 2026-07-17 | Bailey Ober | MIN vs CHC | 7 | 44% O | 28% O | 12% O | 0.534 |
| 2026-07-17 | Brady Singer | CIN vs COL | 6 | 51% O | 40% O | 19% U | 0.214 |
| 2026-07-17 | Bryce Miller | SEA vs SF | 6 | 55% O | 42% O | 30% U | 0.212 |
| 2026-07-17 | Cade Cavalli | WSH vs ATH | 9 | 64% O | 50% O | 37% O | 0.261 |
| 2026-07-17 | Cal Quantrill | TEX vs ATL | 3 | 35% U | 21% U | 10% U | 0.059 |
| 2026-07-17 | Chris Sale | ATL vs TEX | 6 | 66% O | 51% O | 33% U | 0.156 |
| 2026-07-17 | Colin Rea | CHC vs MIN | 6 | 32% O | 18% O | 9% U | 0.381 |
| 2026-07-17 | Dean Kremer | BAL vs HOU | 5 | 64% O | 49% U | 32% U | 0.157 |
| 2026-07-17 | Eduardo Rivera | BOS vs TB | 3 | 38% U | 24% U | 15% U | 0.074 |
| 2026-07-17 | Gabriel Hughes | COL vs CIN | 6 | 68% O | 48% O | 29% U | 0.153 |
| 2026-07-17 | Gage Jump | ATH vs WSH | 8 | 73% O | 51% O | 35% O | 0.246 |
| 2026-07-17 | Gavin Williams | CLE vs PIT | voided: postponed | ||||
| 2026-07-17 | Gerrit Cole | NYY vs LAD | 8 | 56% O | 39% O | 20% O | 0.405 |
| 2026-07-17 | Griffin Jax | TB vs BOS | voided: late | ||||
| 2026-07-17 | Jake Bennett | BOS vs TB | voided: late | ||||
| 2026-07-17 | Jared Jones | PIT vs CLE | voided: postponed | ||||
| 2026-07-17 | Landen Roupp | SF vs SEA | 2 | 65% U | 45% U | 31% U | 0.240 |
| 2026-07-17 | Logan Henderson | MIL vs MIA | 4 | 46% U | 35% U | 24% U | 0.130 |
| 2026-07-17 | Mason Englert | TB vs BOS | 4 | 50% U | 34% U | 16% U | 0.132 |
| 2026-07-17 | Merrill Kelly | AZ vs STL | 3 | 41% U | 29% U | 14% U | 0.090 |
| 2026-07-17 | Michael King | SD vs KC | 4 | 42% U | 27% U | 16% U | 0.092 |
| 2026-07-17 | Michael McGreevy | STL vs AZ | 5 | 32% O | 18% U | 8% U | 0.169 |
| 2026-07-17 | Peter Lambert | HOU vs BAL | 10 | 72% O | 54% O | 37% O | 0.226 |
| 2026-07-17 | Reid Detmers | LAA vs DET | 7 | 66% O | 49% O | 34% O | 0.269 |
| 2026-07-17 | Roki Sasaki | LAD vs NYY | 5 | 58% O | 40% U | 25% U | 0.133 |
| 2026-07-17 | Sandy Alcantara | MIA vs MIL | 7 | 52% O | 37% O | 21% O | 0.416 |
| 2026-07-17 | Seth Lugo | KC vs SD | 3 | 47% U | 28% U | 18% U | 0.112 |
| 2026-07-17 | Spencer Miles | TOR vs CWS | 4 | 33% U | 17% U | 11% U | 0.052 |
| 2026-07-17 | Troy Melton | DET vs LAA | 9 | 53% O | 39% O | 23% O | 0.396 |
Data and methods
Every prediction and grade is downloadable from the archive: dated full-distribution JSON per market, complete grades CSVs, and the raw backtest pairs behind every curve on this page. Files are committed to git at publish time, so timestamps carry a witness.
Sources and caveats: Statcast via pybaseball; grading runs against Statcast play-by-play at grade time. Model features use trailing windows that may span the offseason, backward-only. Openers (nine or fewer batters faced) are excluded from the strikeout model's population. Park HR context blends true park effects with the home roster's power and is used as signal, not sold as a clean park factor.
Probabilities, not picks. This is not betting advice. Page generated 2026-07-20 15:45 UTC.