The public, graded record.

Player-level probabilities for every MLB slate: pitcher strikeouts, batter hits, batter home runs. Published before first pitch, graded against Statcast that night, never edited. The goal is calibration, not beating the market. When this engine says 30%, it should happen about 30% of the time, and this page is the test. Probabilities, not picks. Not betting advice.

3graded markets
601predictions published for 2026-07-19
1366graded, all-time
336voided, shown not deleted
3days live

Void rules, stated in advance: published late voids. Postponed voids. A pitcher who does not start, or a batter outside the starting nine, voids. Voids stay on the record.

Calibration, the whole point

When the engine says 30%, does it happen 30% of the time? Points on the dashed diagonal are perfectly calibrated. Above it the model was too low, below it too high. Bins thinner than 25 samples live in the tables, not the curves.

Strikeouts

Live record

9.5% calibration error, 0.207 Brier, 252 grades

predicted 16%, observed 22% (n=37)predicted 25%, observed 18% (n=38)predicted 35%, observed 44% (n=55)predicted 46%, observed 52% (n=48)predicted 55%, observed 74% (n=34)predicted 65%, observed 77% (n=26)00252550507575100100predicted probability (%)observed frequency (%)
Reliability table
BinPredictedObservedn
0 to 10%9%0%3
10 to 20%16%22%37
20 to 30%25%18%38
30 to 40%35%44%55
40 to 50%46%52%48
50 to 60%55%74%34
60 to 70%65%77%26
70 to 80%74%91%11

Engine change, 2026-07-18 → 2026-07-19. For one day the strikeout model carried an extra feature — the trailing strikeout rate of the opposing starting nine. An adversarial review found it was trained on the lineup each pitcher actually faced but served a posted-or-projected one, and that on the projected path it made the model worse than having no such feature at all; its apparent gain was never distinguishable from zero. It was removed on 2026-07-19 and the model retrained without it. 81 graded line-results (27 starts, all from 7/18) plus the 7/19 board were produced by that version. Those predictions stand exactly as published and stay in the running totals — they were made in good faith before first pitch, and deleting inconvenient history is the one thing this record will not do. Method: docs/K-FEATURE-VERDICT.md.

Backtest, not the live record

2.2% pooled calibration error, per-line 2.2% / 3.9% / 2.6%, 0.205 Brier, 1090 holdout starts

Trained on real starts before 2026-06-01, openers excluded. The three lines per start are nested views of one outcome, so the pooled number can flatter. It earns nothing until the live record proves it.

predicted 15%, observed 14% (n=493)predicted 25%, observed 23% (n=657)predicted 35%, observed 37% (n=606)predicted 45%, observed 47% (n=614)predicted 55%, observed 56% (n=462)predicted 65%, observed 70% (n=315)predicted 74%, observed 75% (n=114)00252550507575100100predicted probability (%)observed frequency (%)
Reliability table
BinPredictedObservedn
0 to 10%10%0%1
10 to 20%15%14%493
20 to 30%25%23%657
30 to 40%35%37%606
40 to 50%45%47%614
50 to 60%55%56%462
60 to 70%65%70%315
70 to 80%74%75%114
80 to 90%81%88%8
ModelLog-lossBrier >4.5Brier >5.5Brier >6.5
model2.2520.22790.21370.1737
marginal2.2880.24720.23670.1921
persistence2.2370.23410.21850.1743

Read it straight: the persistence baseline is ahead of the model on full-distribution log-loss. Stated precisely, because the earlier wording overstated it — the gap is +0.0169 with a 95% paired bootstrap interval of [-0.0073, +0.0410], so persistence leads on the point estimate but the difference is not distinguishable from zero at this sample size. The model does win the P(over) Briers at every published line. The live record is the referee.

Hits

Live record

2.1% calibration error, 0.198 Brier, 1282 grades

predicted 17%, observed 11% (n=161)predicted 24%, observed 23% (n=478)predicted 56%, observed 52% (n=180)predicted 65%, observed 65% (n=440)00252550507575100100predicted probability (%)observed frequency (%)
Reliability table
BinPredictedObservedn
10 to 20%17%11%161
20 to 30%24%23%478
30 to 40%31%0%2
40 to 50%49%62%8
50 to 60%56%52%180
60 to 70%65%65%440
70 to 80%71%38%13

Backtest, not the live record

0.5% pooled calibration error, per-line 0.6% / 0.4%, AUC 0.56 / 0.57, 9946 holdout batter-games

Calibrated but only slightly sharper than the base rate. Projected lineups miss the actual nine about 24% of the time; that scratch rate is declared here in advance. Judge the live record against the live-gradeable subset: log-loss 1.2159, ECE 0.6%. Night-one predictions of 7/17 predate that evening's engine fixes; immutable either way.

predicted 17%, observed 17% (n=3065)predicted 24%, observed 24% (n=6797)predicted 31%, observed 35% (n=84)predicted 49%, observed 59% (n=139)predicted 56%, observed 56% (n=3595)predicted 64%, observed 65% (n=6083)predicted 71%, observed 73% (n=129)00252550507575100100predicted probability (%)observed frequency (%)
Reliability table
BinPredictedObservedn
10 to 20%17%17%3065
20 to 30%24%24%6797
30 to 40%31%35%84
40 to 50%49%59%139
50 to 60%56%56%3595
60 to 70%64%65%6083
70 to 80%71%73%129
ModelLog-lossBrier >0.5Brier >1.5
model1.19920.23390.1706
marginal1.20520.23590.1724
persistence1.21440.23800.1732

On this market the model beats both cheap baselines on every metric. On the K market it does not, and both statements are published with the same straight face.

Home runs

Live record

2.2% calibration error, 0.101 Brier, 641 grades

predicted 8%, observed 2% (n=162)predicted 14%, observed 15% (n=452)predicted 22%, observed 15% (n=26)00252550507575100100predicted probability (%)observed frequency (%)
Reliability table
BinPredictedObservedn
0 to 10%8%2%162
10 to 20%14%15%452
20 to 30%22%15%26
30 to 40%32%0%1

Backtest, not the live record

0.6% calibration error, 0.590 AUC, 10462 holdout batter-games

Aggregate rate bias -0.3%: predicted 12.7% vs observed 13.0%. The mid-summer power surge outran the training window, reported straight. Live-gradeable subset (76%): log-loss 0.4268, ECE 0.4%.

predicted 8%, observed 9% (n=3548)predicted 14%, observed 14% (n=6069)predicted 22%, observed 21% (n=835)00252550507575100100predicted probability (%)observed frequency (%)
Reliability table
BinPredictedObservedn
0 to 10%8%9%3548
10 to 20%14%14%6069
20 to 30%22%21%835
30 to 40%32%20%10
ModelLog-lossBrierAUC
model0.41410.11200.590
marginal0.42020.11340.500
persistence0.41930.11300.572

What the weather features bought (goal HR-2, published 2026-07-19). Split result, reported straight: weather did not measurably improve discrimination — log-loss moved -0.0006 and Brier -0.0001, both smaller than this engine's own run-to-run variation, so they are not evidence of added skill. What it fixed was the level: the summer under-prediction shrank from -0.85pt to -0.31pt, absorbing about 64% of the seasonal bias, and calibration error fell 1.08% → 0.61%. Mostly air density — temperature, humidity and elevation combined into the quantity that actually governs carry. Caveat: the backtest trains on reanalysis (what the weather was) while live serving uses forecasts (what it was predicted to be), and day-ahead wind direction is wrong often enough to flip blowing-out/blowing-in in roughly one game in five. The bias fix rides on density, which forecasts well.

MetricWithout wxWith wxDelta
Log-loss0.41470.4141-0.0006
Brier0.11210.1120-0.0001
AUC0.58830.58980.0015
ECE0.01080.0061-0.0047
Rate bias-0.0085-0.0031+0.0054

Totals (over/under 8.5)

Live record

Accumulating: 0 of 60 grades before the live curve draws.

Scored by CRPS over the whole distribution; the curve here is the 8.5 line only. What clearing the launch gate does and does not mean is stated on the Totals board.

Backtest, not the live record

Holdout 558 games, 2026-06-01 to 2026-07-16
Mean CRPSScore
Model2.5466
Climatology (the gate)2.5678
Null: is_home only2.5604
Persistence2.6783

-0.0212 vs climatology, 95% CI [-0.0413, -0.0007]

Clears the pre-registered floor on 100% of 100 refits. Seed stability is refit reproducibility, not a confidence statement — the interval above is the confidence statement, and it crosses zero.

Mean total 9.06 predicted vs 9.35 actual (-0.29 runs/game, -1.5 sigma): the model tracks the training run environment and the holdout ran hotter. Disclosed because a calibration product that hides a level miss is not a calibration product.

Backtest weather is reanalysis (what actually happened); live serving uses the forecast at publish time, so live lift will be smaller and that gap is not yet measured.

One day, worked

93 graded lines on 2026-07-19

Every probability below was published before first pitch and graded that night against Statcast. Nothing here was edited after the fact; the full log of all 93 lines for this day, and every other day, is in the raw artifacts. Sorted by Brier — lowest is best — so the top of this table is the engine at its most accurate and the bottom is it being wrong in public.

PitcherTeamOpponent LineP(over) Actual KResult Brier
Jacob LopezATHWSH6.513%6under0.017
Germán MárquezSDKC6.513%2under0.018
Andre PallanteSTLAZ6.514%3under0.020
Eduardo RodriguezAZSTL6.517%3under0.030
Grant HolmesATLTEX6.517%2under0.030
Ryan YarbroughNYYLAD6.518%1under0.034
Ryan JohnsonLAADET6.519%5under0.035
Shane McClanahanTBBOS6.519%3under0.036
Zebby MatthewsMINCHC6.519%4under0.037
Logan GilbertSEASF4.579%10over0.045
Noah CameronKCSD6.522%1under0.047
Ryan FeltnerCOLCIN6.522%3under0.047

A Brier score is the squared error of a probability: 0 is perfect, 0.25 is a coin flip, 1 is confidently wrong.

Pre-registered goals

dated, falsifiable, failures shown

Every goal below was written down with its threshold and its deadline before the result was known, in docs/BENCHMARKS.md. A failed goal is marked FAILED and stays visible; it is never deleted or quietly re-scoped. That rule is the whole point — a record that only kept its wins would prove nothing.

GoalStanding
T-1 Totals launch gate: backtest CRPS < 2.5678 MET — 2.5466, clearing on 100 of 100 refits. But the gate is a FLOOR: a model knowing only which team is home scores 2.5604 against it, so clearing it means "not worse than knowing nothing".
W-1 Winners launch gate: backtest Brier < 0.2488 FAILED — not published. Clears on only 73 of 100 refits, and AUC 0.539 against ~0.60 for elite public models. Root cause is structural: deriving win probability from two runs-scored marginals fights the ninth-inning truncation. Held until a v1.1 fix clears it robustly.
K-1 Live K log-loss below the persistence baseline NOT MET in backtest, and the gap is not distinguishable from zero: 2.2587 vs 2.2417, 95% CI [−0.0103, +0.0371]. The honest statement is that we cannot separate the model from persistence, not that persistence wins.
HR-2 Disclose what the weather features bought MET — and the answer was mostly nothing: discrimination moved inside the noise band. What weather bought was calibration, absorbing 64% of the summer under-prediction.

Full table with sample sizes, deadlines and method notes: docs/BENCHMARKS.md in the repository.

Game prep

The scouting packet as a page, one per game: both starters with percentile sliders, arsenal tables, movement and location charts, lineup grids fused with the engine's published board numbers, and bullpen availability. Same data modules as everything above.

Graded log, strikeouts

Every K prediction ever published: probability, outcome, Brier. Batter-market logs live in the data archive as full CSVs.

DatePitcherMatchup KP(K>4.5)P(K>5.5)P(K>6.5)Brier
2026-07-19Alan RangelPHI vs NYM457% U39% U28% U0.183
2026-07-19Andre PallanteSTL vs AZ340% U24% U14% U0.080
2026-07-19Brandon YoungBAL vs HOU751% O31% O16% O0.475
2026-07-19Cam SchlittlerNYY vs LAD873% O61% O50% O0.159
2026-07-19Casey MizeDET vs LAA570% O49% U34% U0.149
2026-07-19Eduardo RodriguezAZ vs STL347% U33% U17% U0.120
2026-07-19Eury PérezMIA vs MIL965% O50% O33% O0.274
2026-07-19Foster GriffinWSH vs ATH267% U50% U32% U0.266
2026-07-19Germán MárquezSD vs KC241% U24% U13% U0.081
2026-07-19Grant HolmesATL vs TEX247% U29% U17% U0.111
2026-07-19Hunter BrownHOU vs BAL465% U47% U27% U0.238
2026-07-19Hunter GreeneCIN vs COL766% O54% O36% O0.248
2026-07-19Jacob LopezATH vs WSH640% O21% O13% U0.333
2026-07-19Joey CantilloCLE vs PIT766% O55% O40% O0.227
2026-07-19Logan GilbertSEA vs SF1079% O67% O51% O0.131
2026-07-19Nathan EovaldiTEX vs ATL274% U62% U44% U0.375
2026-07-19Noah CameronKC vs SD162% U44% U22% U0.206
2026-07-19Nolan McLeanNYM vs PHI1069% O50% O36% O0.255
2026-07-19Paul SkenesPIT vs CLE870% O59% O43% O0.194
2026-07-19Robbie RaySF vs SEA663% O49% O29% U0.160
2026-07-19Robert GasserMIL vs MIA556% O39% U23% U0.133
2026-07-19Ryan FeltnerCOL vs CIN356% U35% U22% U0.162
2026-07-19Ryan JohnsonLAA vs DET557% O39% U19% U0.124
2026-07-19Ryan YarbroughNYY vs LAD140% U25% U18% U0.085
2026-07-19Sean BurkeCWS vs TOR566% O47% U29% U0.142
2026-07-19Shane McClanahanTB vs BOS354% U37% U19% U0.157
2026-07-19Shota ImanagaCHC vs MIN467% U49% U35% U0.269
2026-07-19Sonny GrayBOS vs TB554% O42% U23% U0.147
2026-07-19Trey YesavageTOR vs CWS956% O39% O20% O0.400
2026-07-19Yoshinobu YamamotoLAD vs NYY773% O56% O35% O0.229
2026-07-19Zebby MatthewsMIN vs CHC459% U39% U19% U0.181
2026-07-18Brandon PfaadtAZ vs STL335% U23% U13% U0.065
2026-07-18Bryan WooSEA vs SF760% O45% O30% O0.319
2026-07-18Davis MartinCWS vs TOR546% O29% U16% U0.132
2026-07-18Dustin MaySTL vs AZ648% O33% O19% U0.250
2026-07-18Emmet SheehanLAD vs NYYvoided: postponed
2026-07-18Grayson RodriguezLAA vs DET347% U29% U15% U0.110
2026-07-18Griffin CanningSD vs KC453% U37% U19% U0.151
2026-07-18Ian SeymourTB vs BOS443% U33% U22% U0.114
2026-07-18J.T. GinnATH vs WSH758% O42% O29% O0.341
2026-07-18Jesús LuzardoPHI vs NYM777% O65% O49% O0.147
2026-07-18Logan AllenCLE vs PIT351% U35% U26% U0.152
2026-07-18Logan WebbSF vs SEA549% O31% U21% U0.134
2026-07-18MacKenzie GoreTEX vs ATL764% O48% O35% O0.273
2026-07-18Matthew BoydCHC vs MIN455% U36% U21% U0.159
2026-07-18Max MeyerMIA vs MIL564% O46% U32% U0.151
2026-07-18Patrick SandovalBOS vs TB549% O30% U15% U0.124
2026-07-18Rhett LowderCIN vs COL245% U29% U14% U0.101
2026-07-18Ryan WeathersNYY vs LADvoided: postponed
2026-07-18Sean ManaeaNYM vs PHI758% O39% O18% O0.412
2026-07-18Shane BieberTOR vs CWS645% O30% O17% U0.270
2026-07-18Shane DrohanMIL vs MIA949% O33% O19% O0.452
2026-07-18Spencer ArrighettiHOU vs BAL668% O51% O30% U0.144
2026-07-18Taj BradleyMIN vs CHC674% O51% O39% U0.152
2026-07-18Tarik SkubalDET vs LAA971% O58% O45% O0.191
2026-07-18Tomoyuki SuganoCOL vs CIN354% U32% U20% U0.145
2026-07-18Trevor RogersBAL vs HOU848% O30% O16% O0.493
2026-07-18Zack LittellWSH vs ATH438% U22% U12% U0.070
2026-07-17Anthony KayCWS vs TOR538% O23% U12% U0.152
2026-07-17Bailey OberMIN vs CHC744% O28% O12% O0.534
2026-07-17Brady SingerCIN vs COL651% O40% O19% U0.214
2026-07-17Bryce MillerSEA vs SF655% O42% O30% U0.212
2026-07-17Cade CavalliWSH vs ATH964% O50% O37% O0.261
2026-07-17Cal QuantrillTEX vs ATL335% U21% U10% U0.059
2026-07-17Chris SaleATL vs TEX666% O51% O33% U0.156
2026-07-17Colin ReaCHC vs MIN632% O18% O9% U0.381
2026-07-17Dean KremerBAL vs HOU564% O49% U32% U0.157
2026-07-17Eduardo RiveraBOS vs TB338% U24% U15% U0.074
2026-07-17Gabriel HughesCOL vs CIN668% O48% O29% U0.153
2026-07-17Gage JumpATH vs WSH873% O51% O35% O0.246
2026-07-17Gavin WilliamsCLE vs PITvoided: postponed
2026-07-17Gerrit ColeNYY vs LAD856% O39% O20% O0.405
2026-07-17Griffin JaxTB vs BOSvoided: late
2026-07-17Jake BennettBOS vs TBvoided: late
2026-07-17Jared JonesPIT vs CLEvoided: postponed
2026-07-17Landen RouppSF vs SEA265% U45% U31% U0.240
2026-07-17Logan HendersonMIL vs MIA446% U35% U24% U0.130
2026-07-17Mason EnglertTB vs BOS450% U34% U16% U0.132
2026-07-17Merrill KellyAZ vs STL341% U29% U14% U0.090
2026-07-17Michael KingSD vs KC442% U27% U16% U0.092
2026-07-17Michael McGreevySTL vs AZ532% O18% U8% U0.169
2026-07-17Peter LambertHOU vs BAL1072% O54% O37% O0.226
2026-07-17Reid DetmersLAA vs DET766% O49% O34% O0.269
2026-07-17Roki SasakiLAD vs NYY558% O40% U25% U0.133
2026-07-17Sandy AlcantaraMIA vs MIL752% O37% O21% O0.416
2026-07-17Seth LugoKC vs SD347% U28% U18% U0.112
2026-07-17Spencer MilesTOR vs CWS433% U17% U11% U0.052
2026-07-17Troy MeltonDET vs LAA953% O39% O23% O0.396

Data and methods

Every prediction and grade is downloadable from the archive: dated full-distribution JSON per market, complete grades CSVs, and the raw backtest pairs behind every curve on this page. Files are committed to git at publish time, so timestamps carry a witness.

Sources and caveats: Statcast via pybaseball; grading runs against Statcast play-by-play at grade time. Model features use trailing windows that may span the offseason, backward-only. Openers (nine or fewer batters faced) are excluded from the strikeout model's population. Park HR context blends true park effects with the home roster's power and is used as signal, not sold as a clean park factor.

Probabilities, not picks. This is not betting advice. Page generated 2026-07-20 15:45 UTC.