> > >
On June 11 we locked fourteen call-up probabilities. Today the window closed. Five of the fourteen reached the majors, and the model scored a Brier of 0.2375 against a naive baseline of 0.2296.
The naive baseline is what you get by ignoring the model entirely and predicting the same number, the observed base rate, for every player. Beating it is the minimum bar a forecasting model has to clear. On this cohort, v0.5 did not clear it.
| Player | Org | CUP | Result | What happened |
|---|---|---|---|---|
| Thomas White | MIA | 62% | NO | Shut down for the season in August with a left shoulder capsular sprain. Miami chose to prepare him for 2027. He was on the minor league IL at run date. |
| Charlie Condon | COL | 48% | NO | Never left Triple-A Albuquerque. Hit .337/.462/.814 in June, then fell to a .758 OPS in July and .695 in August as pitchers adjusted. |
| Hagen Smith | CWS | 45% | YES | Two hitless scoreless relief innings in his debut against Cleveland. |
| Owen Caissie | MIA | 33% | YES | Appeared in the majors throughout the window as Miami's primary right fielder. See the data-quality note: he was an active major leaguer at run date and should not have been scored as a call-up candidate. |
| Emmanuel Rodriguez | MIN | 33% | NO | Stayed at Triple-A St. Paul. Minnesota promoted Jenkins ahead of him after the Buxton injury. |
| Kaelen Culpepper | MIN | 33% | YES | Homered in his second major league plate appearance after returning from a recurring glute strain. |
| Walker Jenkins | MIN | 23% | YES | Promoted after a Byron Buxton injury. Singled in his debut. |
| Kade Anderson | SEA | 23% | YES | Started against the Cubs: 5.2 innings, three earned runs, five strikeouts. |
| Ryan Sloan | SEA | 23% | NO | Finished the season at Double-A Arkansas. Nineteen years old. |
| Franklin Arias | BOS | 22% | NO | Strong full season at Double-A but Boston never made the move. |
| Josue De Paula | LAD | 22% | NO | Hit .315 at Double-A Tulsa. Discussed as a September call-up but the Dodgers logjam held, exactly the 0.88 org coefficient at work. |
| Aiva Arquette | MIA | 22% | NO | Two separate IL stints limited him to 49 games between High-A and Double-A. Headed to the Arizona Fall League to make up at-bats. |
| Eduardo Tait | MIN | 12% | NO | Still at High-A. Catcher timelines are the slowest in the model for a reason. |
| Sebastian Walcott | TEX | 5% | NO | UCL surgery announced in February put the whole season in jeopardy. Correctly the lowest number in the cohort. |
Bucket ordering broke. The 30-44 band resolved at a higher rate than the 45-plus band, the first non-monotonic result the model has produced. Mean prediction was 29.0 percent against an observed base rate of 35.7 percent, so v0.5 is still running low overall, the same direction of error the June checkpoint found and the v0.5 patch was meant to fix. The single largest penalty came from underprediction, not overprediction: Jenkins and Anderson both debuted at 23 percent.
Read that carefully, because it is the opposite of the story people expect from a bad checkpoint. The model was not wildly overconfident. Its average prediction was 29.0 percent against an observed rate of 35.7 percent, which means it was still running low, the same direction of error the June checkpoint found and the same one the v0.5 patch was built to fix. The patch moved the number in the right direction and did not move it far enough.
The one place it overshot was the top, and that was a single player. Thomas White at 62 percent contributed the largest overprediction penalty in the cohort, and he was on the injured list on the day we scored him.
Thomas White, 62 percent, injury-status blind spot. Scored 62 percent, the largest number the model has published, while already on the minor league injured list at run date. Miami shut him down in August. The model has no injury-status feature and this is the second time that gap has cost it.
Walker Jenkins, 23 percent, opportunity underweight. Promoted after a Byron Buxton injury opened center field. The model prices org opportunity from season-long context and does not react to in-window injuries to blockers, which is precisely what the live PIX layer was built to catch and the batch model is not.
Kade Anderson, 23 percent, level-proximity underweight. Scored at Double-A and jumped straight to a major league start. The level-proximity feature penalizes non-Triple-A players heavily; for elite arms on a contending club that penalty appears too steep.
Charlie Condon, 48 percent, performance-trend miss. Second-highest number in the cohort. His OPS fell from 1.276 in June to .758 in July and .695 in August. The 30-day performance feature is measured at run date only and cannot see forward decay.
Owen Caissie, 33 percent, data quality. Should never have been in the cohort. He was Miami's primary right fielder during the window. An eligibility check against active MLB rosters at run date would have removed him.
Owen Caissie was scored at 33 percent as a call-up candidate. He was Miami’s primary right fielder for the season and appeared in the majors throughout the window. He was never a prospect awaiting promotion, and the row is a data-entry failure on our side rather than a forecast that went wrong.
We are grading it anyway, because closed runs are never edited and that rule does not get suspended when the error is embarrassing. For transparency: excluding that row, the Brier on the remaining thirteen is 0.2212 against a naive baseline of 0.2130. The model still loses. Removing our own mistake does not rescue the result, which is worth knowing.
Fourteen outcomes is a very small sample and Brier scores on fourteen binary events carry enormous variance. One player moving from NO to YES swings the number materially. So this is a data point, not a verdict on the model.
It is also not nothing. The June checkpoint produced perfect bucket ordering on 28 outcomes. This one produced an inversion between the top two bands and a score worse than a constant guess. Two checkpoints in, the honest summary is that the model sorts players better than it prices them, and that the injury blind spot has now cost it twice.
No refit today. v0.5 stays live exactly as it is. Refitting a mapping table on fourteen outcomes is the small-sample overfitting that the shrinkage step in the June patch exists to prevent, and doing it because the result stung would be the worst possible reason. Sub-Batch C resolves October 8 with 63 open predictions. At that point 79 outcomes across two cohorts inform a single revision.
Candidates already on the v0.6 list, all of which came out of today:
The full run record, including every input and the checkpoint JSON, is published at checkpoint_2026-09-09.json. The running scoreboard is on the record, and the model itself is documented in Inside v0.5.
Outcomes verified against public reporting of each debut. Brier score is the mean squared error of the forecast probabilities against binary outcomes, per Brier (1950). The naive baseline is a constant forecast equal to the observed base rate of the cohort. AUC is the probability that a randomly chosen player who debuted carried a higher CUP than a randomly chosen player who did not.