Every headline number carries a confidence interval, and we track whether our edge is decaying over time instead of quietly hiding it. If the model breaks, this page says so first.
EDGE STILL HOLDINGCUSUM change-point test clean · stat 2.99 · baseline accuracy 65.3%
HEADLINE NUMBERS
95% CI · 1,610 BOOTSTRAP RESAMPLES
66.5%
OVERALL ACCURACY
95% CI 64.1% – 68.8%
Tier
Accuracy
95% CI
S
83.7%
75.6% – 90.7%
A+
75.1%
69.2% – 81.1%
A
68.0%
59.0% – 77.0%
B
66.3%
63.0% – 69.5%
C
58.3%
53.4% – 63.1%
CALIBRATION
PREDICTED VS ACTUAL
The raw ensemble optimizes which side to pick, not the exact win probability — the pick is a hard 3-model vote while this number is the mean of three probabilities, so the mid buckets win a few points more often than stated (underconfident, not wrong). The confidence we display on picks corrects this with a per-season Platt layer (fit on prior seasons only, clamped so it never flips the pick); this chart shows the raw model underneath. The gap column shows it honestly; whiskers are 95% Wilson intervals.