Generated by parity_study.py, 2026-07-08. Luck SD from 200,000 simulated 82-game seasons (p_win=0.5, OT share 0.22): 8.49 points.
Each team's points normalized to an 82-game pace (short seasons 2012-13, 2019-20, 2020-21 included at pace).
| Season | Teams | GP/team | SD(pts/82) |
|---|---|---|---|
| 2010-11 | 30 | 82.0 | 13.27 |
| 2011-12 | 30 | 82.0 | 11.73 |
| 2012-13 | 30 | 48.0 | 16.46 |
| 2013-14 | 30 | 82.0 | 15.26 |
| 2014-15 | 30 | 82.0 | 15.91 |
| 2015-16 | 30 | 82.0 | 12.86 |
| 2016-17 | 30 | 82.0 | 15.13 |
| 2017-18 | 31 | 82.0 | 15.44 |
| 2018-19 | 31 | 82.0 | 13.65 |
| 2019-20 | 31 | 69.8 | 14.12 |
| 2020-21 | 31 | 56.0 | 19.27 |
| 2021-22 | 32 | 82.0 | 20.28 |
| 2022-23 | 32 | 82.0 | 18.90 |
| 2023-24 | 32 | 82.0 | 17.63 |
| 2024-25 | 32 | 82.0 | 14.84 |
| 2025-26 | 32 | 82.0 | 13.18 |
2010-2019 mean SD 14.38 (range 11.73-16.46); 2024-25 14.84, 2025-26 13.18.
signal SD = sqrt(SD(actual)^2 − luck SD^2); luck share = luck var / actual var; r ceiling = signal SD / SD(actual) — the correlation a perfect strength oracle would score; err_sig SD = sqrt(RMSE^2 − luck SD^2) — the model's error on true strength; r expected = corr implied by that signal + that error; r true-inputs = corr of the end-of-season-known-inputs prediction with actual points.
| Season | SD act | SD pred | MAE | RMSE | r (Pearson) | rho (Spearman) | signal SD | luck share | r ceiling | err_sig SD | r expected | r true-inputs |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2022-23 | 18.9 | 8.3 | 11.5 | 14.0 | 0.714 | 0.707 | 16.9 | 20% | 0.893 | 11.1 | 0.746 | 0.908 |
| 2023-24 | 17.6 | 8.9 | 10.7 | 13.2 | 0.671 | 0.626 | 15.5 | 23% | 0.876 | 10.1 | 0.733 | 0.874 |
| 2024-25 | 14.8 | 7.5 | 11.0 | 13.3 | 0.426 | 0.385 | 12.2 | 33% | 0.820 | 10.2 | 0.629 | 0.878 |
| 2025-26 | 13.2 | 6.2 | 10.6 | 12.8 | 0.273 | 0.219 | 10.1 | 41% | 0.765 | 9.6 | 0.554 | 0.844 |
Freeze the model's information at its 2022-23 level (implied info-error SD 12.7 pts, from r = Vs/sqrt((Vs+Ve)(Vs+Vl)) calibrated on 2022-23) and let only the signal SD move:
| Season | signal SD | observed r | compression-only r | SE(r), n=32 |
|---|---|---|---|---|
| 2022-23 | 16.9 | 0.714 | 0.714 | 0.09 |
| 2023-24 | 15.5 | 0.671 | 0.677 | 0.10 |
| 2024-25 | 12.2 | 0.426 | 0.568 | 0.15 |
| 2025-26 | 10.1 | 0.273 | 0.475 | 0.17 |
| Season | RMSE | err_sig SD | mean NNoData | mean RetShare | r( | err | , NNoData) | r( | err | , RetShare) |
|---|---|---|---|---|---|---|---|---|---|---|
| 2022-23 | 14.0 | 11.1 | 0.6 | 0.79 | -0.09 | -0.19 | ||||
| 2023-24 | 13.2 | 10.1 | 0.5 | 0.80 | 0.05 | -0.22 | ||||
| 2024-25 | 13.3 | 10.2 | 0.5 | 0.78 | -0.05 | -0.15 | ||||
| 2025-26 | 12.8 | 9.6 | 0.7 | 0.81 | -0.05 | 0.15 |
Pooled across 128 team-seasons: r(|err|, NNoData) = -0.04, r(|err|, RetShare) = -0.12.
Variant comparison (RMSE / Pearson r):
| Variant | 2022-23 | 2023-24 | 2024-25 | 2025-26 |
|---|---|---|---|---|
| base | 14.0 / 0.71 | 13.2 / 0.67 | 13.3 / 0.43 | 12.8 / 0.27 |
| age | 14.0 / 0.72 | 13.3 / 0.67 | 13.2 / 0.43 | 12.7 / 0.29 |
| ageG | 14.0 / 0.72 | 13.3 / 0.67 | 13.2 / 0.43 | 12.6 / 0.30 |
Mostly parity compression; the model's absolute accuracy did not degrade. The model's error actually improved (RMSE 14.0 -> 12.8, MAE 11.5 -> 10.6), but the spread it is trying to rank shrank dramatically: SD of actual points fell 18.9 -> 13.2, and after removing the 8.5-point schedule-luck floor the true strength SD fell 16.9 -> 10.1; luck now accounts for 41% of actual-points variance vs 20% in 2022-23, and a perfect strength oracle tops out at r = 0.76 (was 0.89). The historical table confirms 2024-26 is the tightest league since at least 2010 (SD82 14.8 and 13.2 vs 2010-2019 mean 14.4, and vs 18.9-20.3 in 2021-23). Attribution: freezing the model's information at its 2022-23 level and shrinking only the signal predicts r = 0.48 for 2025-26 — i.e. compression alone accounts for 0.24 of the 0.44 Pearson-r drop (54%). The residual 0.20 is only ~1.2x the n=32 sampling SE of r (0.17), so evidence of genuine relative degradation is weak — and three independent checks say the model itself held up: the signal-error SD (sqrt(RMSE^2 - luck^2)) fell 11.1 -> 9.6; the true-inputs prediction still ranks well (r 0.84 in 2025-26 vs 0.91 in 2022-23), so the input projections, not the points model, carry what residual there is; and |error| is uncorrelated with roster churn (pooled r vs NNoData -0.04, vs RetShare -0.12), killing the churn/rookie alternative.
2026-27 implications. The 2026-27 predicted-standings spread is narrow — SD of ExpPts is 5.0 vs a recent true-signal SD of ~11.1. That gap is correct behavior, not a bug: a point forecast should be shrunk toward the mean by its own uncertainty (the backtest's SD(pred) was likewise ~6 vs SD(act) 13-19), and in a compressed league the shrinkage is heavier. Do not widen ExpPts to match the actual standings spread. STRENGTH_SD = 10.5 represents strength uncertainty beyond schedule luck; the backtest's realized signal-error SD averaged 9.9 over 2024-26, so 10.5 remains about right, if a shade conservative (implied total per-team spread sqrt(10.5^2 + 8.5^2) = 13.5 vs realized RMSE 12.8; 9.5-10 would match exactly — not worth retuning). The practical consequence is humility on playoff odds: with true-strength SD ~11 points against a luck SD of 8.5 plus ~10 points of projection error, mid-table playoff probabilities belong in the 35-65% band and predicted ranks in the middle two quartiles are close to interchangeable — the sim's wide PlayoffProb bands are a feature of the era, not a model defect.