Distributional Shift as a Cascading Failure Mechanism in Deployed Medical Imaging Classifiers
Main Article Content
Abstract
Distributional shift remains a critical barrier to the safe deployment of medical imaging classifiers because model degradation can propagate beyond predictive error into calibration failure, triage distortion, and downstream clinical decision support risk. This study examines distributional shift as a cascading failure mechanism by evaluating classifier behavior across reference, low covariate-shift, moderate acquisition-shift, moderate label-shift, and high mixed-shift deployment scenarios. The reference-domain classifier achieved strong baseline performance, with AUROC of 0.941, sensitivity of 0.918, specificity of 0.904, F1-score of 0.911, and error rate of 0.087. However, performance declined progressively under deployment shift, with the high mixed-shift condition reducing AUROC to 0.783, sensitivity to 0.721, specificity to 0.766, and F1-score to 0.742, while increasing the error rate to 0.263. Calibration reliability also deteriorated, as Expected Calibration Error increased from 0.031 to 0.142, Brier score rose from 0.071 to 0.184, entropy increased from 0.218 to 0.392, and overconfident error increased from 4.8% to 21.6%. Feature-space analysis further showed that Maximum Mean Discrepancy increased from 0.038 under low covariate shift to 0.137 under high mixed shift, accompanied by prediction drift of 0.158 and entropy shift of 0.174. Cascading simulation revealed that high-shift risk escalated from 0.22 at image acquisition to 0.91 at clinical action, confirming that distributional instability can intensify across workflow stages. The findings indicate that deployment safety requires integrated monitoring of performance, calibration, drift, and propagation risk rather than reliance on aggregate diagnostic accuracy alone.