We examined the limitations of observed course DFWI rates (% of D and F grades, withdrawals, and incompletes) as evaluation metrics, which obscure student characteristics, course design, instruction, structure, and latent factors, posing challenges in identifying courses that need improvement. An artificial neural network (ANN) was trained using student data to model risk, accounting for variations in student characteristics. The model’s predictions on test data were averaged at the course level, producing expected DFWI rates based on student composition. Courses with high observed DFWI rates and large deviations between observed and predicted DFWI rates (the GAP) were ranked and prioritized for review, as they may reflect aspects of course design, structure, or instructional practices warranting further qualitative evaluation. Our predictions are non-causal, and modeling calibration varies across subgroups; therefore, the original GAP rankings, robust to a post-hoc calibration check, are presented as risk-adjusted indicators for prioritizing courses for further review rather than as definitive causal measures of course quality. Rankings based on observed DFWI rates differ substantially from risk-adjusted GAP rankings, indicating that relying on observed DFWI rates alone may misidentify high-risk courses. Our methodology can assist educators and administrators in making fair resource allocation decisions and improving student outcomes.
Joseph et al. (2026) studied this question.