SAFE metrics (Sustainability (S, Sustainability means Robustness), Accuracy (A), Fairness (F), and Explainability (E) metrics) assess model performance via rank change, valued for their intuitive nature and computational simplicity, yet they need to account for the randomness in predicted rankings. To address this, we introduce an entropy-adjusted factor that integrates rank change into the whole evaluation framework, preserving the original benefits while enhancing applicability in uncertain contexts by considering prediction randomness. For models with identical original metric values, lower entropy signifies greater reliability. The general decrease in metric values does not always signify a drop in model performance; rather, it indicates a shift in the evaluation framework. Additionally, we mathematically demonstrate the variations in these SAFE metrics. In our empirical analysis, by using credit ratings as a case study, we took the results of SAFE metrics as an example and compared them with entropy-adjusted SAFE metrics. The results revealed reductions in RGA, RGR, alongside an improvement in RGE, with entropy value plot analysis further validating these findings. RGF is not considered, as it can be replaced by the RGE of the protected variable. We propose that traditional and entropy-adjusted SAFE metrics can be used complementarily, each providing unique advantages at different stages or under varying requirements, thus offering a more comprehensive evaluation framework for model performance in diverse scenarios.
Lunshuai Wu (Tue,) studied this question.