Variable selection is an important step in data envelopment analysis (DEA) when the number of decision making units (DMUs) is insufficient. This research thus proposes a two-stage variable selection method integrating a well-known supervised machine learning technique, the random forest algorithm. In the first stage, a baseline DEA model with full input and output variables is implemented and each DMU can be determined as being efficient or inefficient. In the second stage, a random forest is trained to learn how to classify efficient or inefficient DMUs well in high dimensions. Accordingly, the importance of each variable can be calculated based on permutation importance indices in random forest. This paper further discusses two issues about data augmentation in order to improve the robustness of permutation importance when the number of DMUs is quite small.
Tzu-Pu Chang (Thu,) studied this question.