Randomized trial demonstrates high classification accuracy in microarray cancer datasets, highlighting efficient gene selection.
Gene selection is a critical step in microarray-based cancer classification, where the number of genes greatly exceeds the number of samples. Selecting a small yet informative subset improves classification accuracy, model interpretability, and computational efficiency. While extensive research has focused on population-based metaheuristics for this task, relatively few studies have explored single-solution-based incremental methods, which offer strong potential for efficiency and adaptability in resource-constrained settings. Addressing this gap, we propose a semi-greedy hybrid incremental (SGHI) gene selection method that combines filter-based ranking with wrapper evaluation in a stochastic, stepwise fashion. At each iteration, SGHI probabilistically constructs a small candidate list using softmax-weighted sampling from filter-ranked genes, then greedily selects the best-performing gene based on cross-validated accuracy. This design enables efficient search space exploration while keeping evaluation costs low. Experiments on 11 benchmark microarray datasets using a Random Forest classifier and stratified 10-fold cross-validation demonstrate that SGHI consistently achieves high classification accuracy, even reaching perfect accuracy in several cases, while selecting highly compact gene subsets. On average, SGHI attains a mean accuracy of 0.98 while selecting only 6.4 genes, highlighting its ability to balance predictive performance with minimal feature usage. Notably, under a low-budget configuration, SGHI attains an average accuracy of 0.97 with a mean of just 5.6 selected genes, outperforming state-of-the-art methods in gene reduction while maintaining competitive accuracy. Statistical analyses confirm SGHI’s robust and consistent performance across datasets, offering a computationally efficient solution with substantially reduced computational demand.
No takes yet. Share an insight, caveat, or question.
Osman Gökalp (2026) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: