Automated short-answer grading (ASAG) in educational contexts faces a fundamental trade-off between predictive performance, interpretability, and methodological transparency, particularly under data-constrained educational settings. While recent approaches rely on deep learning architectures, these models require large annotated datasets and offer limited transparency, restricting their applicability in authentic classroom environments. This study proposes a fully specified and interpretable machine learning pipeline for ASAG across multiple educational concepts. The approach is based on a shared TF–IDF representation and evaluates three linear classifiers—Logistic Regression, Multinomial Naïve Bayes, and Linear Support Vector Machines—under a stratified cross-validation framework adapted to small datasets. Model performance is assessed using accuracy, precision, recall, and F1-score. Statistical comparisons using the Wilcoxon signed-rank test indicate exploratory evidence of statistically significant differences between classifiers, although the observed differences remain small in practical magnitude. Additionally, the methodology incorporates token-level analysis to identify discriminative lexical patterns and examine consensus across classifiers. To enhance interpretability, tokens are presented using a bilingual Spanish/English representation while preserving the original feature space. The results across ten concept-specific datasets show consistent performance across models (accuracy ≈ 0.82–0.88) and reveal stable lexical patterns consistently associated with model predictions of correctness. The findings highlight that lightweight, interpretable models can provide consistent and reliable performance under resource-constrained educational conditions. The proposed framework contributes a stability-oriented and interpretable evaluation paradigm for ASAG, offering a practical alternative to data-intensive approaches in educational assessment. It is intended as a methodological reference protocol rather than a performance benchmark. The findings should be interpreted as evidence of within-context consistency instead of broad external generalization.
Maestre et al. (Mon,) studied this question.