• Model performance differs between baseline and working-memory conditions • Aggregate metrics mask structured error differences • Time pressure amplifies cognitive-state-related disparities • Study motivates cognitive-state–aware model evaluation Machine-learning (ML) models that predict human moral decisions in autonomous-vehicle (AV) dilemma tasks are usually evaluated using pooled test performance. This study examines whether such evaluation is sufficient, or whether model behaviour differs depending on the cognitive state in which decisions are made. We treat working-memory (WM) condition as a grouping variable with two states: baseline (no secondary task) and WM load (a concurrent Sternberg letter task). The central question is whether models trained on these decisions perform similarly for baseline and WM trials, or whether performance differs across cognitive states. The dataset contains 4,032 trials from 336 licensed drivers, covering three time-to-collision (TTC) levels and four scenario types. A participant-level train–test split is used (235 training participants, 101 test participants) to avoid information leakage. Six ML classifiers are trained without WM as an input feature and are evaluated separately on baseline and WM trials in the held-out test set. Performance is assessed using balanced accuracy, macro-F1, and true-positive and false-positive rates for utilitarian predictions. Robustness is examined using participant-level bootstrap resampling. Across the models tested, model behaviour differs systematically between baseline and WM trials. This shows that pooled test performance can hide performance disparities linked to the cognitive state under which decisions were formed. These differences show consistent patterns within the evaluated conditions. They are strongest under high time pressure (TTC = 1 s) and are concentrated in specific TTC-by-scenario regions, while many other dilemmas show little or no difference. Bootstrap resampling indicates that the observed global baseline–WM differences are consistent at the aggregate level. The artificial neural network (ANN), which predicted almost the same outcome for nearly all trials, is treated as uninformative rather than as evidence of equal treatment. Overall, predictive accuracy alone is not sufficient to evaluate moral decision prediction models. Reporting cognitive-state–stratified performance provides additional insight into when and where models behave unevenly, particularly in time-constrained moral scenarios. The work contributes a cognitive-state–stratified auditing procedure for safety-critical decision modelling that complements pooled metrics with cognitive-state–stratified and context-stratified evaluation.
Singh et al. (Fri,) studied this question.