Existing machine learning studies on high-entropy alloys (HEAs) compress multi-element compositions into aggregate scalar descriptors thereby discarding element-specific physicochemical information and limiting both predictive accuracy and post-hoc interpretability. This study proposes a composition-aware surrogate model that represents each element as a 13-dimensional token combining 12 physicochemical descriptors with its mole fraction, and learns inter-element interactions through multi-head self-attention. Applied to 232 molecular dynamics simulation samples of the CoCrCuFeNi quinary HEA system, the model achieves ultimate tensile strength R 2 = 0.929 and Young's modulus R 2 = 0.982, outperforming machine learning baselines under identical hyperparameter optimization conditions. Shapley Additive explanations applied over the five-variable composition space reveal model-consistent trends including a positive contribution of Ni to ultimate tensile strength and a negative contribution of Cu to Young's modulus. The framework demonstrates that element representation design is critical for both prediction accuracy and interpretability on small-scale embedded-atom method molecular dynamics data and can be extended to higher-order multi-component HEA composition screening by simply adding element tokens.
Shin et al. (2026) studied this question.