This study compared three deep learning architectures—one-dimensional convolutional neural network (1D-CNN), self-supervised learning (SSL), and Vision Transformer (ViT)—to evaluate their ability to predict carotenoid content from visible–near-infrared (VIS–NIR) spectral reflectance data (400–850 nm) acquired non-destructively from tea leaves. Model performance was evaluated using 10-fold cross-validation and analyzed through the mean SHapley Additive exPlanations values to identify key spectral features. The ViT model achieved the highest predictive accuracy (coefficient of determination R2 = 0.81, root mean square error RMSE = 1.04, ratio of performance to deviation RPD = 2.32), followed by 1D-CNN (R2 = 0.75, RMSE = 1.21, RPD = 1.99), whereas SSL showed substantially lower predictive performance (R2 = 0.30, RMSE = 2.01, RPD = 1.20). Feature importance analysis revealed that ViT focused strongly on the red-edge region around 720 nm, which corresponds to spectral features associated with carotenoids and chlorophyll. The 1D-CNN relied mainly on blue (450–480 nm) and red (670–700 nm) regions, while SSL exhibited a broadly distributed importance pattern across wavelengths. These results indicate that ViT’s self-attention mechanism captures long-range spectral dependencies more effectively than conventional convolutional or self-supervised models. Overall, the study demonstrates that transformer-based architectures provide a powerful and interpretable framework for non-destructive estimation of carotenoid content from VIS–NIR reflectance spectroscopy.
Tsuchiya et al. (Mon,) studied this question.