ABSTRACT The study focuses on proposing a new machine learning approach to classify crude oil samples based on their physicochemical properties, such as sulfur concentration (S), total acidity number (TAN), and American Petroleum Institute (API) gravity. The goal is to overcome the limitations of traditional analysis methods, which are time‐consuming and consume large volumes of samples and solvents, using spectroscopic techniques and machine learning models such as support vector machine ensemble (SVM ensemble). The number of 196 crude oil samples with different sulfur content, different TAN, and API gravity were considered. The SVM ensemble is a powerful approach to improve classification performance because it can reduce the variability of individual models, improve robustness against overfitting, and generalize better than a single model. The parameters of sensitivity, specificity, error rate, Matthews correlation coefficient, and accuracy were considered to compare the SVM ensemble with partial least squares‐discriminant analysis (PLS‐DA) and SVM. The results demonstrated that near infrared spectroscopy (NIR), combined with multivariate classification models, is an efficient and reliable method for simultaneously classifying sulfur content, TAN, and API gravity in crude oils.
Barboza et al. (Wed,) studied this question.