Spectroscopic methods for plant products identification often suffer from batch effects, whereas High-Performance Liquid Chromatography (HPLC) offers superior reproducibility but is rarely used for classification. Here, HPLC quantitative data were integrated with chromaticity values to develop a robust machine learning framework for identifying Panax notoginseng (PN) powder. Using 726 samples across various origins, parts, and grades, a dataset of 19 features (saponins, flavonoids, starch, chromaticity) was compiles. For each classification task (origin, part, grade), the dataset was partitioned into training (80%) and testing (20%) sets; crucially, an independent external validation set was employed to assess the models' real-world performance and generalizability. Four machine learning algorithms, including Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Light Gradient Boosting Machine (LGBM), and Support Vector Machine (SVM), were then used for model development. LGBM optimally distinguished PN parts (validation accuracy 91.09%), while SVM excelled in grade differentiation (85.19%) and origin identification (up to 100%). Feature importance analysis identified specific flavonoids, saponins, and chromaticity values as critical differentiators. These results highlight that integrating chemical and physical data overcomes batch effects, providing a stable, practical solution for real-world quality control of plant products. • A large-scale data matrix of 726 samples × 19 features was established. • Not only train and test sets, but also independent external validation set was used. • The classification models achieved testing set accuracies > 92.86%. • The classification models attained validation set accuracies > 64.58%. • This method offers a robust scheme for real-world identification of plant products.
Ye et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: