Based on the elemental composition of 359 samples of dry wines (Riesling, Chardonnay, Muscat, Cabernet Sauvignon, and Merlot) produced in the Krasnodar Territory (Russia), the influence of data clustering on the generalizing properties of neural network models for solving classification problems was studied. Clustering as a property of predicted classes being compact and separate from each other was evaluated using scatterplots of canonical values from discriminant analysis and the Silhouette Score, Calinski-Harabasz, Davies-Bouldin, and Dunn metrics. With a decrease in data clustering, the generalization properties of neural network models decrease despite an increase in dataset size. Since the predictive properties of neural networks are primarily correlated with the clustering of data, it seems logical to say that their growth will occur when “quantity turns into quality”, that is, clustering will increase with increasing dataset size. The validity of the assumption explains the polarity of trends observed in literature. With the growth of training datasets, some researchers record improvements in the predictive properties of models, while others record deterioration. The results obtained are of great practical importance, as they indicate the unsuitability of unlimited data accumulation and allow optimizing the costs of collecting it, which is important, especially for food quality control.
Khalafyan et al. (Thu,) studied this question.