Background/Objectives: Arboviral diseases share common vectors, geographic distribution, and symptoms. Developing Machine Learning diagnostic tools for co-circulating arboviral diseases faces data-scarcity challenges. This study aimed to demonstrate that proof of concept using synthetic data can establish computational feasibility and guide future real-world validation efforts. Methods: We assembled a synthetic dataset of 28,000 records, with 7000 for each disease—Dengue, Zika, and Chikungunya—plus Influenza as a negative control. These records were obtained from the existing literature. A binary matrix with 67 symptoms was created for detailed statistical analysis using Odds Ratios, Chi-Square, and symptom-specific conditional prevalence to validate the clinical relevance of the simulated data. This dataset was used to train and evaluate various algorithms, including Multi-Layer Perceptron (MLP), Narrow Neural Network (NN), Quadratic Support Vector Machine (QSVM), and Bagged Tree (BT), employing multiple performance metrics: accuracy, precision, sensitivity, specificity, F1-score, AUC-ROC, and Cohen’s kappa coefficient. Results: The dataset aligns with the PAHO guidelines. Similar findings are observed in other arboviral databases, confirming the validity of the synthetic dataset. A notable performance across all evaluated metrics was observed. The NN model achieved an overall accuracy of 0.92 and an AUC above 0.98, with precision, sensitivity, and specificity values exceeding 0.85, and an average Uniform Cohen’s Kappa of 0.89, highlighting its ability to reliably distinguish between Dengue and Influenza, with a slight decrease between Zika and Chikungunya. Conclusions: These models could accelerate early diagnosis of arboviral diseases by leveraging encoded symptom features for Machine Learning and Deep Learning approaches, serving as a support tool in regions with limited healthcare access without replacing clinical medical expertise.
Cruz-Parada et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: