Comparative analysis demonstrates robust nonlinear dependence and distribution testing across real-world datasets, highlighting practical workflows for kernel-based statistical learning.
Classical nonparametric statistical methods often exhibit limited capability when analysing modern high-dimensional datasets containing complex nonlinear relationships and heterogeneous distributions. Kernel-based methods formulated within the framework of reproducing kernel Hilbert spaces (RKHS) provide a flexible alternative by enabling distribution-free inference while implicitly modelling nonlinear structures. This study presents a unified and reproducible framework that integrates several established kernel-based inference techniques, including Maximum Mean Discrepancy (MMD), the Hilbert–Schmidt Independence Criterion (HSIC), Kernel Ridge Regression (KRR), Kernel Principal Component Analysis (Kernel PCA), Kernel Analysis of Variance (K-ANOVA), and Kernel Canonical Correlation Analysis (KCCA), within a common computational workflow. Rather than proposing new kernel algorithms, the contribution lies in providing a systematic implementation and comparative empirical evaluation of complementary kernel methods for distributional comparison, dependence analysis, nonlinear prediction, and feature-space visualization. The framework is evaluated using publicly available healthcare (PIMA Indians Diabetes and UCI Heart Disease) and environmental (Air Quality) datasets. Experimental results demonstrate that kernel-based inference effectively detects nonlinear distributional differences and complex dependence structures that are difficult to identify using conventional approaches. At the same time, the analysis shows that predictive performance remains dependent on data quality and covariate completeness, illustrating that kernel methods cannot fully compensate for missing or uninformative predictors. Overall, the study provides a practical and reproducible reference workflow for applying kernel-based nonparametric inference to real-world statistical learning problems while highlighting important considerations regarding kernel selection, bandwidth sensitivity, computational scalability, and model interpretation.
No takes yet. Share an insight, caveat, or question.
Rasool et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: