Background/objectives Pancreatic ductal adenocarcinoma (PDAC) remains one of the most lethal malignancies, largely due to late diagnosis. Reliable early detection requires diagnostic models that account for biological heterogeneity, including established sex dimorphism in PDAC pathogenesis. This study investigates the structural stability and functional asymmetry of diagnostic models based on a five-predictor urinary biomarker panel (LYVE1, REG1B, TFF1, Age, Creatinine) and assesses diagnostic performance, algorithmic bias, and cross-sex generalisability across asymptomatic (PDAC vs Healthy) and symptomatic (PDAC vs Benign) clinical settings. Methods A retrospective case–control analysis included 199 PDAC patients, 183 healthy controls, and 208 benign cases. Five algorithms—Logistic Regression (LR), Decision Tree, Support Vector Machine, Random Forest, and Synolitic Graph Neural Network (SGNN)—were evaluated using sex-stratified 5-fold cross-validation with cross-sex testing (f → f, f → m, m → m, m → f). The SGNN method was utilized in addition to the classical approaches to explore the interaction topology of biomarkers. Performance was assessed using AUC, sensitivity, specificity, predictive values, and equalised odds differences. Interpretability analyses (SHAP and full LR coefficients) and adversarial robustness testing were conducted for the selected LR model. Results Logistic Regression demonstrated the most stable and interpretable performance across all validation scenarios. For PDAC vs Healthy, LR achieved AUCs ranging from 0.896 to 0.931; for PDAC vs Benign, AUCs ranged from 0.772 to 0.874. Cross-sex generalisation was asymmetric: male-trained models transferred well to females (e.g., m → f AUC = 0.931 in PDAC vs Healthy; 0.863 in PDAC vs Benign), whereas female-trained models showed reduced performance on males, particularly in the symptomatic setting (f → m AUC = 0.772). Equalised odds differences indicated minimal bias in the asymptomatic setting (0.023) and modest differences in the symptomatic setting (0.064). REG1B was the dominant predictor across models, with LYVE1 contributing consistently and creatinine showing a negative association. Adversarial testing demonstrated moderate performance degradation under perturbation, supporting overall model robustness. SGNN yielded a small improvement in mean AUC relative to Logistic Regression; however, this gain was accompanied by increased model complexity, whereas LR provided stable, interpretable, and robust performance across validation scenarios. Conclusions Urinary biomarker-based PDAC diagnostic signatures exhibit sex-dimorphic and asymmetric generalisation patterns. Sex-stratified Logistic Regression provides a stable, interpretable, and fair diagnostic framework across clinical contexts. Incorporating sex-specific modelling may improve equitable and reliable early detection of PDAC and support future clinical deployment of biomarker-driven screening tools.
No takes yet. Share an insight, caveat, or question.
Krivonosov et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: