Online Health Communities (OHCs) connect patients for peer support, but users face a discovery challenge when they have minimal prior interactions to guide personalization. We address recommendation under extreme interaction sparsity in a survey-driven setting where each user provides a 16-dimensional intake vector at registration and each support group has a structured feature profile. We augment three Neural Collaborative Filtering (NCF) architectures—Matrix Factorization (MF), Multi-Layer Perceptron (MLP), and NeuMF—with an auxiliary pseudo-label objective derived from survey–group feature alignment using cosine similarity mapped to 0, 1. This yields Pseudo-Label NCF (PL-NCF), which learns dual embedding spaces: main embeddings optimized for ranking, and PL-specific embeddings intended to capture semantic alignment between users and support groups. On a dataset of 165 users, 498 support groups, and 498 observed memberships (three per user), we evaluate ranking quality (HR@5, NDCG@5) primarily under a leave-one-out protocol—one validation and one test interaction per user, yielding approximately one training positive per user—which best reflects our target cold-start regime, with a train-validation-test (70/15/15) split for supplementary analysis. We analyze embedding structure in the leave-one-out setting using spherical 𝑘-means silhouette scores computed in the original high-dimensional embedding space, alongside 2D t-SNE visualizations for qualitative inspection. Under leave-one-out evaluation, all three PL variants improve ranking: MLP-PL improves HR@5 from 2. 65 % to 5. 30 %, NeuMF-PL from 4. 46 % to 5. 18 %, and MF-PL from 4. 58 % to 5. 42 %. Critically, PL-specific embedding spaces exhibit higher cosine silhouette scores than baseline main embeddings under per-seed optimal clustering with 𝑘 ∈ 3, 4, 5, 6, 7, 8, 10: MF-PL improves from 0. 0394 to 0. 0684; NeuMF-PL improves from 0. 0263 to 0. 0653. We also observe a separability–accuracy trade-off: main-embedding clusterability is negatively correlated with ranking accuracy (Spearman 𝜌 ≈ −0. 38 on leave-one-out; 𝜌 ≈ −0. 59 on the train-validation-test split), suggesting that embeddings optimized for ranking may sacrifice interpretability. These findings demonstrate that survey-derived pseudo-labels can regularize sparse-data training for improved ranking, while producing interpretable task-specialized embedding spaces via dedicated PL-specific representations
Barman et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: