Graph-Augmented Machine Learning Framework for Residue-Level Protein–Peptide Interface Classification Using Structural and Physicochemical Features This project presents a modular computational framework for residue-level protein–peptide interface classification using structural bioinformatics, graph-based molecular representation, classical machine learning, and graph neural network (GNN)-ready datasets. The framework integrates: Residue-level structural feature extraction Protein–peptide interface labeling Residue contact graph construction Graph topology analysis Graph-augmented machine learning Grouped cross-validation by protein complex GNN dataset preparation using PyTorch Geometric-compatible graph structures Structural and physicochemical descriptors include: residue–peptide distance, local atomic contact density, contact atom counts, hydrophobicity, residue polarity, aromaticity, charge, and B-factor-derived flexibility. Graph-derived descriptors include: node degree, clustering coefficient, betweenness centrality, closeness centrality, eigenvector centrality, PageRank, graph communities, and interface-neighborhood enrichment. A pilot benchmark dataset containing 1,450 residues from four experimentally solved protein–peptide complexes was assembled. Interface residues were defined using a 5 Å contact threshold. Under grouped cross-validation, Random Forest achieved the strongest overall performance: Mean F1-score: 0.959 ± 0.049 Mean ROC-AUC: 0.999 ± 0.002 Graph neural network experiments were subsequently initiated using residue contact graphs and graph convolutional architectures.
Taufia Hussain (Sat,) studied this question.