The study of protein–protein interactions (PPIs) is of significant importance for elucidating biological processes, clarifying pathological mechanisms, and promoting drug development. In this study, we proposed a method to predict PPIs based on protein sequence and gene sequence information, combined with convolutional neural networks (CNNs). First, we extracted three types of features from protein sequence: global physicochemical properties features of the protein sequence, local same type of amino acid position variation features, and protein evolutionary conservation features; simultaneously, we extracted single nucleotide frequency and positional features, dinucleotide frequency features, and trinucleotide frequency features from the corresponding gene sequence. During the feature extraction process, we employed the amphiphilic pseudo amino acid composition (APAAC) method to extract the global hydrophobicity and hydrophilicity features of the protein sequence; we defined a new mathematical descriptor—θ interval deviation product factor—to extract protein evolutionary conservation features from Position Specific Scoring Matrix (PSSM); we also defined a mapping function to map all nucleotides in the gene sequence onto a unit circle, and then extracted nucleotide positional features from the mapped points. Second, based on extracted features, we constructed a 36 × 32 sample feature grayscale map to represent a protein pair sample. Finally, we developed a CNN model to predict PPIs. Our method achieved superior results on four species test sets: an accuracy of 99.28% on the Saccharomyces cerevisiae dataset, 98.15% on the Drosophila melanogaster dataset, 98.62% on the Homo sapiens dataset, and 96.84% on the Mus musculus dataset, outperforming existing computational methods. Furthermore, we extended the application of this method to the prediction of protein–protein interaction networks and non-interaction networks, and also achieved promising results.
Shi et al. (Mon,) studied this question.