Key points are not available for this paper at this time.
Protein–DNA and protein–RNA interactions are central to gene regulation and genetic disease, yet experimental identification remains costly and complex. Machine learning (ML) offers an efficient alternative, though challenges persist in representing protein sequences due to residue variability, dimensionality issues, and the risk of losing biological context. Traditional approaches such as k-mer counting or neural network encodings provide standardized sequence representations but often demand high computational resources and may obscure functional information. To address these limitations, a novel encoding method based on interpolation of physicochemical properties (PCPs) is introduced. Discrete PCPs values are transformed into continuous functions using logarithmic enhancement, highlighting residues that contribute most to nucleic acid interactions while preserving biological relevance across variable sequence lengths. Statistical features extracted from the resulting spectra via Tsfresh are then used for binary classification of DNA- and RNA-binding proteins. Six classifiers were evaluated, and the proposed method achieved up to 99% accuracy, precision, recall, and F1 score when amino acid highlighting was applied, compared with 66% without highlighting. Benchmarking against k-mer and neural network approaches confirmed superior efficiency and reliability, underscoring the potential of this method for protein interaction prediction. Our framework may be extended to multiclass problems and applied to the study of protein variants, offering a scalable tool for broader protein interaction prediction.
Cabello-Lima et al. (Sat,) studied this question.