Nguyễn Tương Tri * , Mai Xuân Văn , Trần Võ Hoàng Nguyên Trần Hoài Nhân

* Người chịu trách nhiệm về bài viết: (ntuongtri@hueuni.edu.vn)

Abstract

Determining protein-protein interactions plays an important role in understanding the structure and biological function of cells. Accurate prediction of interactions remains challenging. In this study, we introduce a method that combine the Hilbert transform on evolutionary information for the feature extraction step and the Deep Forest classification model for the prediction step. The use of Hilbert transformation aims to enhance the evolutionary information stored in the protein sequence. Additionally, Deep Forest, a powerful ensemble classification model, is used to improve prediction performance. The proposed method is evaluated on the benchmark Yeast dataset and compared with other existing methods. The results show that our method is highly effective in predicting PPI. The datasets and source code of the method are provided at https://github.com/thnhanhub/FeatPSSM.git
Keywords: Protein Interactions, Ensemble Learning, Hilbert Transform

Tóm tắt

Xác định tương tác protein-protein (PPI) có vai trò quan trọng trong việc hiểu biết về cấu trúc và chức năng sinh học của tế bào. Việc dự đoán chính xác các tương tác vẫn đang là một thách thức. Trong bài báo này, chúng tôi giới thiệu một phương pháp kết hợp giữa biến đổi Hilbert trên thông tin tiến hóa cho bước trích xuất đặc trưng và mô hình phân loại Deep Forest (DF) cho bước dự đoán tương tác. Việc sử dụng biến đổi Hilbert là nhằm tăng cường thông tin tiến hóa được lưu trữ trong trình tự protein. Bên cạnh đó, DF, một mô hình phân loại tổng hợp mạnh, được sử dụng để cải thiện hiệu suất dự đoán. Phương pháp đề xuất đã được đánh giá trên bộ dữ liệu chuẩn Yeast và so sánh với các phương pháp hiện có khác. Kết quả cho thấy phương pháp của chúng tôi đã đạt hiệu quả cao trong dự đoán PPI. Các bộ dữ liệu và mã nguồn của phương pháp được cung cấp tại https://github.com/thnhanhub/FeatPSSM.git.
Từ khóa: Tương tác protein-protein, Học tổng hợp, Biến đổi Hilbert

Article Details

Tài liệu tham khảo

An, J.-Y., Zhou, Y., Zhao, Y.-J., & Yan, Z.-J. (2019). An efficient feature extraction technique based on local coding PSSM and multifeatures fusion for predicting protein–protein interactions. Evolutionary Bioinformatics, 15, 1176934319879920. https://doi.org/10.1177/1176934319879920

Carnes, R. M., Kesterson, R. A., Korf, B. R., Mobley, J. A., & Wallis, D. (2019). Affinity purification of NF1 protein–protein interactors identifies keratins and neurofibromin itself as binding partners. Genes, 10(9), 650. https://doi.org/10.3390/genes10090650

Castel, P., Holtz-Morris, A., Kwon, Y., Suter, B. P., & McCormick, F. (2021). DoMY-Seq: A yeast two-hybrid–based technique for precision mapping of protein–protein interaction motifs. Journal of Biological Chemistry, 296, 100023. https://doi.org/10.1074/jbc.RA120.014284

Chakraborty, A., Mitra, S., Bhattacharjee, M., De, D., & Pal, A. J. (2021). Determining protein–protein interaction using support vector machine: A review. IEEE Access, 9, 12473–12490. https://doi.org/10.1109/ACCESS.2021.3051006

Chakraborty, A., Mitra, S., Bhattacharjee, M., De, D., & Pal, A. J. (2023). Determining human-coronavirus protein–protein interaction using machine intelligence. Medicine in Novel Technology and Devices, 18, 100228. https://doi.org/10.1016/j.medntd.2023.100228

Chen, Z.-H., You, Z.-H., Li, L.-P., Wang, Y.-B., Wong, L., & Yi, H.-C. (2019). Prediction of self-interacting proteins from protein sequence information based on random projection model and fast Fourier transform. International Journal of Molecular Sciences, 20(4), 930. https://doi.org/10.3390/ijms20040930

Copeland, R. A. (1994). Introduction to protein structure. In J. P. J. M. van Os (Ed.), Methods for protein analysis (pp. 1–16). Springer. https://doi.org/10.1007/978-1-4757-1505-7_1

De Las Rivas, J., & Fontanillo, C. (2010). Protein–protein interactions essentials: Key concepts to building and analyzing interactome networks. PLoS Computational Biology, 6(6), e1000807. https://doi.org/10.1371/journal.pcbi.1000807

Fu, L., Niu, B., Zhu, Z., Wu, S., & Li, W. (2012). CD-HIT: Accelerated for clustering the next-generation sequencing data. Bioinformatics, 28(23), 3150–3152. https://doi.org/10.1093/bioinformatics/bts565

Gelfand, M. S. (1995). Prediction of function in DNA sequence analysis. Journal of Computational Biology, 2(1), 87–115. https://doi.org/10.1089/cmb.1995.2.87

Gribskov, M., McLachlan, A. D., & Eisenberg, D. (1987). Profile analysis: Detection of distantly related proteins. Proceedings of the National Academy of Sciences, 84(13), 4355–4358. https://doi.org/10.1073/pnas.84.13.4355

Guo, Y., Yu, L., Wen, Z., & Li, M. (2008). Using support vector machine combined with auto covariance to predict protein–protein interactions from protein sequences. Nucleic Acids Research, 36(9), 3025–3030. https://doi.org/10.1093/nar/gkn159

Huang, Y.-A., You, Z.-H., Gao, X., Wong, L., & Wang, L. (2015). Using weighted sparse representation model combined with discrete cosine transformation to predict protein–protein interactions from protein sequence. BioMed Research International, 2015, 902198. https://doi.org/10.1155/2015/902198

Kido, K. (2015). Discrete Fourier transform. In Digital Fourier analysis: Fundamentals (pp. 77–105). Springer. https://doi.org/10.1007/978-1-4614-9260-3_4

Li, J.-Q., You, Z.-H., Li, X., Ming, Z., & Chen, X. (2017). PSPEL: In silico prediction of self-interacting proteins from amino acid sequences using ensemble learning. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 14(5), 1165–1172. https://doi.org/10.1109/TCBB.2017.2649529

Lin, X., & Chen, X. (2013). Heterogeneous data integration by tree-augmented naïve Bayes for protein–protein interactions prediction. Proteomics, 13(2), 261–268. https://doi.org/10.1002/pmic.201200326

Pan, J., Li, L.-P., Yu, C.-Q., You, Z.-H., Ren, Z.-H., & Tang, J.-Y. (2021). FWHT-RF: A novel computational approach to predict plant protein–protein interactions via an ensemble learning method. Scientific Programming, 2021, Article 1607946. https://doi.org/10.1155/2021/1607946

Pan, J., Wang, S., Yu, C., Li, L., You, Z., & Sun, Y. (2022). A novel ensemble learning-based computational method to predict protein–protein interactions from protein primary sequences. Biology, 11(5), 775. https://doi.org/10.3390/biology11050775

Rodriguez, J. J., Kuncheva, L. I., & Alonso, C. J. (2006). Rotation forest: A new classifier ensemble method. IEEE Transactions on Pattern Analysis and Machine Intelligence, 28(10), 1619–1630. https://doi.org/10.1109/TPAMI.2006.211

Romero-Molina, S., Ruiz-Blanco, Y. B., Harms, M., Münch, J., & Sanchez-Garcia, E. (2019). PPI-Detect: A support vector machine model for sequence-based prediction of protein–protein interactions. Journal of Computational Chemistry, 40(11), 1233–1242. https://doi.org/10.1002/jcc.25780

Shen, J., Zhang, J., Luo, X., Zhu, W., Yu, K., Chen, K., Li, Y., & Jiang, H. (2007). Predicting protein–protein interactions based only on sequences information. Proceedings of the National Academy of Sciences, 104(11), 4337–4341. https://doi.org/10.1073/pnas.0607879104

Uetz, P., Titz, B., & Cagney, G. (2008). Experimental methods for protein interaction identification and characterization. In Protein–protein interactions: Methods and applications (pp. 1–32). Humana Press. https://doi.org/10.1007/978-1-84800-125-1_1

Wang, L., Zhang, M., Yang, C., Li, Y., & Wang, Y. (2018). Using two-dimensional principal component analysis and rotation forest for prediction of protein–protein interactions. Scientific Reports, 8(1), 12874. https://doi.org/10.1038/s41598-018-30694-1

Wang, Y., Cheng, J., Liu, Y., & Chen, Y. (2016). Prediction of protein secondary structure using support vector machine with PSSM profiles. In 2016 IEEE Information Technology, Networking, Electronic and Automation Control Conference (ITNEC) (pp. 502–505). IEEE. https://doi.org/10.1109/ITNEC.2016.7560411

Xenarios, I., Salwínski, Ł., Duan, X. J., Higney, P., Kim, S.-M., & Eisenberg, D. (2002). DIP, the database of interacting proteins: A research tool for studying cellular networks of protein interactions. Nucleic Acids Research, 30(1), 303–305. https://doi.org/10.1093/nar/30.1.303

Yakubu, R. R., Nieves, E., & Weiss, L. M. (2019). The methods employed in mass spectrometric analysis of posttranslational modifications (PTMs) and protein–protein interactions (PPIs). In Post-translational modifications in health and disease (pp. 169–198). Springer. https://doi.org/10.1007/978-3-030-15950-4_10

Yang, X., Yang, S., Li, Q., Wuchty, S., & Zhang, Z. (2020). Prediction of human-virus protein–protein interactions through a sequence embedding-based machine learning method. Computational and Structural Biotechnology Journal, 18, 153–161. https://doi.org/10.1016/j.csbj.2019.12.005

You, Z.-H., Lei, Y.-K., Zhu, L., Xia, J., & Wang, B. (2013). Prediction of protein–protein interactions from amino acid sequences with ensemble extreme learning machines and principal component analysis. BMC Bioinformatics, 14(Suppl. 8), S10. https://doi.org/10.1186/1471-2105-14-S8-S10

Zeng, M., Zhang, F., Wu, F.-X., Li, Y., Wang, J., & Li, M. (2020). Protein–protein interaction site prediction through combining local and global features with deep neural networks. Bioinformatics, 36(4), 1114–1120. https://doi.org/10.1093/bioinformatics/btz699

Zhao, T.-H., Shi, X.-H., Wang, L., Zhang, Y.-N., & cộng sự. (2013). A novel method of predicting protein disordered regions based on sequence features. BioMed Research International, 2013, Article 414327. https://doi.org/10.1155/2013/414327

Zhou, Y. Z., Gao, Y., & Zheng, Y. Y. (2011). Prediction of protein–protein interactions using local description of amino acid sequence. In Advanced Data Mining and Applications (pp. 254–262). Springer. https://doi.org/10.1007/978-3-642-22456-0_37

Zhou, Z.-H., & Feng, J. (2019). Deep forest. National Science Review, 6(1), 74–86. https://doi.org/10.1093/nsr/nwy108