C400KNN: Predictive Model for Enhancing Anti-HIV Peptide Classification

Authors

  • Jaru Nikom Department of Mathematics and Computer Science, Faculty of Science and Technology, Prince of Songkla University, Pattani Campus, Pattani 94000, Thailand https://orcid.org/0009-0002-8147-6782
  • Watshara Shoombuatong Center for Research Innovation and Biomedical Informatics, Faculty of Medical Technology, Mahidol University, Bangkok 10700, Thailand https://orcid.org/0000-0002-3394-8709
  • Phasit Charoenkwan Modern Management and Information Technology, College of Arts, Media and Technology, Chiang Mai University, Chiang Mai 50200, Thailand https://orcid.org/0000-0002-5656-2796
  • Salang Musikasuwan Department of Mathematics and Computer Science, Faculty of Science and Technology, Prince of Songkla University, Pattani Campus, Pattani 94000, Thailand https://orcid.org/0000-0002-8733-9626

DOI:

https://doi.org/10.59796/jcst.V16N2.2026.191

Keywords:

predictive model, anti-HIV peptide, machine learning, chi-square

Abstract

The design of peptide sequences to perform anti-HIV activities is a particularly time-consuming stage in the production of AIDS medications. One way to solve this issue is by computer modeling. Predictive models for anti-HIV peptides help decrease the time and cost of producing anti-HIV peptide drugs. This paper introduces the C400KNN model, developed by a machine learning approach that uses non-antimicrobial peptides as the negative dataset and incorporates a feature selection procedure of 400 variables using the chi-square test, followed by training with the KNN algorithm. The amino acid sequences were gathered from databases containing anti-HIV, antiviral, and antimicrobial information. Next, all features were extracted from 12 descriptors, encompassing both the chemical structure and physicochemical properties. Ten classifiers were used to create the models. The models were evaluated to discover descriptors using AUC and ACC. Then, the features of those descriptors were combined, and feature selection methods were used to choose the most important features. The best model was selected based on its performance. The findings indicate that the descriptor groups that improve model efficiency are AAC, DPC, PAAC, and APAAC. The Chi-square method was used in the feature selection step.   Additionally, the accuracy of the model was found to be 0.83. The C400KNN model is thought to be effective and might help researchers in creating anti-HIV peptides for use in pharmaceutical manufacturing.

References

Charoenkwan, P., Anuwongcharoen, N., Nantasenamat, C., Hasan, M. M., & Shoombuatong, W. (2021). In silico approaches for the prediction and analysis of antiviral peptides: A review. Current Pharmaceutical Design, 27(18), 2180-2188. https://doi.org/10.2174/1381612826666201102105827

Charoenkwan, P., Nantasenamat, C., Hasan, M. M., & Shoombuatong, W. (2020). Meta-iPVP: A sequence-based meta-predictor for improving the prediction of phage virion proteins using effective feature representation. Journal of Computer-Aided Molecular Design, 34(10), 1105-1116. https://doi.org/10.1007/s10822-020-00323-z

Charoenkwan, P., Schaduangrat, N., & Shoombuatong, W. (2023). StackTTCA: A stacking ensemble learning-based framework for accurate and high-throughput identification of tumor T cell antigens. BMC Bioinformatics, 24(1), Article 301. https://doi.org/10.1186/s12859-023-05421-x

Chen, Z., Zhao, P., Li, F., Leier, A., Marquez-Lago, T. T., Wang, Y., ... & Song, J. (2018). iFeature: A python package and web server for features extraction and selection from protein and peptide sequences. Bioinformatics, 34(14), 2499-2502. https://doi.org/10.1093/bioinformatics/bty140

Chicco, D., & Jurman, G. (2023). The Matthews correlation coefficient (MCC) should replace the ROC AUC as the standard metric for assessing binary classification. BioData Mining, 16(1), Article 4. https://doi.org/10.1186/s13040-023-00322-4

Chou, K. C., Wu, Z. C., & Xiao, X. (2012). iLoc-Hum: Using the accumulation-label scale to predict subcellular locations of human proteins with both single and multiple sites. Molecular Biosystems, 8(2), 629-641. https://doi.org/10.1039/C1MB05420A

Deng, H., Ding, M., Wang, Y., Li, W., Liu, G., & Tang, Y. (2023). ACP-MLC: A two-level prediction engine for identification of anticancer peptides and multi-label classification of their functional types. Computers in Biology and Medicine, 158, Article 106844. https://doi.org/10.1016/j.compbiomed.2023.106844

Esmaeili, M., Mohabatkar, H., & Mohsenzadeh, S. (2010). Using the concept of Chou's pseudo amino acid composition for risk type prediction of human papillomaviruses. Journal of Theoretical Biology, 263(2), 203-209. https://doi.org/10.1016/j.jtbi.2009.11.016

Gen, M., & Lin, L. (2023). Genetic algorithms and their applications. Springer handbook of engineering statistics. London, UK: Springer London. https://doi.org/10.1007/978-1-4471-7503-2_33

Greenberg, M. L., & Cammack, N. (2004). Resistance to enfuvirtide, the first HIV fusion inhibitor. Journal of Antimicrobial Chemotherapy, 54(2), 333-340. https://doi.org/10.1093/jac/dkh330

Hatta, M., Wahid, W. N., Yusuf, F., Hidayat, F., Santoso, N. A., & Aini, Q. (2024). Enhancing predictive models in system development using machine learning algorithms. International Journal of Cyber and IT Service Management, 4(2), 80-87. https://doi.org/10.34306/ijcitsm.v4i2.159

Hemelaar, J. (2012). The origin and diversity of the HIV-1 pandemic. Trends in Molecular Medicine, 18(3), 182-192. https://doi.org/10.1016/j.molmed.2011.12.001

Huang, Y., Niu, B., Gao, Y., Fu, L., & Li, W. (2010). CD-HIT Suite: A web server for clustering and comparing biological sequences. Bioinformatics, 26(5), 680-682. https://doi.org/10.1093/bioinformatics/btq003

Kumar, V., & Dogra, N. (2022). A comprehensive review on deep synergistic drug prediction techniques for cancer. Archives of Computational Methods in Engineering, 29(3), 1443-1461. https://doi.org/10.1007/s11831-021-09617-3

Kwong, P. D., Wyatt, R., Robinson, J., Sweet, R. W., Sodroski, J., & Hendrickson, W. A. (1998). Structure of an HIV gp120 envelope glycoprotein in complex with the CD4 receptor and a neutralizing human antibody. Nature, 393(6686), 648-659. https://doi.org/10.1038/31405

Lama, J., & Planelles, V. (2007). Host factors influencing susceptibility to HIV infection and AIDS progression. Retrovirology, 4(1), Article 52. https://doi.org/10.1186/1742-4690-4-52

Lertampaiporn, S., Wattanapornprom, W., Thammarongtham, C., & Hongsthong, A. (2025). EnsembleNPPred: A robust approach to neuropeptide prediction and recognition using ensemble machine learning and deep learning methods. Life, 15(7), Article 1010. https://doi.org/10.3390/life15071010

Li, W., & Godzik, A. (2006). Cd-hit: A fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics, 22(13), 1658-1659. https://doi.org/10.1093/bioinformatics/btl158

Lissabet, J. F. B., Belén, L. H., & Farias, J. G. (2019). AntiVPP 1.0: A portable tool for prediction of antiviral peptides. Computers in Biology and Medicine, 107, 127-130. https://doi.org/10.1016/j.compbiomed.2019.02.011

Liu, Y., Ouyang, X. H., Xiao, Z. X., Zhang, L., & Cao, Y. (2020). A review on the methods of peptide-MHC binding prediction. Current Bioinformatics, 15(8), 878-888. https://doi.org/10.2174/1574893615999200429122801

Liu, Y., Zhu, Y., Sun, X., Ma, T., Lao, X., & Zheng, H. (2023). DRAVP: A comprehensive database of antiviral peptides and proteins. Viruses, 15(4), Article 820. https://doi.org/10.3390/v15040820

Mathiyazhagan, B., Liyaskar, J., Azar, A. T., Inbarani, H. H., Javed, Y., Kamal, N. A., & Fouad, K. M. (2022). Rough set based classification and feature selection using improved harmony search for peptide analysis and prediction of anti-HIV-1 activities. Applied Sciences, 12(4), Article 2020. https://doi.org/10.3390/app12042020

Mori, T., O'Keefe, B. R., Sowder, R. C., Bringans, S., Gardella, R., Berg, S., ... & Boyd, M. R. (2005). Isolation and characterization of griffithsin, a novel HIV-inactivating protein, from the red alga Griffithsia sp. Journal of Biological Chemistry, 280(10), 9345-9353. https://doi.org/10.1074/jbc.M411122200

Pang, Y., Yao, L., Jhong, J. H., Wang, Z., & Lee, T. Y. (2021). AVPIden: A new scheme for identification and functional prediction of antiviral peptides based on machine learning approaches. Briefings in Bioinformatics, 22(6), Article bbab263. https://doi.org/10.1093/bib/bbab263

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., ... & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. The Journal of Machine Learning Research, 12, 2825-2830.

Poorinmohammad, N., & Mohabatkar, H. (2015). A comparison of different machine learning algorithms for the prediction of anti-HIV-1 peptides based on their sequence-related properties. International Journal of Peptide Research and Therapeutics, 21(1), 57-62. https://doi.org/10.1007/s10989-014-9432-x

Poorinmohammad, N., Mohabatkar, H., Behbahani, M., & Biria, D. (2015). Computational prediction of anti HIV‐1 peptides and in vitro evaluation of anti HIV‐1 activity of HIV‐1 P24‐derived peptides. Journal of Peptide Science, 21(1), 10-16. https://doi.org/10.1002/psc.2712

Prabakaran, P., Dimitrov, A. S., Fouts, T. R., & Dimitrov, D. S. (2007). Structure and function of the HIV envelope glycoprotein as entry mediator, vaccine immunogen, and target for inhibitors. Advances in Pharmacology, 55, 33-97. https://doi.org/10.1016/S1054-3589(07)55002-7

Quaranta, L., Calefato, F., & Lanubile, F. (2021). Kgtorrent: A dataset of python jupyter notebooks from kaggle [Conference presentation]. 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR). IEEE, Madrid, Spain. https://doi.org/10.1109/MSR52588.2021.00072

Qureshi, A., Thakur, N., & Kumar, M. (2013). HIPdb: A database of experimentally validated HIV inhibiting peptides. PloS One, 8(1), Article e54908. https://doi.org/10.1371/journal.pone.0054908

Rambaut, A., Posada, D., Crandall, K. A., & Holmes, E. C. (2004). The causes and consequences of HIV evolution. Nature Reviews Genetics, 5(1), 52-61. https://doi.org/10.1038/nrg1246

Roberts, D. R., Bahn, V., Ciuti, S., Boyce, M. S., Elith, J., Guillera‐Arroita, G., ... & Dormann, C. F. (2017). Cross‐validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913-929. https://doi.org/10.1111/ecog.02881

Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., & Müller, K. R. (Eds.). (2019). Explainable AI: Interpreting, explaining and visualizing deep learning. Switzerland: Springer Nature. https://doi.org/10.1007/978-3-030-28954-6

Schaduangrat, N., Nantasenamat, C., Prachayasittikul, V., & Shoombuatong, W. (2019). Meta-iAVP: A sequence-based meta-predictor for improving the prediction of antiviral peptides using effective feature representation. International Journal of Molecular Sciences, 20(22), Article 5743. https://doi.org/10.3390/ijms20225743

Sharma, R., Shrivastava, S., Singh, S. K., Kumar, A., Singh, A. K., & Saxena, S. (2021). Deep-AVPpred: Artificial intelligence driven discovery of peptide drugs for viral infections. IEEE Journal of Biomedical and Health Informatics, 26(10), 5067-5074. https://doi.org/10.1109/JBHI.2021.3130825

Steffler, M., Li, Y., Weir, S., Shaikh, S., Murtada, F., Wright, J. G., & Kantarevic, J. (2021). Trends in prevalence of chronic disease and multimorbidity in Ontario, Canada. Canadian Medical Association Journal, 193(8), E270-E277. https://doi.org/10.1503/cmaj.201473

Thakur, N., Qureshi, A., & Kumar, M. (2012). AVPpred: Collection and prediction of highly effective antiviral peptides. Nucleic Acids Research, 40(W1), W199-W204. https://doi.org/10.1093/nar/gks450

Vamathevan, J., Clark, D., Czodrowski, P., Dunham, I., Ferran, E., Lee, G., ... & Zhao, S. (2019). Applications of machine learning in drug discovery and development. Nature Reviews Drug Discovery, 18(6), 463-477. https://doi.org/10.1038/s41573-019-0024-5

Van den Goorbergh, R., Van Smeden, M., Timmerman, D., & Van Calster, B. (2022). The harm of class imbalance corrections for risk prediction models: Illustration and simulation using logistic regression. Journal of the American Medical Informatics Association, 29(9), 1525-1534. https://doi.org/10.1093/jamia/ocac093

Wang, G., Li, X., & Wang, Z. (2016). APD3: The antimicrobial peptide database as a tool for research and education. Nucleic Acids Research, 44(D1), D1087-D1093. https://doi.org/10.1093/nar/gkv1278

Wang, W., Owen, S. M., Rudolph, D. L., Cole, A. M., Hong, T., Waring, A. J., ... & Lehrer, R. I. (2004). Activity of α-and θ-defensins against primary isolates of HIV-1. The Journal of Immunology, 173(1), 515-520. https://doi.org/10.4049/jimmunol.173.1.515

Wei, L., Zhou, C., Su, R., & Zou, Q. (2019). PEPred-Suite: Improved and robust prediction of therapeutic peptides using adaptive feature representation learning. Bioinformatics, 35(21), 4272-4280. https://doi.org/10.1093/bioinformatics/btz246

Xu, J., Li, F., Leier, A., Xiang, D., Shen, H. H., Marquez Lago, T. T., ... & Song, J. (2021). Comprehensive assessment of machine learning-based methods for predicting antimicrobial peptides. Briefings in Bioinformatics, 22(5), Article bbab083. https://doi.org/10.1093/bib/bbab083

Zhang, W., Ding, Y., Wei, L., Guo, X., & Ni, F. (2024). Therapeutic peptides identification via kernel risk sensitive loss-based k-nearest neighbor model and multi-Laplacian regularization. Briefings in Bioinformatics, 25(6), Article bbae534. https://doi.org/10.1093/bib/bbae534

Zhou, Y., Fujikura, K., Mkrtchian, S., & Lauschke, V. M. (2018). Computational methods for the pharmacogenetic interpretation of next generation sequencing data. Frontiers in Pharmacology, 9, Article 1437. https://doi.org/10.3389/fphar.2018.01437

Downloads

Published

2026-07-21

How to Cite

Nikom, J., Shoombuatong, W., Charoenkwan, P., & Musikasuwan, S. (2026). C400KNN: Predictive Model for Enhancing Anti-HIV Peptide Classification. Journal of Current Science and Technology, 16(3), 191. https://doi.org/10.59796/jcst.V16N2.2026.191

Issue

Section

Research Article