LARGE-SCALE COMPLEX NETWORK ANALYSIS TO INVESTIGATE ASSOCIATIONS OF DISORDERED PROTEINS AND SCALE-FREE NETWORK
Keywords:
Protein-protein interaction network, disordered protein, scale-free network, degree distribution, power-law formAbstract
In biomedical engineering studies, a protein-protein interaction network is widely used to identify biological processes and characterize associations among proteins. Typically, its structure forms a scale-free network which contains a few hub proteins and a lot of low-degree proteins. Thus, its degree distribution follows the power-law distribution. Investigating such a protein that might affect this structure may reveal an important protein for the whole network. In this research, we investigated disordered proteins whose removal results in a non-scale-free structure that may cause severe diseases like cancers. Later, a simple measure to identify these proteins was developed as an alternative method avoiding fitting a power-law distribution. A high performance for identifying disordered proteins was achieved by using the area under the ROC curve value greater than 0.9. In addition, the measure yielded a superior performance to the random selection procedure.
References
Amberger, J.S., Bocchini, C.A., Schiettecatte, F., Scott, A.F., and Hamosh, A. (2015). OMIM.org: Online Mendelian Inheritance in Man (OMIM ®), an online catalog of human genes and genetic disorders. Nucleic Acids Res., 43:789-798.
Bonacich, P. (2017). Power and centrality: a family of measures. Am. J. Sociol., 92(5):1,170-1,182.
Chawla, N.V., Bowyer, K.W., Hall, L.O., and Kegelmeyer, W.P. (2002). SMOTE: Synthetic minority over-sampling technique. J. Artif. Intell. Res., (16):321-357.
Haynes, C., Oldfield, C.J., Ji, F., Klitgord, N., Cusick, M.E., Radivojac, P., Uversky, V.N., Vidal, M., and Iakoucheva, L.M. (2006). Intrinsic disorder is a common feature of hub proteins from four eukaryotic interactomes. PLoS Comput. Biol., 2(8):0890-0901.
Kosch, D. and Schreiber, F. (2004). Comparison of centralities for biological network. Proc. Ger. Conf. Bioinform., p. 199–206.
Li, J., Feng, Y., Wang, X., Liu, W., Rong, L., and Bao, J. (2015). An overview of predictors for intrinsically disordered proteins over 2010-2014. Int. J. Mol. Sci., 16(1):23,446-23,462.
Plaimas, K., Eils, R., and König, R. (2010). Identifying essential genes in bacterial metabolic networks with machine learning methods. BMC Syst. Biol., l(4):56.
Province, J. (2015). Yeast protein-protein interaction network model based on biological experimental data. Appl. Math. Mech., 36:827-834.
Sickmeier, M., Hamilton, J.A., LeGall, T., Vacic, V., Cortese, M.S., Tantos, A., Szabo, B., Tompa, P., Chen, J., Uversky, V.N., Obradovic, Z., and Dunker, A.K. (2007). DisProt: The database of disordered proteins. Nucleic Acids Res., 35(Database Issue):D786-D793.
Soffer, S.N. and Vázquez, A. (2005). Network clustering coefficient without degree-correlation biases. Phys. Rev. E, 71(5):2-5.
Suratanee, A. and Plaimas, K. (2014). Identification of inflammatory bowel disease-related proteins using a reverse k-nearest neighbor search. J. Bioinf. Comput. Biol., 12(4):1450017.
Szklarczyk, D., Franceschini, A., Wyder, S., Forslund, K., Heller, D., Huerta-Cepas, J., Simonovic, M., Roth, A., Santos, A., Tsafou, K.P., Kuhn, M., Bork, P., Jensen, L.J., and Von Mering, C. (2015). STRING v10: Protein-protein interaction networks, integrated over the tree of life. Nucleic Acids Res., 43(D1): D447-D452.








