CANCER CLASSIFICATION BASED ON MICROARRAY DATA USING BINARY-TUBE TREE
Keywords:
Classification, microarray data, binary-tube tree, decision treeAbstract
Analysis of large gene expression data via machine learning has been widely acknowledged to accurately predict cancer classes from the microarray data which are cast as the numerical matrix from a collection of complex RNAs. This research proposes a novel continuous value binary-tube tree constructed from the core vector of multi-features and the tolerance radius of the specific cancer types. To validate this classifier, the accuracy measure and Matthews correlation coefficient measure are compared with the support vector machine and the decision tree classifier using k-folds cross-validation. The result shows that the binary-tube tree outperforms the support vector machine and the decision tree classifier for leukemia, colon cancer, lymphoma, and breast cancer; only the lung cancer data set shows no significant improvement.
References
Ben-Dor, A., Chor, B., Karp, R., and Yakhini, Z. (2003). Discovering local structure in gene expression data: the order-preserving submatrix problem. J. Comput. Biol., 10:373-384.
Bennet, J., Ganaprakasam, C., and Kumar, N. (2015). A hybrid approach for gene selection and classification using support vector machine. Int. Arab J. Inf. Techn., 12, 695-700.
Cho, S.-B. and Won, H.-H. (2003). Machine learning in DNA microarray analysis for cancer classification. Proceedings of the First Asia-Pacific Bioinformatics Conference on Bioinformatics 2003, 19:189-198.
Furey, T.S., Cristianini, N., Duffy, N., Bednarski, D.W., Schummer, M., and Haussler, D. (2000). Support vector machine classification and validation of cancer tissue samples using microarray expression data. Bioinformatics, 16:906.
Hong, J-H. and Cho, B-S. (2006). Multi-class cancer classification with OVR-support vector machines selected by naïve Bayes classifier. Lect. Notes Comput. Sci., 2:155-164.
Hornik, K., Buchta, C., and Zeileis, A. (2009). Open-source machine learning: R meets Weka. Computation. Stat., 24:225-232.
Hoshida, Y., Brunet, J-P., Tamayo, P., Golub, T.R., and Mesirov, J.P. (2007). Subclass mapping: Identifying common subtypes in independent disease data sets. PloS One, 2:e1195.
Kanchanasuk, S. and Sinapiromsaran, K. (2016). Recursive binary tube partitioning for classification. Proc. Adapt. Learn. Opt., 5:99-107.
Lewis, D.P., Jebara, T., and Noble, W.S. (2006). Support vector machine learning from heterogeneous data: an empirical analysis using protein sequence and structure. Bioinformatics, 22:2,753-2,760.
Lotte, F., Bougrain, L., Cichocki, A., Clerc, M., Congedo, M., Rakotomamonjy, A., and Yger, F. (2018). A review of classification algorithms for EEG-based brain-computer interfaces: a 10 year update. J. Neural Eng., 15:031005.
Matthews, B.W. (1975). Comparison of the predicted and observed secondary structure of T4 phage lysozyme. BBA-Protein Structure, 405:442-451.
Meyer, D., Dimitriadou, E., Hornik, K., Weingessel, A., Leisch, F., Chang, C.C., and Lin, C.C. (2017). e1071: Misc Functions of the Department of Statistics, Probability Theory Group. R Foundation for Statistical Computing, Vienna, Austria.
Mohamad, M.S. and Deris, S. (2005). A hybrid of genetic algorithm and support vector machine for features selection and classification of gene expression microarray. Int. J. Comput. Intel. Appl., 5:91-106.
Mukherjee, S. (2003). Classifying microarray data using support vector machines. In: A Practical Approach to Microarray Data Analysis. Berrar, D.P., Dubitzky, W., and Granzow, M. (eds). Springer Science+Business Media, Boston, MA, USA, p. 166-185.
Nguyen, D.V. and Rocke, D.M. (2002). Tumor classification by partial least squares using microarray gene expression data. Bioinformatics, 18:39.
Painsky, A. and Rosset, S. (2014). Optimal set cover formulation for exclusive row biclustering of gene expression. J. Comput. Sci. Technol., 29:423-435.
Quinlan, J.R. (1993). C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers, Burlington, MA, USA, 302p.
R Core Team. (2015). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria., 55:275-286.
Shipp, M.A., Ross, K.N., Tamayo, P., Weng, A.P., Kutok, J.L., Aguiar, R.C., Gaasenbeek, M., Angelo, M., Reich, M., Pinkus, G.S., Ray, T.S., Koval, M.A., Last, K.W., Norton, A., Lister, T.A., Mesirov, J., Neuberg, D.S., Lander, E.S., Aster, J.C., and Golub, T.R. (2002). Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning. Nat. Med., 8:68-74.
Tarca, A.L., Romero, R., and Draghici, S. (2006). Analysis of microarray experiments of gene expression profiling. Am. J. Obstet. Gynecol., 195:373-388.
van 't Veer, L.J., Dai, H., van de Vijver, M.J., He, Y.D., Hart, A.A., Mao, M., Peterse, H.L., van der Kooy, K., Marton, M.J., Witteveen, A.T., Schreiber, G.J., Kerkhoven, R.M., Roberts, C., Linsley, P.S., Bernards, R., and Friend S.H. (2002). Gene expression profiling predicts clinical outcome of breast cancer. Nature, 415:530-536.
Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics Bull., 1:80-83.
Witten, I.H., Frank, E., Hall, M.A., and Pal, C.J. (2016). Data Mining: Practical Machine Learning Tools and Techniques. 4th ed. Morgan Kaufmann Publishers, Burlington, MA, USA, 654p.
Wu, X. and Kumar, V. (2009). The Top Ten Algorithms in Data Mining. CRC Press, Boca Raton, FL, USA, 232p.
Zhao, Y-H., Wang, G-R., Yin, Y., and Xu, G-Y. (2007). A novel approach to revealing positive and negative co-regulated genes. J. Comput. Sci. Technol., 22:261-272.








