CODE SLICING TO IMPROVE CASE CLASSIFICATION ACCURACY

Authors

  • Omar Abdalgani Shiba
  • Mohamed Nasir Sulaiman Faculty of Computer Science and Information Technology University Putra Malaysia
  • Fatimah Ahmad Faculty of Computer Science and Information Technology University Putra Malaysia
  • Ali Mamat Faculty of Computer Science and Information Technology University Putra Malaysia

Keywords:

Data mining, case-slicing technique, classification accuracy, case classification

Abstract

Finding a good classification algorithm is an important component of many data mining projects. Data mining researchers often use classifiers to identify important classes of objects within a data repository. The goal of this paper is to improve the case classification accuracy in data mining. The paper achieves this goal by introducing a new approach of similarity-based retrieval based on program slicing techniques and is called Case Slicing Technique (CST). The proposed approach helps identify the subset of features used to compute the similarity measures needed by the classification algorithms. The idea is based on slicing cases with respect to the slicing criterion. Likewise, the paper presents the experimental results of the CST using five real-world datasets which are; Australian Credit Application (AUS), Cleveland Heart Disease (CLEV), Breast Cancer (BCO), German Credit Card (GERM) and Hepatitis Domain (HEPA). The paper compares CST with other selected approaches. The results obtained showed that the classification accuracy can be improved when we use CST.

References

Batchelor, B. (1978). Pattern recognition: ideas in practice. Plenum Press, New York, p. 71-72.

Biberman, Y. (1994). A Context similarity measure. Proceedings of the European Conference on Machine Learning (ECML-94); April 6-8, 1994. Catalina, Italy. Springer Verlag, p. 49-63.

Ching, J.Y., Wong, A.K., and Chan, K.C. (1995). Class-dependent discretization for inductive learning from continuous and mixed mode data. IEEE Transactions on Pattern Analysis and Machine Intelligence, 17(7):641-651.

Cost, S., and Salzberg, S. (1993). A weighted nearest neighbor algorithm for learning with symbolic features. Machine Learning, 10:57-78.

Domingos, P. (1995). Rule induction and instance-based learning: a unified approach. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI-95); August 20-11, 1995. Montreal, Canada; Morgan Kaufmann, p. 1,226-1,232.

Dougherty, J., Kohavi, R., and Sahami, M. (1995). Supervised and unsupervised discretization of continuous features. Proceedings of the 12th International Conference on Machine Learning; July 9-12, 1995. San Francisco, Tahoe City; CA: Morgan Kaufmann, p. 194-202.

Edwin, D. (1974). Recent progress in distance and similarity measures in pattern recognition. Proceedings of 2nd International Joint Conference on Pattern Recognition; July 4, 1974. Tokyo, p. 534-539.

Fayyad, U.M., and Irani, K.B. (1992). On the handling of continuous-valued attributes in decision tree generation. Machine Learning, p. 8:87-102.

Gallagher, K., and Lyle, J. (1991). Using program slicing in software maintenance. IEEE Transaction on Software Engineering, 17(8):751-761.

Horwitz, S., Reps, T., and Binkley, D. (1990). Interprocedural slicing using dependence graphs. ACM Transactions on Programming Languages and Systems, 12(1):35-46.

Kohavi, R. (1995). A Study of cross-validation and bootstrap for accuracy estimation and model selection. Proceedings of 14th International Joint Conference on Artificial Intelligence (IJCAI). San Mateo CA: Morgan Kaufmann, p. 1,137-1,145.

Michalski, R., Robert, E., and Edwin, D. (1981). A Recent advance in data analysis: clustering objects into classes characterized by conjunctive concepts. Progress in Pattern Recognition, 1, Kanal, L.N., and Rosenfeld, A. (eds.). North- Holland, New York, p. 33-56.

Mitchell, T.M. (1997). Machine learning. McGraw-Hill, New York.

Mohri, T., and Tanaka, H. (1994). An optimalweighting criterion of case indexing for both numeric and symbolic attributes. David, W.A (ed.), Case-Based Reasoning: Papers from the 1994 Workshop (Report No. WS-94-01). Menlo Park, CA: AIII Press, p.123-127.

Murphy, P.M. (2001). UCI Repositories of Machine Learning and Domain Theories. University of California, Irvine UCI. Available from:http://www.isc.uci.edu/ ~mlearn/MLRepository.html. Accessed November 12, 2001.

Nadler, M., and Eric, P. (1993). Pattern recognition engineering. Wiley, New York. p. 293-294.

Pfahringer, B. (1995). Compression-based discretization of continuous attributes. Proceedings of the 12th International Conference on Machine Learning. US, Lake Tahoe, p. 456-463.

Quinlan, J.R. (1993). C4.5: programs for machine learning. CA: Morgan Kaufmann Publishers, Inc.

Quinlan, J.R. (1986). Induction of decision trees. Machine Learning, 1(1):81-106.

Rachlin, J., Simon, K., Salzberg, S., and David, W. (1994). Towards a better understanding of memory-based and bayesian classifiers. Proceedings of the Eleventh International Machine Learning Conference. New Brunswick, NJ: Morgan Kaufmann, p. 242-250.

Salzberg, S. (1991). A nearest hyperrectangle learning method. Machine Learning, 6:277-309.

Tapia, R., and Thompson, J. (1978). Nonparametric probability density estimation, Baltimore. MD: The Johns Hopkins University Press.

Thamar, S., and Olac, F. (2001). Improving classifier accuracy using unlabeled data. Proceedings of the IASTED International Conference on Artificial Intelligence and Applications (AIA2001). Marbella, Spain.

Thamar, S., and Olac, F. (2002). Improving classification accuracy of large test sets using the ordered classification algorithm. Proceedings of IBERAMIA-02, Lecture Notes in Computer Science, Sevilla Spain, Springer-Verlag Heidelberg, p. 70-79.

Tip, F. (1995). A Survey of program slicing techniques. Journal of Programming Languages, 3:121-189.

Tversky, A. (1977). Features of similarity. Psychological Review, 84(4):327-352.

Vasconcelos, W.W. (2000). Slicing knowledgebased systems techniques and applications. Knowledge-Based Systems Journal, 13:177-198.

Weiser, M. (1984). Program slicing. IEEE Trans. Software Engineering, 10(4):352-357.

Wettschereck, D., and Aha, D.W. (1995). Weighting features. Proceedings of the 1st. International Conference on CBR (ICCBR-95). Sesimbra, Portugal. Springer-Verlag, p. 347-358.

Wettschereck, D., and Thomas, G.D. (1995). An experimental comparison of nearest neighbor and nearest-hyperrectangle algorithms. Machine Learning, 19(1): 5-28.

Wilson, D.R., and Martinez, T.R. (1996). Value difference metrics for continuously valued attributes. Proceedings of the International Conference on Artificial Intelligence, Expert Systems and Neural Networks, p. 11-14.

Xiaoli, Q.A. (1999). Case-based reasoning system for bearing design, [MSc. thesis]. Faculty of Computer Science, Drexel University, Philadelphia, PA.

Downloads

Published

2026-08-27

How to Cite

Abdalgani Shiba, O., Nasir Sulaiman, M., Ahmad, F., & Mamat, A. (2026). CODE SLICING TO IMPROVE CASE CLASSIFICATION ACCURACY. Suranaree Journal of Science and Technology, 11(2), 107–114. retrieved from https://ph04.tci-thaijo.org/index.php/SUJST/article/view/13141

Issue

Section

Research Article