COSHAP: An Explainable AI Framework Integrating Conformal Prediction and SHAP for Uncertainty-Aware Consumer Segmentation
DOI:
https://doi.org/10.59796/jcst.V16N4.2026.213Keywords:
conformal prediction, consumer segmentation, explainable artificial intelligence, machine learning, SHAP, RFM analysis, uncertainty quantificationAbstract
Data-driven decision-making in digital marketing is often constrained by the lack of transparency and uncertainty estimation in conventional machine learning models. This study proposes COSHAP (Conformal Segmentation with Hierarchical Attribution Paths), a unified framework that integrates Conformal Prediction (CP) and SHAP to model uncertainty while enhancing interpretability in customer segmentation. Using the Online Retail II dataset, segmentation was performed based on Recency, Frequency, Monetary (RFM), Tenure, and Average Basket Size features using K-Means, which was subsequently approximated using supervised models (Random Forest, XGBoost, Naïve Bayes, and SVM). CP was applied to generate prediction sets with guaranteed statistical coverage, while SHAP analysis was extended by conditioning feature attributions on prediction set membership through Hierarchical Attribution Paths. Experimental results demonstrated that the XGBoost model achieved the highest approximation accuracy of 97.36%. At a 95% confidence level, the framework achieved empirical coverage consistent with the statistical target for the ensemble models, with an average prediction set size close to 1.0, indicating high certainty in the majority of cases. However, a subset of customers at decision boundaries yielded multi-label predictions, accurately reflecting behavioral ambiguity. SHAP analysis identified Recency and Frequency as the primary drivers of segmentation, whereas COSHAP successfully uncovered the push-pull feature dynamics that trigger prediction uncertainty. The COSHAP framework effectively transforms point-based segmentation into a reliable, calibrated, and transparent decision-support system.
References
Al-Kerboly, D. M. A., Hamad, M. M., & Dawood, O. A. (2023). Clustering algorithms comparison for University of Anbar researchers’ Google Scholar profiles. In 2023 Al-Sadiq International Conference on Communication and Information Technology (AICCIT). IEEE. https://doi.org/10.1109/AICCIT57614.2023.10217951
Alves Gomes, M., & Meisen, T. (2023). A review on customer segmentation methods for personalized customer targeting in e-commerce use cases. Information Systems and e-Business Management, 21, 527–570. https://doi.org/10.1007/s10257-023-00640-4
Battur, S., & Totad, S. G. (2026). Data clustering algorithms for incremental datasets: A comparative study. In International Conference on Data Science and Applications (pp. 421–433). Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-10940-8_33
Chaudhuri, S. E., Ben Chaouch, Z., Hauber, B., Mange, B., Zhou, M., Christopher, S., et al. (2025). Use of Bayesian decision analysis to maximize value in patient-centered randomized clinical trials in Parkinson’s disease. Journal of Biopharmaceutical Statistics, 35(5), 981–1000. https://doi.org/10.1080/10543406.2023.2170400
Chen, D. (2012). Online Retail II [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5CG6D
Cohen, K. M., Park, S., Simeone, O., & Shitz, S. S. (2023). Calibrating AI models for wireless communications via conformal prediction. IEEE Transactions on Machine Learning in Communications and Networking, 1, 296-312. https://doi.org/10.1109/TMLCN.2023.3319282
Correia, A. H., Massoli, F. V., Louizos, C., & Behboodi, A. (2024). An information theoretic perspective on conformal prediction. Advances in Neural Information Processing Systems, 37, 101000-101041. https://doi.org/10.52202/079017-3203
Dracup, C. (2007). Confidence intervals. Encyclopedia of Statistics in Quality and Reliability. Wiley. https://doi.org/10.1002/9780470061572.eqr218
Efendioğlu, İ. (2023). The change of digital marketing with artificial intelligence [Conference presentation]. The 7th International Conference on Applied Research in Management, Economics and Accounting. Dublin Republic of Ireland. https://doi.org/10.33422/7th.iarmea.2023.07.101
Farag, M., Emam, A., Leonhardt, J., & Roscher, R. (2025). Enhancing decision support in crop production: Analyzing conformal prediction for uncertainty quantification. Computers and Electronics in Agriculture, 237, Article 110559. https://doi.org/10.1016/j.compag.2025.110559
Hu, X., Liu, A., Li, X., Dai, Y., & Nakao, M. (2023). Explainable AI for customer segmentation in product development. CIRP Annals - Manufacturing Technology, 72(1), 89-92. https://doi.org/10.1016/j.cirp.2023.03.004
Karahan, S. N., Güllü, M., Karhan, D., Cimen, S., Osmanca, M. S., & Barışçı, N. (2025). Realistic performance assessment of machine learning algorithms for 6G network slicing: A dual-methodology approach with explainable AI integration. Electronics, 14(19), Article 3841. https://doi.org/10.3390/electronics14193841
Kendall, A., & Gal, Y. (2017). What uncertainties do we need in bayesian deep learning for computer vision? [Conference presentation]. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA USA.
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., & Viegas, F. (2018). Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV) [Conference presentation]. International Conference on Machine Learning. PMLR. Stockholm, Sweden.
Kumar, N. (2025). Intelligent customer segmentation: Unveiling consumer patterns with machine learning. Journal of Umm Al-Qura University for Engineering and Architecture, 16(3), 774-783. https://doi.org/10.1007/s43995-025-00180-7
Levenstien, M. A., Yang, Y., & Ott, J. (2003). Statistical significance for hierarchical clustering in genetic association and microarray expression studies. BMC Bioinformatics, 4(1), Article 62. https://doi.org/10.1186/1471-2105-4-62
Ljepava, N. (2022). AI-enabled marketing solutions in marketing decision making: AI application in different stages of marketing process. TEM Journal, 11(3), 1308-1315. https://doi.org/10.18421/TEM113-40
Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions [Conference presentation]. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.
Lv, M. (2022). Application of an K-means improved clustering analysis algorithm in the design of resource management information system [Conference presentation]. 2022 World Automation Congress (WAC). IEEE. San Antonio, TX, USA. https://doi.org/10.23919/WAC55640.2022.9934387
Molnar, C. (2019). Interpretable machine learning, a guide for making black box models explainable. Retrieved from https://christophm.github.io/interpretable-ml-book/shap.html
Msaddi, S., & Kumbasar, T. (2023). Generating high-quality prediction intervals for regression tasks via fuzzy C-means clustering-based conformal prediction. In International Conference on Intelligent and Fuzzy Systems (pp. 728–739). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-39777-6_63
Pakpahan, S. I., Mukid, M. A., Saputra, B. A., & Rochayani, M. Y. (2025). Simulation study and application to compare the performance of the K-means and K-medoids algorithms [Conference presentation]. 2025 International Conference on Computer Sciences, Engineering, and Technology Innovation (ICoCSETI). IEEE. Jakarta, Indonesia. https://doi.org/10.1109/ICoCSETI63724.2025.11019227
Portela, A., Banga, J. R., & Matabuena, M. (2025). Conformal prediction for uncertainty quantification in dynamic biological systems. PLOS Computational Biology, 21(5), e1013098. https://doi.org/10.1371/journal.pcbi.1013098
Pradeepa, P., & Jeyakumar, M. K. (2022). Data redundancy removal using KMAD-based self-tuning spectral clustering and CKD prediction using ML techniques. Journal of Current Science and Technology, 12(3), 517–537. https://doi.org/10.14456/jcst.2022.40
Resheff, Y. S., Rotics, S., Harel, R., Spiegel, O., & Nathan, R. (2014). AcceleRater: A web application for supervised learning of behavioral modes from acceleration measurements. Movement Ecology, 2(1), Article 27. https://doi.org/10.1186/s40462-014-0027-0
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should i trust you?” Explaining the predictions of any classifier [Conference presentation]. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York, US. https://doi.org/10.1145/2939672.2939778
Romano, Y., Sesia, M., & Candes, E. (2020). Classification with valid and adaptive coverage. Advances in Neural Information Processing Systems, 33, 3581-3591.
Salih, A. M., Raisi‐Estabragh, Z., Galazzo, I. B., Radeva, P., Petersen, S. E., Lekadir, K., & Menegaz, G. (2025). A perspective on explainable artificial intelligence methods: SHAP and LIME. Advanced Intelligent Systems, 7(1), Article 2400304. https://doi.org/10.1002/aisy.202400304
Sharma, P., Mirzan, S. R., Bhandari, A., Pimpley, A., Eswaran, A., Srinivasan, S., & Shao, L. (2020). Evaluating tree explanation methods for anomaly reasoning: A case study of SHAP TreeExplainer and TreeInterpreter [Conference presentation]. International Conference on Conceptual Modeling. Vienna Austria. https://doi.org/10.1007/978-3-030-65847-2_4
Szepannek, G., & Lübke, K. (2022). Explaining artificial intelligence with care: Analyzing the explainability of black box multiclass machine learning models in forensics. KI-Künstliche Intelligenz, 36(2), 125-134. https://doi.org/10.1007/s13218-022-00764-8
Takaew, S., & Romsaiyud, W. (2025). An explainable approach to sentiment analysis of Thai hotel reviews using a fine-tuned language model and SHAP. Journal of Current Science and Technology, 15(4), Article 149. https://doi.org/10.59796/jcst.v15n4.2025.149
Thakur, Y., & Mittal, N. (2024). Customer segmentation using K-mean clustering [Conference presentation]. Biennial International Conference on Future Learning Aspects of Mechanical Engineering. Singapore: Springer Nature Singapore. https://doi.org/10.1007/978-981-96-7214-1_32
Vovk, V., Gammerman, A., & Shafer, G. (2005). Algorithmic learning in a random world. Boston, MA: Springer US. https://doi.org/10.1007/b106715
Wang, Q. (2022). Support vector machine algorithm in machine learning [Conference presentation]. 2022 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA). IEEE. Dalian, China. https://doi.org/10.1109/ICAICA54878.2022.9844516
Downloads
Published
How to Cite
Issue
Section
Categories
License
Copyright (c) 2026 Journal of Current Science and Technology

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.


