COSHAP: An Explainable AI Framework Integrating Conformal Prediction and SHAP for Uncertainty-Aware Consumer Segmentation

Authors

DOI:

https://doi.org/10.59796/jcst.V16N4.2026.213

Keywords:

conformal prediction, consumer segmentation, explainable artificial intelligence, machine learning, SHAP, RFM analysis, uncertainty quantification

Abstract

Data-driven decision-making in digital marketing is often constrained by the lack of transparency and uncertainty estimation in conventional machine learning models. This study proposes COSHAP (Conformal Segmentation with Hierarchical Attribution Paths), a unified framework that integrates Conformal Prediction (CP) and SHAP to model uncertainty while enhancing interpretability in customer segmentation. Using the Online Retail II dataset, segmentation was performed based on Recency, Frequency, Monetary (RFM), Tenure, and Average Basket Size features using K-Means, which was subsequently approximated using supervised models (Random Forest, XGBoost, Naïve Bayes, and SVM). CP was applied to generate prediction sets with guaranteed statistical coverage, while SHAP analysis was extended by conditioning feature attributions on prediction set membership through Hierarchical Attribution Paths. Experimental results demonstrated that the XGBoost model achieved the highest approximation accuracy of 97.36%. At a 95% confidence level, the framework achieved empirical coverage consistent with the statistical target for the ensemble models, with an average prediction set size close to 1.0, indicating high certainty in the majority of cases. However, a subset of customers at decision boundaries yielded multi-label predictions, accurately reflecting behavioral ambiguity. SHAP analysis identified Recency and Frequency as the primary drivers of segmentation, whereas COSHAP successfully uncovered the push-pull feature dynamics that trigger prediction uncertainty. The COSHAP framework effectively transforms point-based segmentation into a reliable, calibrated, and transparent decision-support system.

References

Al-Kerboly, D. M. A., Hamad, M. M., & Dawood, O. A. (2023). Clustering algorithms comparison for University of Anbar researchers’ Google Scholar profiles. In 2023 Al-Sadiq International Conference on Communication and Information Technology (AICCIT). IEEE. https://doi.org/10.1109/AICCIT57614.2023.10217951

Alves Gomes, M., & Meisen, T. (2023). A review on customer segmentation methods for personalized customer targeting in e-commerce use cases. Information Systems and e-Business Management, 21, 527–570. https://doi.org/10.1007/s10257-023-00640-4

Battur, S., & Totad, S. G. (2026). Data clustering algorithms for incremental datasets: A comparative study. In International Conference on Data Science and Applications (pp. 421–433). Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-10940-8_33

Chaudhuri, S. E., Ben Chaouch, Z., Hauber, B., Mange, B., Zhou, M., Christopher, S., et al. (2025). Use of Bayesian decision analysis to maximize value in patient-centered randomized clinical trials in Parkinson’s disease. Journal of Biopharmaceutical Statistics, 35(5), 981–1000. https://doi.org/10.1080/10543406.2023.2170400

Chen, D. (2012). Online Retail II [Data set]. UCI Machine Learning Repository. https://doi.org/10.24432/C5CG6D

Cohen, K. M., Park, S., Simeone, O., & Shitz, S. S. (2023). Calibrating AI models for wireless communications via conformal prediction. IEEE Transactions on Machine Learning in Communications and Networking, 1, 296-312. https://doi.org/10.1109/TMLCN.2023.3319282

Correia, A. H., Massoli, F. V., Louizos, C., & Behboodi, A. (2024). An information theoretic perspective on conformal prediction. Advances in Neural Information Processing Systems, 37, 101000-101041. https://doi.org/10.52202/079017-3203

Dracup, C. (2007). Confidence intervals. Encyclopedia of Statistics in Quality and Reliability. Wiley. https://doi.org/10.1002/9780470061572.eqr218

Efendioğlu, İ. (2023). The change of digital marketing with artificial intelligence [Conference presentation]. The 7th International Conference on Applied Research in Management, Economics and Accounting. Dublin Republic of Ireland. https://doi.org/10.33422/7th.iarmea.2023.07.101

Farag, M., Emam, A., Leonhardt, J., & Roscher, R. (2025). Enhancing decision support in crop production: Analyzing conformal prediction for uncertainty quantification. Computers and Electronics in Agriculture, 237, Article 110559. https://doi.org/10.1016/j.compag.2025.110559

Hu, X., Liu, A., Li, X., Dai, Y., & Nakao, M. (2023). Explainable AI for customer segmentation in product development. CIRP Annals - Manufacturing Technology, 72(1), 89-92. https://doi.org/10.1016/j.cirp.2023.03.004

Karahan, S. N., Güllü, M., Karhan, D., Cimen, S., Osmanca, M. S., & Barışçı, N. (2025). Realistic performance assessment of machine learning algorithms for 6G network slicing: A dual-methodology approach with explainable AI integration. Electronics, 14(19), Article 3841. https://doi.org/10.3390/electronics14193841

Kendall, A., & Gal, Y. (2017). What uncertainties do we need in bayesian deep learning for computer vision? [Conference presentation]. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA USA.

Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., & Viegas, F. (2018). Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV) [Conference presentation]. International Conference on Machine Learning. PMLR. Stockholm, Sweden.

Kumar, N. (2025). Intelligent customer segmentation: Unveiling consumer patterns with machine learning. Journal of Umm Al-Qura University for Engineering and Architecture, 16(3), 774-783. https://doi.org/10.1007/s43995-025-00180-7

Levenstien, M. A., Yang, Y., & Ott, J. (2003). Statistical significance for hierarchical clustering in genetic association and microarray expression studies. BMC Bioinformatics, 4(1), Article 62. https://doi.org/10.1186/1471-2105-4-62

Ljepava, N. (2022). AI-enabled marketing solutions in marketing decision making: AI application in different stages of marketing process. TEM Journal, 11(3), 1308-1315. https://doi.org/10.18421/TEM113-40

Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions [Conference presentation]. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA.

Lv, M. (2022). Application of an K-means improved clustering analysis algorithm in the design of resource management information system [Conference presentation]. 2022 World Automation Congress (WAC). IEEE. San Antonio, TX, USA. https://doi.org/10.23919/WAC55640.2022.9934387

Molnar, C. (2019). Interpretable machine learning, a guide for making black box models explainable. Retrieved from https://christophm.github.io/interpretable-ml-book/shap.html

Msaddi, S., & Kumbasar, T. (2023). Generating high-quality prediction intervals for regression tasks via fuzzy C-means clustering-based conformal prediction. In International Conference on Intelligent and Fuzzy Systems (pp. 728–739). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-39777-6_63

Pakpahan, S. I., Mukid, M. A., Saputra, B. A., & Rochayani, M. Y. (2025). Simulation study and application to compare the performance of the K-means and K-medoids algorithms [Conference presentation]. 2025 International Conference on Computer Sciences, Engineering, and Technology Innovation (ICoCSETI). IEEE. Jakarta, Indonesia. https://doi.org/10.1109/ICoCSETI63724.2025.11019227

Portela, A., Banga, J. R., & Matabuena, M. (2025). Conformal prediction for uncertainty quantification in dynamic biological systems. PLOS Computational Biology, 21(5), e1013098. https://doi.org/10.1371/journal.pcbi.1013098

Pradeepa, P., & Jeyakumar, M. K. (2022). Data redundancy removal using KMAD-based self-tuning spectral clustering and CKD prediction using ML techniques. Journal of Current Science and Technology, 12(3), 517–537. https://doi.org/10.14456/jcst.2022.40

Resheff, Y. S., Rotics, S., Harel, R., Spiegel, O., & Nathan, R. (2014). AcceleRater: A web application for supervised learning of behavioral modes from acceleration measurements. Movement Ecology, 2(1), Article 27. https://doi.org/10.1186/s40462-014-0027-0

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should i trust you?” Explaining the predictions of any classifier [Conference presentation]. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York, US. https://doi.org/10.1145/2939672.2939778

Romano, Y., Sesia, M., & Candes, E. (2020). Classification with valid and adaptive coverage. Advances in Neural Information Processing Systems, 33, 3581-3591.

Salih, A. M., Raisi‐Estabragh, Z., Galazzo, I. B., Radeva, P., Petersen, S. E., Lekadir, K., & Menegaz, G. (2025). A perspective on explainable artificial intelligence methods: SHAP and LIME. Advanced Intelligent Systems, 7(1), Article 2400304. https://doi.org/10.1002/aisy.202400304

Sharma, P., Mirzan, S. R., Bhandari, A., Pimpley, A., Eswaran, A., Srinivasan, S., & Shao, L. (2020). Evaluating tree explanation methods for anomaly reasoning: A case study of SHAP TreeExplainer and TreeInterpreter [Conference presentation]. International Conference on Conceptual Modeling. Vienna Austria. https://doi.org/10.1007/978-3-030-65847-2_4

Szepannek, G., & Lübke, K. (2022). Explaining artificial intelligence with care: Analyzing the explainability of black box multiclass machine learning models in forensics. KI-Künstliche Intelligenz, 36(2), 125-134. https://doi.org/10.1007/s13218-022-00764-8

Takaew, S., & Romsaiyud, W. (2025). An explainable approach to sentiment analysis of Thai hotel reviews using a fine-tuned language model and SHAP. Journal of Current Science and Technology, 15(4), Article 149. https://doi.org/10.59796/jcst.v15n4.2025.149

Thakur, Y., & Mittal, N. (2024). Customer segmentation using K-mean clustering [Conference presentation]. Biennial International Conference on Future Learning Aspects of Mechanical Engineering. Singapore: Springer Nature Singapore. https://doi.org/10.1007/978-981-96-7214-1_32

Vovk, V., Gammerman, A., & Shafer, G. (2005). Algorithmic learning in a random world. Boston, MA: Springer US. https://doi.org/10.1007/b106715

Wang, Q. (2022). Support vector machine algorithm in machine learning [Conference presentation]. 2022 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA). IEEE. Dalian, China. https://doi.org/10.1109/ICAICA54878.2022.9844516

Downloads

Published

2026-09-15

How to Cite

Asfi, M., Warsito, B., & Wibowo, A. (2026). COSHAP: An Explainable AI Framework Integrating Conformal Prediction and SHAP for Uncertainty-Aware Consumer Segmentation. Journal of Current Science and Technology, 16(4), 213. https://doi.org/10.59796/jcst.V16N4.2026.213