ANALYZING AND CLUSTERING THAILAND’S MULTIDIMENSIONAL POVERTY DATA USING AN IMPROVED K-MEANS ALGORITHM
Keywords:
Clustering, Density Peak, Improved K-Means, Multidimensional Poverty, PCA, ThailandAbstract
This study proposes an Improved K-Means (IKM) algorithm to enhance the clustering of Thailand’s multidimensional poverty data. The framework integrates Principal Component Analysis (PCA) for dimensionality reduction, the Elbow Method for optimal cluster selection, and Density Peak Clustering (DPC) for centroid initialization within an automated workflow. Provincial-level data from the National Statistical Office (NSO) for 2019, covering all 76 provinces, were analyzed. The performance of IKM was compared with standard K-Means and Gaussian Mixture Models (GMM) using the Silhouette coefficient as the internal validation metric. IKM achieved the highest Silhouette scores in six of seven indicators, indicating superior cluster compactness and separation. To verify robustness, the framework was further evaluated using district-level data from the Thai People Map and Analytics Platform (TPMAP). Paired bootstrap and paired permutation tests confirmed that the observed improvements were statistically significant rather than attributable to random variation. For example, comparison with Hierarchical Clustering using 2,000 bootstrap resamples produced a 95% confidence interval excluding zero, while both statistical tests yielded p-values below 0.05. Comparable results were obtained for comparisons with K-Means and GMM, confirming the consistency and statistical reliability of IKM. Overall, the proposed framework delivers more accurate, stable, and interpretable clustering, effectively capturing multidimensional socioeconomic disparities across Thailand. These findings provide a rigorous analytical basis for evidence-based, spatially targeted poverty reduction policies and offer a scalable framework for future district-level and longitudinal analyses.
References
Aldahdooh, R. T., & Ashour, W. (2013). DIMK-means: Distance-based initialization method for K-means clustering algorithm. International Journal of Intelligent Systems and Applications, 5(2), 41–51. https://doi.org/10.5815/ijisa.2013.02.05
Alkire, S., Nogales, R., Quinn, N. N., & Suppa, N. (2021). Global multidimensional poverty and COVID-19: A decade of progress at risk? Social Science & Medicine, 291, Article 114457. https://doi.org/10.1016/j.socscimed.2021.114457
Clausen, C., & Wechsler, H. (2000). Color image compression using PCA and backpropagation learning.
Pattern Recognition, 33(9), 1555–1560. https://doi.org/10.1016/S0031-3203(99)00126-0
Efron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap. Chapman & Hall/CRC. https://doi.org/10.1201/9780429246593
Fonta, C. L., Yameogo, T. B., Tinto, H., van Huysen, T., Natama, H. M., Compaore, A., & Fonta, W. M. (2020). Decomposing multidimensional child poverty and its drivers in the Mouhoun region of Burkina Faso, West Africa. BMC Public Health, 20, Article 149. https://doi.org/10.1186/s12889-020-8254-3
Good, P. I. (2005). Permutation, parametric, and bootstrap tests of hypotheses (3rd ed.). Springer.
Gottumukkal, R., & Asari, V. K. (2004). An improved face recognition technique based on modular PCA
approach. Pattern Recognition Letters, 25(4), 429–436. https://doi.org/10.1016/j.patrec.2003.11.005
He, M., Zhang, Y., Wen, D., & Wang, Y. (2021). Forecasting crude oil prices: A scaled PCA approach. Energy Economics, 97, Article 105189. https://doi.org/10.1016/j.eneco.2021.105189
Hotelling, H. (1933). Analysis of a complex of statistical variables into principal components. Journal
of Educational Psychology, 24(6), 417–441. https://doi.org/10.1037/h0071325
Hubert, L., & Arabie, P. (1985). Comparing partitions. Journal of Classification, 2(1), 193–218. https://doi.org/10.1007/BF01908075
Ikotun, A. M., Ezugwu, A. E., Abualigah, L., Abuhaija, B., & Heming, J. (2023). K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data. Information Sciences, 622, 178–210. https://doi.org/10.1016/j.ins.2022.11.139
Khan, S. S., & Ahmad, A. (2004). Cluster center initialization algorithm for K-means clustering. Pattern Recognition Letters, 25(11), 1293–1302. https://doi.org/10.1016/j.patrec.2004.04.007
Lloyd, S. (1982). Least squares quantization in PCM. IEEE Transactions on Information Theory, 28(2), 129–137. https://doi.org/10.1109/TIT.1982.1056489
MacQueen, J. B. (1967). Some methods for classification and analysis of multivariate observations. In Le Cam, L. M. and Neyman, J. (Eds.), Proceedings of the fifth Berkeley symposium on mathematical statistics and probability: Vol. 1. Statistics (p. 281–297). University of California Press.
Mahikul, W., Aiyasuwan, O., Thanartthanaboon, P., Chancharoen, W., Achararit, P., Sirisombat, T., & Singkham, P. (2022). Factors affecting bus accident severity in Thailand: A multinomial logit
model. PLOS ONE, 17, Article e0277318. https://doi.org/10.1371/journal.pone.0277318
Mahmud, M. S., Rahman, M. M., & Akhtar, M. N. (2012). Improvement of K-means clustering algorithm with better initial centroids based on weighted average. In 2012 7th International Conference on Electrical and Computer Engineering (pp. 647–650). IEEE. https://doi.org/10.1109/ICECE.2012.6471633
Maitra, R. (2009). Initializing partition-optimization algorithms. IEEE/ACM Transactions on Computational Biology and Bioinformatics, 6(1), 144–157. https://doi.org/10.1109/TCBB.2007.70244
Mohanty, S. K., Agrawal, N. K., Mahapatra, B., Choudhury, D., Tuladhar, S., & Holmgren, E. V. (2017). Multidimensional poverty and catastrophic health spending in the mountainous regions of Myanmar, Nepal and India. International Journal for Equity in Health, 16, Article 21. https://doi.org/10.1186/s12939-016-0514-6
Nikolopoulos, K., Punia, S., Schäfers, A., Tsinopoulos, C., & Vasilakis, C. (2021). Forecasting and planning during a pandemic: COVID-19 growth rates, supply chain disruptions, and governmental decisions. European Journal of Operational Research, 290(1), 99–115. https://doi.org/10.1016/j.ejor.2020.08.001
Oxford Poverty and Human Development Initiative (OPHI) and United Nations Development Programme (UNDP) (2023). 2023 global multidimensional poverty index: Unstacking global poverty: Data for high-impact action. United Nations Development Programme.
Pinilla-Roncancio, M., Mactaggart, I., Kuper, H., Dionicio, C., Naber, J., Murthy, G. V. S., & Polack, S. (2020). Multidimensional poverty and disability: A case control study in India, Cameroon, and Guatemala. SSM–Population Health, 11, Article 100591. https://doi.org/10.1016/j.ssmph.2020.100591
Rodriguez, A., & Laio, A. (2014). Clustering by fast search and find of density peaks. Science, 344(6191), 1492–1496. https://doi.org/10.1126/science.1242072
Sano, A. V. D., & Nindito, H. (2016). Application of K-means algorithm for cluster analysis on poverty of provinces in Indonesia. ComTech: Computer, Mathematics and Engineering Applications, 7(2), 141–150. https://doi.org/10.21512/comtech.v7i2.2254
Strehl, A., & Ghosh, J. (2002). Cluster ensembles—A knowledge reuse framework for combining multiple partitions. Journal of Machine Learning Research, 3, 583–617. https://doi.org/10.1162/153244303321897735
Wiraphanphong, A. (2020). A study of multidimensional poverty management in Thailand according to the Thai People Map and Analytics Platform (TPMAP) through provincial budget allocation and provincial groups, annual budget 2017–2019 [Doctoral dissertation, National Institute of Development Administration].
Yu, H., Chen, R., & Zhang, G. (2014). A SVM stock selection model within PCA. Procedia Computer Science, 31, 406–412. https://doi.org/10.1016/j.procs.2014.05.284
Zelasky, S., Martin, C. L., Weaver, C., Baxter, L. K., & Rappazzo, K. M. (2023). Identifying groups of children’s social mobility opportunity for public health applications using K-means clustering. Heliyon, 9(9), Article e20250. https://doi.org/10.1016/j.heliyon.2023.e20250
Zhou, X., Tang, X., & Zhang, R. (2020). Impact of green finance on economic development and environmental quality: A study based on provincial panel data from China. Environmental Science and Pollution Research, 27, 19915–19932. https://doi.org/10.1007/s11356-020-08383-2








