CFP last date
20 August 2026
Reseach Article

K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan

by Aolia Ikhwanudin, Tubagus Toifur, Muhamad Yusuf
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Number 125
Year of Publication: 2026
Authors: Aolia Ikhwanudin, Tubagus Toifur, Muhamad Yusuf
10.5120/ijca44904f79e07d

Aolia Ikhwanudin, Tubagus Toifur, Muhamad Yusuf . K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan. International Journal of Computer Applications. 187, 125 ( Jul 2026), 45-53. DOI=10.5120/ijca44904f79e07d

@article{ 10.5120/ijca44904f79e07d,
author = { Aolia Ikhwanudin, Tubagus Toifur, Muhamad Yusuf },
title = { K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan },
journal = { International Journal of Computer Applications },
issue_date = { Jul 2026 },
volume = { 187 },
number = { 125 },
month = { Jul },
year = { 2026 },
issn = { 0975-8887 },
pages = { 45-53 },
numpages = {9},
url = { https://ijcaonline.org/archives/volume187/number125/k-means-clustering-for-regional-segmentation-of-municipal-permit-services-a-case-study-of-tangerang-selatan/ },
doi = { 10.5120/ijca44904f79e07d },
publisher = {Foundation of Computer Science (FCS), NY, USA},
address = {New York, USA}
}
%0 Journal Article
%1 2026-07-29T00:34:13.747472+05:30
%A Aolia Ikhwanudin
%A Tubagus Toifur
%A Muhamad Yusuf
%T K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan
%J International Journal of Computer Applications
%@ 0975-8887
%V 187
%N 125
%P 45-53
%D 2026
%I Foundation of Computer Science (FCS), NY, USA
Abstract

Effective allocation of public service resources in Indonesian cities requires understanding the spatial heterogeneity of permit-type demand at the sub-district level. This paper presents a multi-method unsupervised clustering framework applied to 56 sub-districts in Tangerang Selatan City, using proportion vectors for eight permit categories combined with log-transformed application volume. Four algorithms were evaluated — K-Means, Agglomerative (Ward), Gaussian Mixture Model (GMM), and Spectral Clustering — across k=2..6. Four outlier sub-districts were identified and excluded prior to clustering (kel_0, Alue Bagok, Pondok Benda, Pamulang Timur). On the remaining 52 sub-district, K-Means at k=4 (collapsed to 3 interpretable clusters) achieved the highest silhouette score (0.2456), outperforming Agglomerative (0.2347), GMM (0.1596), and Spectral (0.1401). The three final clusters represent: Cluster 0 (5 sub-districts) — cemetery/land-use permit specialists with high tariffed field-inspection demand (p_Field=0.448, mean 8,106 applications); Cluster 1 (28 sub-districts) — residential service generalists with high free-field inspection demand (p_FreeF=0.413, mean 3,644 applications); and Cluster 2 (19 sub-districts) — high-volume administrative-review centers with elevated inter-agency permit activity (p_Admin=0.440, mean 7,618 applications). These findings provide a data-driven basis for differentiated surveyor allocation, digital service channel design, and spatial planning prioritization at DPMPTSP Tangerang Selatan. Beyond silhouette alone, a comprehensive multi-metric evaluation (Davies–Bouldin and Calinski–Harabasz indices across k=2..6 for all four algorithms), a 500-iteration bootstrap stability analysis, and one-way ANOVA significance tests on the resulting cluster profiles were conducted to validate the robustness of the chosen segmentation.

References
  1. MacQueen, J. 1967. Some methods for classification and analysis of multivariate observations. Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability, vol. 1, pp. 281-297. University of California Press, Berkeley.
  2. Rousseeuw, P.J. 1987. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20,53-65. https://doi.org/10.1016/0377-0427(87)90125-7
  3. Ikotun, A.M., Ezugwu, A.E., Abualigah, L., Abuhaija, B., and Heming, J. 2023. K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data. Information Sciences, 622,178-210. https://doi.org/10.1016/j.ins.2022.11.139
  4. Ezugwu, A.E., Ikotun, A.M., Oyelade, O.O., Abualigah, L., Agushaka, J.O., Eke, C.I., and Akinyelu, A.A. 2022. A comprehensive survey of clustering algorithms: State-of-the-art machine learning applications, taxonomy, challenges, and future research prospects. Engineering Applications of Artificial Intelligence, 110, Article 104743. https://doi.org/10.1016/j.engappai.2022.104743
  5. Ward, J.H. 1963. Hierarchical grouping to optimize an objective function. Journal of the American Statistical Association, 58(301), 236-244. https://doi.org/10.1080/01621459.1963.10500845
  6. Jain, A.K. 2010. Data clustering: 50 years beyond K-Means. Pattern Recognition Letters, 31(8), 651-666. https://doi.org/10.1016/j.patrec.2009.09.011
  7. Halkidi, M., Batistakis, Y., and Vazirgiannis, M. 2001. On clustering validation techniques. Journal of Intelligent Information Systems, 17(2-3), 107-145. https://doi.org/10.1023/A:1012801612483
  8. Liu, F.T., Ting, K.M., and Zhou, Z.H. 2012. Isolation-based anomaly detection. ACM Transactions on Knowledge Discovery from Data, 6(1), Article 3, pp. 1-39. https://doi.org/10.1145/2133360.2133363
  9. Warrens, M.J. and van der Hoef, H. 2022. Understanding the adjusted Rand index and other partition comparison indices based on counting object pairs. Journal of Classification, 39,487-509. https://doi.org/10.1007/s00357-022-09413-z
  10. Pal, S. and Heumann, C. 2022. Clustering compositional data using Dirichlet mixture model. PLOS ONE, 17(5), e0268438. https://doi.org/10.1371/journal.pone.0268438
  11. Ng, A.Y., Jordan, M.I., and Weiss, Y. 2001. On spectral clustering: Analysis and an algorithm. Advances in Neural Information Processing Systems (NIPS), 14,849-856.
  12. Arthur, D. and Vassilvitskii, S. 2007. k-means++: the advantages of careful seeding. Proceedings of the 18th ACM-SIAM Symposium on Discrete Algorithms (SODA), pp. 1027-1035.
  13. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12,2825-2830.
  14. Greenacre, M., Blanco, V., Bove, G., Cannari, M., Carlier, A., Castellan, G., Chavent, M., and Greenacre, L. 2023. Compositional data analysis. Nature Reviews Methods Primers, 3,14. https://doi.org/10.1038/s43586-023-00195-3
  15. Xu, D. and Tian, Y. 2015. A comprehensive survey of clustering algorithms. Annals of Data Science, 2(2), 165-193. https://doi.org/10.1007/s40745-015-0040-1
  16. Soffan, S., Bramantoro, A., and Alzahrani, A.A. 2025. Combination of machine learning and data envelopment analysis to measure the efficiency of the Tax Service Office. PeerJ Computer Science, 11, e2672. https://doi.org/10.7717/peerj-cs.2672
  17. Batty, M. 2013. The New Science of Cities. MIT Press, Cambridge, MA.
  18. Anselin, L. 1995. Local indicators of spatial association: LISA. Geographical Analysis, 27(2), 93-115. https://doi.org/10.1111/j.1538-4632.1995.tb00338.x
  19. Fadhel, M.A., Duhaim, A.M., Saihood, A., Sewify, A., Al-Hamadani, M.N.A., and colleagues. 2024. Assessing urban vulnerability to emergencies: a spatiotemporal approach using K-Means clustering. Land, 13(11), 1744. https://doi.org/10.3390/land13111744
  20. Lenssen, L. and Schubert, E. 2023. Medoid silhouette clustering with automatic cluster number selection. Data Mining and Knowledge Discovery, 38,1523-1564. https://doi.org/10.1007/s10618-023-00979-9
  21. Shao, C., Du, X., Yu, J., and Chen, J. 2022. Cluster-based improved Isolation Forest. Entropy, 24(5), 611. https://doi.org/10.3390/e24050611
  22. Schubert, E. 2023. Stop using the elbow criterion for k-means and how to choose the number of clusters instead. ACM SIGKDD Explorations Newsletter, 25(1), 36-42. https://doi.org/10.1145/3606274.3606278
  23. Gracias, J.S., Parnell, G.S., Specking, E., Pohl, E.A., and Buchanan, R. 2023. Smart cities: A structured literature review. Smart Cities, 6(4), 1719-1743. https://doi.org/10.3390/smartcities6040079
  24. Alamsyah, A. and Muhammad, M.A. 2024. Leveraging generative AI for public service innovation: a path to smart government in Indonesia. Digital Government: Research and Practice, 5(4), Article 32. https://doi.org/10.1145/3761820
  25. Xia, J., Zhang, Y., Song, J., Chen, Y., Wang, Y., and Liu, S. 2022. Revisiting dimensionality reduction techniques for visual cluster analysis: An empirical study. IEEE Transactions on Visualization and Computer Graphics, 28(1), 529-539. https://doi.org/10.1109/TVCG.2021.3114694
  26. Rousseeuw, P.J. and Kaufman, L. 1990. Finding Groups in Data: An Introduction to Cluster Analysis. Wiley, New York.
  27. [AUTHOR TO COMPLETE] - Reference on clustering for Indonesian regional development or fiscal capacity analysis.
  28. [AUTHOR TO COMPLETE] - Reference on data-driven public service analytics in Southeast Asian municipalities.
Index Terms

Computer Science
Information Sciences

Keywords

K-Means clustering silhouette coefficient permit segmentation sub-district Tangerang Selatan DPMPTSP outlier detection method comparison StandardScaler