| International Journal of Computer Applications |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 187 - Number 125 |
| Year of Publication: 2026 |
| Authors: Aolia Ikhwanudin, Agianto Syam Halim, Tubagus Toifur, Mumahamad Yusuf, Ibnu Masud, Tb Ai Munandar |
10.5120/ijcaebdb557ce0b6
|
Aolia Ikhwanudin, Agianto Syam Halim, Tubagus Toifur, Mumahamad Yusuf, Ibnu Masud, Tb Ai Munandar . Machine Learning-based Prediction of Building Permit Levy: A Comparative Study of Regression Algorithms. International Journal of Computer Applications. 187, 125 ( Jul 2026), 33-44. DOI=10.5120/ijcaebdb557ce0b6
The building permit levy is a critical source of regional own-source revenue for local governments in Indonesia. Accurate prediction of levy amounts is essential for fiscal planning, transparency, and administrative efficiency. This study applies seven machine learning regression algorithms — Elastic Net, Random Forest, XGBoost, LightGBM, Gradient Boosting, Decision Tree, and k-Nearest Neighbors — to predict IMB levy values using building-related features derived from 14,971 permit application records. A comprehensive exploratory data analysis (EDA) pipeline was implemented, including missing-value audit, outlier analysis (IQR and Z-score), Pearson and Spearman correlation analysis, and data-leakage auditing. The evaluation combines a held-out test set, 5-fold cross-validation, bootstrap confidence intervals, feature ablation, raw-versus-log target comparison, randomized hyperparameter tuning, and a scenario-based analysis across levy magnitude, building function, designation, and building size segments. Random Forest achieves R² = 0.6875, MAE = Rp 5,002,478, and MAPE = 11.18% on the held-out set, while Gradient Boosting attains the best hold-out R² (0.8490); cross-validation shows the two ensembles are statistically close under the heavy-tailed target, and Elastic Net fails catastrophically (R² = −160.65), confirming the non-linear levy–feature relationship. Feature importance and ablation identify total building area (total_luas) as the overwhelmingly dominant predictor (83.76% importance), consistent with the multiplicative structure of the IMB levy formula. Scenario analysis reveals that relative errors concentrate in non-residential permits and at the extremes of the levy distribution, motivating a human-in-the-loop deployment for those segments. These findings establish tree-based ensembles as robust candidates for automated levy estimation in regional government digital services.