KLASIFIKASI KELAYAKAN KREDIT MENGGUNAKAN RANDOM FOREST DENGAN OPTIMISASI SMOTENC DAN GRIDSEARCHCV

Authors

  • Nandito Diaz Vannesa Universitas Dian Nuswantoro
  • Fikri Budiman

DOI:

https://doi.org/10.36080/skanika.v9i2.3829

Keywords:

Random Forest, SMOTENC, GridSearchCV, SHAP, Credit Classification

Abstract

Creditworthiness is critical to financial system stability, yet conventional methods struggle with imbalanced data and mixed numeric-categorical features. This study develops the CreditRF-OptSHAP pipeline, integrating Random Forest, SMOTENC, GridSearchCV, and SHAP to improve credit classification performance and interpretability. The dataset used, Statlog German Credit Data (UCI), comprises 1,000 instances, 20 mixed features, and a 70:30 imbalance ratio. The pipeline comprises five phases: data exploration; preprocessing via label encoding and StandardScaler; SMOTENC-based training-set balancing; hyperparameter optimization via GridSearchCV with Stratified 10-Fold Cross Validation; and model interpretation via SHAP TreeExplainer so that each feature's contribution to prediction is explained, making the model no longer a black box. The significance of this recall gain over RF Default was validated using McNemar's Test to rule out statistical coincidence. The main model (RF+SMOTENC+GridSearchCV) achieved a recall of 0.5833 for the bad credit class and an AUC-ROC of 0.7638, up from 0.5000 for RF Default. SHAP analysis identified checking account status as the most dominant feature (importance 0.1239). The McNemar test confirmed a statistically significant difference (p=0.041), confirming its validity. The CreditRF-OptSHAP pipeline yields a model that is more sensitive to non-performing loans, transparent, and statistically validated, thus supporting accountable credit decisions in financial institutions.

Downloads

Download data is not yet available.

References

[1] “NPL Kredit Konsumsi Bulanan Hingga Februari 2025.” Accessed: May 03, 2026. [Online]. Available: https://pusatdata.kontan.co.id/infografik/94/NPL-Kredit-Konsumsi-Bulanan-Hingga-Februari-2025

[2] N. Suhadolnik and J. Ueyama, “Machine Learning for Enhanced Credit Risk Assessment : An Empirical Approach,” 2023.

[3] T. Wongvorachan, S. He, and O. Bulut, “A Comparison of Undersampling, Oversampling, and SMOTE Methods for Dealing with Imbalanced Classification in Educational Data Mining,” vol. 14, no. 1, p.54, 2023, doi: 10.3390/info14010054.

[4] M. S. Uddin, “Leveraging random forest in micro-enterprises credit risk modelling for accuracy and interpretability,” vol. 27, no. 23, pp. 1–31, 2020, doi: 10.1002/ijfe.2346.

[5] Y. Zhou, L. Shen, and L. Ballester, “A two-stage credit scoring model based on random forest : Evidence from Chinese small firms,” International Review of Financial Analysis, vol. 89, no. 4, p. 102755, 2023, doi: 10.1016/j.irfa.2023.102755.

[6] O. Pahlevi and Y. Handrianto, “Implementasi Algoritma Klasifikasi Random Forest Untuk Penilaian Kelayakan Kredit,” Jurnal Infortech, vol. 5, no. 1, pp. 71–76, 2023, doi: 10.31294/infortech.v5i1.15829.

[7] W. Wulansari, and D. Purwitasari, “Algoritma Random Forest pada Prediksi Status Kredit Usaha Rakyat untuk Mengurangi Nonperforming Loan Rate,” Journal of Intelligent System and Computation, vol. 5, no. 2, pp. 109–114, 2023, doi: 10.52985/insyst.v5i2.358.

[8] B. Prasojo and E. Haryatmi, “Analisa Prediksi Kelayakan Pemberian Kredit Pinjaman dengan Metode Random,” Jurnal Nasional Teknologi dan Sistem Informasi, vol. 7, no. 2, pp. 79–89, 2021, doi: 10.25077/TEKNOSI.v7i2.2021.79-89.

[9] V. Chang, et al., “Credit Risk Prediction Using Machine Learning and Deep Learning : A Study on Credit Card Customers,” Risks, vol. 12, no. 11, pp. 1-33, 2024, doi: 10.3390/risks12110174.

[10] A. O. Kuyoro, O. A. Ogunyolu, T. G. Ayanwola, and F. Yetunde, “Ingénierie des Systèmes d ’ Information Dynamic Effectiveness of Random Forest Algorithm in Financial Credit Risk Management for Improving Output Accuracy and Loan Classification Prediction,” Ingénierie des systèmes d information, vol. 27, no. 5, pp. 815–821, 2022, doi: 10.18280/isi.270515.

[11] “Statlog (German Credit Data) - UCI Machine Learning Repository.” Accessed: May 03, 2026. [Online]. Available: https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data

[12] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” vol. 1923, pp. 1895–1923, 1998.

[13] M. Eltawil, et al., “Comment on Iacobescu et al. Evaluating Binary Classifiers for Cardiovascular Disease Prediction: Enhancing Early Diagnostic Capabilities,” J. Cardiovasc. Dev. Dis., vol. 11, no. 12, pp. 4–9, 2026, doi: 10.3390/jcdd13010046.

[14] Y. Elor, and H. A.-Elor, "To SMOTE, or not to SMOTE?," Balance, vol. 1, no. 1. Association for Computing Machinery, 2022.

[15] F. Pedregosa, R. Weiss, and M. Brucher, “Scikit-learn: Machine Learning in Python,” vol. 12, no. 85, pp. 2825–2830, 2011.

16] M. S. Santos, P. H. Abreu, N. Japkowicz, and A. Fernández, “A unifying view of class overlap and imbalance : Key concepts , multi-view panorama, and open avenues for research,” vol. 89, no. 2, 2022, pp. 228–253, 2023.

[17] M. Loecher, “Debiasing SHAP scores in random forests,” AStA Adv. Stat. Anal., vol. 108, no. 2, pp. 427–440, 2024, doi: 10.1007/s10182-023-00479-7.

[18] S. M. Lundberg and S. Lee, “Predictions,” no. Section 2, pp. 1–10, 2017.

[19] Y. Zhong and H. Wang, “Methods,” IEEE Access, vol. PP, p. 1, 2023, doi: 10.1109/ACCESS.2023.3239889.

[20] M. R. Forest, “Analisa Rekomendasi Fitur Persetujuan Pinjaman Perusahaan,” JATISI: Jurnal Teknik Informatika dan Sistem Informasi, vol. 9, no. 3, pp. 2055-2070, 2022, doi: 10.35957/jatisi.v9i3.2258.

Downloads

Published

2026-08-19

How to Cite

[1]
Nandito Diaz Vannesa and Fikri Budiman, “KLASIFIKASI KELAYAKAN KREDIT MENGGUNAKAN RANDOM FOREST DENGAN OPTIMISASI SMOTENC DAN GRIDSEARCHCV”, SKANIKA, vol. 9, no. 2, Aug. 2026.