Explainable Phishing Website Detection Using Comparative Machine Learning and SHAP
DOI:
https://doi.org/10.36080/idealis.v9i2.3810Keywords:
Explainable Artificial Intelligence, Machine Learning, Phishing URL Detection, SHAP, XGboostAbstract
Phishing attacks distributed through fraudulent URLs remain one of the most damaging cyber threats, including in Indonesia where malicious links spread widely through messaging applications and e-mail. Blacklist-based approaches cannot recognize newly created phishing URLs, while accurate machine learning models are often difficult to interpret. This study aims to compare six machine learning algorithms, namely XGBoost, Random Forest, Support Vector Machine, K-Nearest Neighbors, Logistic Regression, and Naive Bayes, for phishing website detection based on lexical URL, page content, and external reputation features, while providing model transparency through Explainable Artificial Intelligence (XAI). Experiments were conducted on the Web Page Phishing Detection Dataset containing 11,430 URLs with 87 features and a balanced class distribution. The research stages include exploratory data analysis, feature selection analysis using Chi-Square, Mutual Information, and Recursive Feature Elimination, an 80:20 data split, model training, hyperparameter optimization using RandomizedSearchCV, and interpretation of the best model using SHapley Additive exPlanations (SHAP). The results show that XGBoost delivers the best performance with 96.50% accuracy, 96.26% precision, 96.76% recall, 96.51% F1-score, and an AUC of 0.9943. SHAP analysis identifies google_index, page_rank, and nb_hyperlinks as the most influential features, dominated by external reputation-based features. Under this experimental setting, the findings indicate that an accurate phishing detection model can be equipped with interpretable explanations of its feature contributions. The main contribution of this study is an integrated comparative evaluation that combines six-algorithm benchmarking, leakage-free hyperparameter optimization, and SHAP-based interpretation on a public phishing dataset, offering practical guidance for security analysts.
Downloads
References
[1] A. Safi and S. Singh, "A systematic literature review on phishing website detection techniques," J. King Saud Univ. - Comput. Inf. Sci., vol. 35, no. 2, pp. 590-611, 2023, doi: 10.1016/j.jksuci.2023.01.004.
[2] A. F. Mahmud and S. Wirawan, "Deteksi Phishing Website Menggunakan Machine Learning Metode Klasifikasi," Sistemasi: J. Sist. Inf., vol. 13, no. 4, pp. 1368-1380, 2024, doi: 10.32520/stmsi.v13i4.3456.
[3] A. S. Y. Irawan, N. Heryana, H. S. Hopipah, and D. Rahma, "Identifikasi Website Phishing dengan Perbandingan Algoritma Klasifikasi," Syntax: J. Inform., vol. 10, no. 1, pp. 57-67, 2021, doi: 10.35706/syji.v10i01.5292.
[4] Y. Muliono, M. A. Ma’ruf, and Z. M. Azzahra, "Phishing Site Detection Classification Model Using Machine Learning Approach," J. EMACS (Eng. Math. Comput. Sci.), vol. 5, no. 2, pp. 63-67, 2023, doi: 10.21512/emacsjournal.v5i2.9951.
[5] D. Komalasari, T. B. Kurniawan, D. A. Dewi, M. Z. Zakaria, Z. Abdullah, and A. Alanda, "Phishing Domain Detection Using Machine Learning Algorithms," Int. J. Adv. Sci. Eng. Inf. Technol., vol. 15, no. 1, pp. 318-327, 2025, doi: 10.18517/ijaseit.15.1.12553.
[6] K. Adane, B. Beyene, and M. Abebe, "ML and DL-based Phishing Website Detection: The Effects of Varied Size Datasets and Informative Feature Selection Techniques," J. Artif. Intell. Technol., vol. 4, no. 1, pp. 18-30, 2024, doi: 10.37965/jait.2023.0269.
[7] P. Ponni and P. Dhandayudam, "An Optimized Bagging Learning with Ensemble Feature Selection Method for URL Phishing Detection," J. Electr. Eng. Technol., vol. 19, no. 3, pp. 1881-1889, 2023, doi: 10.1007/s42835-023-01680-z.
[8] E. Sangra, R. Agrawal, P. R. Gundalwar, K. Sharma, D. Bangri, and D. Nandi, "Malicious Website Detection Using Random Forest and Pearson Correlation for Effective Feature Selection," Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 8, pp. 772-780, 2024, doi: 10.14569/IJACSA.2024.0150876.
[9] O. K. Sahingoz, E. Buber, O. Demir, and B. Diri, "Machine learning based phishing detection from URLs," Expert Syst. Appl., vol. 117, pp. 345-357, 2019, doi: 10.1016/j.eswa.2018.09.029.
[10] S. S. Shafin, "An explainable feature selection framework for web phishing detection with machine learning," Data Sci. Manag., vol. 8, no. 2, pp. 127-136, 2025, doi: 10.1016/j.dsm.2024.08.004.
[11] A. Oest, Y. Safaei, P. Zhang, B. Wardman, K. Tyers, Y. Shoshitaishvili, A. Doupe, and G.-J. Ahn, "PhishTime: Continuous longitudinal measurement of the effectiveness of anti-phishing blacklists," in Proc. 29th USENIX Security Symposium, 2020, pp. 379-396.
[12] S. R. Alotaibi et al., "Explainable artificial intelligence in web phishing classification on secure IoT with cloud-based cyber-physical systems," Alexandria Eng. J., vol. 110, pp. 490-505, 2025, doi: 10.1016/j.aej.2024.09.115.
[13] S. M. Lundberg and S.-I. Lee, "A Unified Approach to Interpreting Model Predictions," in Proc. 31st Int. Conf. Neural Information Processing Systems (NeurIPS), 2017, pp. 4765-4774.
[14] S. Y. Nailendra, W. Witanti, and G. Abdillah, "Optimasi Prediksi Penjualan Retail Online Menggunakan LightGBM dan Hyperparameter Tuning," J. Algoritm., vol. 22, no. 2, pp. 1931-1942, 2025, doi: 10.33364/algoritma/v.22-2.2551.
[15] A. Hannousse and S. Yahiouche, "Towards benchmark datasets for machine learning based website phishing detection: An experimental study," Eng. Appl. Artif. Intell., vol. 104, p. 104347, 2021, doi: 10.1016/j.engappai.2021.104347.
[16] T. Chen and C. Guestrin, "XGBoost: A Scalable Tree Boosting System," in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 785-794, doi: 10.1145/2939672.2939785.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Juni Ismail, Raja Anan Nasution, Muhammad Nasri Gea

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.










