KOMPARASI MNB, CNB, DAN SVM UNTUK DETEKSI UJARAN KEBENCIAN BAHASA INDONESIA
DOI:
https://doi.org/10.36080/skanika.v9i2.3830Keywords:
hate speech detection, Complement Naive Bayes, TF-IDF, imbalanced class, lexicon-based calibrationAbstract
Indonesia had 235.26 million internet users and 81.72% internet penetration in 2026, while hate speech threatens social cohesion. Based on the literature reviewed in this study, no prior study has been found that systematically conducted a three-way benchmark of MNB, CNB, and SVM with preprocessing ablation, McNemar testing, and Wilson CI error analysis in Indonesian hate speech detection, using a dataset of 13,169 tweets, with per-algorithm imbalanced-class handling (SMOTE for MNB/CNB; CSL and Lexicon-Based Coefficient Override for SVM), and ablation of six preprocessing configurations. MNB achieves F1 Macro 72.10% and highest Abusive Recall (84.30%), while SVM achieves Abusive Recall 80.20% with inference latency 0.10 ms on CPU; a direct comparison with IndoBERT on GPU T4 is not fully equivalent across hardware, though an additional CPU-only measurement shows IndoBERT at ~102.4 ms (~1,020× ratio). McNemar tests confirm SVM differs significantly from MNB/CNB (p<0.0001), CNB vs MNB non-significant (p=0.0704), indicating a negative SMOTE×CNB interaction specific to this study’s dataset and configuration. SVM is selected as the deployment model based on Utility Score (81.39) under the weighting scenarios tested, and may serve as a first-pass screening aid (not a substitute for human verification) for platforms operating under Indonesian ITE Law No. 1/2024.
Downloads
References
[1] Asosiasi Penyelenggara Jasa Internet Indonesia (APJII), “Survei Profil Internet Indonesia 2026,” 2026. [Online]. Available: https://apjii.or.id
[2] Republik Indonesia, “Undang-Undang Nomor 1 Tahun 2024 tentang Perubahan Kedua atas Undang-Undang Nomor 11 Tahun 2008 tentang Informasi dan Transaksi Elektronik,” 2024. [Online]. Available: https://peraturan.bpk.go.id/details/274494/uu-no-1-tahun2024
[3] A. Gandhi, P. Ahir, K. Adhvaryu, P. Shah, R. Lohiya, E. Cambria, S. Poria, and A. Hussain, “Hate speech detection: A comprehensive review of recent works,” Expert Systems, vol. 41, no. 8, Art. e13562, 2024, doi: 10.1111/exsy.13562.
[4] M. O. Ibrohim and I. Budi, “Multi-label Hate Speech and Abusive Language Detection in Indonesian Twitter,” in Proceedings of the 3rd Workshop on Abusive Language Online (ALW3), 2019, pp. 46–57, doi: 10.18653/v1/w19-3506.
[5] I. Putu Widiarta Nandana Githa, A. Syananda, R. Faustine, I. S. Edbert, and D. Suhartono, “Hate Speech Classification in Indonesian Tweets Using TF-IDF and Data Augmentation,” in 2024 International Conference on Green Energy, Computing and Sustainable Technology, GECOST 2024, 2024, pp. 61–65, doi: 10.1109/GECOST60902.2024.10474781.
[6] I. I. R. Difandana and I. Imaduddin, “Comparative Analysis of Naive Bayes and Support Vector Machine for Hate Speech Classification,” MALCOM Indones. J. Mach. Learn. Comput. Sci., vol. 6, no. 1, pp. 414–422, Feb. 2026, doi: 10.57152/MALCOM.V6I1.2571.
[7] S. A. Zikrina and F. Fitriyani, “Advancing Hate Speech Detection in Indonesian Language Using Graph Neural Networks and TF-IDF,” J. RESTI, vol. 9, no. 1, pp. 137–145, 2025, doi: 10.29207/resti.v9i1.6179.
[8] S. D. A. Putri, M. O. Ibrohim, and I. Budi, “Abusive Language and Hate Speech Detection for Indonesian-Local Language in Social Media Text,” in Lecture Notes in Networks and Systems, 2021, pp. 88–98, doi: 10.1007/978-3-030-79757-7_9.
[9] N. Khoirunnisaa, K. Nabila Nastiti Kesuma, S. Setiawan, and A. Yunizar Pratama Yusuf, “Klasifikasi teks ulasan aplikasi Netflix pada Google Play Store menggunakan algoritma Naïve Bayes dan SVM,” SKANIKA Sist. Komput. dan Tek. Inform., vol. 7, no. 1, pp. 64–73, 2024, doi: 10.36080/skanika.v7i1.3138.
[10] M. O. Ibrohim and I. Budi, “Hate speech and abusive language detection in Indonesian social media: Progress and challenges,” Heliyon, vol. 9, no. 8, Art. e18647, 2023, doi: 10.1016/j.heliyon.2023.e18647.
[11] A. N. A. Saputra, R. E. Saputro, and D. I. S. Saputra, “Labeling Optimization and Hybrid CNN Model in Sentiment Analysis of Movie Reviews with Slang Handling,” Jurnal Teknik Informatika (JUTIF), vol. 6, no. 6, 2025, doi: 10.52436/1.jutif.2025.6.6.4465.
[12] X. Xiang, “Application of an Improved TF-IDF Method in Literary Text Classification,” Advances in Multimedia, 2022, doi: 10.1155/2022/9285324.
[13] D. Rosadi and S. Rakasiwi, “Anti-Data Leakage Pipeline for Differentiated Thyroid Cancer Recurrence Prediction: Integrating SMOTE, Optuna-based Optimization, and Bootstrap BCa Validation,” Journal of Applied Informatics and Computing, vol. 10, no. 3, 2026, doi: 10.30871/jaic.v10i3.12922.
[14] M. Koren, O. Peretz, and O. Koren, “Naive Bayes classifier – An ensemble procedure for recall and precision enrichment,” Engineering Applications of Artificial Intelligence, vol. 136, Part B, Art. 108972, 2024, doi: 10.1016/j.engappai.2024.108972.
[15] D. Elreedy, A. F. Atiya, and F. Kamalov, “A theoretical distribution analysis of synthetic minority oversampling technique (SMOTE) for imbalanced learning,” Machine Learning, vol. 113, pp. 4903–4923, 2024, doi: 10.1007/s10994-022-06296-4.
[16] A. M. Mardiana, I. F. Rozi, and R. Arianto, “Support Vector Machine with FastText Word Embedding for Hate Speech Aspect Categorization,” Paradigma - Jurnal Komputer dan Informatika, vol. 27, no. 2, pp. 92–98, 2025, doi: 10.31294/p.v27i2.5127.
[17] W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, “A survey on imbalanced learning: latest research, applications and future directions,” Artif. Intell. Rev., vol. 57, no. 6, p. 137, 2024, doi: 10.1007/s10462-024-10759-6.
[18] A. S. Aribowo, Y. Fauziah, Y. Bantulu, S. Saifullah, and A. M. A. Fubalo, “Hate Speech Analysis Using IndoBERT in YouTube Comments on the 2024 Indonesian Presidential Debate Video,” Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control, vol. 11, no. 3, 2026, doi: 10.22219/kinetik.v11i3.2604.
[19] O. Rainio, J. Teuho, and R. Klén, “Evaluation metrics and statistical tests for machine learning,” Scientific Reports, vol. 14, Art. 6086, 2024, doi: 10.1038/s41598-024-56706-x.
[20] T. Sugihartono and R. R. C. Putra, “Penerapan metode Support Vector Machine dalam klasifikasi ulasan pengguna aplikasi Mobile JKN,” SKANIKA Sist. Komput. dan Tek. Inform., vol. 7, no. 2, pp. 144–153, 2024, doi: 10.36080/skanika.v7i2.3193.
[21] L. Andersson and A. Nerman, “A Note on Confidence Intervals for a Binomial p: Andersson–Nerman vs. Wilson,” Stat, vol. 13, Art. e70027, 2024, doi: 10.1002/sta4.70027.
[22] G. Rau and Y.-S. Shih, “Evaluation of Cohen’s kappa and other measures of inter-rater agreement for genre analysis and other nominal data,” Journal of English for Academic Purposes, vol. 53, Art. 101026, 2021, doi: 10.1016/j.jeap.2021.101026.
[23] scikit-learn 1.3.0, Zenodo, 30 June 2023, doi: 10.5281/zenodo.8098905. [Online]. Available: https://zenodo.org/record/8098905
[24] M. I. Wijanarko, L. Susanto, P. A. Pratama, I. Idris, T. Hong, and D. Wijaya, “Monitoring Hate Speech in Indonesia: An NLP-based Classification of Social Media Texts,” in Proceedings of EMNLP 2024: System Demonstrations, pp. 142–152, 2024, doi: 10.18653/v1/2024.emnlp-demo.15.
[25] F. Koto, J. H. Lau, and T. Baldwin, “INDOBERTWEET: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” in EMNLP 2021 - 2021 Conference on Empirical Methods in Natural Language Processing, Proceedings, 2021, pp. 10660–10668, doi: 10.18653/v1/2021.emnlp-main.833.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Fisco Maulana Ikhwan, Edi Sugiarto

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
CC BY-SA 4.0
Creative Commons Attribution-ShareAlike 4.0 International
This license requires that reusers give credit to the creator. It allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, even for commercial purposes. If others remix, adapt, or build upon the material, they must license the modified material under identical terms.
BY: Credit must be given to you, the creator.
SA: Adaptations must be shared under the same terms.ng







