IMPLEMENTASI METODE NER STATISTIK BERBASIS WEB UNTUK EKSTRAKSI ENTITAS PADA BERITA BENCANA ALAM TVRI
DOI:
https://doi.org/10.36080/skanika.v9i2.3897Keywords:
Naive Bayes, Named Entity Recognition, Natural Disaster News, Random Undersampling, Manual LabelingAbstract
Disaster news published on TVRI's portal is typically presented in unstructured text containing important details such as disaster type, location, time, and related organizations, making manual information extraction inefficient. This study implements a web-based Statistical Named Entity Recognition (NER) using the Naive Bayes algorithm to extract disaster, location, time, date, and organization entities from 200 TVRI disaster news articles. Training data was created through expert-validated manual labeling using the BIO (Begin-Inside-Outside) scheme, producing 40,667 labeled tokens converted into CurrentWord, Token Type, CurrentTag, Bef1Tag, and Class features, with Laplace Smoothing applied to address zero probability. Testing was conducted before and after applying Random Undersampling to assess its effect on class balance. Before Random Undersampling, the model achieved 87.91% accuracy, 55.73% precision, 58.66% recall, and 56.63% F1-score, after application, accuracy dropped to 83.17% and recall to 55.50%, while precision rose to 62.74% and F1-score to 57.45%. These results show that Random Undersampling improved the precision-recall balance despite lower accuracy, proving that Naive Bayes-based Statistical NER can automatically extract entities, though the relatively small F1-score improvement (about 0.82 percentage points) indicates a need for further testing.
Downloads
References
[1] G. Fajar et al., “Journal of Open Innovation : Technology , Market , and Complexity Indonesian disaster named entity recognition from multi source information using bidirectional LSTM ( BiLSTM ),” J. Open Innov. Technol. Mark. Complex., vol. 10, no. 3, p. 100358, 2024, doi: 10.1016/j.joitmc.2024.100358.
[2] Badan Nasional Penanggulangan Bencana (BNPB), “Indonesia Tangguh Menghadapi Bencana 2025.” [Online]. Available: https://indonesiabaik.id/infografis/indonesia-tangguh-menghadapi-bencana-2025
[3] I. Budi and R. R. Suryono, “Application of named entity recognition method for Indonesian datasets: a review,” Bull. Electr. Eng. Informatics, vol. 12, no. 2, pp. 969–978, 2023, doi: 10.11591/eei.v12i2.4529.
[4] Z. Hu, W. Hou, and X. Liu, “Deep learning for named entity recognition: a survey,” Neural Comput. Appl., vol. 36, no. 16, pp. 8995–9022, 2024, doi: 10.1007/s00521-024-09646-6.
[5] E. Yulianti, et al., “Named entity recognition on Indonesian legal documents: a dataset and study using transformer-based models,” Int. J. Electr. Comput. Eng., vol. 14, no. 5, pp. 5489–5501, 2024, doi: 10.11591/ijece.v14i5.pp5489-5501.
[6] P. J. B. Pajila, et al., “A Comprehensive Survey on Naive Bayes Algorithm: Advantages, Limitations and Applications,” Proc. 4th Int. Conf. Smart Electron. Commun. ICOSEC 2023, no. Icosec, pp. 1228–1234, 2023, doi: 10.1109/ICOSEC58147.2023.10276274.
[7] S. O. Khairunnisa, Z. Chen, and M. Komachi, “Improving Domain-Specific NER in the Indonesian Language Through Domain Transfer and Data Augmentation,” J. Adv. Comput. Intell. Intell. Informatics, vol. 28, no. 6, pp. 1299–1312, 2024, doi: 10.20965/jaciii.2024.p1299.
[8] I. Indra, et al., “Implementation of Named Entity Recognition (NER) for Spatio-Temporal Detection and Sentence Modeling of Natural Disasters Using Naive Bayes and C.45,” in 2024 International Seminar on Intelligent Technology and Its Applications (ISITIA), 2024, pp. 166–171. doi: 10.1109/ISITIA63062.2024.10668211.
[9] I. M. Karo Karo, S. Dewi, and A. Syahrin, “Ekstraksi Informasi Bencana Banjir Dari Berita Online Berbasis Named Enitity Recognition,” Multinetics, vol. 11, no. 1, pp. 35–42, 2025, doi: 10.32722/multinetics.v11i1.7499.
[10] N. Istiqomah and F. Novika, “Perbandingan Kinerja Model NER IndoBERT dan IndoLEM dalam Ekstraksi Informasi Kesehatan Pascabencana dari Berita Daring di Indonesia,” J. Comput. Sci. Informatics Eng., vol. 04, no. 3, pp. 158–174, 2025, doi: 10.55537/cosie.v4i3.1174
[11] M. Carvalho, A. J. Pinho, and S. Brás, “Resampling approaches to handle class imbalance: a review from a data perspective,” J. Big Data, vol. 12, no. 1, p. 71, Mar. 2025, doi: 10.1186/s40537-025-01119-4.
[12] C. P. Chai, “Comparison of text preprocessing methods,” Nat. Lang. Eng., vol. 29, no. 3, pp. 509–553, 2023, doi: 10.1017/S1351324922000213.
[13] M. Siino, I. Tinnirello, and M. La Cascia, “Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers,” Inf. Syst., vol. 121, no. March 2023, p. 102342, 2024, doi: 10.1016/j.is.2023.102342.
[14] N. Lestari et al., “Implementation Of Text Mining And Pattern Discovery With Naive Bayes Algorithm For Classification Of Text Documents,” J. Teknol. Inf. dan Komun., vol. 14, no. 1, pp. 88–102, 2023.
[15] R. Hidayat and W. Windarto, “Klasifikasi Komentar Netizen X Tentang Pemecatan Shin Tae Yong Dari PSSI Dengan Algoritma Naive Bayes,” SKANIKA: Sistem Komputer dan Teknik Informatika, vol. 9, no. 1, pp. 171–181, 2026, doi: 10.36080/skanika.v9i1.3605.
[16] Y. Feng, M. Zhou, and X. Tong, “Imbalanced classification: A paradigm‐based review,” Stat. Anal. Data Min. ASA Data Sci. J., vol. 14, no. 5, pp. 383–406, Oct. 2021, doi: 10.1002/sam.11538.
[17] K. Ghosh, C. Bellinger, R. Corizzo, P. Branco, B. Krawczyk, and N. Japkowicz, “The class imbalance problem in deep learning,” Mach. Learn., vol. 113, no. 7, pp. 4845–4901, 2024, doi: 10.1007/s10994-022-06268-8.
[18] H. Alamro, T. Gojobori, M. Essack, and X. Gao, “BioBBC: a multi-feature model that enhances the detection of biomedical entities,” Sci. Rep., vol. 14, no. 1, pp. 1–14, 2024, doi: 10.1038/s41598-024-58334-x.
[19] I. M. Widi, A. Ari, and I. W. Supriana, “Dampak Penggunaan Anotasi Penamaan yang Berbeda Pada Kinerja NER,” J. Nas. Teknol. Inf. dan Apl., vol. 1, no. 4, pp. 1141–1148, 2023.
[20] G. Syahrani, S. Sevira, and A. Yunizar Pratama Yusuf, “Rancangan Chatbot Rekomendasi Coffee Shop Jabodetabek Dengan Menggunakan Dialogflow Natural Language Processing,” SKANIKA: Sistem Komputer dan Teknik Informormatika, vol. 7, no. 1, pp. 74–84, 2024, doi: 10.36080/skanika.v7i1.3139.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Muhammad Iqbal Shiddiq, Indra Indra

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
CC BY-SA 4.0
Creative Commons Attribution-ShareAlike 4.0 International
This license requires that reusers give credit to the creator. It allows reusers to distribute, remix, adapt, and build upon the material in any medium or format, even for commercial purposes. If others remix, adapt, or build upon the material, they must license the modified material under identical terms.
BY: Credit must be given to you, the creator.
SA: Adaptations must be shared under the same terms.ng








