Comparison of Yolo26 and RT-DETR-L for Manga Visual Element Detection on The Manga109 Dataset
DOI:
https://doi.org/10.36080/idealis.v9i2.3842Keywords:
Manga Image Analysis, Manga109, Object Detection, RT-DETR, YOLO26Abstract
Automatic comic element detection is a critical prerequisite for manga digital analysis applications such as indexing, translation, and accessibility enhancement. Manga109 is a widely used benchmark for this task, yet the rapid progress of real-time object detectors has not been matched by a systematic comparison among state-of-the-art models in this domain, particularly between YOLO-family detectors employing Small-Target-Aware Label assignment (STAL) and transformer-based detectors such as RT-DETR. This study benchmarks YOLO26 (nano, small, and medium variants) against RT-DETR-L for detecting four comic element classes (panel, character, text, and face) on Manga109. All models were trained for 100 epochs at 1024×1024 pixels and evaluated using mAP and per-size AP following COCO conventions. YOLO26m achieves the best performance (mAP@0.5:0.95 of 0.7471), while RT-DETR-L obtains the lowest (0.7001) despite the largest parameter count and slowest inference. For small text detection, YOLO26m outperforms RT-DETR-L by 42.7% relatively, and the AP gap across object sizes narrows monotonically as model size grows, supporting the STAL design claim. RT-DETR-L also exhibits significant training instability. The main contribution is a systematic, multi-dimensional benchmark of YOLO26 against RT-DETR-L on Manga109 that evidence STAL effectiveness for small-text detection and offers practical guidance for selecting real-time detectors in manga image analysis.
Downloads
References
[1] K. Aizawa et al., "Building a Manga Dataset 'Manga109' with Annotations for Multimedia Applications," IEEE MultiMedia, vol. 27, no. 2, pp. 8-18, 2020, doi: 10.1109/MMUL.2020.2987895.
[2] Y. Matsui, K. Ito, Y. Aramaki, A. Fujimoto, T. Ogawa, T. Yamasaki, and K. Aizawa, "Sketch-based Manga Retrieval using Manga109 Dataset," Multimedia Tools and Applications, vol. 76, no. 20, pp. 21811-21838, 2017, doi: 10.1007/s11042-016-4020-z.
[3] T. Ogawa, A. Otsubo, R. Narita, Y. Matsui, T. Yamasaki, and K. Aizawa, "Object Detection for Comics using Manga109 Annotations," arXiv:1803.08670, 2018.
[4] E. Vivoli, M. Bertini, and D. Karatzas, "CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding," in Proc. NeurIPS Datasets and Benchmarks Track, 2024.
[5] S. Paval, P. Meißner, and I. P. Yamshchikov, "ComicScene154: A Scene Dataset for Comic Analysis," inProc. 2025 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2025, pp. 31574-31580, doi: 10.18653/v1/2025.emnlp-main.1608.
[6] Q. Feng, X. Xu, and Z. Wang, "Deep learning-based small object detection: A survey,"Mathematical Biosciences and Engineering, vol. 20, no. 4, pp. 6551-6590, 2023, doi: 10.3934/mbe.2023282.
[7] N.-T. Le, N. T. Thai, and C. V. Bui, "Benchmarking Real-Time Object Detection: Evaluating YOLO and RT-DETR on Speed, Accuracy, and Efficiency," inInformation and Communication Technology (SOICT 2024), Communications in Computer and Information Science. Springer, 2025, pp. 224-234, doi: 10.1007/978-981-96-4285-4_19.
[8] N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, "End-to-End Object Detection with Transformers," in Proc. ECCV, 2020.
[9] Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and J. Chen, "DETRs Beat YOLOs on Real-time Object Detection," in Proc. IEEE/CVF CVPR, 2024.
[10] R. Sapkota, R. H. Cheppally, A. Sharda, and M. Karkee, "YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection," arXiv:2509.25164, 2026.
[11] A. J. González Hernández, J. P. Sánchez Hernández, D. L. Hernández Rabadán, J. Frausto Solis, and J. González Barbosa, "An analysis of YOLO models versus RT-DETR applied to multi-object detection in images,"International Journal of Combinatorial Optimization Problems and Informatics, vol. 16, no. 3, pp. 360-377, 2025, doi: 10.61467/2007.1558.2025.v16i3.778.
[12] T.-Y. Lin et al., "Microsoft COCO: Common Objects in Context," in Proc. ECCV, 2014.
[13] W. Lv, Y. Zhao, Q. Chang, K. Huang, G. Wang, and Y. Liu, "RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer," arXiv:2407.17140, 2024.
[14] S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, "CBAM: Convolutional Block Attention Module," in Proc. ECCV, 2018.
[15] N.-V. Nguyen, C. Rigaud, and J.-C. Burie, "Comic MTL: Optimized Multi-task Learning for Comic Book Image Analysis," International Journal on Document Analysis and Recognition (IJDAR), vol. 22, pp. 265-284, 2019.
[16] J. Baek, A. Miyai, S. Onohara, H. Ikuta, and K. Aizawa, "Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding," in Proc. Culture × AI Workshop at ICML, 2026.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Daniel Sande Bona, Tindia Febriyati

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.










