Yayın: Turkish Scene Text Recognition with a Lightweight and Robust Transformer
Yükleniyor...
Tarih
Yazarlar
Danışman
item.page.editor
Editör
Bölüm / Program
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
IEEE
DOI
10.1109/siu66497.2025.11111830
Özet
In this study, we propose two lightweight vision transformers, ViT-TR-Tiny and ViT-TR-Nano, for scene text recognition. These models achieve the optimal balance of recognition accuracy and computational efficiency by significantly reducing overall network complexity. Experimental results show that the proposed models achieve competitive word accuracy with only minor accuracy degradation when compared to wellknown approaches in the literature. Remarkably, the TensorRT-optimized ViT-TR-Tiny achieved 93.44% word accuracy on STRIT and 92.78% on TS-TR while processing 2264 images per second. These findings highlight the promise of efficient transformer-based architectures for tackling complex scene text recognition tasks, particularly in Turkish.
Tanım
Dergi veya Seri
2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU
ISSN
2165-0608
ISBN
979-8-3315-6656-2; 979-8-3315-6655-5