Yayın:
Turkish Scene Text Recognition with a Lightweight and Robust Transformer

dc.contributor.authorYildiz, Serdar
dc.date.accessioned2026-06-27T15:31:04Z
dc.date.issued2025
dc.description.abstractIn this study, we propose two lightweight vision transformers, ViT-TR-Tiny and ViT-TR-Nano, for scene text recognition. These models achieve the optimal balance of recognition accuracy and computational efficiency by significantly reducing overall network complexity. Experimental results show that the proposed models achieve competitive word accuracy with only minor accuracy degradation when compared to wellknown approaches in the literature. Remarkably, the TensorRT-optimized ViT-TR-Tiny achieved 93.44% word accuracy on STRIT and 92.78% on TS-TR while processing 2264 images per second. These findings highlight the promise of efficient transformer-based architectures for tackling complex scene text recognition tasks, particularly in Turkish.en
dc.description.urihttps://doi.org/10.1109/siu66497.2025.11111830
dc.identifier.doi10.1109/siu66497.2025.11111830
dc.identifier.isbn979-8-3315-6656-2; 979-8-3315-6655-5
dc.identifier.issn2165-0608
dc.identifier.urihttps://hdl.handle.net/20.500.14981/71439
dc.identifier.wos001575462500048
dc.language.isotur
dc.publisherIEEE
dc.relation.conference33rd Conference on Signal Processing and Communications Applications-SIU-Annual
dc.relation.ispartof2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU
dc.subjectscene text recognition
dc.subjectoptical character recognition
dc.subjectvision transformer
dc.subjectComputer Science
dc.subjectEngineering
dc.subjectTelecommunications
dc.titleTurkish Scene Text Recognition with a Lightweight and Robust Transformer
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar