Yayın:
Turkish scene text recognition: Introducing extensive real and synthetic datasets and a novel recognition model

dc.contributor.authorYildiz, Serdar
dc.date.accessioned2026-06-27T14:59:59Z
dc.date.issued2024
dc.description.abstractIn the advancing field of computer vision, scene text recognition (STR) has been progressively gaining prominence. Despite this progress, the lack of a comprehensive study or a suitable dataset for STR, particularly for languages like Turkish, stands out. Existing datasets, regardless of the language, tend to grapple with issues such as limited sample quantity and high noise levels, which considerably restrict the progression and overall efficacy of STR research and applications. Addressing these shortcomings, we introduce the Turkish Scene Text Recognition (TS-TR) dataset, one of the most substantial STR datasets to date, comprising 7288 text instances. In addition, we propose the Synthetic Turkish Scene Text Recognition (STS-TR) dataset, an enormous collection of 12 million samples created using a novel histogram-based method, more efficient than common synthetic data generation methods. Moreover, we present a novel recognition model, the Masked Vision Transformer for Text Recognition (MViT-TR), which achieves a word accuracy of 94.42% on the challenging TS-TR test dataset, underlining its robustness and performance efficacy. We extend our investigation to the influence of synthetic datasets, the utilization of patch masking, and the function of the position attention module on recognition performance. To foster future STR research, we have made all datasets and source codes publicly available.en
dc.description.urihttps://doi.org/10.1016/j.jestch.2024.101881
dc.identifier.doi10.1016/j.jestch.2024.101881
dc.identifier.issn2215-0986
dc.identifier.urihttps://hdl.handle.net/20.500.14981/66958
dc.identifier.volume60
dc.identifier.wos001355199700001
dc.language.isoeng
dc.publisherELSEVIER - DIVISION REED ELSEVIER INDIA PVT LTD
dc.relation.ispartofENGINEERING SCIENCE AND TECHNOLOGY-AN INTERNATIONAL JOURNAL-JESTECH
dc.rightsopenAccess
dc.subjectScene text recognition dataset
dc.subjectSynthetic scene text recognition dataset
dc.subjectPatch masking
dc.subjectPosition attention
dc.subjectVision transformers
dc.subjectNETWORK
dc.subjectEngineering
dc.titleTurkish scene text recognition: Introducing extensive real and synthetic datasets and a novel recognition model
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar