Yayın:
Tiny TR-CAP: A novel small-scale benchmark dataset for general-purpose image captioning tasks

dc.contributor.authorMemis, Abbas
dc.contributor.authorYildiz, Serdar
dc.date.accessioned2026-06-27T15:14:03Z
dc.date.issued2025
dc.description.abstractIn the last decade, the outstanding performance of deep learning has also led to a rapid and inevitable rise in automatic image captioning, as well as the need for large amounts of data. Although well-known, conventional and publicly available datasets have been proposed for the image captioning task, the lack of ground-truth caption data still remains a major challenge in the generation of accurate image captions. To address this issue, in this paper we introduced a novel image captioning benchmark dataset called Tiny TRCAP, which consists of 1076 original images and 5380 handwritten captions (5 captions for each image with high diversity). The captions, which were translated into English using two web-based language translation APIs and a novel multilingual deep machine translation model, were tested against 11 state-of-the-art and prominent deep learning-based models, including CLIPCap, BLIP, BLIP2, FUSECAP, OFA, PromptCap, Kosmos2, MiniGPT4, LlaVA, BakLlaVA, and GIT. In the experimental studies, the accuracy statistics of the captions generated by the related models were reported in terms of the BLEU, METEOR, ROUGE-L, CIDEr, SPICE, and WMD captioning metrics, and their performance was evaluated comparatively. In the performance analysis, quite promising captioning performances were observed, and the best success rates were achieved with the OFA model with scores of 0.7097 BLEU-1, 0.5389 BLEU-2, 0.3940 BLEU-3, 0.2875 BLEU-4, 0.1797 METEOR, 0.4627 ROUGE-L, 0.2938 CIDEr, 0.0626 SPICE, and 0.4605 WMD. To support research studies in the field of image captioning, the image and caption sets of Tiny TR-CAP will also be publicly available on GitHub (https://github.com/abbasmemis/tiny_TR-CAP) for academic research purposes.en
dc.description.urihttps://doi.org/10.1016/j.jestch.2025.102009
dc.identifier.doi10.1016/j.jestch.2025.102009
dc.identifier.issn2215-0986
dc.identifier.urihttps://hdl.handle.net/20.500.14981/69282
dc.identifier.volume64
dc.identifier.wos001439491400001
dc.language.isoeng
dc.publisherELSEVIER - DIVISION REED ELSEVIER INDIA PVT LTD
dc.relation.ispartofENGINEERING SCIENCE AND TECHNOLOGY-AN INTERNATIONAL JOURNAL-JESTECH
dc.rightsopenAccess
dc.subjectImage captioning
dc.subjectImage captioning benchmark dataset
dc.subjectDataset construction
dc.subjectImage caption generation
dc.subjectDeep learning based image captioning
dc.subjectCaption generation
dc.subjectEngineering
dc.titleTiny TR-CAP: A novel small-scale benchmark dataset for general-purpose image captioning tasks
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar