Yayın:
TRCaptionNet plus plus : A high-performance encoder-decoder based deep Turkish image captioning model fine-tuned with a large-scale set of pretrain data

dc.contributor.authorYildiz, Serdar
dc.contributor.authorMemis, Abbas
dc.contributor.authorVarli, Songul
dc.date.accessioned2026-06-27T15:25:20Z
dc.date.issued2025
dc.description.abstractThis paper introduces a novel and high-performance encoder-decoder-based deep model called TRCaption-Net++ for generic Turkish image captioning tasks. The proposed model is an improved and refined version of TRCaptionNet, which essentially employs a CLIP (contrastive language-image pretraining) image encoder, a feature projection layer and a BERT (bidirectional encoder representations from transformers) text decoder. Within the scope of the study, the regular TRCaptionNet model was trained and specifically fine-tuned with a massive set of image data. In this respect, approximately 2,000,000 random images representing the words in the MS COCO and Flickr caption sets were retrieved through web crawling in the initial stage. Then, nearly 8,000,000 caption texts were generated for each image via 4 different image captioning models. Finally, the text decoder module of the proposed model was improved by using the image-caption features of these crawled images. The performance evaluation test of the TRCaptionNet++ model was carried out on two Turkish caption datasets (TasvirEt and Turkish MS COCO) and two machine-translated caption sets (MS COCO and Flickr30K) by measuring common image captioning metrics such as BLEU, METEOR, ROUGE-L, CIDEr and SPICE. As a result of the performance tests, quite remarkable captioning success rates were achieved and it is observed that the proposed model has a superior performance outperforming all the related works. Project details and demo links of TRCaptionNet++ will also be available on the project's page https://serdaryildiz.com/TRCaptionNetppen
dc.description.urihttps://doi.org/10.55730/1300-0632.4150
dc.identifier.doi10.55730/1300-0632.4150
dc.identifier.eissn1303-6203
dc.identifier.issn1300-0632
dc.identifier.issue5
dc.identifier.urihttps://hdl.handle.net/20.500.14981/70782
dc.identifier.volume33
dc.identifier.wos001585775800007
dc.language.isoeng
dc.publisherTubitak Scientific & Technological Research Council Turkey
dc.relation.ispartofTURKISH JOURNAL OF ELECTRICAL ENGINEERING AND COMPUTER SCIENCES
dc.rightsopenAccess
dc.subjectImage captioning
dc.subjectTurkish image captioning
dc.subjectimage encoders
dc.subjecttext decoders
dc.subjectdeep learning
dc.subjectComputer Science
dc.subjectEngineering
dc.titleTRCaptionNet plus plus : A high-performance encoder-decoder based deep Turkish image captioning model fine-tuned with a large-scale set of pretrain data
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar