Yayın:
LLaVa-RS: A Unified Model for Image and Change Captioning in Remote Sensing

dc.contributor.authorYildirim, Turabi
dc.contributor.authorAmasyali, Mehmet Fatih
dc.contributor.authorKaraca, Ali Can
dc.date.accessioned2026-06-27T15:30:07Z
dc.date.issued2025
dc.description.abstractVisual instruction tuning has significantly advanced the functionality of Multimodal Large Language Models (MLMMs), enabling enhanced comprehension and response capabilities in many forms of data. However, existing open MLLMs remain largely focused on tasks with general types of images, with little exploration into remote sensing applications and captioning differences at two images. To adress these issues we propose LLaVa-RS, a MLLM that is focused on remote images with features such as image captioning, and change captioning. Code available at https://github.com/ChangeCapsInRS/LLAVA-RS.en
dc.description.urihttps://doi.org/10.1109/siu66497.2025.11112506
dc.identifier.doi10.1109/siu66497.2025.11112506
dc.identifier.isbn979-8-3315-6656-2; 979-8-3315-6655-5
dc.identifier.issn2165-0608
dc.identifier.urihttps://hdl.handle.net/20.500.14981/71240
dc.identifier.wos001575462500402
dc.language.isotur
dc.publisherIEEE
dc.relation.conference33rd Conference on Signal Processing and Communications Applications-SIU-Annual
dc.relation.ispartof2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU
dc.subjectMultimodal large language models
dc.subjectlarge vision models
dc.subjectremote sensing change captioning
dc.subjectimage captioning
dc.subjectComputer Science
dc.subjectEngineering
dc.subjectTelecommunications
dc.titleLLaVa-RS: A Unified Model for Image and Change Captioning in Remote Sensing
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar