Publication: LLaVa-RS: A Unified Model for Image and Change Captioning in Remote Sensing
Loading...
Date
Advisor
item.page.editor
Editor
Department
Journal Title
Journal ISSN
Volume Title
Publisher
IEEE
DOI
10.1109/siu66497.2025.11112506
Abstract
Visual instruction tuning has significantly advanced the functionality of Multimodal Large Language Models (MLMMs), enabling enhanced comprehension and response capabilities in many forms of data. However, existing open MLLMs remain largely focused on tasks with general types of images, with little exploration into remote sensing applications and captioning differences at two images. To adress these issues we propose LLaVa-RS, a MLLM that is focused on remote images with features such as image captioning, and change captioning. Code available at https://github.com/ChangeCapsInRS/LLAVA-RS.
Description
Journal or Series
2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU
ISSN
2165-0608
ISBN
979-8-3315-6656-2; 979-8-3315-6655-5