Yayın: LLaVa-RS: A Unified Model for Image and Change Captioning in Remote Sensing
Yükleniyor...
Tarih
Danışman
item.page.editor
Editör
Bölüm / Program
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
IEEE
DOI
10.1109/siu66497.2025.11112506
Özet
Visual instruction tuning has significantly advanced the functionality of Multimodal Large Language Models (MLMMs), enabling enhanced comprehension and response capabilities in many forms of data. However, existing open MLLMs remain largely focused on tasks with general types of images, with little exploration into remote sensing applications and captioning differences at two images. To adress these issues we propose LLaVa-RS, a MLLM that is focused on remote images with features such as image captioning, and change captioning. Code available at https://github.com/ChangeCapsInRS/LLAVA-RS.
Tanım
Dergi veya Seri
2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU
ISSN
2165-0608
ISBN
979-8-3315-6656-2; 979-8-3315-6655-5