Yayın: LLaVa-RS: A Unified Model for Image and Change Captioning in Remote Sensing
| dc.contributor.author | Yildirim, Turabi | |
| dc.contributor.author | Amasyali, Mehmet Fatih | |
| dc.contributor.author | Karaca, Ali Can | |
| dc.date.accessioned | 2026-06-27T15:30:07Z | |
| dc.date.issued | 2025 | |
| dc.description.abstract | Visual instruction tuning has significantly advanced the functionality of Multimodal Large Language Models (MLMMs), enabling enhanced comprehension and response capabilities in many forms of data. However, existing open MLLMs remain largely focused on tasks with general types of images, with little exploration into remote sensing applications and captioning differences at two images. To adress these issues we propose LLaVa-RS, a MLLM that is focused on remote images with features such as image captioning, and change captioning. Code available at https://github.com/ChangeCapsInRS/LLAVA-RS. | en |
| dc.description.uri | https://doi.org/10.1109/siu66497.2025.11112506 | |
| dc.identifier.doi | 10.1109/siu66497.2025.11112506 | |
| dc.identifier.isbn | 979-8-3315-6656-2; 979-8-3315-6655-5 | |
| dc.identifier.issn | 2165-0608 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/71240 | |
| dc.identifier.wos | 001575462500402 | |
| dc.language.iso | tur | |
| dc.publisher | IEEE | |
| dc.relation.conference | 33rd Conference on Signal Processing and Communications Applications-SIU-Annual | |
| dc.relation.ispartof | 2025 33RD SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, SIU | |
| dc.subject | Multimodal large language models | |
| dc.subject | large vision models | |
| dc.subject | remote sensing change captioning | |
| dc.subject | image captioning | |
| dc.subject | Computer Science | |
| dc.subject | Engineering | |
| dc.subject | Telecommunications | |
| dc.title | LLaVa-RS: A Unified Model for Image and Change Captioning in Remote Sensing | |
| dc.type | Proceedings Paper | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |