Yayın: Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework
| dc.contributor.author | Karaca, Ali Can | |
| dc.contributor.author | Ozelbas, Enes | |
| dc.contributor.author | Berber, Saadettin | |
| dc.contributor.author | Karimli, Orkhan | |
| dc.contributor.author | Yildirim, Turabi | |
| dc.contributor.author | Amasyali, Mehmet Fatih | |
| dc.date.accessioned | 2026-06-27T15:22:59Z | |
| dc.date.issued | 2025 | |
| dc.description.abstract | Existing remote sensing image change captioning (RSICC) methods often fail under challenges, such as illumination differences, viewpoint changes, and blur effects, leading to inaccuracies, especially in no-change regions. Moreover, images acquired at different spatial resolutions and with registration errors tend to affect the captions. To address these issues, we introduce SECOND-CC, a novel RSICC dataset featuring high-resolution RGB image pairs, semantic segmentation maps, and diverse real-world scenarios. SECOND-CC contains 6041 pairs of bitemporal remote sensing images and 30 205 sentences describing the differences between the images. In addition, we propose MModalCC, a multimodal framework that integrates semantic and visual data using advanced attention mechanisms, including cross-modal cross attention and multimodal gated cross attention. In addition, we adapt MModalCC to handle noisy semantic inputs by integrating a semantic change detector, improving its robustness for real-world applications. Detailed ablation studies and attention visualizations further demonstrate its effectiveness and ability to address the challenges of RSICC. Comprehensive experiments show that MModalCC outperforms state-of-the-art RSICC methods, including RSICCformer, Chg2Cap, and PSNet with +4.6% improvement on BLEU4 score and +9.6% improvement on CIDEr score in SECOND-CC dataset. MModalCC was further validated on the LEVIR-MCI benchmark, where it achieved an average S*m score of 83.51, significantly outperforming previous state-of-the-art methods. | en |
| dc.description.sponsorship | Scientific and Technological Research Council of Turkey (TUBITAK) [3501, 122E666] | |
| dc.description.uri | https://doi.org/10.1109/jstars.2025.3600613 | |
| dc.identifier.doi | 10.1109/jstars.2025.3600613 | |
| dc.identifier.eissn | 2151-1535 | |
| dc.identifier.endpage | 21513 | |
| dc.identifier.issn | 1939-1404 | |
| dc.identifier.startpage | 21494 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/70302 | |
| dc.identifier.volume | 18 | |
| dc.identifier.wos | 001566927000013 | |
| dc.language.iso | eng | |
| dc.publisher | IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC | |
| dc.relation.ispartof | IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING | |
| dc.rights | openAccess | |
| dc.subject | Semantics | |
| dc.subject | Feature extraction | |
| dc.subject | Remote sensing | |
| dc.subject | Decoding | |
| dc.subject | Buildings | |
| dc.subject | Adaptation models | |
| dc.subject | Visualization | |
| dc.subject | Vegetation mapping | |
| dc.subject | Logic gates | |
| dc.subject | Earth | |
| dc.subject | Change captioning | |
| dc.subject | multimodal change captioning (MModalCC) | |
| dc.subject | remote sensing (RS) images | |
| dc.subject | Engineering | |
| dc.subject | Physical Geography | |
| dc.subject | Imaging Science & Photographic Technology | |
| dc.title | Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework | |
| dc.type | Article | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |