Yayın:
Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework

dc.contributor.authorKaraca, Ali Can
dc.contributor.authorOzelbas, Enes
dc.contributor.authorBerber, Saadettin
dc.contributor.authorKarimli, Orkhan
dc.contributor.authorYildirim, Turabi
dc.contributor.authorAmasyali, Mehmet Fatih
dc.date.accessioned2026-06-27T15:22:59Z
dc.date.issued2025
dc.description.abstractExisting remote sensing image change captioning (RSICC) methods often fail under challenges, such as illumination differences, viewpoint changes, and blur effects, leading to inaccuracies, especially in no-change regions. Moreover, images acquired at different spatial resolutions and with registration errors tend to affect the captions. To address these issues, we introduce SECOND-CC, a novel RSICC dataset featuring high-resolution RGB image pairs, semantic segmentation maps, and diverse real-world scenarios. SECOND-CC contains 6041 pairs of bitemporal remote sensing images and 30 205 sentences describing the differences between the images. In addition, we propose MModalCC, a multimodal framework that integrates semantic and visual data using advanced attention mechanisms, including cross-modal cross attention and multimodal gated cross attention. In addition, we adapt MModalCC to handle noisy semantic inputs by integrating a semantic change detector, improving its robustness for real-world applications. Detailed ablation studies and attention visualizations further demonstrate its effectiveness and ability to address the challenges of RSICC. Comprehensive experiments show that MModalCC outperforms state-of-the-art RSICC methods, including RSICCformer, Chg2Cap, and PSNet with +4.6% improvement on BLEU4 score and +9.6% improvement on CIDEr score in SECOND-CC dataset. MModalCC was further validated on the LEVIR-MCI benchmark, where it achieved an average S*m score of 83.51, significantly outperforming previous state-of-the-art methods.en
dc.description.sponsorshipScientific and Technological Research Council of Turkey (TUBITAK) [3501, 122E666]
dc.description.urihttps://doi.org/10.1109/jstars.2025.3600613
dc.identifier.doi10.1109/jstars.2025.3600613
dc.identifier.eissn2151-1535
dc.identifier.endpage21513
dc.identifier.issn1939-1404
dc.identifier.startpage21494
dc.identifier.urihttps://hdl.handle.net/20.500.14981/70302
dc.identifier.volume18
dc.identifier.wos001566927000013
dc.language.isoeng
dc.publisherIEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC
dc.relation.ispartofIEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING
dc.rightsopenAccess
dc.subjectSemantics
dc.subjectFeature extraction
dc.subjectRemote sensing
dc.subjectDecoding
dc.subjectBuildings
dc.subjectAdaptation models
dc.subjectVisualization
dc.subjectVegetation mapping
dc.subjectLogic gates
dc.subjectEarth
dc.subjectChange captioning
dc.subjectmultimodal change captioning (MModalCC)
dc.subjectremote sensing (RS) images
dc.subjectEngineering
dc.subjectPhysical Geography
dc.subjectImaging Science & Photographic Technology
dc.titleRobust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC Framework
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar