Leveraging Textual Compositional Reasoning for Robust Change Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Kyu Ri, Park, Jiyoung, Kim, Seong Tae, Lee, Hong Joo, Kim, Jung Uk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917111191830528
author Park, Kyu Ri
Park, Jiyoung
Kim, Seong Tae
Lee, Hong Joo
Kim, Jung Uk
author_facet Park, Kyu Ri
Park, Jiyoung
Kim, Seong Tae
Lee, Hong Joo
Kim, Jung Uk
contents Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and compositional semantics. To alleviate this, we present CORTEX (COmpositional Reasoning-aware TEXt-guided), a novel framework that integrates complementary textual cues to enhance change understanding. In addition to capturing cues from pixel-level differences, CORTEX utilizes scene-level textual knowledge provided by Vision Language Models (VLMs) to extract richer image text signals that reveal underlying compositional reasoning. CORTEX consists of three key modules: (i) an Image-level Change Detector that identifies low-level visual differences between paired images, (ii) a Reasoning-aware Text Extraction (RTE) module that use VLMs to generate compositional reasoning descriptions implicit in visual features, and (iii) an Image-Text Dual Alignment (ITDA) module that aligns visual and textual features for fine-grained relational reasoning. This enables CORTEX to reason over visual and textual features and capture changes that are otherwise ambiguous in visual features alone.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22903
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Textual Compositional Reasoning for Robust Change Captioning
Park, Kyu Ri
Park, Jiyoung
Kim, Seong Tae
Lee, Hong Joo
Kim, Jung Uk
Computer Vision and Pattern Recognition
Artificial Intelligence
Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to represent explicitly structured information such as object relationships and compositional semantics. To alleviate this, we present CORTEX (COmpositional Reasoning-aware TEXt-guided), a novel framework that integrates complementary textual cues to enhance change understanding. In addition to capturing cues from pixel-level differences, CORTEX utilizes scene-level textual knowledge provided by Vision Language Models (VLMs) to extract richer image text signals that reveal underlying compositional reasoning. CORTEX consists of three key modules: (i) an Image-level Change Detector that identifies low-level visual differences between paired images, (ii) a Reasoning-aware Text Extraction (RTE) module that use VLMs to generate compositional reasoning descriptions implicit in visual features, and (iii) an Image-Text Dual Alignment (ITDA) module that aligns visual and textual features for fine-grained relational reasoning. This enables CORTEX to reason over visual and textual features and capture changes that are otherwise ambiguous in visual features alone.
title Leveraging Textual Compositional Reasoning for Robust Change Captioning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.22903