Imagine How To Change: Explicit Procedure Modeling for Change Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Jiayang, Guo, Zixin, Cao, Min, Zhu, Guibo, Laaksonen, Jorma |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hybrid Deep Learning for Hyperspectral Single Image Super-Resolution
by: Muhammad, Usman, et al.
Published: (2025)
by: Muhammad, Usman, et al.
Published: (2025)
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Figuring out Figures: Using Textual References to Caption Scientific Figures
by: Cao, Stanley, et al.
Published: (2024)
by: Cao, Stanley, et al.
Published: (2024)
Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning
by: Chen, Fuhai, et al.
Published: (2026)
by: Chen, Fuhai, et al.
Published: (2026)
OSCBench: Benchmarking Object State Change in Text-to-Video Generation
by: Han, Xianjing, et al.
Published: (2026)
by: Han, Xianjing, et al.
Published: (2026)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
SAM Guided Semantic and Motion Changed Region Mining for Remote Sensing Change Captioning
by: Wang, Futian, et al.
Published: (2025)
by: Wang, Futian, et al.
Published: (2025)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
by: Bai, Longju, et al.
Published: (2024)
by: Bai, Longju, et al.
Published: (2024)
Decoding fMRI Data into Captions using Prefix Language Modeling
by: Shen, Vyacheslav, et al.
Published: (2025)
by: Shen, Vyacheslav, et al.
Published: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
Unveiling the Invisible: Captioning Videos with Metaphors
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
by: Chen, Pingyi, et al.
Published: (2023)
by: Chen, Pingyi, et al.
Published: (2023)
Mask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change Captioning
by: Sun, Dongwei, et al.
Published: (2024)
by: Sun, Dongwei, et al.
Published: (2024)
Updating CLIP to Prefer Descriptions Over Captions
by: Zur, Amir, et al.
Published: (2024)
by: Zur, Amir, et al.
Published: (2024)
Vision-Language Agents for Interactive Forest Change Analysis
by: Brock, James, et al.
Published: (2026)
by: Brock, James, et al.
Published: (2026)
AC-Lite : A Lightweight Image Captioning Model for Low-Resource Assamese Language
by: Choudhury, Pankaj, et al.
Published: (2025)
by: Choudhury, Pankaj, et al.
Published: (2025)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
by: Ghaboura, Sara, et al.
Published: (2024)
by: Ghaboura, Sara, et al.
Published: (2024)
A Fusion-Guided Inception Network for Hyperspectral Image Super-Resolution
by: Muhammad, Usman, et al.
Published: (2025)
by: Muhammad, Usman, et al.
Published: (2025)
StableI2I: Spotting Unintended Changes in Image-to-Image Transition
by: Li, Jiayang, et al.
Published: (2026)
by: Li, Jiayang, et al.
Published: (2026)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
CIC: A Framework for Culturally-Aware Image Captioning
by: Yun, Youngsik, et al.
Published: (2024)
by: Yun, Youngsik, et al.
Published: (2024)
HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments
by: Zhou, Qinhong, et al.
Published: (2024)
by: Zhou, Qinhong, et al.
Published: (2024)
HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
by: Dönmez, Esra, et al.
Published: (2026)
by: Dönmez, Esra, et al.
Published: (2026)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
by: Lee, Yebin, et al.
Published: (2024)
by: Lee, Yebin, et al.
Published: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
by: Kim, Si-Woo, et al.
Published: (2025)
by: Kim, Si-Woo, et al.
Published: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
by: Kim, Hyunjong, et al.
Published: (2025)
by: Kim, Hyunjong, et al.
Published: (2025)
Transfer Learning from ImageNet for MEG-Based Decoding of Imagined Speech
by: Jhilal, Soufiane, et al.
Published: (2026)
by: Jhilal, Soufiane, et al.
Published: (2026)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
by: Bucciarelli, Davide, et al.
Published: (2024)
by: Bucciarelli, Davide, et al.
Published: (2024)
MAMI: Multi-Attentional Mutual-Information for Long Sequence Neuron Captioning
by: Fauzulhaq, Alfirsa Damasyifa, et al.
Published: (2024)
by: Fauzulhaq, Alfirsa Damasyifa, et al.
Published: (2024)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
by: Matsuda, Kazuki, et al.
Published: (2024)
by: Matsuda, Kazuki, et al.
Published: (2024)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
by: Gondal, Moazzam Umer, et al.
Published: (2025)
by: Gondal, Moazzam Umer, et al.
Published: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
by: Du, Yang, et al.
Published: (2025)
by: Du, Yang, et al.
Published: (2025)
Similar Items
-
Hybrid Deep Learning for Hyperspectral Single Image Super-Resolution
by: Muhammad, Usman, et al.
Published: (2025) -
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024) -
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026) -
Figuring out Figures: Using Textual References to Caption Scientific Figures
by: Cao, Stanley, et al.
Published: (2024) -
Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning
by: Chen, Fuhai, et al.
Published: (2026)