CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Guanghao, Zhong, Tao, Xia, Yan, Liu, Mushui, Yu, Zhelun, Li, Haoyuan, He, Wanggui, Shu, Fangxun, She, Dong, Wang, Yi, Jiang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
by: Xiao, Wenyi, et al.
Published: (2024)
by: Xiao, Wenyi, et al.
Published: (2024)
MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
by: He, Wanggui, et al.
Published: (2024)
by: He, Wanggui, et al.
Published: (2024)
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
by: Huang, Ziwei, et al.
Published: (2024)
by: Huang, Ziwei, et al.
Published: (2024)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
by: Shu, Fangxun, et al.
Published: (2024)
by: Shu, Fangxun, et al.
Published: (2024)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
by: Liu, Jiang, et al.
Published: (2024)
by: Liu, Jiang, et al.
Published: (2024)
Enhancing Retrieval Augmentation via Adversarial Collaboration
by: Zhang, Letian, et al.
Published: (2025)
by: Zhang, Letian, et al.
Published: (2025)
CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers
by: She, D., et al.
Published: (2025)
by: She, D., et al.
Published: (2025)
SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought
by: Li, Guanghao, et al.
Published: (2025)
by: Li, Guanghao, et al.
Published: (2025)
PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning
by: Liu, Jinlong, et al.
Published: (2026)
by: Liu, Jinlong, et al.
Published: (2026)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
by: She, Dong, et al.
Published: (2025)
by: She, Dong, et al.
Published: (2025)
MemoPhishAgent: Memory-Augmented Multi-Modal LLM Agent for Phishing URL Detection
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
by: Zhang, Wenqiao, et al.
Published: (2024)
by: Zhang, Wenqiao, et al.
Published: (2024)
Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning
by: Kim, Dahun, et al.
Published: (2026)
by: Kim, Dahun, et al.
Published: (2026)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Interleaved-Modal Chain-of-Thought
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
High-Dimensional Multi-Study Multi-Modality Covariate-Augmented Generalized Factor Model
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning
by: Chen, Yiqun, et al.
Published: (2026)
by: Chen, Yiqun, et al.
Published: (2026)
SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification
by: Zhang, Yunkai, et al.
Published: (2025)
by: Zhang, Yunkai, et al.
Published: (2025)
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition
by: Lin, Tianwei, et al.
Published: (2024)
by: Lin, Tianwei, et al.
Published: (2024)
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
by: Wang, Xierui, et al.
Published: (2024)
by: Wang, Xierui, et al.
Published: (2024)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
by: Wang, Yaoting, et al.
Published: (2025)
by: Wang, Yaoting, et al.
Published: (2025)
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
by: Shao, Yifei, et al.
Published: (2026)
by: Shao, Yifei, et al.
Published: (2026)
MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
by: Nguyen, Thang, et al.
Published: (2025)
by: Nguyen, Thang, et al.
Published: (2025)
Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation
by: Xiong, Siheng, et al.
Published: (2025)
by: Xiong, Siheng, et al.
Published: (2025)
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback
by: Byun, Ju-Seung, et al.
Published: (2024)
by: Byun, Ju-Seung, et al.
Published: (2024)
Chain-of-Thought Augmentation with Logit Contrast for Enhanced Reasoning in Language Models
by: Shim, Jay, et al.
Published: (2024)
by: Shim, Jay, et al.
Published: (2024)
LLMs can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought
by: Jiang, Zhuoxuan, et al.
Published: (2024)
by: Jiang, Zhuoxuan, et al.
Published: (2024)
Degeneration of the archimedean height pairing of algebraically trivial cycles
by: Chen, Zhelun
Published: (2025)
by: Chen, Zhelun
Published: (2025)
Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities
by: Song, Peibo, et al.
Published: (2026)
by: Song, Peibo, et al.
Published: (2026)
Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems
by: Feng, Tao, et al.
Published: (2026)
by: Feng, Tao, et al.
Published: (2026)
SAG: Style-Aligned Article Generation via Model Collaboration
by: Xu, Chenning, et al.
Published: (2024)
by: Xu, Chenning, et al.
Published: (2024)
Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration
by: Shen, Zhixuan, et al.
Published: (2024)
by: Shen, Zhixuan, et al.
Published: (2024)
Understanding Before Reasoning: Enhancing Chain-of-Thought with Iterative Summarization Pre-Prompting
by: Zhu, Dong-Hai, et al.
Published: (2025)
by: Zhu, Dong-Hai, et al.
Published: (2025)
CoTKR: Chain-of-Thought Enhanced Knowledge Rewriting for Complex Knowledge Graph Question Answering
by: Wu, Yike, et al.
Published: (2024)
by: Wu, Yike, et al.
Published: (2024)
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
by: Liu, Zhenghao, et al.
Published: (2025)
by: Liu, Zhenghao, et al.
Published: (2025)
Observing Schrödinger's Cat with Artificial Intelligence: Emergent Classicality from Information Bottleneck
by: Zhang, Zhelun, et al.
Published: (2023)
by: Zhang, Zhelun, et al.
Published: (2023)
Similar Items
-
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025) -
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025) -
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI Feedback
by: Xiao, Wenyi, et al.
Published: (2024) -
MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
by: He, Wanggui, et al.
Published: (2024) -
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
by: Huang, Ziwei, et al.
Published: (2024)