MCoT-MVS: Multi-level Vision Selection by Multi-modal Chain-of-Thought Reasoning for Composed Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Xuri, Wang, Chunhao, Wang, Xindi, Qin, Zheyun, Chen, Zhumin, Xin, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval
by: Park, Jeong-Woo, et al.
Published: (2025)
by: Park, Jeong-Woo, et al.
Published: (2025)
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026)
by: Zhang, Shutong, et al.
Published: (2026)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
by: Sun, Zelong, et al.
Published: (2025)
by: Sun, Zelong, et al.
Published: (2025)
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval
by: Tang, Yuanmin, et al.
Published: (2024)
by: Tang, Yuanmin, et al.
Published: (2024)
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
by: Zhu, Shaojie, et al.
Published: (2023)
by: Zhu, Shaojie, et al.
Published: (2023)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
by: Lin, Weihuang, et al.
Published: (2025)
by: Lin, Weihuang, et al.
Published: (2025)
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
by: Ma, Qihang, et al.
Published: (2025)
by: Ma, Qihang, et al.
Published: (2025)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Composed Multi-modal Retrieval: A Survey of Approaches and Applications
by: Zhang, Kun, et al.
Published: (2025)
by: Zhang, Kun, et al.
Published: (2025)
MGRR-Net: Multi-level Graph Relational Reasoning Network for Facial Action Units Detection
by: Ge, Xuri, et al.
Published: (2022)
by: Ge, Xuri, et al.
Published: (2022)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View Stereo
by: Yuan, Zhenlong, et al.
Published: (2024)
by: Yuan, Zhenlong, et al.
Published: (2024)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
by: Liu, Zheyuan, et al.
Published: (2023)
by: Liu, Zheyuan, et al.
Published: (2023)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
IA-MVS: Instance-Focused Adaptive Depth Sampling for Multi-View Stereo
by: Wang, Yinzhe, et al.
Published: (2025)
by: Wang, Yinzhe, et al.
Published: (2025)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
by: Chen, Qiguang, et al.
Published: (2024)
by: Chen, Qiguang, et al.
Published: (2024)
Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
by: Du, Yifan, et al.
Published: (2025)
by: Du, Yifan, et al.
Published: (2025)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
by: Wang, Yabing, et al.
Published: (2025)
by: Wang, Yabing, et al.
Published: (2025)
Multi-modal Attribute Prompting for Vision-Language Models
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
by: Zhang, Zhenguo, et al.
Published: (2025)
by: Zhang, Zhenguo, et al.
Published: (2025)
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View Stereo
by: Yuan, Zhenlong, et al.
Published: (2024)
by: Yuan, Zhenlong, et al.
Published: (2024)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
TSAR-MVS: Textureless-aware Segmentation and Correlative Refinement Guided Multi-View Stereo
by: Yuan, Zhenlong, et al.
Published: (2023)
by: Yuan, Zhenlong, et al.
Published: (2023)
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
by: Serussi, Gabriele, et al.
Published: (2026)
by: Serussi, Gabriele, et al.
Published: (2026)
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
by: Tu, Rong-Cheng, et al.
Published: (2025)
by: Tu, Rong-Cheng, et al.
Published: (2025)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
by: Wen, Junwei, et al.
Published: (2026)
by: Wen, Junwei, et al.
Published: (2026)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
Multi-level Cross-modal Alignment for Image Clustering
by: Qiu, Liping, et al.
Published: (2024)
by: Qiu, Liping, et al.
Published: (2024)
CFIR: Fast and Effective Long-Text To Image Retrieval for Large Corpora
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
DeliCIR: Deliberative Test-Time Evolutionary Hierarchical Multi-Agents for Composed Image Retrieval
by: Pei, Xingtian, et al.
Published: (2026)
by: Pei, Xingtian, et al.
Published: (2026)
Similar Items
-
MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval
by: Park, Jeong-Woo, et al.
Published: (2025) -
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026) -
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
by: Sun, Zelong, et al.
Published: (2025) -
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval
by: Tang, Yuanmin, et al.
Published: (2024) -
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
by: Zhu, Shaojie, et al.
Published: (2023)