Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xia, Jiaer, Tong, Bingkui, Zang, Yuhang, Shao, Rui, Zhou, Kaiyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
von: Tong, Bingkui, et al.
Veröffentlicht: (2025)
von: Tong, Bingkui, et al.
Veröffentlicht: (2025)
Measuring Epistemic Humility in Multimodal Large Language Models
von: Tong, Bingkui, et al.
Veröffentlicht: (2025)
von: Tong, Bingkui, et al.
Veröffentlicht: (2025)
Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
Streaming Video Instruction Tuning
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
von: Xia, Jiaer, et al.
Veröffentlicht: (2025)
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
Contextual Object Detection with Multimodal Large Language Models
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
von: Zang, Yuhang, et al.
Veröffentlicht: (2023)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
SHAN: Object-Level Privacy Detection via Inference on Scene Heterogeneous Graph
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2024)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2024)
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
von: Abdelrahman, Eslam, et al.
Veröffentlicht: (2023)
Beyond Visual Appearances: Privacy-sensitive Objects Identification via Hybrid Graph Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2024)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2024)
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
von: Ma, Ji, et al.
Veröffentlicht: (2026)
von: Ma, Ji, et al.
Veröffentlicht: (2026)
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
von: Wu, Jiahe, et al.
Veröffentlicht: (2026)
von: Wu, Jiahe, et al.
Veröffentlicht: (2026)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
TumorChain: Interleaved Multimodal Chain-of-Thought Reasoning for Traceable Clinical Tumor Analysis
von: Li, Sijing, et al.
Veröffentlicht: (2026)
von: Li, Sijing, et al.
Veröffentlicht: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
von: Kancheti, Sai Srinivas, et al.
Veröffentlicht: (2026)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
Unified Reward Model for Multimodal Understanding and Generation
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
Multimodal Chain-of-Thought Reasoning in Language Models
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits
von: Jin, Leyang, et al.
Veröffentlicht: (2026)
von: Jin, Leyang, et al.
Veröffentlicht: (2026)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
von: Gao, Timin, et al.
Veröffentlicht: (2024)
von: Gao, Timin, et al.
Veröffentlicht: (2024)
Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
von: Wang, Jiaze, et al.
Veröffentlicht: (2025)
von: Wang, Jiaze, et al.
Veröffentlicht: (2025)
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
von: Qian, Rui, et al.
Veröffentlicht: (2025)
von: Qian, Rui, et al.
Veröffentlicht: (2025)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
von: Li, Yunheng, et al.
Veröffentlicht: (2026)
Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
von: Zhang, Honglu, et al.
Veröffentlicht: (2025)
von: Zhang, Honglu, et al.
Veröffentlicht: (2025)
Fuel Gauge: Estimating Chain-of-Thought Length Ahead of Time in Large Multimodal Models
von: Yang, Yuedong, et al.
Veröffentlicht: (2026)
von: Yang, Yuedong, et al.
Veröffentlicht: (2026)
mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
von: Wu, Peixi, et al.
Veröffentlicht: (2026)
von: Wu, Peixi, et al.
Veröffentlicht: (2026)
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models
von: Yu, Yongsheng, et al.
Veröffentlicht: (2024)
von: Yu, Yongsheng, et al.
Veröffentlicht: (2024)
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
von: Qian, Jiahe, et al.
Veröffentlicht: (2025)
von: Qian, Jiahe, et al.
Veröffentlicht: (2025)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
von: Tong, Bingkui, et al.
Veröffentlicht: (2025) -
Measuring Epistemic Humility in Multimodal Large Language Models
von: Tong, Bingkui, et al.
Veröffentlicht: (2025) -
Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
von: Xia, Jiaer, et al.
Veröffentlicht: (2025) -
Streaming Video Instruction Tuning
von: Xia, Jiaer, et al.
Veröffentlicht: (2025) -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
von: Wang, Yibin, et al.
Veröffentlicht: (2025)