Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Corbière, Charles, Roburin, Simon, Montariol, Syrielle, Bosselut, Antoine, Alahi, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Helvipad: A Real-World Dataset for Omnidirectional Stereo Depth Estimation
by: Zayene, Mehdi, et al.
Published: (2024)
by: Zayene, Mehdi, et al.
Published: (2024)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
by: Gao, Silin, et al.
Published: (2025)
by: Gao, Silin, et al.
Published: (2025)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025)
by: Bhagwatkar, Rishika, et al.
Published: (2025)
Interleaved-Modal Chain-of-Thought
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
by: Zhang, Zhenguo, et al.
Published: (2025)
by: Zhang, Zhenguo, et al.
Published: (2025)
Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model
by: Endres, Jannik, et al.
Published: (2025)
by: Endres, Jannik, et al.
Published: (2025)
PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection
by: Mamooler, Sepideh, et al.
Published: (2024)
by: Mamooler, Sepideh, et al.
Published: (2024)
ConGeo: Robust Cross-view Geo-localization across Ground View Variations
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
by: Loison, António, et al.
Published: (2026)
by: Loison, António, et al.
Published: (2026)
Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought
by: Guo, Yuchen, et al.
Published: (2026)
by: Guo, Yuchen, et al.
Published: (2026)
PRIMEDrive-CoT: A Precognitive Chain-of-Thought Framework for Uncertainty-Aware Object Interaction in Driving Scene Scenario
by: Mandalika, Sriram, et al.
Published: (2025)
by: Mandalika, Sriram, et al.
Published: (2025)
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
by: Butsanets, Léo, et al.
Published: (2025)
by: Butsanets, Léo, et al.
Published: (2025)
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
by: Guo, Guangfu, et al.
Published: (2026)
by: Guo, Guangfu, et al.
Published: (2026)
CODE: Confident Ordinary Differential Editing
by: van Delft, Bastien, et al.
Published: (2024)
by: van Delft, Bastien, et al.
Published: (2024)
CoT-Drive: Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
by: Wu, Yike, et al.
Published: (2025)
by: Wu, Yike, et al.
Published: (2025)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
by: Kancheti, Sai Srinivas, et al.
Published: (2026)
Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
by: Jain, Riddhi, et al.
Published: (2025)
by: Jain, Riddhi, et al.
Published: (2025)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
by: Lim, Byeonggeuk, et al.
Published: (2026)
by: Lim, Byeonggeuk, et al.
Published: (2026)
CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
by: Yi, Shixin, et al.
Published: (2025)
by: Yi, Shixin, et al.
Published: (2025)
Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
by: Kim, Dong-Hee, et al.
Published: (2026)
by: Kim, Dong-Hee, et al.
Published: (2026)
Long-term Traffic Simulation with Interleaved Autoregressive Motion and Scenario Generation
by: Yang, Xiuyu, et al.
Published: (2025)
by: Yang, Xiuyu, et al.
Published: (2025)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
by: Li, Zhang, et al.
Published: (2026)
by: Li, Zhang, et al.
Published: (2026)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
UniTalk: Towards Universal Active Speaker Detection in Real World Scenarios
by: Nguyen, Le Thien Phuc, et al.
Published: (2025)
by: Nguyen, Le Thien Phuc, et al.
Published: (2025)
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
by: Lu, Wenting, et al.
Published: (2026)
by: Lu, Wenting, et al.
Published: (2026)
On the Real-World Adversarial Robustness of Real-Time Semantic Segmentation Models for Autonomous Driving
by: Rossolini, Giulio, et al.
Published: (2022)
by: Rossolini, Giulio, et al.
Published: (2022)
Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
by: Wu, Yuan, et al.
Published: (2026)
by: Wu, Yuan, et al.
Published: (2026)
A Multi-Loss Strategy for Vehicle Trajectory Prediction: Combining Off-Road, Diversity, and Directional Consistency Losses
by: Rahimi, Ahmad, et al.
Published: (2024)
by: Rahimi, Ahmad, et al.
Published: (2024)
DS MYOLO: A Reliable Object Detector Based on SSMs for Driving Scenarios
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Research on Driving Scenario Technology Based on Multimodal Large Lauguage Model Optimization
by: Mengjie, Wang, et al.
Published: (2025)
by: Mengjie, Wang, et al.
Published: (2025)
Reinforcing Structured Chain-of-Thought for Video Understanding
by: Wang, Peiyao, et al.
Published: (2026)
by: Wang, Peiyao, et al.
Published: (2026)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
by: Gao, Hong, et al.
Published: (2025)
by: Gao, Hong, et al.
Published: (2025)
Similar Items
-
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
by: Mi, Li, et al.
Published: (2024) -
Helvipad: A Real-World Dataset for Omnidirectional Stereo Depth Estimation
by: Zayene, Mehdi, et al.
Published: (2024) -
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
by: Gao, Silin, et al.
Published: (2025) -
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
by: Bhagwatkar, Rishika, et al.
Published: (2025) -
Interleaved-Modal Chain-of-Thought
by: Gao, Jun, et al.
Published: (2024)