Rethinking Chain-of-Thought Reasoning for Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Yiwu, Hu, Zi-Yuan, Li, Yin, Wang, Liwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
Enhancing Temporal Modeling of Video LLMs via Time Gating
von: Hu, Zi-Yuan, et al.
Veröffentlicht: (2024)
von: Hu, Zi-Yuan, et al.
Veröffentlicht: (2024)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024)
PAVE: Patching and Adapting Video Large Language Models
von: Liu, Zhuoming, et al.
Veröffentlicht: (2025)
von: Liu, Zhuoming, et al.
Veröffentlicht: (2025)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
von: Yerramilli, Sahiti, et al.
Veröffentlicht: (2025)
von: Yerramilli, Sahiti, et al.
Veröffentlicht: (2025)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
von: Li, Shaoxuan, et al.
Veröffentlicht: (2026)
von: Li, Shaoxuan, et al.
Veröffentlicht: (2026)
GeoRC: A Benchmark for Geolocation Reasoning Chains
von: Talreja, Mohit, et al.
Veröffentlicht: (2026)
von: Talreja, Mohit, et al.
Veröffentlicht: (2026)
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
Multimodal Chain-of-Thought Reasoning in Language Models
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
von: Lin, Yuanze, et al.
Veröffentlicht: (2024)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
von: Parmar, Paritosh, et al.
Veröffentlicht: (2025)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2025)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
Thought Flow Nets: From Single Predictions to Trains of Model Thought
von: Schuff, Hendrik, et al.
Veröffentlicht: (2021)
von: Schuff, Hendrik, et al.
Veröffentlicht: (2021)
Interleaved-Modal Chain-of-Thought
von: Gao, Jun, et al.
Veröffentlicht: (2024)
von: Gao, Jun, et al.
Veröffentlicht: (2024)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
von: Acuna, David, et al.
Veröffentlicht: (2025)
von: Acuna, David, et al.
Veröffentlicht: (2025)
ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
von: Li, Yifan, et al.
Veröffentlicht: (2025)
von: Li, Yifan, et al.
Veröffentlicht: (2025)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
von: Zhang, Jianrui, et al.
Veröffentlicht: (2024)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
CapsFusion: Rethinking Image-Text Data at Scale
von: Yu, Qiying, et al.
Veröffentlicht: (2023)
von: Yu, Qiying, et al.
Veröffentlicht: (2023)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
von: Han, Songhao, et al.
Veröffentlicht: (2024)
von: Han, Songhao, et al.
Veröffentlicht: (2024)
A Survey of Reasoning with Foundation Models
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
von: Sun, Jiankai, et al.
Veröffentlicht: (2023)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
von: Gu, Hexiang, et al.
Veröffentlicht: (2025)
von: Gu, Hexiang, et al.
Veröffentlicht: (2025)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2025)
von: Chen, Qiguang, et al.
Veröffentlicht: (2025)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
Rethinking Genomic Modeling Through Optical Character Recognition
von: Xiang, Hongxin, et al.
Veröffentlicht: (2026)
von: Xiang, Hongxin, et al.
Veröffentlicht: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024) -
Enhancing Temporal Modeling of Video LLMs via Time Gating
von: Hu, Zi-Yuan, et al.
Veröffentlicht: (2024) -
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
von: Zhong, Yiwu, et al.
Veröffentlicht: (2024) -
PAVE: Patching and Adapting Video Large Language Models
von: Liu, Zhuoming, et al.
Veröffentlicht: (2025) -
Fine-grained Spatiotemporal Grounding on Egocentric Videos
von: Liang, Shuo, et al.
Veröffentlicht: (2025)