Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tong, Jingqi, Mou, Yurong, Li, Hangcheng, Li, Mingzhe, Yang, Yongzhuo, Zhang, Ming, Chen, Qiguang, Liang, Tianyi, Hu, Xiaomeng, Zheng, Yining, Chen, Xinchi, Zhao, Jun, Huang, Xuanjing, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AI Can Learn Scientific Taste
von: Tong, Jingqi, et al.
Veröffentlicht: (2026)
von: Tong, Jingqi, et al.
Veröffentlicht: (2026)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
von: Zhao, Jun, et al.
Veröffentlicht: (2024)
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
von: Wang, Weiyi, et al.
Veröffentlicht: (2026)
von: Wang, Weiyi, et al.
Veröffentlicht: (2026)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
von: Li, Bowen, et al.
Veröffentlicht: (2026)
von: Li, Bowen, et al.
Veröffentlicht: (2026)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
von: Luo, Qi, et al.
Veröffentlicht: (2025)
von: Luo, Qi, et al.
Veröffentlicht: (2025)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
Multi-hop Reasoning via Early Knowledge Alignment
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering
von: Wang, Yuxin, et al.
Veröffentlicht: (2026)
von: Wang, Yuxin, et al.
Veröffentlicht: (2026)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval
von: Luo, Qi, et al.
Veröffentlicht: (2025)
von: Luo, Qi, et al.
Veröffentlicht: (2025)
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models
von: Liu, Peiju, et al.
Veröffentlicht: (2026)
von: Liu, Peiju, et al.
Veröffentlicht: (2026)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning with Focused Thinking
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
von: Wang, Yuxin, et al.
Veröffentlicht: (2025)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
Thinking with Spatial Code for Physical-World Video Reasoning
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
von: Zong, Yi, et al.
Veröffentlicht: (2024)
von: Zong, Yi, et al.
Veröffentlicht: (2024)
Code2Video: A Code-centric Paradigm for Educational Video Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
von: Liu, Xu, et al.
Veröffentlicht: (2026)
von: Liu, Xu, et al.
Veröffentlicht: (2026)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
Scaling Laws for Fact Memorization of Large Language Models
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
von: Sun, Yuhong, et al.
Veröffentlicht: (2025)
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
von: Wang, Lu, et al.
Veröffentlicht: (2026)
von: Wang, Lu, et al.
Veröffentlicht: (2026)
A New Geometric Representation for 3D Bijective Mappings and Applications
von: Chen, Qiguang, et al.
Veröffentlicht: (2023)
von: Chen, Qiguang, et al.
Veröffentlicht: (2023)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
von: Wu, Jianzong, et al.
Veröffentlicht: (2024)
Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
Text-Video Multi-Grained Integration for Video Moment Montage
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AI Can Learn Scientific Taste
von: Tong, Jingqi, et al.
Veröffentlicht: (2026) -
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
von: Zhao, Jun, et al.
Veröffentlicht: (2024) -
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
von: Wang, Weiyi, et al.
Veröffentlicht: (2026) -
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025) -
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
von: Li, Bowen, et al.
Veröffentlicht: (2026)