VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yan, Zeng, Yawen, Zheng, Jingsheng, Xing, Xiaofen, Xu, Jin, Xu, Xiangmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
von: Wang, Yanan, et al.
Veröffentlicht: (2025)
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
von: Lee, Daeun, et al.
Veröffentlicht: (2025)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
von: Han, Songhao, et al.
Veröffentlicht: (2024)
von: Han, Songhao, et al.
Veröffentlicht: (2024)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
Rethinking Chain-of-Thought Reasoning for Videos
von: Zhong, Yiwu, et al.
Veröffentlicht: (2025)
von: Zhong, Yiwu, et al.
Veröffentlicht: (2025)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
von: Qi, Ji, et al.
Veröffentlicht: (2023)
von: Qi, Ji, et al.
Veröffentlicht: (2023)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
von: Gao, Timin, et al.
Veröffentlicht: (2024)
von: Gao, Timin, et al.
Veröffentlicht: (2024)
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
von: Dai, Xuanlang, et al.
Veröffentlicht: (2026)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
von: Wang, Yunnan, et al.
Veröffentlicht: (2026)
von: Wang, Yunnan, et al.
Veröffentlicht: (2026)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
von: Chen, Qiguang, et al.
Veröffentlicht: (2024)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
von: Liang, Jiafeng, et al.
Veröffentlicht: (2024)
From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2025)
von: Wang, Xiangfeng, et al.
Veröffentlicht: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
Towards Patronizing and Condescending Language in Chinese Videos: A Multimodal Dataset and Detector
von: Wang, Hongbo, et al.
Veröffentlicht: (2024)
von: Wang, Hongbo, et al.
Veröffentlicht: (2024)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
Compact Model Training by Low-Rank Projection with Energy Transfer
von: Guo, Kailing, et al.
Veröffentlicht: (2022)
von: Guo, Kailing, et al.
Veröffentlicht: (2022)
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
von: Xu, Yichen, et al.
Veröffentlicht: (2026)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
von: Zheng, Duo, et al.
Veröffentlicht: (2024)
von: Zheng, Duo, et al.
Veröffentlicht: (2024)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
von: Lin, Jingyang, et al.
Veröffentlicht: (2026)
Video-guided Machine Translation with Global Video Context
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
von: Cheng, Ming, et al.
Veröffentlicht: (2025)
von: Cheng, Ming, et al.
Veröffentlicht: (2025)
Reinforcing Structured Chain-of-Thought for Video Understanding
von: Wang, Peiyao, et al.
Veröffentlicht: (2026)
von: Wang, Peiyao, et al.
Veröffentlicht: (2026)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
von: Qi, Ji, et al.
Veröffentlicht: (2025)
von: Qi, Ji, et al.
Veröffentlicht: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
von: Du, Yang, et al.
Veröffentlicht: (2025)
von: Du, Yang, et al.
Veröffentlicht: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
von: Wang, Yanan, et al.
Veröffentlicht: (2025) -
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
von: Lee, Daeun, et al.
Veröffentlicht: (2025) -
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
von: Han, Songhao, et al.
Veröffentlicht: (2024) -
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025) -
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)