Gespeichert in:
| Hauptverfasser: | Xie, Yuan, Chen, Tianshui, Ge, Zheng, Ni, Lionel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2508.20478 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning with Focused Thinking
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
von: Yang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Yang, Zhenyu, et al.
Veröffentlicht: (2025)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
von: Pan, Yaning, et al.
Veröffentlicht: (2025)
von: Pan, Yaning, et al.
Veröffentlicht: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2025)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
von: Su, Yuhao, et al.
Veröffentlicht: (2025)
von: Su, Yuhao, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
von: Pu, Bowei, et al.
Veröffentlicht: (2025)
von: Pu, Bowei, et al.
Veröffentlicht: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
Zero-Shot Long-Form Video Understanding through Screenplay
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
Hallucination Mitigation Prompts Long-term Video Understanding
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
von: Qi, Ji, et al.
Veröffentlicht: (2025)
von: Qi, Ji, et al.
Veröffentlicht: (2025)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025) -
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
von: Fu, Shenghao, et al.
Veröffentlicht: (2025) -
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2026) -
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025) -
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)