Agentic Very Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rege, Aniket, Sadhu, Arka, Li, Yuliang, Li, Kejie, Vinayak, Ramya Korlakai, Chai, Yuning, Lee, Yong Jae, Kim, Hyo Jin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
von: Rege, Aniket, et al.
Veröffentlicht: (2025)
von: Rege, Aniket, et al.
Veröffentlicht: (2025)
PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences
von: Chen, Daiwei, et al.
Veröffentlicht: (2024)
von: Chen, Daiwei, et al.
Veröffentlicht: (2024)
DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
von: Kumar, Raja, et al.
Veröffentlicht: (2025)
von: Kumar, Raja, et al.
Veröffentlicht: (2025)
GRASP: Graph Agentic Search over Propositions for Multi-hop Question Answering
von: Jenkins, Stockton, et al.
Veröffentlicht: (2026)
von: Jenkins, Stockton, et al.
Veröffentlicht: (2026)
An Efficient Streaming Video Understanding Framework with Agentic Control
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
von: Cai, Mu, et al.
Veröffentlicht: (2023)
von: Cai, Mu, et al.
Veröffentlicht: (2023)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
AffectSeek: Agentic Affective Understanding in Long Videos under Vague User Queries
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
von: Li, Keliang, et al.
Veröffentlicht: (2026)
von: Li, Keliang, et al.
Veröffentlicht: (2026)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
von: Xu, Weili, et al.
Veröffentlicht: (2025)
von: Xu, Weili, et al.
Veröffentlicht: (2025)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
von: Liu, Keliang, et al.
Veröffentlicht: (2025)
von: Liu, Keliang, et al.
Veröffentlicht: (2025)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
von: Durante, Zane, et al.
Veröffentlicht: (2026)
von: Durante, Zane, et al.
Veröffentlicht: (2026)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
MVLight: Relightable Text-to-3D Generation via Light-conditioned Multi-View Diffusion
von: Shim, Dongseok, et al.
Veröffentlicht: (2024)
von: Shim, Dongseok, et al.
Veröffentlicht: (2024)
Bridging Lifelong and Multi-Task Representation Learning via Algorithm and Complexity Measure
von: Wang, Zhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhi, et al.
Veröffentlicht: (2025)
Taming False Positives in Out-of-Distribution Detection with Human Feedback
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
von: Vishwakarma, Harit, et al.
Veröffentlicht: (2024)
Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
von: Yamada, Daisuke, et al.
Veröffentlicht: (2025)
von: Yamada, Daisuke, et al.
Veröffentlicht: (2025)
Metric Learning from Limited Pairwise Preference Comparisons
von: Wang, Zhi, et al.
Veröffentlicht: (2024)
von: Wang, Zhi, et al.
Veröffentlicht: (2024)
DrVideo: Document Retrieval Based Long Video Understanding
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
von: Yu, Xueqing, et al.
Veröffentlicht: (2026)
von: Yu, Xueqing, et al.
Veröffentlicht: (2026)
Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
von: Rege, Aniket, et al.
Veröffentlicht: (2025) -
PAL: Pluralistic Alignment Framework for Learning from Heterogeneous Preferences
von: Chen, Daiwei, et al.
Veröffentlicht: (2024) -
DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
von: Kumar, Raja, et al.
Veröffentlicht: (2025) -
GRASP: Graph Agentic Search over Propositions for Multi-hop Question Answering
von: Jenkins, Stockton, et al.
Veröffentlicht: (2026) -
An Efficient Streaming Video Understanding Framework with Agentic Control
von: Liu, Jinming, et al.
Veröffentlicht: (2026)