Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Hong, Bao, Yiming, Tu, Xuezhen, Xu, Yutong, Jin, Yue, Mu, Yiyang, Zhong, Bin, Yue, Linan, Zhang, Min-Ling |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
por: Gao, Hong, et al.
Publicado: (2025)
por: Gao, Hong, et al.
Publicado: (2025)
Training Multimodal Large Reasoning Models Needs Better Thoughts: A Three-Stage Framework for Long Chain-of-Thought Synthesis and Selection
por: Wang, Yizhi, et al.
Publicado: (2025)
por: Wang, Yizhi, et al.
Publicado: (2025)
An Efficient Streaming Video Understanding Framework with Agentic Control
por: Liu, Jinming, et al.
Publicado: (2026)
por: Liu, Jinming, et al.
Publicado: (2026)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
por: Zhang, Yue, et al.
Publicado: (2026)
por: Zhang, Yue, et al.
Publicado: (2026)
Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models
por: Wang, Yizhi, et al.
Publicado: (2026)
por: Wang, Yizhi, et al.
Publicado: (2026)
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
por: Zheng, Mingzhe, et al.
Publicado: (2026)
por: Zheng, Mingzhe, et al.
Publicado: (2026)
Agentic Very Long Video Understanding
por: Rege, Aniket, et al.
Publicado: (2026)
por: Rege, Aniket, et al.
Publicado: (2026)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
por: Sun, Yiming, et al.
Publicado: (2025)
por: Sun, Yiming, et al.
Publicado: (2025)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
por: Liu, Wenqi, et al.
Publicado: (2026)
por: Liu, Wenqi, et al.
Publicado: (2026)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
LensWalk: Agentic Video Understanding by Planning How You See in Videos
por: Li, Keliang, et al.
Publicado: (2026)
por: Li, Keliang, et al.
Publicado: (2026)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
por: Zhang, Xiaoyi, et al.
Publicado: (2025)
por: Zhang, Xiaoyi, et al.
Publicado: (2025)
Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding
por: Zhong, Yutong
Publicado: (2025)
por: Zhong, Yutong
Publicado: (2025)
LumiVideo: An Intelligent Agentic System for Video Color Grading
por: Guo, Yuchen, et al.
Publicado: (2026)
por: Guo, Yuchen, et al.
Publicado: (2026)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
por: Fan, Yue, et al.
Publicado: (2024)
por: Fan, Yue, et al.
Publicado: (2024)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
por: Zhao, Yiming, et al.
Publicado: (2026)
por: Zhao, Yiming, et al.
Publicado: (2026)
OwlSight: A Robust Illumination Adaptation Framework for Dark Video Human Action Recognition
por: Cheng, Shihao, et al.
Publicado: (2025)
por: Cheng, Shihao, et al.
Publicado: (2025)
A Unified Framework for Human-centric Point Cloud Video Understanding
por: Xu, Yiteng, et al.
Publicado: (2024)
por: Xu, Yiteng, et al.
Publicado: (2024)
ContextFlow: Training-Free Video Object Editing via Adaptive Context Enrichment
por: Chen, Yiyang, et al.
Publicado: (2025)
por: Chen, Yiyang, et al.
Publicado: (2025)
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding
por: Tu, Xuezhen, et al.
Publicado: (2026)
por: Tu, Xuezhen, et al.
Publicado: (2026)
Video2LoRA: Unified Semantic-Controlled Video Generation via Per-Reference-Video LoRA
por: Wu, Zexi, et al.
Publicado: (2026)
por: Wu, Zexi, et al.
Publicado: (2026)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
por: Wang, Jiapeng, et al.
Publicado: (2024)
por: Wang, Jiapeng, et al.
Publicado: (2024)
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
por: Liu, Penghui, et al.
Publicado: (2025)
por: Liu, Penghui, et al.
Publicado: (2025)
The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
por: Wu, Zhuoyuan, et al.
Publicado: (2025)
por: Wu, Zhuoyuan, et al.
Publicado: (2025)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
por: Liu, Dongyang, et al.
Publicado: (2025)
por: Liu, Dongyang, et al.
Publicado: (2025)
Hybrid 3D Human Pose Estimation with Monocular Video and Sparse IMUs
por: Bao, Yiming, et al.
Publicado: (2024)
por: Bao, Yiming, et al.
Publicado: (2024)
Apollo: An Exploration of Video Understanding in Large Multimodal Models
por: Zohar, Orr, et al.
Publicado: (2024)
por: Zohar, Orr, et al.
Publicado: (2024)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
por: Wang, Ziyang, et al.
Publicado: (2025)
por: Wang, Ziyang, et al.
Publicado: (2025)
VideoNSA: Native Sparse Attention Scales Video Understanding
por: Song, Enxin, et al.
Publicado: (2025)
por: Song, Enxin, et al.
Publicado: (2025)
Preacher: Paper-to-Video Agentic System
por: Liu, Jingwei, et al.
Publicado: (2025)
por: Liu, Jingwei, et al.
Publicado: (2025)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
por: Wang, Ziyi, et al.
Publicado: (2025)
por: Wang, Ziyi, et al.
Publicado: (2025)
VideoCoF: Unified Video Editing with Temporal Reasoner
por: Yang, Xiangpeng, et al.
Publicado: (2025)
por: Yang, Xiangpeng, et al.
Publicado: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
por: Yin, Yufei, et al.
Publicado: (2025)
por: Yin, Yufei, et al.
Publicado: (2025)
Code2MCP: Transforming Code Repositories into MCP Services
por: Ouyang, Chaoqian, et al.
Publicado: (2025)
por: Ouyang, Chaoqian, et al.
Publicado: (2025)
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
por: Zhang, Xingjian, et al.
Publicado: (2025)
por: Zhang, Xingjian, et al.
Publicado: (2025)
EEA: Exploration-Exploitation Agent for Long Video Understanding
por: Yang, Te, et al.
Publicado: (2025)
por: Yang, Te, et al.
Publicado: (2025)
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
por: Zhou, Yiyang, et al.
Publicado: (2025)
por: Zhou, Yiyang, et al.
Publicado: (2025)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
por: Li, Handong, et al.
Publicado: (2025)
por: Li, Handong, et al.
Publicado: (2025)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
por: Qiu, Chenhao, et al.
Publicado: (2026)
por: Qiu, Chenhao, et al.
Publicado: (2026)
DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos
por: Mu, Juncheng, et al.
Publicado: (2026)
por: Mu, Juncheng, et al.
Publicado: (2026)
Ejemplares similares
-
APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
por: Gao, Hong, et al.
Publicado: (2025) -
Training Multimodal Large Reasoning Models Needs Better Thoughts: A Three-Stage Framework for Long Chain-of-Thought Synthesis and Selection
por: Wang, Yizhi, et al.
Publicado: (2025) -
An Efficient Streaming Video Understanding Framework with Agentic Control
por: Liu, Jinming, et al.
Publicado: (2026) -
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
por: Zhang, Yue, et al.
Publicado: (2026) -
Bridging Efficiency and Transparency: Explainable CoT Compression in Multimodal Large Reasoning Models
por: Wang, Yizhi, et al.
Publicado: (2026)