Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Jiahua, Zhang, Zhanhe, Xu, Chenghao, Xu, Zhe, Wei, Kun, Yang, Xu, Deng, Cheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
por: Wang, Zheng, et al.
Publicado: (2026)
por: Wang, Zheng, et al.
Publicado: (2026)
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
por: Su, Zibo, et al.
Publicado: (2025)
por: Su, Zibo, et al.
Publicado: (2025)
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
por: Xu, Zhe, et al.
Publicado: (2024)
por: Xu, Zhe, et al.
Publicado: (2024)
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
por: Chen, Hanmo, et al.
Publicado: (2026)
por: Chen, Hanmo, et al.
Publicado: (2026)
Active Perception Agent for Omnimodal Audio-Video Understanding
por: Tao, Keda, et al.
Publicado: (2025)
por: Tao, Keda, et al.
Publicado: (2025)
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
por: Li, Wenjun, et al.
Publicado: (2024)
por: Li, Wenjun, et al.
Publicado: (2024)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
por: Zhang, Huaxin, et al.
Publicado: (2024)
por: Zhang, Huaxin, et al.
Publicado: (2024)
Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
por: Lyu, Guangtao, et al.
Publicado: (2025)
por: Lyu, Guangtao, et al.
Publicado: (2025)
Long-Video Audio Synthesis with Multi-Agent Collaboration
por: Zhang, Yehang, et al.
Publicado: (2025)
por: Zhang, Yehang, et al.
Publicado: (2025)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
por: Yang, Zhongyu, et al.
Publicado: (2026)
por: Yang, Zhongyu, et al.
Publicado: (2026)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
por: Yan, Haiyang, et al.
Publicado: (2026)
por: Yan, Haiyang, et al.
Publicado: (2026)
Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning
por: Chen, Hanmo, et al.
Publicado: (2026)
por: Chen, Hanmo, et al.
Publicado: (2026)
VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree
por: Li, Wenlong, et al.
Publicado: (2025)
por: Li, Wenlong, et al.
Publicado: (2025)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
por: Chen, Jiahua, et al.
Publicado: (2026)
por: Chen, Jiahua, et al.
Publicado: (2026)
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
por: Xu, Yichang, et al.
Publicado: (2026)
por: Xu, Yichang, et al.
Publicado: (2026)
AStF: Motion Style Transfer via Adaptive Statistics Fusor
por: Chen, Hanmo, et al.
Publicado: (2025)
por: Chen, Hanmo, et al.
Publicado: (2025)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
por: Ma, Martin Q., et al.
Publicado: (2026)
por: Ma, Martin Q., et al.
Publicado: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
por: Gao, Zhe, et al.
Publicado: (2026)
por: Gao, Zhe, et al.
Publicado: (2026)
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
por: Li, Zizhong, et al.
Publicado: (2026)
por: Li, Zizhong, et al.
Publicado: (2026)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
por: Wang, Ziyang, et al.
Publicado: (2025)
por: Wang, Ziyang, et al.
Publicado: (2025)
Generative Frame Sampler for Long Video Understanding
por: Yao, Linli, et al.
Publicado: (2025)
por: Yao, Linli, et al.
Publicado: (2025)
LiSD: An Efficient Multi-Task Learning Framework for LiDAR Segmentation and Detection
por: Xu, Jiahua, et al.
Publicado: (2024)
por: Xu, Jiahua, et al.
Publicado: (2024)
Towards Arbitrary Motion Completing via Hierarchical Continuous Representation
por: Xu, Chenghao, et al.
Publicado: (2025)
por: Xu, Chenghao, et al.
Publicado: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
por: Hu, Pengfei, et al.
Publicado: (2025)
por: Hu, Pengfei, et al.
Publicado: (2025)
A Turn Toward Better Alignment: Few-Shot Generative Adaptation with Equivariant Feature Rotation
por: Xu, Chenghao, et al.
Publicado: (2025)
por: Xu, Chenghao, et al.
Publicado: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
por: Xu, Jiaqi, et al.
Publicado: (2023)
por: Xu, Jiaqi, et al.
Publicado: (2023)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
por: Chen, Boyu, et al.
Publicado: (2025)
por: Chen, Boyu, et al.
Publicado: (2025)
SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
por: Xu, Zhiyu, et al.
Publicado: (2025)
por: Xu, Zhiyu, et al.
Publicado: (2025)
VCA: Video Curious Agent for Long Video Understanding
por: Yang, Zeyuan, et al.
Publicado: (2024)
por: Yang, Zeyuan, et al.
Publicado: (2024)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
por: Jin, Hongbo, et al.
Publicado: (2025)
por: Jin, Hongbo, et al.
Publicado: (2025)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
por: Zhi, Zhuo, et al.
Publicado: (2025)
por: Zhi, Zhuo, et al.
Publicado: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
por: Wang, Xiaodong, et al.
Publicado: (2026)
por: Wang, Xiaodong, et al.
Publicado: (2026)
Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing
por: Chen, Xi, et al.
Publicado: (2026)
por: Chen, Xi, et al.
Publicado: (2026)
EEA: Exploration-Exploitation Agent for Long Video Understanding
por: Yang, Te, et al.
Publicado: (2025)
por: Yang, Te, et al.
Publicado: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
por: Yin, Yufei, et al.
Publicado: (2026)
por: Yin, Yufei, et al.
Publicado: (2026)
HOI-M3:Capture Multiple Humans and Objects Interaction within Contextual Environment
por: Zhang, Juze, et al.
Publicado: (2024)
por: Zhang, Juze, et al.
Publicado: (2024)
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
por: Zhang, Yidan, et al.
Publicado: (2025)
por: Zhang, Yidan, et al.
Publicado: (2025)
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
por: Guo, Weiyu, et al.
Publicado: (2025)
por: Guo, Weiyu, et al.
Publicado: (2025)
Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering
por: Lyu, Guangtao, et al.
Publicado: (2026)
por: Lyu, Guangtao, et al.
Publicado: (2026)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
por: Fu, Shenghao, et al.
Publicado: (2025)
por: Fu, Shenghao, et al.
Publicado: (2025)
Ejemplares similares
-
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
por: Wang, Zheng, et al.
Publicado: (2026) -
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
por: Su, Zibo, et al.
Publicado: (2025) -
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
por: Xu, Zhe, et al.
Publicado: (2024) -
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
por: Chen, Hanmo, et al.
Publicado: (2026) -
Active Perception Agent for Omnimodal Audio-Video Understanding
por: Tao, Keda, et al.
Publicado: (2025)