Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jiahua, Zhang, Zhanhe, Xu, Chenghao, Xu, Zhe, Wei, Kun, Yang, Xu, Deng, Cheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
von: Su, Zibo, et al.
Veröffentlicht: (2025)
von: Su, Zibo, et al.
Veröffentlicht: (2025)
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
Active Perception Agent for Omnimodal Audio-Video Understanding
von: Tao, Keda, et al.
Veröffentlicht: (2025)
von: Tao, Keda, et al.
Veröffentlicht: (2025)
Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
von: Zhang, Huaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Huaxin, et al.
Veröffentlicht: (2024)
Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
von: Lyu, Guangtao, et al.
Veröffentlicht: (2025)
von: Lyu, Guangtao, et al.
Veröffentlicht: (2025)
Long-Video Audio Synthesis with Multi-Agent Collaboration
von: Zhang, Yehang, et al.
Veröffentlicht: (2025)
von: Zhang, Yehang, et al.
Veröffentlicht: (2025)
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
von: Yang, Zhongyu, et al.
Veröffentlicht: (2026)
von: Yang, Zhongyu, et al.
Veröffentlicht: (2026)
Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding
von: Yan, Haiyang, et al.
Veröffentlicht: (2026)
von: Yan, Haiyang, et al.
Veröffentlicht: (2026)
Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
von: Chen, Hanmo, et al.
Veröffentlicht: (2026)
VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree
von: Li, Wenlong, et al.
Veröffentlicht: (2025)
von: Li, Wenlong, et al.
Veröffentlicht: (2025)
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
von: Xu, Yichang, et al.
Veröffentlicht: (2026)
von: Xu, Yichang, et al.
Veröffentlicht: (2026)
AStF: Motion Style Transfer via Adaptive Statistics Fusor
von: Chen, Hanmo, et al.
Veröffentlicht: (2025)
von: Chen, Hanmo, et al.
Veröffentlicht: (2025)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
LiSD: An Efficient Multi-Task Learning Framework for LiDAR Segmentation and Detection
von: Xu, Jiahua, et al.
Veröffentlicht: (2024)
von: Xu, Jiahua, et al.
Veröffentlicht: (2024)
Towards Arbitrary Motion Completing via Hierarchical Continuous Representation
von: Xu, Chenghao, et al.
Veröffentlicht: (2025)
von: Xu, Chenghao, et al.
Veröffentlicht: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
A Turn Toward Better Alignment: Few-Shot Generative Adaptation with Equivariant Feature Rotation
von: Xu, Chenghao, et al.
Veröffentlicht: (2025)
von: Xu, Chenghao, et al.
Veröffentlicht: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
von: Xu, Zhiyu, et al.
Veröffentlicht: (2025)
von: Xu, Zhiyu, et al.
Veröffentlicht: (2025)
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
von: Zhi, Zhuo, et al.
Veröffentlicht: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
von: Wang, Xiaodong, et al.
Veröffentlicht: (2026)
von: Wang, Xiaodong, et al.
Veröffentlicht: (2026)
Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
EEA: Exploration-Exploitation Agent for Long Video Understanding
von: Yang, Te, et al.
Veröffentlicht: (2025)
von: Yang, Te, et al.
Veröffentlicht: (2025)
HOI-M3:Capture Multiple Humans and Objects Interaction within Contextual Environment
von: Zhang, Juze, et al.
Veröffentlicht: (2024)
von: Zhang, Juze, et al.
Veröffentlicht: (2024)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
von: Zhang, Yidan, et al.
Veröffentlicht: (2025)
von: Zhang, Yidan, et al.
Veröffentlicht: (2025)
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
von: Guo, Weiyu, et al.
Veröffentlicht: (2025)
von: Guo, Weiyu, et al.
Veröffentlicht: (2025)
Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
von: Lyu, Guangtao, et al.
Veröffentlicht: (2026)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
von: Wang, Zheng, et al.
Veröffentlicht: (2026) -
A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
von: Su, Zibo, et al.
Veröffentlicht: (2025) -
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons
von: Xu, Zhe, et al.
Veröffentlicht: (2024) -
Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
von: Chen, Hanmo, et al.
Veröffentlicht: (2026) -
Active Perception Agent for Omnimodal Audio-Video Understanding
von: Tao, Keda, et al.
Veröffentlicht: (2025)