Active Perception Agent for Omnimodal Audio-Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Tao, Keda, Du, Wenjie, Yu, Bohan, Wang, Weiqiang, Liu, Jian, Wang, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
by: Tao, Keda, et al.
Published: (2026)
by: Tao, Keda, et al.
Published: (2026)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
by: Jin, Xinqi, et al.
Published: (2025)
by: Jin, Xinqi, et al.
Published: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
by: Chen, Xueyi, et al.
Published: (2025)
by: Chen, Xueyi, et al.
Published: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)
by: Wang, Xiaodong, et al.
Published: (2026)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
by: Ying, Kaining, et al.
Published: (2025)
by: Ying, Kaining, et al.
Published: (2025)
Is Oracle Pruning the True Oracle?
by: Feng, Sicheng, et al.
Published: (2024)
by: Feng, Sicheng, et al.
Published: (2024)
Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
HoliTom: Holistic Token Merging for Fast Video Large Language Models
by: Shao, Kele, et al.
Published: (2025)
by: Shao, Kele, et al.
Published: (2025)
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
by: Tao, Keda, et al.
Published: (2024)
by: Tao, Keda, et al.
Published: (2024)
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
PhotoArtAgent: Intelligent Photo Retouching with Language Model-Based Artist Agents
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
by: Li, Bingzhou, et al.
Published: (2026)
by: Li, Bingzhou, et al.
Published: (2026)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025)
by: Zhong, Hao, et al.
Published: (2025)
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
by: Li, Jiahua, et al.
Published: (2025)
by: Li, Jiahua, et al.
Published: (2025)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
by: Liang, Yongyuan, et al.
Published: (2025)
by: Liang, Yongyuan, et al.
Published: (2025)
EarlyTom: Early Token Compression Completes Fast Video Understanding
by: Wang, Hesong, et al.
Published: (2026)
by: Wang, Hesong, et al.
Published: (2026)
VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning
by: Tan, Hao, et al.
Published: (2026)
by: Tan, Hao, et al.
Published: (2026)
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
by: Yu, Xueqing, et al.
Published: (2026)
by: Yu, Xueqing, et al.
Published: (2026)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
by: Bai, Purui, et al.
Published: (2026)
by: Bai, Purui, et al.
Published: (2026)
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
by: Hong, Jack, et al.
Published: (2025)
by: Hong, Jack, et al.
Published: (2025)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Overcoming False Illusions in Real-World Face Restoration with Multi-Modal Guided Diffusion Model
by: Tao, Keda, et al.
Published: (2024)
by: Tao, Keda, et al.
Published: (2024)
3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis
by: Wang, Ziyue, et al.
Published: (2026)
by: Wang, Ziyue, et al.
Published: (2026)
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025)
by: Zhang, Yehang, et al.
Published: (2025)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
by: Ma, Junxian, et al.
Published: (2025)
by: Ma, Junxian, et al.
Published: (2025)
Understanding Attention Mechanism in Video Diffusion Models
by: Liu, Bingyan, et al.
Published: (2025)
by: Liu, Bingyan, et al.
Published: (2025)
Scaling Language-Centric Omnimodal Representation Learning
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
VABench: A Comprehensive Benchmark for Audio-Video Generation
by: Hua, Daili, et al.
Published: (2025)
by: Hua, Daili, et al.
Published: (2025)
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
by: Jin, Xin, et al.
Published: (2025)
by: Jin, Xin, et al.
Published: (2025)
Edge-guided and Class-balanced Active Learning for Semantic Segmentation of Aerial Images
by: Shan, Lianlei, et al.
Published: (2024)
by: Shan, Lianlei, et al.
Published: (2024)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
by: Xun, Shuhang, et al.
Published: (2025)
by: Xun, Shuhang, et al.
Published: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Similar Items
-
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
by: Tao, Keda, et al.
Published: (2025) -
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
by: Tao, Keda, et al.
Published: (2026) -
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
by: Jin, Xinqi, et al.
Published: (2025) -
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
by: Chen, Xueyi, et al.
Published: (2025) -
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)