Active Perception Agent for Omnimodal Audio-Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tao, Keda, Du, Wenjie, Yu, Bohan, Wang, Weiqiang, Liu, Jian, Wang, Huan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910012643737600
author Tao, Keda
Du, Wenjie
Yu, Bohan
Wang, Weiqiang
Liu, Jian
Wang, Huan
author_facet Tao, Keda
Du, Wenjie
Yu, Bohan
Wang, Weiqiang
Liu, Jian
Wang, Huan
contents Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address these limitations, we introduce OmniAgent, to our best knowledge, the first fully active perception agent that dynamically orchestrates specialized unimodal tools to achieve more fine-grained omnimodal reasoning. Unlike previous works that rely on rigid, static workflows and dense frame-captioning, we demonstrate a paradigm shift from passive response generation to active multimodal inquiry. OmniAgent employs dynamic planning to autonomously orchestrate tool invocation on demand, strategically concentrating perceptual attention on task-relevant cues. Central to our approach is a novel coarse-to-fine audio-guided perception paradigm, which leverages audio cues to localize temporal events and guide subsequent reasoning. Extensive empirical evaluations on three audio-video understanding benchmarks demonstrate that OmniAgent achieves state-of-the-art performance, surpassing leading open-source and closed-source models by substantial margins of 10% - 20% accuracy without training.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Active Perception Agent for Omnimodal Audio-Video Understanding
Tao, Keda
Du, Wenjie
Yu, Bohan
Wang, Weiqiang
Liu, Jian
Wang, Huan
Computer Vision and Pattern Recognition
Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address these limitations, we introduce OmniAgent, to our best knowledge, the first fully active perception agent that dynamically orchestrates specialized unimodal tools to achieve more fine-grained omnimodal reasoning. Unlike previous works that rely on rigid, static workflows and dense frame-captioning, we demonstrate a paradigm shift from passive response generation to active multimodal inquiry. OmniAgent employs dynamic planning to autonomously orchestrate tool invocation on demand, strategically concentrating perceptual attention on task-relevant cues. Central to our approach is a novel coarse-to-fine audio-guided perception paradigm, which leverages audio cues to localize temporal events and guide subsequent reasoning. Extensive empirical evaluations on three audio-video understanding benchmarks demonstrate that OmniAgent achieves state-of-the-art performance, surpassing leading open-source and closed-source models by substantial margins of 10% - 20% accuracy without training.
title Active Perception Agent for Omnimodal Audio-Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.23646