Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Ke, Wang, Yuhao, Cheng, Ziyang, Liu, Hongcheng, Wang, Yanfeng, Wang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
von: Li, Caorui, et al.
Veröffentlicht: (2025)
von: Li, Caorui, et al.
Veröffentlicht: (2025)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
von: Ou, Siqu, et al.
Veröffentlicht: (2025)
von: Ou, Siqu, et al.
Veröffentlicht: (2025)
M$^3$-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering
von: Xie, Peijin, et al.
Veröffentlicht: (2026)
von: Xie, Peijin, et al.
Veröffentlicht: (2026)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
von: Xie, Tianyu, et al.
Veröffentlicht: (2026)
von: Xie, Tianyu, et al.
Veröffentlicht: (2026)
PAR$^2$-RAG: Planned Active Retrieval and Reasoning for Multi-Hop Question Answering
von: Li, Xingyu, et al.
Veröffentlicht: (2026)
von: Li, Xingyu, et al.
Veröffentlicht: (2026)
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
von: Wan, Zhen, et al.
Veröffentlicht: (2026)
von: Wan, Zhen, et al.
Veröffentlicht: (2026)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Miner:Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models
von: Jiang, Shuyang, et al.
Veröffentlicht: (2026)
von: Jiang, Shuyang, et al.
Veröffentlicht: (2026)
Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought
von: Li, Xuanchen, et al.
Veröffentlicht: (2026)
von: Li, Xuanchen, et al.
Veröffentlicht: (2026)
KG-Reasoner: A Reinforced Model for End-to-End Multi-Hop Knowledge Graph Reasoning
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
von: Zhang, Guohui, et al.
Veröffentlicht: (2026)
von: Zhang, Guohui, et al.
Veröffentlicht: (2026)
CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
von: Shou, Zhou-Peng, et al.
Veröffentlicht: (2025)
von: Shou, Zhou-Peng, et al.
Veröffentlicht: (2025)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhangquan, et al.
Veröffentlicht: (2026)
Audio-Guided Dynamic Modality Fusion with Stereo-Aware Attention for Audio-Visual Navigation
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning
von: Han, Henry, et al.
Veröffentlicht: (2026)
von: Han, Henry, et al.
Veröffentlicht: (2026)
Enhancing Multi-Hop Knowledge Graph Reasoning through Reward Shaping Techniques
von: Li, Chen, et al.
Veröffentlicht: (2024)
von: Li, Chen, et al.
Veröffentlicht: (2024)
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
von: Wang, Shenzhi, et al.
Veröffentlicht: (2026)
von: Wang, Shenzhi, et al.
Veröffentlicht: (2026)
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework
von: Du, Yuexi, et al.
Veröffentlicht: (2026)
von: Du, Yuexi, et al.
Veröffentlicht: (2026)
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
von: Zeng, Yu, et al.
Veröffentlicht: (2025)
von: Zeng, Yu, et al.
Veröffentlicht: (2025)
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2026)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2026)
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench
von: Xu, Ke, et al.
Veröffentlicht: (2026)
von: Xu, Ke, et al.
Veröffentlicht: (2026)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
Agentic DAG-Orchestrated Planner Framework for Multi-Modal, Multi-Hop Question Answering in Hybrid Data Lakes
von: B, Kirushikesh D, et al.
Veröffentlicht: (2026)
von: B, Kirushikesh D, et al.
Veröffentlicht: (2026)
Agentic RAG with Knowledge Graphs for Complex Multi-Hop Reasoning in Real-World Applications
von: Lelong, Jean, et al.
Veröffentlicht: (2025)
von: Lelong, Jean, et al.
Veröffentlicht: (2025)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
Verifiable Process Rewards for Agentic Reasoning
von: Yuan, Huining, et al.
Veröffentlicht: (2026)
von: Yuan, Huining, et al.
Veröffentlicht: (2026)
OmniFuser: Adaptive Multimodal Fusion for Service-Oriented Predictive Maintenance
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
von: Wang, Ziqi, et al.
Veröffentlicht: (2025)
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation
von: Sun, Jiashuo, et al.
Veröffentlicht: (2026)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2026)
ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
von: Zhao, Xinjie, et al.
Veröffentlicht: (2025)
von: Zhao, Xinjie, et al.
Veröffentlicht: (2025)
OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery
von: Prabhakar, Vignesh, et al.
Veröffentlicht: (2025)
von: Prabhakar, Vignesh, et al.
Veröffentlicht: (2025)
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
von: Wang, Daoyu, et al.
Veröffentlicht: (2025)
von: Wang, Daoyu, et al.
Veröffentlicht: (2025)
AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025) -
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
von: Li, Caorui, et al.
Veröffentlicht: (2025) -
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
von: Wang, Yuhao, et al.
Veröffentlicht: (2025) -
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
von: Ou, Siqu, et al.
Veröffentlicht: (2025) -
M$^3$-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering
von: Xie, Peijin, et al.
Veröffentlicht: (2026)