OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Yifan, Mu, Xinyu, Feng, Tao, Ou, Zhonghong, Gong, Yuning, Luo, Haoran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Active Perception Agent for Omnimodal Audio-Video Understanding
von: Tao, Keda, et al.
Veröffentlicht: (2025)
von: Tao, Keda, et al.
Veröffentlicht: (2025)
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search
von: Luo, Haoran, et al.
Veröffentlicht: (2025)
von: Luo, Haoran, et al.
Veröffentlicht: (2025)
LongAudio-RAG: Event-Grounded Question Answering over Multi-Hour Long Audio
von: Vakada, Naveen, et al.
Veröffentlicht: (2026)
von: Vakada, Naveen, et al.
Veröffentlicht: (2026)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
von: Tao, Keda, et al.
Veröffentlicht: (2026)
von: Tao, Keda, et al.
Veröffentlicht: (2026)
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
von: Tao, Keda, et al.
Veröffentlicht: (2025)
von: Tao, Keda, et al.
Veröffentlicht: (2025)
From RAG to Agentic RAG for Faithful Islamic Question Answering
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2026)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
von: Zhong, Hao, et al.
Veröffentlicht: (2025)
von: Zhong, Hao, et al.
Veröffentlicht: (2025)
MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
von: Gong, Ziyu, et al.
Veröffentlicht: (2025)
von: Gong, Ziyu, et al.
Veröffentlicht: (2025)
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
Agentic Keyframe Search for Video Question Answering
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
von: Fan, Lin, et al.
Veröffentlicht: (2024)
von: Fan, Lin, et al.
Veröffentlicht: (2024)
Dynin-Omni: Omnimodal Unified Large Diffusion Language Model
von: Kim, Jaeik, et al.
Veröffentlicht: (2026)
von: Kim, Jaeik, et al.
Veröffentlicht: (2026)
A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
von: Kurpath, Mohammed Irfan, et al.
Veröffentlicht: (2025)
von: Kurpath, Mohammed Irfan, et al.
Veröffentlicht: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models
von: Deng, Yuchen, et al.
Veröffentlicht: (2026)
von: Deng, Yuchen, et al.
Veröffentlicht: (2026)
GlobalRAG: Enhancing Global Reasoning in Multi-hop Question Answering via Reinforcement Learning
von: Luo, Jinchang, et al.
Veröffentlicht: (2025)
von: Luo, Jinchang, et al.
Veröffentlicht: (2025)
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
von: Li, Bingzhou, et al.
Veröffentlicht: (2026)
OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
Open-Ended Multi-Modal Relational Reasoning for Video Question Answering
von: Luo, Haozheng, et al.
Veröffentlicht: (2020)
von: Luo, Haozheng, et al.
Veröffentlicht: (2020)
An Entity Linking Agent for Question Answering
von: Luo, Yajie, et al.
Veröffentlicht: (2025)
von: Luo, Yajie, et al.
Veröffentlicht: (2025)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
von: Xu, Ke, et al.
Veröffentlicht: (2026)
von: Xu, Ke, et al.
Veröffentlicht: (2026)
PAR$^2$-RAG: Planned Active Retrieval and Reasoning for Multi-Hop Question Answering
von: Li, Xingyu, et al.
Veröffentlicht: (2026)
von: Li, Xingyu, et al.
Veröffentlicht: (2026)
Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering
von: Kobeissi, Amine, et al.
Veröffentlicht: (2026)
von: Kobeissi, Amine, et al.
Veröffentlicht: (2026)
Spatial Audio Question Answering and Reasoning on Dynamic Source Movements
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2026)
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2026)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
Efficient Multimodal Planning Agent for Visual Question-Answering
von: Chen, Zhuo, et al.
Veröffentlicht: (2026)
von: Chen, Zhuo, et al.
Veröffentlicht: (2026)
Structured RAG for Answering Aggregative Questions
von: Koshorek, Omri, et al.
Veröffentlicht: (2025)
von: Koshorek, Omri, et al.
Veröffentlicht: (2025)
OmniAudio: Generating Spatial Audio from 360-Degree Video
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
Extending Embodied Question Answering from Perception to Decision
von: Gong, Xicheng, et al.
Veröffentlicht: (2026)
von: Gong, Xicheng, et al.
Veröffentlicht: (2026)
GROUNDEDKG-RAG: Grounded Knowledge Graph Index for Long-document Question Answering
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
von: Li, Yangning, et al.
Veröffentlicht: (2025)
von: Li, Yangning, et al.
Veröffentlicht: (2025)
FoRAG: Factuality-optimized Retrieval Augmented Generation for Web-enhanced Long-form Question Answering
von: Cai, Tianchi, et al.
Veröffentlicht: (2024)
von: Cai, Tianchi, et al.
Veröffentlicht: (2024)
Narrative Aligned Long Form Video Question Answering
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
TSVC:Tripartite Learning with Semantic Variation Consistency for Robust Image-Text Retrieval
von: Lyu, Shuai, et al.
Veröffentlicht: (2025)
von: Lyu, Shuai, et al.
Veröffentlicht: (2025)
Traceable Cross-Source RAG for Chinese Tibetan Medicine Question Answering
von: Chen, Fengxian, et al.
Veröffentlicht: (2026)
von: Chen, Fengxian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Active Perception Agent for Omnimodal Audio-Video Understanding
von: Tao, Keda, et al.
Veröffentlicht: (2025) -
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026) -
KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search
von: Luo, Haoran, et al.
Veröffentlicht: (2025) -
LongAudio-RAG: Event-Grounded Question Answering over Multi-Hour Long Audio
von: Vakada, Naveen, et al.
Veröffentlicht: (2026) -
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)