Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Tao, ding, yiming, Chai, Shenghua, Zhang, Minghui, Luo, Zhongtian, Wang, Xinming, Chen, Xinlong, Kang, Zhaolu, Gong, Junhao, Zhou, Yuxuan, Jin, Haopeng, Cui, Zhiqing, Yang, Jiabing, Zhang, YiFan, Yi, Hongzhu, He, Zheqi, Yang, Xi, Huang, Yan, Wang, Liang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
von: Yu, Tao, et al.
Veröffentlicht: (2026)
von: Yu, Tao, et al.
Veröffentlicht: (2026)
ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
von: Yu, Tao, et al.
Veröffentlicht: (2026)
von: Yu, Tao, et al.
Veröffentlicht: (2026)
PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
von: Yu, Tao, et al.
Veröffentlicht: (2026)
von: Yu, Tao, et al.
Veröffentlicht: (2026)
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
von: Yang, Yujia, et al.
Veröffentlicht: (2026)
von: Yang, Yujia, et al.
Veröffentlicht: (2026)
Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization
von: Yu, Tao, et al.
Veröffentlicht: (2026)
von: Yu, Tao, et al.
Veröffentlicht: (2026)
LightSearcher: Efficient DeepSearch via Experiential Memory
von: Lan, Hengzhi, et al.
Veröffentlicht: (2025)
von: Lan, Hengzhi, et al.
Veröffentlicht: (2025)
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
von: Wu, Fang, et al.
Veröffentlicht: (2025)
von: Wu, Fang, et al.
Veröffentlicht: (2025)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
von: Yu, Tao, et al.
Veröffentlicht: (2025)
von: Yu, Tao, et al.
Veröffentlicht: (2025)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
von: Ding, Yue, et al.
Veröffentlicht: (2026)
von: Ding, Yue, et al.
Veröffentlicht: (2026)
OmniBench: Towards The Future of Universal Omni-Language Models
von: Li, Yizhi, et al.
Veröffentlicht: (2024)
von: Li, Yizhi, et al.
Veröffentlicht: (2024)
OmniSearchSage: Multi-Task Multi-Entity Embeddings for Pinterest Search
von: Agarwal, Prabhat, et al.
Veröffentlicht: (2024)
von: Agarwal, Prabhat, et al.
Veröffentlicht: (2024)
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
von: Zeng, Weixuan, et al.
Veröffentlicht: (2026)
von: Zeng, Weixuan, et al.
Veröffentlicht: (2026)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
von: Li, Caorui, et al.
Veröffentlicht: (2025)
von: Li, Caorui, et al.
Veröffentlicht: (2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
OmniRe: Omni Urban Scene Reconstruction
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Deep Omni-supervised Learning for Rib Fracture Detection from Chest Radiology Images
von: Chai, Zhizhong, et al.
Veröffentlicht: (2023)
von: Chai, Zhizhong, et al.
Veröffentlicht: (2023)
To Search or Not to Search: Aligning the Decision Boundary of Deep Search Agents via Causal Intervention
von: Zhang, Wenlin, et al.
Veröffentlicht: (2026)
von: Zhang, Wenlin, et al.
Veröffentlicht: (2026)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
Qwen3-Omni Technical Report
von: Xu, Jin, et al.
Veröffentlicht: (2025)
von: Xu, Jin, et al.
Veröffentlicht: (2025)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
von: Wang, Kun, et al.
Veröffentlicht: (2026)
von: Wang, Kun, et al.
Veröffentlicht: (2026)
OmniGAIA: Towards Native Omni-Modal AI Agents
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
von: Liu, Jing, et al.
Veröffentlicht: (2025)
von: Liu, Jing, et al.
Veröffentlicht: (2025)
OmniStyle2: Learning to Stylize by Learning to Destylize
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization
von: Han, Minghao, et al.
Veröffentlicht: (2026)
von: Han, Minghao, et al.
Veröffentlicht: (2026)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
von: Xie, Tianyu, et al.
Veröffentlicht: (2026)
von: Xie, Tianyu, et al.
Veröffentlicht: (2026)
Cross-Domain Deep Code Search with Meta Learning
von: Chai, Yitian, et al.
Veröffentlicht: (2022)
von: Chai, Yitian, et al.
Veröffentlicht: (2022)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
OmniDrones: An Efficient and Flexible Platform for Reinforcement Learning in Drone Control
von: Xu, Botian, et al.
Veröffentlicht: (2023)
von: Xu, Botian, et al.
Veröffentlicht: (2023)
Logics-Parsing-Omni Technical Report
von: An, Xin, et al.
Veröffentlicht: (2026)
von: An, Xin, et al.
Veröffentlicht: (2026)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
von: Yu, Tao, et al.
Veröffentlicht: (2026) -
ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
von: Yu, Tao, et al.
Veröffentlicht: (2026) -
PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
von: Yu, Tao, et al.
Veröffentlicht: (2026) -
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
von: Yang, Yujia, et al.
Veröffentlicht: (2026) -
Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization
von: Yu, Tao, et al.
Veröffentlicht: (2026)