OmniGAIA: Towards Native Omni-Modal AI Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiaoxi, Jiao, Wenxiang, Jin, Jiarui, Wang, Shijian, Dong, Guanting, Jin, Jiajie, Wang, Hao, Wang, Yinuo, Wen, Ji-Rong, Lu, Yuan, Dou, Zhicheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents
von: Jin, Jiarui, et al.
Veröffentlicht: (2026)
von: Jin, Jiarui, et al.
Veröffentlicht: (2026)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
DeepAgent: A General Reasoning Agent with Scalable Toolsets
von: Li, Xiaoxi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2025)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
von: Chen, Yijing, et al.
Veröffentlicht: (2025)
von: Chen, Yijing, et al.
Veröffentlicht: (2025)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
von: Chen, Yuheng, et al.
Veröffentlicht: (2026)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
Ola: Pushing the Frontiers of Omni-Modal Language Model
von: Liu, Zuyan, et al.
Veröffentlicht: (2025)
von: Liu, Zuyan, et al.
Veröffentlicht: (2025)
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
GAIA: Zero-shot Talking Avatar Generation
von: He, Tianyu, et al.
Veröffentlicht: (2023)
von: He, Tianyu, et al.
Veröffentlicht: (2023)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
von: Liu, Che, et al.
Veröffentlicht: (2026)
von: Liu, Che, et al.
Veröffentlicht: (2026)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
von: Li, Shufan, et al.
Veröffentlicht: (2024)
von: Li, Shufan, et al.
Veröffentlicht: (2024)
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
LongCat-Flash-Omni Technical Report
von: Meituan LongCat Team, et al.
Veröffentlicht: (2025)
von: Meituan LongCat Team, et al.
Veröffentlicht: (2025)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
von: Lau, Kin Wai, et al.
Veröffentlicht: (2026)
von: Lau, Kin Wai, et al.
Veröffentlicht: (2026)
OmniSense: Towards Edge-Assisted Online Analytics for 360-Degree Videos
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
von: Hao, Jing, et al.
Veröffentlicht: (2025)
von: Hao, Jing, et al.
Veröffentlicht: (2025)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
Improving Gloss-free Sign Language Translation by Reducing Representation Density
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
von: Ye, Jinhui, et al.
Veröffentlicht: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
von: Liu, Jing, et al.
Veröffentlicht: (2023)
von: Liu, Jing, et al.
Veröffentlicht: (2023)
Explore the Limits of Omni-modal Pretraining at Scale
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents
von: Jin, Jiarui, et al.
Veröffentlicht: (2026) -
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024) -
DeepAgent: A General Reasoning Agent with Scalable Toolsets
von: Li, Xiaoxi, et al.
Veröffentlicht: (2025) -
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026) -
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)