Ola: Pushing the Frontiers of Omni-Modal Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zuyan, Dong, Yuhao, Wang, Jiahui, Liu, Ziwei, Hu, Winston, Lu, Jiwen, Rao, Yongming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
von: Yue, Xianghu, et al.
Veröffentlicht: (2024)
von: Yue, Xianghu, et al.
Veröffentlicht: (2024)
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
von: Ragano, Alessandro, et al.
Veröffentlicht: (2024)
von: Ragano, Alessandro, et al.
Veröffentlicht: (2024)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
von: Cao, Yuqin, et al.
Veröffentlicht: (2024)
von: Cao, Yuqin, et al.
Veröffentlicht: (2024)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
von: Mingote, Victoria, et al.
Veröffentlicht: (2024)
von: Mingote, Victoria, et al.
Veröffentlicht: (2024)
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
von: Premananth, Gowtham, et al.
Veröffentlicht: (2023)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2023)
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
von: Shimada, Kazuki, et al.
Veröffentlicht: (2025)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2025)
Two Web Toolkits for Multimodal Piano Performance Dataset Acquisition and Fingering Annotation
von: Park, Junhyung, et al.
Veröffentlicht: (2025)
von: Park, Junhyung, et al.
Veröffentlicht: (2025)
Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
von: Wen, Liuyuan
Veröffentlicht: (2024)
von: Wen, Liuyuan
Veröffentlicht: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
von: Sulun, Serkan, et al.
Veröffentlicht: (2025)
von: Sulun, Serkan, et al.
Veröffentlicht: (2025)
KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
von: Zhou, Huali, et al.
Veröffentlicht: (2024)
von: Zhou, Huali, et al.
Veröffentlicht: (2024)
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
von: Salaj, Ina, et al.
Veröffentlicht: (2025)
von: Salaj, Ina, et al.
Veröffentlicht: (2025)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
von: Li, Shufan, et al.
Veröffentlicht: (2024)
von: Li, Shufan, et al.
Veröffentlicht: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation
von: Chan, Nolan, et al.
Veröffentlicht: (2026)
von: Chan, Nolan, et al.
Veröffentlicht: (2026)
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
von: Gong, Jingyao
Veröffentlicht: (2026)
von: Gong, Jingyao
Veröffentlicht: (2026)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
von: Cai, Zhuojiang, et al.
Veröffentlicht: (2024)
von: Cai, Zhuojiang, et al.
Veröffentlicht: (2024)
A Unified Framework for Modality-Agnostic Deepfakes Detection
von: Yu, Cai, et al.
Veröffentlicht: (2023)
von: Yu, Cai, et al.
Veröffentlicht: (2023)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
von: Li, Sifei, et al.
Veröffentlicht: (2025)
von: Li, Sifei, et al.
Veröffentlicht: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
von: Park, Jay, et al.
Veröffentlicht: (2025)
von: Park, Jay, et al.
Veröffentlicht: (2025)
Improvement Of Audiovisual Quality Estimation Using A Nonlinear Autoregressive Exogenous Neural Network And Bitstream Parameters
von: Kossi, Koffi, et al.
Veröffentlicht: (2024)
von: Kossi, Koffi, et al.
Veröffentlicht: (2024)
Efficient Face Detection with Audio-Based Region Proposals for Human-Robot Interactions
von: Aris, William, et al.
Veröffentlicht: (2023)
von: Aris, William, et al.
Veröffentlicht: (2023)
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
Deep Learning for Steganalysis of Diverse Data Types: A review of methods, taxonomy, challenges and future directions
von: Kheddar, Hamza, et al.
Veröffentlicht: (2023)
von: Kheddar, Hamza, et al.
Veröffentlicht: (2023)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
von: Wu, Shu, et al.
Veröffentlicht: (2025)
von: Wu, Shu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
von: Yue, Xianghu, et al.
Veröffentlicht: (2024) -
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
von: Ragano, Alessandro, et al.
Veröffentlicht: (2024) -
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
von: Cao, Yuqin, et al.
Veröffentlicht: (2024) -
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
von: Mingote, Victoria, et al.
Veröffentlicht: (2024) -
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2024)