Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
Fuente:
arXiv
Saved in:
| Main Authors: | Lau, Kin Wai, Rehman, Yasar Abbas Ur, Po, Lai-Man, de Gusmão, Pedro Porto Buarque |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
by: Lau, Kin Wai, et al.
Published: (2024)
by: Lau, Kin Wai, et al.
Published: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription
by: Hamberger, Anna, et al.
Published: (2025)
by: Hamberger, Anna, et al.
Published: (2025)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
by: Niu, Xinlei, et al.
Published: (2025)
by: Niu, Xinlei, et al.
Published: (2025)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
by: He, Jianfeng, et al.
Published: (2023)
by: He, Jianfeng, et al.
Published: (2023)
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
by: Wang, Junyu, et al.
Published: (2025)
by: Wang, Junyu, et al.
Published: (2025)
V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation
by: Chan, Nolan, et al.
Published: (2026)
by: Chan, Nolan, et al.
Published: (2026)
Dance-to-Music Generation with Encoder-based Textual Inversion
by: Li, Sifei, et al.
Published: (2024)
by: Li, Sifei, et al.
Published: (2024)
Benchmarking Cross-Domain Audio-Visual Deception Detection
by: Guo, Xiaobao, et al.
Published: (2024)
by: Guo, Xiaobao, et al.
Published: (2024)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
by: Gong, Jingyao
Published: (2026)
by: Gong, Jingyao
Published: (2026)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
by: Oh, Hyunwoo, et al.
Published: (2025)
by: Oh, Hyunwoo, et al.
Published: (2025)
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
by: Liu, Che, et al.
Published: (2025)
by: Liu, Che, et al.
Published: (2025)
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
by: Li, Junjie, et al.
Published: (2024)
by: Li, Junjie, et al.
Published: (2024)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024)
by: Cai, Zhuojiang, et al.
Published: (2024)
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
by: Huang, Lian, et al.
Published: (2024)
by: Huang, Lian, et al.
Published: (2024)
A Unified Framework for Modality-Agnostic Deepfakes Detection
by: Yu, Cai, et al.
Published: (2023)
by: Yu, Cai, et al.
Published: (2023)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
by: Wang, Le, et al.
Published: (2025)
by: Wang, Le, et al.
Published: (2025)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
by: Sudarsanam, Parthasaarathy, et al.
Published: (2025)
by: Sudarsanam, Parthasaarathy, et al.
Published: (2025)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
by: Li, Sifei, et al.
Published: (2025)
by: Li, Sifei, et al.
Published: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
by: Lin, Yueqian, et al.
Published: (2025)
by: Lin, Yueqian, et al.
Published: (2025)
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
by: Kim, Minyoung, et al.
Published: (2025)
by: Kim, Minyoung, et al.
Published: (2025)
Ola: Pushing the Frontiers of Omni-Modal Language Model
by: Liu, Zuyan, et al.
Published: (2025)
by: Liu, Zuyan, et al.
Published: (2025)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
by: Wang, Haoxu, et al.
Published: (2024)
by: Wang, Haoxu, et al.
Published: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
by: Su, Fei, et al.
Published: (2026)
by: Su, Fei, et al.
Published: (2026)
LaunchpadGPT: Language Model as Music Visualization Designer on Launchpad
by: Xu, Siting, et al.
Published: (2023)
by: Xu, Siting, et al.
Published: (2023)
AudioSetMix: Enhancing Audio-Language Datasets with LLM-Assisted Augmentations
by: Xu, David
Published: (2024)
by: Xu, David
Published: (2024)
MMSD-Net: Towards Multi-modal Stuttering Detection
by: Nie, Liangyu, et al.
Published: (2024)
by: Nie, Liangyu, et al.
Published: (2024)
Fine-Tuning MIDI-to-Audio Alignment using a Neural Network on Piano Roll and CQT Representations
by: Murgul, Sebastian, et al.
Published: (2025)
by: Murgul, Sebastian, et al.
Published: (2025)
Gender Representation in TV and Radio: Automatic Information Extraction methods versus Manual Analyses
by: Doukhan, David, et al.
Published: (2024)
by: Doukhan, David, et al.
Published: (2024)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
by: Wu, Shu, et al.
Published: (2025)
by: Wu, Shu, et al.
Published: (2025)
Double Mixture: Towards Continual Event Detection from Speech
by: Kang, Jingqi, et al.
Published: (2024)
by: Kang, Jingqi, et al.
Published: (2024)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
by: He, Jiajun, et al.
Published: (2024)
by: He, Jiajun, et al.
Published: (2024)
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
by: Guan, Yiwen, et al.
Published: (2025)
by: Guan, Yiwen, et al.
Published: (2025)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Similar Items
-
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
by: Lau, Kin Wai, et al.
Published: (2024) -
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024) -
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025) -
Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024) -
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024)