Towards Multi-Modal Mastery: A 4.5B Parameter Truly Multi-Modal Small Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Koska, Ben, Horváth, Mojmír |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
by: Liu, Lei, et al.
Published: (2024)
by: Liu, Lei, et al.
Published: (2024)
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
by: Guan, Yiwen, et al.
Published: (2024)
by: Guan, Yiwen, et al.
Published: (2024)
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
by: Zhang, Haomin, et al.
Published: (2025)
by: Zhang, Haomin, et al.
Published: (2025)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
by: Chen, Xueyuan, et al.
Published: (2024)
by: Chen, Xueyuan, et al.
Published: (2024)
Exploring Multi-Modal Control in Music-Driven Dance Generation
by: Li, Ronghui, et al.
Published: (2024)
by: Li, Ronghui, et al.
Published: (2024)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
by: Kuan, Chun-Yi, et al.
Published: (2024)
by: Kuan, Chun-Yi, et al.
Published: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
by: Xue, Jinlong, et al.
Published: (2024)
by: Xue, Jinlong, et al.
Published: (2024)
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
by: Lu, Ke-Han, et al.
Published: (2025)
by: Lu, Ke-Han, et al.
Published: (2025)
Ola: Pushing the Frontiers of Omni-Modal Language Model
by: Liu, Zuyan, et al.
Published: (2025)
by: Liu, Zuyan, et al.
Published: (2025)
Closing the Modality Reasoning Gap for Speech Large Language Models
by: Wang, Chaoren, et al.
Published: (2026)
by: Wang, Chaoren, et al.
Published: (2026)
SSR: Alignment-Aware Modality Connector for Speech Language Models
by: Tan, Weiting, et al.
Published: (2024)
by: Tan, Weiting, et al.
Published: (2024)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
by: Wang, Xuechen, et al.
Published: (2024)
by: Wang, Xuechen, et al.
Published: (2024)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
by: Cheng, Yifan, et al.
Published: (2025)
by: Cheng, Yifan, et al.
Published: (2025)
Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
by: Jiang, Xilin, et al.
Published: (2025)
by: Jiang, Xilin, et al.
Published: (2025)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
by: Wang, Hankun, et al.
Published: (2024)
by: Wang, Hankun, et al.
Published: (2024)
Dual Audio-Centric Modality Coupling for Talking Head Generation
by: Fu, Ao, et al.
Published: (2025)
by: Fu, Ao, et al.
Published: (2025)
Self-Powered LLM Modality Expansion for Large Speech-Text Models
by: Yu, Tengfei, et al.
Published: (2024)
by: Yu, Tengfei, et al.
Published: (2024)
Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
by: Chen, Maximillian, et al.
Published: (2024)
by: Chen, Maximillian, et al.
Published: (2024)
Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
Segment Beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation
by: Wu, Renjie, et al.
Published: (2023)
by: Wu, Renjie, et al.
Published: (2023)
CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling
by: Yang, Ruihan, et al.
Published: (2023)
by: Yang, Ruihan, et al.
Published: (2023)
Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
by: Dao, Alan, et al.
Published: (2024)
by: Dao, Alan, et al.
Published: (2024)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
by: Zhong, Jinzuomu, et al.
Published: (2023)
by: Zhong, Jinzuomu, et al.
Published: (2023)
Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
by: Di Pierno, Andrea, et al.
Published: (2025)
by: Di Pierno, Andrea, et al.
Published: (2025)
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
by: Lemerle, Théodor, et al.
Published: (2024)
by: Lemerle, Théodor, et al.
Published: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2026)
by: Cappellazzo, Umberto, et al.
Published: (2026)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
by: Li, Yingting, et al.
Published: (2024)
by: Li, Yingting, et al.
Published: (2024)
Qwen2.5-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
Dynamic Modality and View Selection for Multimodal Emotion Recognition with Missing Modalities
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
by: Wang, Junyu, et al.
Published: (2025)
by: Wang, Junyu, et al.
Published: (2025)
Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
by: Lau, Kin Wai, et al.
Published: (2026)
by: Lau, Kin Wai, et al.
Published: (2026)
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Cross-Modal Denoising: A Novel Training Paradigm for Enhancing Speech-Image Retrieval
by: Zhou, Lifeng, et al.
Published: (2024)
by: Zhou, Lifeng, et al.
Published: (2024)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
by: Kolehmainen, Jari, et al.
Published: (2024)
by: Kolehmainen, Jari, et al.
Published: (2024)
Similar Items
-
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition
by: Liu, Lei, et al.
Published: (2024) -
Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
by: Guan, Yiwen, et al.
Published: (2024) -
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
by: Zhang, Haomin, et al.
Published: (2025) -
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
by: Chen, Xueyuan, et al.
Published: (2024) -
Exploring Multi-Modal Control in Music-Driven Dance Generation
by: Li, Ronghui, et al.
Published: (2024)