MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Qian, Chen, Yafeng, Chen, Yanni, Chen, Mengzhe, Chen, Yingda, Deng, Chong, Du, Zhihao, Gao, Ruize, Gao, Changfeng, Gao, Zhifu, Li, Yabin, Lv, Xiang, Liu, Jiaqing, Luo, Haoneng, Ma, Bin, Ni, Chongjia, Shi, Xian, Tang, Jialong, Wang, Hui, Wang, Hao, Wang, Wen, Wang, Yuxuan, Xu, Yunlan, Yu, Fan, Yan, Zhijie, Yang, Yexin, Yang, Baosong, Yang, Xian, Yang, Guanrou, Zhao, Tianyu, Zhang, Qinglin, Zhang, Shiliang, Zhao, Nan, Zhang, Pei, Zhang, Chong, Zhou, Jinren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
di: Du, Zhihao, et al.
Pubblicazione: (2025)
di: Du, Zhihao, et al.
Pubblicazione: (2025)
CTC-Assisted LLM-Based Contextual ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
Differentiable Reward Optimization for LLM based TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
di: An, Keyu, et al.
Pubblicazione: (2025)
di: An, Keyu, et al.
Pubblicazione: (2025)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
di: Chen, Yafeng, et al.
Pubblicazione: (2023)
di: Chen, Yafeng, et al.
Pubblicazione: (2023)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
Spatially covariant gravity with nonmetricity
di: Yu, Yang, et al.
Pubblicazione: (2024)
di: Yu, Yang, et al.
Pubblicazione: (2024)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
A Novel Adaptive Control Scheme for Accuracy‐Preassigned Finite‐Time Synchronization of Delayed Chaotic Neural Networks
di: Yantao Wang, et al.
Pubblicazione: (2025)
di: Yantao Wang, et al.
Pubblicazione: (2025)
Fun-ASR Technical Report
di: An, Keyu, et al.
Pubblicazione: (2025)
di: An, Keyu, et al.
Pubblicazione: (2025)
Impact of mass transfer on the orbital evolution of a white dwarf close to an intermediate-mass black hole
di: Yang, Yang, et al.
Pubblicazione: (2025)
di: Yang, Yang, et al.
Pubblicazione: (2025)
AFIRE: Accurate and Fast Image Reconstruction Algorithm for Geometric-inconsistent Multispectral CT
di: Gao, Yu, et al.
Pubblicazione: (2025)
di: Gao, Yu, et al.
Pubblicazione: (2025)
Convergence Analysis of Nonlinear Kaczmarz Method for Systems of Nonlinear Equations with Component-wise Convex Mapping
di: Gao, Yu, et al.
Pubblicazione: (2023)
di: Gao, Yu, et al.
Pubblicazione: (2023)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
di: Cheng, Luyao, et al.
Pubblicazione: (2023)
di: Cheng, Luyao, et al.
Pubblicazione: (2023)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
Interplay of Lyapunov exponents, phase transitions and chaos bound in nonlinear electrodynamics black hole
di: Gao, Chuanhong, et al.
Pubblicazione: (2025)
di: Gao, Chuanhong, et al.
Pubblicazione: (2025)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
di: Sun, Weisong, et al.
Pubblicazione: (2025)
di: Sun, Weisong, et al.
Pubblicazione: (2025)
Layer by Layer Spraying Fabrication of Aggregation‐Induced Emission Metal‐Organic Frameworks Thin Film
di: Xue‐Xian Yang, et al.
Pubblicazione: (2024)
di: Xue‐Xian Yang, et al.
Pubblicazione: (2024)
Effects of particle overall regularity and surface roughness on fabric evolution of granular materials: DEM simulations
di: Jing Chen, et al.
Pubblicazione: (2024)
di: Jing Chen, et al.
Pubblicazione: (2024)
DEMO: A Statistical Perspective for Efficient Image-Text Matching
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
di: Zhang, Chong, et al.
Pubblicazione: (2025)
di: Zhang, Chong, et al.
Pubblicazione: (2025)
Enhancing low-temperature quantum thermometry via sequential measurements
di: Zhang, Ning, et al.
Pubblicazione: (2024)
di: Zhang, Ning, et al.
Pubblicazione: (2024)
Pelvic Radiotherapy in Rectal Cancer Patients With Synchronous Potentially Treatable Liver Metastases
di: Yayu Huang, et al.
Pubblicazione: (2025)
di: Yayu Huang, et al.
Pubblicazione: (2025)
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
Interacting hypersurfaces and multiple scalar-tensor theories
di: Yu, Yang, et al.
Pubblicazione: (2024)
di: Yu, Yang, et al.
Pubblicazione: (2024)
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
di: Song, Yakun, et al.
Pubblicazione: (2024)
di: Song, Yakun, et al.
Pubblicazione: (2024)
Meta-Reasoning: Semantics-Symbol Deconstruction for Large Language Models
di: Wang, Yiming, et al.
Pubblicazione: (2023)
di: Wang, Yiming, et al.
Pubblicazione: (2023)
Cross pathogenicity, host range and molecular characteristics of Fusarium oxysporum species complex populations isolated from tobacco in Jilin Province, China
di: Zhao Xie, et al.
Pubblicazione: (2024)
di: Zhao Xie, et al.
Pubblicazione: (2024)
Transient Receptor Potential Ankyrin 1 (TRPA1) Mediated LPS‐Induced Inflammation in Periodontal Ligament Stem Cells by Inhibiting the Phosphorylation of JNK
di: Xian Wang, et al.
Pubblicazione: (2024)
di: Xian Wang, et al.
Pubblicazione: (2024)
Dynamic stacking ensemble learning with investor knowledge representations for stock market index prediction based on multi-source financial data
di: Gao, Ruize, et al.
Pubblicazione: (2025)
di: Gao, Ruize, et al.
Pubblicazione: (2025)
Overexpression of Anthocyanidin Reductase Increases Flavonoids Content to Combat Fusarium Wilt in the Root Xylem of Vernicia montana
di: Jia Wang, et al.
Pubblicazione: (2026)
di: Jia Wang, et al.
Pubblicazione: (2026)
Multi-category Angle-based Classifier Refit
di: Yau, Guo Xian, et al.
Pubblicazione: (2016)
di: Yau, Guo Xian, et al.
Pubblicazione: (2016)
Documenti analoghi
-
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
di: Du, Zhihao, et al.
Pubblicazione: (2025) -
CTC-Assisted LLM-Based Contextual ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024) -
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024) -
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
di: Yang, Guanrou, et al.
Pubblicazione: (2024) -
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
di: Ma, Ziyang, et al.
Pubblicazione: (2024)