MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Qian, Chen, Yafeng, Chen, Yanni, Chen, Mengzhe, Chen, Yingda, Deng, Chong, Du, Zhihao, Gao, Ruize, Gao, Changfeng, Gao, Zhifu, Li, Yabin, Lv, Xiang, Liu, Jiaqing, Luo, Haoneng, Ma, Bin, Ni, Chongjia, Shi, Xian, Tang, Jialong, Wang, Hui, Wang, Hao, Wang, Wen, Wang, Yuxuan, Xu, Yunlan, Yu, Fan, Yan, Zhijie, Yang, Yexin, Yang, Baosong, Yang, Xian, Yang, Guanrou, Zhao, Tianyu, Zhang, Qinglin, Zhang, Shiliang, Zhao, Nan, Zhang, Pei, Zhang, Chong, Zhou, Jinren |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
por: Du, Zhihao, et al.
Publicado: (2025)
por: Du, Zhihao, et al.
Publicado: (2025)
CTC-Assisted LLM-Based Contextual ASR
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
por: Ma, Ziyang, et al.
Publicado: (2024)
por: Ma, Ziyang, et al.
Publicado: (2024)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
por: Du, Zhihao, et al.
Publicado: (2024)
por: Du, Zhihao, et al.
Publicado: (2024)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
por: An, Keyu, et al.
Publicado: (2024)
por: An, Keyu, et al.
Publicado: (2024)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
por: Gao, Changfeng, et al.
Publicado: (2025)
por: Gao, Changfeng, et al.
Publicado: (2025)
Differentiable Reward Optimization for LLM based TTS system
por: Gao, Changfeng, et al.
Publicado: (2025)
por: Gao, Changfeng, et al.
Publicado: (2025)
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
por: An, Keyu, et al.
Publicado: (2025)
por: An, Keyu, et al.
Publicado: (2025)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2023)
por: Chen, Yafeng, et al.
Publicado: (2023)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
por: An, Keyu, et al.
Publicado: (2024)
por: An, Keyu, et al.
Publicado: (2024)
Spatially covariant gravity with nonmetricity
por: Yu, Yang, et al.
Publicado: (2024)
por: Yu, Yang, et al.
Publicado: (2024)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
por: Yang, Guanrou, et al.
Publicado: (2025)
por: Yang, Guanrou, et al.
Publicado: (2025)
A Novel Adaptive Control Scheme for Accuracy‐Preassigned Finite‐Time Synchronization of Delayed Chaotic Neural Networks
por: Yantao Wang, et al.
Publicado: (2025)
por: Yantao Wang, et al.
Publicado: (2025)
Fun-ASR Technical Report
por: An, Keyu, et al.
Publicado: (2025)
por: An, Keyu, et al.
Publicado: (2025)
Impact of mass transfer on the orbital evolution of a white dwarf close to an intermediate-mass black hole
por: Yang, Yang, et al.
Publicado: (2025)
por: Yang, Yang, et al.
Publicado: (2025)
AFIRE: Accurate and Fast Image Reconstruction Algorithm for Geometric-inconsistent Multispectral CT
por: Gao, Yu, et al.
Publicado: (2025)
por: Gao, Yu, et al.
Publicado: (2025)
Convergence Analysis of Nonlinear Kaczmarz Method for Systems of Nonlinear Equations with Component-wise Convex Mapping
por: Gao, Yu, et al.
Publicado: (2023)
por: Gao, Yu, et al.
Publicado: (2023)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
por: Cheng, Luyao, et al.
Publicado: (2023)
por: Cheng, Luyao, et al.
Publicado: (2023)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
Interplay of Lyapunov exponents, phase transitions and chaos bound in nonlinear electrodynamics black hole
por: Gao, Chuanhong, et al.
Publicado: (2025)
por: Gao, Chuanhong, et al.
Publicado: (2025)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
por: Du, Zhihao, et al.
Publicado: (2024)
por: Du, Zhihao, et al.
Publicado: (2024)
Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
por: Sun, Weisong, et al.
Publicado: (2025)
por: Sun, Weisong, et al.
Publicado: (2025)
Layer by Layer Spraying Fabrication of Aggregation‐Induced Emission Metal‐Organic Frameworks Thin Film
por: Xue‐Xian Yang, et al.
Publicado: (2024)
por: Xue‐Xian Yang, et al.
Publicado: (2024)
Effects of particle overall regularity and surface roughness on fabric evolution of granular materials: DEM simulations
por: Jing Chen, et al.
Publicado: (2024)
por: Jing Chen, et al.
Publicado: (2024)
DEMO: A Statistical Perspective for Efficient Image-Text Matching
por: Zhang, Fan, et al.
Publicado: (2024)
por: Zhang, Fan, et al.
Publicado: (2024)
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
por: Zhang, Chong, et al.
Publicado: (2025)
por: Zhang, Chong, et al.
Publicado: (2025)
Enhancing low-temperature quantum thermometry via sequential measurements
por: Zhang, Ning, et al.
Publicado: (2024)
por: Zhang, Ning, et al.
Publicado: (2024)
Pelvic Radiotherapy in Rectal Cancer Patients With Synchronous Potentially Treatable Liver Metastases
por: Yayu Huang, et al.
Publicado: (2025)
por: Yayu Huang, et al.
Publicado: (2025)
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
Interacting hypersurfaces and multiple scalar-tensor theories
por: Yu, Yang, et al.
Publicado: (2024)
por: Yu, Yang, et al.
Publicado: (2024)
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
por: Song, Yakun, et al.
Publicado: (2024)
por: Song, Yakun, et al.
Publicado: (2024)
Meta-Reasoning: Semantics-Symbol Deconstruction for Large Language Models
por: Wang, Yiming, et al.
Publicado: (2023)
por: Wang, Yiming, et al.
Publicado: (2023)
Cross pathogenicity, host range and molecular characteristics of Fusarium oxysporum species complex populations isolated from tobacco in Jilin Province, China
por: Zhao Xie, et al.
Publicado: (2024)
por: Zhao Xie, et al.
Publicado: (2024)
Transient Receptor Potential Ankyrin 1 (TRPA1) Mediated LPS‐Induced Inflammation in Periodontal Ligament Stem Cells by Inhibiting the Phosphorylation of JNK
por: Xian Wang, et al.
Publicado: (2024)
por: Xian Wang, et al.
Publicado: (2024)
Dynamic stacking ensemble learning with investor knowledge representations for stock market index prediction based on multi-source financial data
por: Gao, Ruize, et al.
Publicado: (2025)
por: Gao, Ruize, et al.
Publicado: (2025)
Overexpression of Anthocyanidin Reductase Increases Flavonoids Content to Combat Fusarium Wilt in the Root Xylem of Vernicia montana
por: Jia Wang, et al.
Publicado: (2026)
por: Jia Wang, et al.
Publicado: (2026)
Multi-category Angle-based Classifier Refit
por: Yau, Guo Xian, et al.
Publicado: (2016)
por: Yau, Guo Xian, et al.
Publicado: (2016)
Ejemplares similares
-
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
por: Du, Zhihao, et al.
Publicado: (2025) -
CTC-Assisted LLM-Based Contextual ASR
por: Yang, Guanrou, et al.
Publicado: (2024) -
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
por: Yang, Guanrou, et al.
Publicado: (2024) -
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
por: Yang, Guanrou, et al.
Publicado: (2024) -
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
por: Ma, Ziyang, et al.
Publicado: (2024)