MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Bowen, Guo, Congchao, Yang, Geng, Yu, Hang, Zhang, Haozhe, Lei, Heidi, Mai, Jialong, Yan, Junjie, Yang, Kaiyue, Yang, Mingqi, Huang, Peikai, Jin, Ruiyang, Jiang, Sitan, Cheng, Weihua, Li, Yawei, Xiao, Yichen, Zhou, Yiying, Zhang, Yongmao, Lu, Yuan, He, Yucen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiniMax-01: Scaling Foundation Models with Lightning Attention
by: MiniMax, et al.
Published: (2025)
by: MiniMax, et al.
Published: (2025)
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
by: MiniMax, et al.
Published: (2026)
by: MiniMax, et al.
Published: (2026)
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
by: MiniMax, et al.
Published: (2025)
by: MiniMax, et al.
Published: (2025)
MiniMax Entropy Network: Learning Category-Invariant Features for Domain Adaptation
by: Tao, Chaofan, et al.
Published: (2019)
by: Tao, Chaofan, et al.
Published: (2019)
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
by: Zi, Bojia, et al.
Published: (2025)
by: Zi, Bojia, et al.
Published: (2025)
Mitigating Backdoor Attacks in Federated Learning Using PPA and MiniMax Game Theory
by: Wehbi, Osama, et al.
Published: (2026)
by: Wehbi, Osama, et al.
Published: (2026)
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
by: Yang, Xiaoyu, et al.
Published: (2024)
by: Yang, Xiaoyu, et al.
Published: (2024)
CUPE: Contextless Universal Phoneme Encoder for Language-Agnostic Speech Processing
by: Rehman, Abdul, et al.
Published: (2025)
by: Rehman, Abdul, et al.
Published: (2025)
FM3Q: Factorized Multi-Agent MiniMax Q-Learning for Two-Team Zero-Sum Markov Game
by: Hu, Guangzheng, et al.
Published: (2024)
by: Hu, Guangzheng, et al.
Published: (2024)
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
by: Guo, Dake, et al.
Published: (2024)
by: Guo, Dake, et al.
Published: (2024)
Speaker Anonymisation for Speech-based Suicide Risk Detection
by: Cui, Ziyun, et al.
Published: (2025)
by: Cui, Ziyun, et al.
Published: (2025)
Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
by: Wu, Suhang, et al.
Published: (2025)
by: Wu, Suhang, et al.
Published: (2025)
LMD: A Learnable Mask Network to Detect Adversarial Examples for Speaker Verification
by: Chen, Xing, et al.
Published: (2022)
by: Chen, Xing, et al.
Published: (2022)
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
by: Yang, Fei, et al.
Published: (2026)
by: Yang, Fei, et al.
Published: (2026)
PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen Speakers
by: Zhang, Xiangyue, et al.
Published: (2026)
by: Zhang, Xiangyue, et al.
Published: (2026)
USAT: A Universal Speaker-Adaptive Text-to-Speech Approach
by: Wang, Wenbin, et al.
Published: (2024)
by: Wang, Wenbin, et al.
Published: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
by: Zhang, Shaolei, et al.
Published: (2024)
by: Zhang, Shaolei, et al.
Published: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
by: Kim, Ji-Hoon, et al.
Published: (2024)
by: Kim, Ji-Hoon, et al.
Published: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
by: Peng, Yifan, et al.
Published: (2024)
by: Peng, Yifan, et al.
Published: (2024)
Optimizing Speech-Input Length for Speaker-Independent Depression Classification
by: Rutowski, Tomasz, et al.
Published: (2024)
by: Rutowski, Tomasz, et al.
Published: (2024)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
by: Fang, Qingkai, et al.
Published: (2024)
by: Fang, Qingkai, et al.
Published: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation
by: Fang, Fengyi, et al.
Published: (2025)
by: Fang, Fengyi, et al.
Published: (2025)
FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
by: Guo, Shoutao, et al.
Published: (2025)
by: Guo, Shoutao, et al.
Published: (2025)
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
by: Falai, Alessio, et al.
Published: (2025)
by: Falai, Alessio, et al.
Published: (2025)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
by: Shan, Weiqiao, et al.
Published: (2025)
by: Shan, Weiqiao, et al.
Published: (2025)
DiariST: Streaming Speech Translation with Speaker Diarization
by: Yang, Mu, et al.
Published: (2023)
by: Yang, Mu, et al.
Published: (2023)
SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation
by: Huang, Ziling, et al.
Published: (2025)
by: Huang, Ziling, et al.
Published: (2025)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
by: Jin, Jiawei, et al.
Published: (2025)
by: Jin, Jiawei, et al.
Published: (2025)
Iwasawa $λ$ invariant and Massey product
by: Qi, Peikai
Published: (2024)
by: Qi, Peikai
Published: (2024)
DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
by: Wang, Kyra, et al.
Published: (2024)
by: Wang, Kyra, et al.
Published: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
by: Ashihara, Takanori, et al.
Published: (2024)
by: Ashihara, Takanori, et al.
Published: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
by: Fang, Qingkai, et al.
Published: (2024)
by: Fang, Qingkai, et al.
Published: (2024)
Understanding the Nuances: The Quantity and Quality of Learner Feedback in Second‐Language Pronunciation
by: Congchao Hua
Published: (2026)
by: Congchao Hua
Published: (2026)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
by: Yang, Runyan, et al.
Published: (2024)
by: Yang, Runyan, et al.
Published: (2024)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
by: Yang, Yufeng, et al.
Published: (2025)
by: Yang, Yufeng, et al.
Published: (2025)
Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
Paralleling and Accelerating Arc Consistency Enforcement with Recurrent Tensor Computations
by: Yang, Mingqi
Published: (2024)
by: Yang, Mingqi
Published: (2024)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
by: Ma, Zhengrui, et al.
Published: (2024)
by: Ma, Zhengrui, et al.
Published: (2024)
Similar Items
-
MiniMax-01: Scaling Foundation Models with Lightning Attention
by: MiniMax, et al.
Published: (2025) -
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
by: MiniMax, et al.
Published: (2026) -
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
by: MiniMax, et al.
Published: (2025) -
MiniMax Entropy Network: Learning Category-Invariant Features for Domain Adaptation
by: Tao, Chaofan, et al.
Published: (2019) -
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
by: Zi, Bojia, et al.
Published: (2025)