Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yuancheng, Zheng, Jiachen, Zhang, Junan, Zhang, Xueyao, Liao, Huan, Wu, Zhizheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Metric Preference Alignment for Generative Speech Restoration
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
Overview of the Amphion Toolkit (v0.2)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
von: Ni, Qinke, et al.
Veröffentlicht: (2026)
Aliasing-Free Neural Audio Synthesis
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization
von: Zhang, Ellie L., et al.
Veröffentlicht: (2025)
von: Zhang, Ellie L., et al.
Veröffentlicht: (2025)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Speech Enhancement Based on Drifting Models
von: Xu, Liang, et al.
Veröffentlicht: (2026)
von: Xu, Liang, et al.
Veröffentlicht: (2026)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
Solid State Bus-Comp: A Large-Scale and Diverse Dataset for Dynamic Range Compressor Virtual Analog Modeling
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
von: Bae, Hanbin, et al.
Veröffentlicht: (2024)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2025)
Constant Directivity Loudspeaker Beamforming
von: Luo, Yuancheng
Veröffentlicht: (2024)
von: Luo, Yuancheng
Veröffentlicht: (2024)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
Deep Active Speech Cancellation with Mamba-Masking Network
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
AI-Generated Music Detection in Broadcast Monitoring
von: López-Ayala, David, et al.
Veröffentlicht: (2026)
von: López-Ayala, David, et al.
Veröffentlicht: (2026)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
von: Kim, Minje, et al.
Veröffentlicht: (2024)
von: Kim, Minje, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Predicting Heart Activity from Speech using Data-driven and Knowledge-based features
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
von: Elbanna, Gasser, et al.
Veröffentlicht: (2024)
Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2026)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2026)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
von: Huang, Kuan-Tang, et al.
Veröffentlicht: (2026)
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
Leveraging Diverse Semantic-based Audio Pretrained Models for Singing Voice Conversion
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2023)
Recent Advances in Discrete Speech Tokens: A Review
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
von: Guo, Yiwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Metric Preference Alignment for Generative Speech Restoration
von: Zhang, Junan, et al.
Veröffentlicht: (2025) -
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
von: Wang, Yuancheng, et al.
Veröffentlicht: (2024) -
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025) -
Overview of the Amphion Toolkit (v0.2)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025) -
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)