MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baas, Matthew, Scholtz, Pieter, Mehta, Arnav, Dyson, Elliott, Prakash, Akshat, Kamper, Herman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Disentanglement in a GAN for Unconditional Speech Synthesis
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
von: Visser, Nicol, et al.
Veröffentlicht: (2025)
von: Visser, Nicol, et al.
Veröffentlicht: (2025)
Visually Grounded Speech Models have a Mutual Exclusivity Bias
von: Nortje, Leanne, et al.
Veröffentlicht: (2024)
von: Nortje, Leanne, et al.
Veröffentlicht: (2024)
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
von: van Rensburg, Kyle Janse, et al.
Veröffentlicht: (2026)
von: van Rensburg, Kyle Janse, et al.
Veröffentlicht: (2026)
Translating speech with just images
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
Visually grounded few-shot word learning in low-resource settings
von: Nortje, Leanne, et al.
Veröffentlicht: (2023)
von: Nortje, Leanne, et al.
Veröffentlicht: (2023)
Towards few-shot isolated word reading assessment
von: Smit, Reuben, et al.
Veröffentlicht: (2025)
von: Smit, Reuben, et al.
Veröffentlicht: (2025)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
von: Visser, Nicol, et al.
Veröffentlicht: (2026)
von: Visser, Nicol, et al.
Veröffentlicht: (2026)
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
The mutual exclusivity bias of bilingual visually grounded speech models
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
von: Oneata, Dan, et al.
Veröffentlicht: (2025)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
von: Malan, Simon, et al.
Veröffentlicht: (2024)
von: Malan, Simon, et al.
Veröffentlicht: (2024)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
von: Malan, Simon, et al.
Veröffentlicht: (2025)
von: Malan, Simon, et al.
Veröffentlicht: (2025)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
von: Jacobs, Christiaan, et al.
Veröffentlicht: (2025)
von: Jacobs, Christiaan, et al.
Veröffentlicht: (2025)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
von: Kamper, Herman, et al.
Veröffentlicht: (2025)
von: Kamper, Herman, et al.
Veröffentlicht: (2025)
Spoken-Term Discovery using Discrete Speech Units
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
Speech Codec Probing from Semantic and Phonetic Perspectives
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
von: Abdullah, Badr M., et al.
Veröffentlicht: (2025)
von: Abdullah, Badr M., et al.
Veröffentlicht: (2025)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
von: Louw, Retief, et al.
Veröffentlicht: (2025)
von: Louw, Retief, et al.
Veröffentlicht: (2025)
Feature-based analysis of oral narratives from Afrikaans and isiXhosa children
von: Sharratt, Emma, et al.
Veröffentlicht: (2025)
von: Sharratt, Emma, et al.
Veröffentlicht: (2025)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
Probing the Robustness Properties of Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
A Neural Speech Codec for Noise Robust Speech Coding
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
von: Carbonneau, Marc-André, et al.
Veröffentlicht: (2025)
von: Carbonneau, Marc-André, et al.
Veröffentlicht: (2025)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
von: Yang, Leyan, et al.
Veröffentlicht: (2026)
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
von: Song, Yakun, et al.
Veröffentlicht: (2024)
von: Song, Yakun, et al.
Veröffentlicht: (2024)
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
von: Wu, Haibin, et al.
Veröffentlicht: (2025)
von: Wu, Haibin, et al.
Veröffentlicht: (2025)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
von: Wang, Guansu, et al.
Veröffentlicht: (2025)
von: Wang, Guansu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Disentanglement in a GAN for Unconditional Speech Synthesis
von: Baas, Matthew, et al.
Veröffentlicht: (2023) -
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
von: Visser, Nicol, et al.
Veröffentlicht: (2025) -
Visually Grounded Speech Models have a Mutual Exclusivity Bias
von: Nortje, Leanne, et al.
Veröffentlicht: (2024) -
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
von: van Rensburg, Kyle Janse, et al.
Veröffentlicht: (2026) -
Translating speech with just images
von: Oneata, Dan, et al.
Veröffentlicht: (2024)