Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Semin, Chung, Seungjun, Moon, Taehong, Lee, Sangheon, Ahn, Minyoung, Lee, Keon, Kim, Nam Soo, Cho, Jaewoong, Schmidt, Ludwig, Lee, Kangwook, Park, Dongmin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
by: Kim, Jaehyeon, et al.
Published: (2024)
by: Kim, Jaehyeon, et al.
Published: (2024)
Raon-Speech Technical Report
by: Kim, Beomsoo, et al.
Published: (2026)
by: Kim, Beomsoo, et al.
Published: (2026)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
by: Lee, Keon, et al.
Published: (2024)
by: Lee, Keon, et al.
Published: (2024)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
by: Kim, Jaehyeon, et al.
Published: (2024)
by: Kim, Jaehyeon, et al.
Published: (2024)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
by: Kim, Minchan, et al.
Published: (2024)
by: Kim, Minchan, et al.
Published: (2024)
Image Clustering Conditioned on Text Criteria
by: Kwon, Sehyun, et al.
Published: (2023)
by: Kwon, Sehyun, et al.
Published: (2023)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
by: Kim, Nam-Gyu, et al.
Published: (2025)
by: Kim, Nam-Gyu, et al.
Published: (2025)
A Simple Early Exiting Framework for Accelerated Sampling in Diffusion Models
by: Moon, Taehong, et al.
Published: (2024)
by: Moon, Taehong, et al.
Published: (2024)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
by: Park, Dongmin, et al.
Published: (2025)
by: Park, Dongmin, et al.
Published: (2025)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
by: Kim, Haechan, et al.
Published: (2026)
by: Kim, Haechan, et al.
Published: (2026)
How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects
by: Lee, Wonkwang, et al.
Published: (2025)
by: Lee, Wonkwang, et al.
Published: (2025)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026)
by: Park, Jaehyun, et al.
Published: (2026)
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
by: Lee, Joun Yeop, et al.
Published: (2024)
by: Lee, Joun Yeop, et al.
Published: (2024)
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
by: Kim, Minchan, et al.
Published: (2024)
by: Kim, Minchan, et al.
Published: (2024)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
by: Ahn, Hyebin, et al.
Published: (2025)
by: Ahn, Hyebin, et al.
Published: (2025)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
by: Cho, Deok-Hyeon, et al.
Published: (2024)
by: Cho, Deok-Hyeon, et al.
Published: (2024)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
by: Lee, Seokgi, et al.
Published: (2025)
by: Lee, Seokgi, et al.
Published: (2025)
MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
by: Kim, Semin, et al.
Published: (2024)
by: Kim, Semin, et al.
Published: (2024)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
by: Lee, Yoonhyung, et al.
Published: (2026)
by: Lee, Yoonhyung, et al.
Published: (2026)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
by: Cho, Deok-Hyeon, et al.
Published: (2025)
by: Cho, Deok-Hyeon, et al.
Published: (2025)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
by: Lee, Hyeonseung, et al.
Published: (2024)
by: Lee, Hyeonseung, et al.
Published: (2024)
OVS Meets Continual Learning: Towards Sustainable Open-Vocabulary Segmentation
by: Hwang, Dongjun, et al.
Published: (2024)
by: Hwang, Dongjun, et al.
Published: (2024)
Adaptive Tracking of a Single-Rigid-Body Character in Various Environments
by: Kwon, Taesoo, et al.
Published: (2023)
by: Kwon, Taesoo, et al.
Published: (2023)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
by: Kim, Hyeongju, et al.
Published: (2025)
by: Kim, Hyeongju, et al.
Published: (2025)
Poly(butylene terephthalate)/poly(ethylene glycol) blends with compatibilizers
by: Hayeong Lee, et al.
Published: (2024)
by: Hayeong Lee, et al.
Published: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
by: Ren, Yong, et al.
Published: (2026)
by: Ren, Yong, et al.
Published: (2026)
High‐Resolution Microlens‐Assisted Tunable n‐Type Optical Doping in Monolayer MoS 2
by: Junil Kim, et al.
Published: (2026)
by: Junil Kim, et al.
Published: (2026)
Le classement des réglementations du travail à l’épreuve de la diversité des salaires minima
by: Sangheon Lee
Published: (2012)
by: Sangheon Lee
Published: (2012)
Varieties of minimum wage system: through the dubious lens of indicator-based country rankings
by: Sangheon Lee
Published: (2012)
by: Sangheon Lee
Published: (2012)
Indicadores clasificatorios de normativas laborales: el caso del salario mínimo demuestra su ineficacia
by: Sangheon Lee
Published: (2012)
by: Sangheon Lee
Published: (2012)
Perceptual Cue Weighting Matters in Real‐Time Integration of Acoustic Information During Spoken Word Recognition
by: Hyoju Kim, et al.
Published: (2024)
by: Hyoju Kim, et al.
Published: (2024)
Ancient Korean Archive Translation: Comparison Analysis on Statistical phrase alignment, LLM in-context learning, and inter-methodological approach
by: Kim, Sojung Lucia, et al.
Published: (2024)
by: Kim, Sojung Lucia, et al.
Published: (2024)
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
by: Hong, Changi, et al.
Published: (2026)
by: Hong, Changi, et al.
Published: (2026)
TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model Agents
by: Jeong, Jongwon, et al.
Published: (2026)
by: Jeong, Jongwon, et al.
Published: (2026)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
by: Jang, Kangwook, et al.
Published: (2023)
by: Jang, Kangwook, et al.
Published: (2023)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
by: Park, Jongho, et al.
Published: (2024)
by: Park, Jongho, et al.
Published: (2024)
Similar Items
-
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
by: Kim, Jaehyeon, et al.
Published: (2024) -
Raon-Speech Technical Report
by: Kim, Beomsoo, et al.
Published: (2026) -
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
by: Lee, Keon, et al.
Published: (2024) -
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024) -
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
by: Kim, Jaehyeon, et al.
Published: (2024)