Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Hussain, Shehzeen, Neekhara, Paarth, Yang, Xuesong, Casanova, Edresson, Ghosh, Subhankar, Fejgin, Roy, Langman, Ryan, Desta, Mikyas, Tavabi, Leili, Li, Jason |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
di: Hussain, Shehzeen, et al.
Pubblicazione: (2025)
di: Hussain, Shehzeen, et al.
Pubblicazione: (2025)
Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
di: Fejgin, Roy, et al.
Pubblicazione: (2025)
di: Fejgin, Roy, et al.
Pubblicazione: (2025)
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
di: Langman, Ryan, et al.
Pubblicazione: (2025)
di: Langman, Ryan, et al.
Pubblicazione: (2025)
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
di: Casanova, Edresson, et al.
Pubblicazione: (2025)
di: Casanova, Edresson, et al.
Pubblicazione: (2025)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
di: Casanova, Edresson, et al.
Pubblicazione: (2024)
di: Casanova, Edresson, et al.
Pubblicazione: (2024)
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
di: Neekhara, Paarth, et al.
Pubblicazione: (2024)
di: Neekhara, Paarth, et al.
Pubblicazione: (2024)
REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
di: Zhang, Ruisi, et al.
Pubblicazione: (2023)
di: Zhang, Ruisi, et al.
Pubblicazione: (2023)
SelfVC: Voice Conversion With Iterative Refinement using Self Transformations
di: Neekhara, Paarth, et al.
Pubblicazione: (2023)
di: Neekhara, Paarth, et al.
Pubblicazione: (2023)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
di: Hu, Ke, et al.
Pubblicazione: (2025)
di: Hu, Ke, et al.
Pubblicazione: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
A guide to american film directors : the sound era : 1929 - 1979 / Larry Langman
di: Langman, Larry
Pubblicazione: (1929)
di: Langman, Larry
Pubblicazione: (1929)
The Transformation of Mohamed Atta: The Relevance of Personality in Radicalization
di: Peter Langman
Pubblicazione: (2025)
di: Peter Langman
Pubblicazione: (2025)
LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization
di: Chary, Luis Felipe, et al.
Pubblicazione: (2025)
di: Chary, Luis Felipe, et al.
Pubblicazione: (2025)
A CNN Based Framework for Unistroke Numeral Recognition in Air-Writing
di: Roy, Prasun, et al.
Pubblicazione: (2023)
di: Roy, Prasun, et al.
Pubblicazione: (2023)
Dampening Long-Period Doppler Shift Oscillations using Deep Machine Learning Techniques in the Solar Network and Internetwork
di: Sadeghi, Rayhaneh, et al.
Pubblicazione: (2024)
di: Sadeghi, Rayhaneh, et al.
Pubblicazione: (2024)
Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
di: Fejgin, Daniel, et al.
Pubblicazione: (2025)
di: Fejgin, Daniel, et al.
Pubblicazione: (2025)
IRIS observational approach to the oscillatory and damping nature of network and internetwork chromosphere small-scale brightening (SSBs) and their unusual dynamical and morphological differences in different regions on the solar disk
di: Sadeghi, Rayhane, et al.
Pubblicazione: (2024)
di: Sadeghi, Rayhane, et al.
Pubblicazione: (2024)
Coherence-Based Frequency Subset Selection For Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
di: Fejgin, Daniel, et al.
Pubblicazione: (2022)
di: Fejgin, Daniel, et al.
Pubblicazione: (2022)
Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting a Calibrated External Microphone Array
di: Fejgin, Daniel, et al.
Pubblicazione: (2022)
di: Fejgin, Daniel, et al.
Pubblicazione: (2022)
Exploring Damping Properties of IRIS Bright Points using Deep Learning Techniques
di: Tavabi, E., et al.
Pubblicazione: (2024)
di: Tavabi, E., et al.
Pubblicazione: (2024)
Characterizing Solar Spicules and their Role in Solar Wind Production using Machine Learning and the Hough Transform
di: Sadeghi, R., et al.
Pubblicazione: (2024)
di: Sadeghi, R., et al.
Pubblicazione: (2024)
Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs
di: Langman, Ryan, et al.
Pubblicazione: (2024)
di: Langman, Ryan, et al.
Pubblicazione: (2024)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR
di: Ogun, Sewade, et al.
Pubblicazione: (2025)
di: Ogun, Sewade, et al.
Pubblicazione: (2025)
A2TTS: TTS for Low Resource Indian Languages
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
di: Bhadoriya, Ayush Singh, et al.
Pubblicazione: (2025)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
di: Chou, Cheng-Kang, et al.
Pubblicazione: (2025)
STEFANN: Scene Text Editor using Font Adaptive Neural Network
di: Roy, Prasun, et al.
Pubblicazione: (2019)
di: Roy, Prasun, et al.
Pubblicazione: (2019)
Effects of Degradations on Deep Neural Network Architectures
di: Roy, Prasun, et al.
Pubblicazione: (2018)
di: Roy, Prasun, et al.
Pubblicazione: (2018)
Multi-scale Attention Guided Pose Transfer
di: Roy, Prasun, et al.
Pubblicazione: (2022)
di: Roy, Prasun, et al.
Pubblicazione: (2022)
BRUDEX Database: Binaural Room Impulse Responses with Uniformly Distributed External Microphones
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
Aligning Generative Music AI with Human Preferences: Methods and Challenges
di: Herremans, Dorien, et al.
Pubblicazione: (2025)
di: Herremans, Dorien, et al.
Pubblicazione: (2025)
Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS
di: Al-Sabbagh, Rania
Pubblicazione: (2026)
di: Al-Sabbagh, Rania
Pubblicazione: (2026)
Preference Alignment Improves Language Model-Based TTS
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
Pathway‐Based Similarity Measurement to Quantify Transcriptomics Similarity Between Human Tissues and Preclinical Models
di: Paarth Parekh, et al.
Pubblicazione: (2024)
di: Paarth Parekh, et al.
Pubblicazione: (2024)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
di: Zhang, Jiacheng, et al.
Pubblicazione: (2024)
di: Zhang, Jiacheng, et al.
Pubblicazione: (2024)
Unification and Texture Universality: The Essence of Hermiticity
di: Chakraborty, Pralay, et al.
Pubblicazione: (2025)
di: Chakraborty, Pralay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance
di: Hussain, Shehzeen, et al.
Pubblicazione: (2025) -
Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation
di: Fejgin, Roy, et al.
Pubblicazione: (2025) -
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
di: Langman, Ryan, et al.
Pubblicazione: (2025) -
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
di: Casanova, Edresson, et al.
Pubblicazione: (2025) -
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
di: Casanova, Edresson, et al.
Pubblicazione: (2024)