Super Monotonic Alignment Search
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Junhyeok, Kim, Hyeongju |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Length-Aware Rotary Position Embedding for Text-Speech Alignment
por: Kim, Hyeongju, et al.
Publicado: (2025)
por: Kim, Hyeongju, et al.
Publicado: (2025)
Training Flow Matching Models with Reliable Labels via Self-Purification
por: Kim, Hyeongju, et al.
Publicado: (2025)
por: Kim, Hyeongju, et al.
Publicado: (2025)
Improving Test-Time Performance of RVQ-based Neural Codecs
por: Kim, Hyeongju, et al.
Publicado: (2025)
por: Kim, Hyeongju, et al.
Publicado: (2025)
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
por: Neekhara, Paarth, et al.
Publicado: (2024)
por: Neekhara, Paarth, et al.
Publicado: (2024)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
por: Cho, Hyunjae, et al.
Publicado: (2024)
por: Cho, Hyunjae, et al.
Publicado: (2024)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
por: Yang, Jinhyeok, et al.
Publicado: (2024)
por: Yang, Jinhyeok, et al.
Publicado: (2024)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
por: Lee, Junhyeok, et al.
Publicado: (2025)
por: Lee, Junhyeok, et al.
Publicado: (2025)
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
por: Lee, Yongjoon, et al.
Publicado: (2024)
por: Lee, Yongjoon, et al.
Publicado: (2024)
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
por: Lee, Junhyeok, et al.
Publicado: (2026)
por: Lee, Junhyeok, et al.
Publicado: (2026)
Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting
por: Jung, Youngmoon, et al.
Publicado: (2025)
por: Jung, Youngmoon, et al.
Publicado: (2025)
MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
por: Kim, Semin, et al.
Publicado: (2024)
por: Kim, Semin, et al.
Publicado: (2024)
Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders
por: Kim, Seungbae, et al.
Publicado: (2025)
por: Kim, Seungbae, et al.
Publicado: (2025)
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
por: Oh, Hyung-Seok, et al.
Publicado: (2024)
por: Oh, Hyung-Seok, et al.
Publicado: (2024)
Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition
por: Kim, Byunggun, et al.
Publicado: (2024)
por: Kim, Byunggun, et al.
Publicado: (2024)
TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs
por: Peng, Jing, et al.
Publicado: (2026)
por: Peng, Jing, et al.
Publicado: (2026)
Spatiotemporal Emotional Synchrony in Dyadic Interactions: The Role of Speech Conditions in Facial and Vocal Affective Alignment
por: Herbuela, Von Ralph Dane Marquez, et al.
Publicado: (2025)
por: Herbuela, Von Ralph Dane Marquez, et al.
Publicado: (2025)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
por: Wang, Helin, et al.
Publicado: (2025)
por: Wang, Helin, et al.
Publicado: (2025)
Inference-time Scaling for Diffusion-based Audio Super-resolution
por: Jin, Yizhu, et al.
Publicado: (2025)
por: Jin, Yizhu, et al.
Publicado: (2025)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
por: Choi, Woongjib, et al.
Publicado: (2025)
por: Choi, Woongjib, et al.
Publicado: (2025)
CONMOD: Controllable Neural Frame-based Modulation Effects
por: Lee, Gyubin, et al.
Publicado: (2024)
por: Lee, Gyubin, et al.
Publicado: (2024)
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
por: Hong, Changi, et al.
Publicado: (2026)
por: Hong, Changi, et al.
Publicado: (2026)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
por: Lee, Seo-Hyun, et al.
Publicado: (2023)
por: Lee, Seo-Hyun, et al.
Publicado: (2023)
Adaptive Duration Model for Text Speech Alignment
por: Cao, Junjie
Publicado: (2025)
por: Cao, Junjie
Publicado: (2025)
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
por: Yun, Jun-Hak, et al.
Publicado: (2025)
por: Yun, Jun-Hak, et al.
Publicado: (2025)
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
por: Kim, Jaeyeon, et al.
Publicado: (2024)
por: Kim, Jaeyeon, et al.
Publicado: (2024)
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
por: Kang, Ju Yeon, et al.
Publicado: (2025)
por: Kang, Ju Yeon, et al.
Publicado: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
por: Kim, Nam-Gyu, et al.
Publicado: (2025)
por: Kim, Nam-Gyu, et al.
Publicado: (2025)
Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
por: Rhyu, Seungyeon, et al.
Publicado: (2024)
por: Rhyu, Seungyeon, et al.
Publicado: (2024)
LeVo: High-Quality Song Generation with Multi-Preference Alignment
por: Lei, Shun, et al.
Publicado: (2025)
por: Lei, Shun, et al.
Publicado: (2025)
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
por: Lee, Do Hyun, et al.
Publicado: (2024)
por: Lee, Do Hyun, et al.
Publicado: (2024)
Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
por: Jang, Jaehyuk, et al.
Publicado: (2026)
por: Jang, Jaehyuk, et al.
Publicado: (2026)
DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
por: Du, Zongcai, et al.
Publicado: (2025)
por: Du, Zongcai, et al.
Publicado: (2025)
Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition
por: Tzeng, Jing-Tong, et al.
Publicado: (2025)
por: Tzeng, Jing-Tong, et al.
Publicado: (2025)
MR-RawNet: Speaker verification system with multiple temporal resolutions for variable duration utterances using raw waveforms
por: Kim, Seung-bin, et al.
Publicado: (2024)
por: Kim, Seung-bin, et al.
Publicado: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
por: Kim, Jaeyeon, et al.
Publicado: (2024)
por: Kim, Jaeyeon, et al.
Publicado: (2024)
DroneAudioset: An Audio Dataset for Drone-based Search and Rescue
por: Gupta, Chitralekha, et al.
Publicado: (2025)
por: Gupta, Chitralekha, et al.
Publicado: (2025)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
por: Liu, Yisi, et al.
Publicado: (2025)
por: Liu, Yisi, et al.
Publicado: (2025)
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
por: Kim, Minsu, et al.
Publicado: (2025)
por: Kim, Minsu, et al.
Publicado: (2025)
Probing Cross-modal Information Hubs in Audio-Visual LLMs
por: Jung, Jihoo, et al.
Publicado: (2026)
por: Jung, Jihoo, et al.
Publicado: (2026)
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
por: Tamiti, Tarikul Islam, et al.
Publicado: (2025)
por: Tamiti, Tarikul Islam, et al.
Publicado: (2025)
Ejemplares similares
-
Length-Aware Rotary Position Embedding for Text-Speech Alignment
por: Kim, Hyeongju, et al.
Publicado: (2025) -
Training Flow Matching Models with Reliable Labels via Self-Purification
por: Kim, Hyeongju, et al.
Publicado: (2025) -
Improving Test-Time Performance of RVQ-based Neural Codecs
por: Kim, Hyeongju, et al.
Publicado: (2025) -
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
por: Neekhara, Paarth, et al.
Publicado: (2024) -
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
por: Cho, Hyunjae, et al.
Publicado: (2024)