Training Articulatory Inversion Models for Interspeaker Consistency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | McGhee, Charles, Gales, Mark J. F., Knill, Kate M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
von: Loweimi, Erfan, et al.
Veröffentlicht: (2025)
von: Loweimi, Erfan, et al.
Veröffentlicht: (2025)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Data Augmentation for Spoken Grammatical Error Correction
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
RCT: Random Consistency Training for Semi-supervised Sound Event Detection
von: Shao, Nian, et al.
Veröffentlicht: (2021)
von: Shao, Nian, et al.
Veröffentlicht: (2021)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Towards End-to-End Spoken Grammatical Error Correction
von: Bannò, Stefano, et al.
Veröffentlicht: (2023)
von: Bannò, Stefano, et al.
Veröffentlicht: (2023)
Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
von: Manor, Hila, et al.
Veröffentlicht: (2024)
von: Manor, Hila, et al.
Veröffentlicht: (2024)
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2024)
von: Saito, Koichi, et al.
Veröffentlicht: (2024)
Post-Training Embedding Alignment for Decoupling Enrollment and Runtime Speaker Recognition Models
von: Gao, Chenyang, et al.
Veröffentlicht: (2024)
von: Gao, Chenyang, et al.
Veröffentlicht: (2024)
Watermarking Training Data of Music Generation Models
von: Epple, Pascal, et al.
Veröffentlicht: (2024)
von: Epple, Pascal, et al.
Veröffentlicht: (2024)
Score-Based Training for Energy-Based TTS Models
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
Deep Speech Synthesis from Multimodal Articulatory Representations
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Music2Latent: Consistency Autoencoders for Latent Audio Compression
von: Pasini, Marco, et al.
Veröffentlicht: (2024)
von: Pasini, Marco, et al.
Veröffentlicht: (2024)
Subtractive Training for Music Stem Insertion using Latent Diffusion Models
von: Villa-Renteria, Ivan, et al.
Veröffentlicht: (2024)
von: Villa-Renteria, Ivan, et al.
Veröffentlicht: (2024)
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2024)
von: de Oliveira, Danilo, et al.
Veröffentlicht: (2024)
SoundMorpher: Perceptually-Uniform Sound Morphing with Diffusion Model
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
von: Vora, Jayneel, et al.
Veröffentlicht: (2024)
von: Vora, Jayneel, et al.
Veröffentlicht: (2024)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
von: Ko, Myeongjin, et al.
Veröffentlicht: (2023)
von: Ko, Myeongjin, et al.
Veröffentlicht: (2023)
Model as Loss: A Self-Consistent Training Paradigm
von: Phaye, Saisamarth Rajesh, et al.
Veröffentlicht: (2025)
von: Phaye, Saisamarth Rajesh, et al.
Veröffentlicht: (2025)
Test-Time Training for Speech Enhancement
von: Behera, Avishkar, et al.
Veröffentlicht: (2025)
von: Behera, Avishkar, et al.
Veröffentlicht: (2025)
Test-Time Training for Depression Detection
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Evaluation of Neural Surrogates for Physical Modelling Synthesis of Nonlinear Elastic Plates
von: Martin, Carlos De La Vega, et al.
Veröffentlicht: (2025)
von: Martin, Carlos De La Vega, et al.
Veröffentlicht: (2025)
Music Genre Classification: Training an AI model
von: Mogonediwa, Keoikantse
Veröffentlicht: (2024)
von: Mogonediwa, Keoikantse
Veröffentlicht: (2024)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
Sound Tagging in Infant-centric Home Soundscapes
von: Khan, Mohammad Nur Hossain, et al.
Veröffentlicht: (2024)
von: Khan, Mohammad Nur Hossain, et al.
Veröffentlicht: (2024)
From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
CORN: Co-Trained Full- And No-Reference Speech Quality Assessment
von: Manocha, Pranay, et al.
Veröffentlicht: (2023)
von: Manocha, Pranay, et al.
Veröffentlicht: (2023)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
von: Janiczek, John, et al.
Veröffentlicht: (2024)
von: Janiczek, John, et al.
Veröffentlicht: (2024)
Computational music analysis from first principles
von: Tymoczko, Dmitri, et al.
Veröffentlicht: (2024)
von: Tymoczko, Dmitri, et al.
Veröffentlicht: (2024)
Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
von: Chung, Woo-Jin, et al.
Veröffentlicht: (2024)
von: Chung, Woo-Jin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024) -
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
von: Loweimi, Erfan, et al.
Veröffentlicht: (2025) -
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024) -
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025) -
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)