GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
Fuente:
arXiv
Salvato in:
| Autori principali: | Watanabe, Chihiro, Kameoka, Hirokazu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024)
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024)
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2026)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2026)
Vocoder-Projected Feature Discriminator
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025)
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025)
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025)
Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
di: Raimon, Athul, et al.
Pubblicazione: (2024)
di: Raimon, Athul, et al.
Pubblicazione: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
di: Ravi, Nagarathna, et al.
Pubblicazione: (2024)
di: Ravi, Nagarathna, et al.
Pubblicazione: (2024)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
di: Cui, Jianwei, et al.
Pubblicazione: (2024)
di: Cui, Jianwei, et al.
Pubblicazione: (2024)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2025)
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2025)
Rethinking Mean Opinion Scores in Speech Quality Assessment: Aggregation through Quantized Distribution Fitting
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
Selecting N-lowest scores for training MOS prediction models
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
End-to-end Piano Performance-MIDI to Score Conversion with Transformers
di: Beyer, Tim, et al.
Pubblicazione: (2024)
di: Beyer, Tim, et al.
Pubblicazione: (2024)
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications
di: Wee, Marcus Yu Zhe, et al.
Pubblicazione: (2025)
di: Wee, Marcus Yu Zhe, et al.
Pubblicazione: (2025)
End-to-End Spoken Grammatical Error Correction
di: Qian, Mengjie, et al.
Pubblicazione: (2025)
di: Qian, Mengjie, et al.
Pubblicazione: (2025)
FunnelNet: An End-to-End Deep Learning Framework to Monitor Digital Heart Murmur in Real-Time
di: Jobayer, Md, et al.
Pubblicazione: (2024)
di: Jobayer, Md, et al.
Pubblicazione: (2024)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
di: Yu, Yifeng, et al.
Pubblicazione: (2024)
di: Yu, Yifeng, et al.
Pubblicazione: (2024)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
di: Dissen, Yehoshua, et al.
Pubblicazione: (2024)
di: Dissen, Yehoshua, et al.
Pubblicazione: (2024)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2020)
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2020)
End-to-End Multi-Task Learning for Adjustable Joint Noise Reduction and Hearing Loss Compensation
di: Gonzalez, Philippe, et al.
Pubblicazione: (2026)
di: Gonzalez, Philippe, et al.
Pubblicazione: (2026)
Content Adaptive Front End For Audio Classification
di: Verma, Prateek, et al.
Pubblicazione: (2023)
di: Verma, Prateek, et al.
Pubblicazione: (2023)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
di: Labrak, Yanis, et al.
Pubblicazione: (2024)
di: Labrak, Yanis, et al.
Pubblicazione: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
di: Ahmad, Hawraz A., et al.
Pubblicazione: (2024)
di: Ahmad, Hawraz A., et al.
Pubblicazione: (2024)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
di: Raissi, Tina, et al.
Pubblicazione: (2025)
di: Raissi, Tina, et al.
Pubblicazione: (2025)
GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting
di: Zhu, Pai, et al.
Pubblicazione: (2024)
di: Zhu, Pai, et al.
Pubblicazione: (2024)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
di: Chen, Jinming, et al.
Pubblicazione: (2024)
di: Chen, Jinming, et al.
Pubblicazione: (2024)
Music Genre Classification: Training an AI model
di: Mogonediwa, Keoikantse
Pubblicazione: (2024)
di: Mogonediwa, Keoikantse
Pubblicazione: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
di: Rufai, Amina Mardiyyah, et al.
Pubblicazione: (2020)
di: Rufai, Amina Mardiyyah, et al.
Pubblicazione: (2020)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
Automatic Identification of Samples in Hip-Hop Music via Multi-Loss Training and an Artificial Dataset
di: Cheston, Huw, et al.
Pubblicazione: (2025)
di: Cheston, Huw, et al.
Pubblicazione: (2025)
Patient-Aware Feature Alignment for Robust Lung Sound Classification:Cohesion-Separation and Global Alignment Losses
di: Jeong, Seung Gyu, et al.
Pubblicazione: (2025)
di: Jeong, Seung Gyu, et al.
Pubblicazione: (2025)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
di: Kushwaha, Saksham Singh, et al.
Pubblicazione: (2024)
di: Kushwaha, Saksham Singh, et al.
Pubblicazione: (2024)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2025)
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
di: Shakeel, Muhammad, et al.
Pubblicazione: (2024)
di: Shakeel, Muhammad, et al.
Pubblicazione: (2024)
Speaker Adaptation for Quantised End-to-End ASR Models
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
di: Palzer, David, et al.
Pubblicazione: (2025)
di: Palzer, David, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024) -
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2024) -
MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2026) -
Vocoder-Projected Feature Discriminator
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025) -
FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
di: Kaneko, Takuhiro, et al.
Pubblicazione: (2025)