Salvato in:
| Autori principali: | Uddin, Majbah, Huynh, Nathan, Vidal, Jose M, Taaffe, Kevin M, Fredendall, Lawrence D, Greenstein, Joel S |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.03369 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications
di: de Groot, Dimme, et al.
Pubblicazione: (2025)
di: de Groot, Dimme, et al.
Pubblicazione: (2025)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
di: Togootogtokh, Enkhtogtokh, et al.
Pubblicazione: (2025)
di: Togootogtokh, Enkhtogtokh, et al.
Pubblicazione: (2025)
TidyVoice 2026 Challenge Evaluation Plan
di: Farhadipour, Aref, et al.
Pubblicazione: (2026)
di: Farhadipour, Aref, et al.
Pubblicazione: (2026)
Descriptor:: Extended-Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD-SVDSR)
di: Vijaykumar, Rahul, et al.
Pubblicazione: (2025)
di: Vijaykumar, Rahul, et al.
Pubblicazione: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
di: Lee, Jaejun, et al.
Pubblicazione: (2025)
Development of the Listening in Spatialized Noise-Sentences (LiSN-S) Test in Brazilian Portuguese: Presentation Software, Speech Stimuli, and Sentence Equivalence
di: Masiero, Bruno S., et al.
Pubblicazione: (2024)
di: Masiero, Bruno S., et al.
Pubblicazione: (2024)
PERSONA: An Application for Emotion Recognition, Gender Recognition and Age Estimation
di: Koshal, Devyani, et al.
Pubblicazione: (2024)
di: Koshal, Devyani, et al.
Pubblicazione: (2024)
An Extensive Analysis of the Singing Voice Conversion Challenge 2025 Evaluation Results
di: Violeta, Lester Phillip, et al.
Pubblicazione: (2025)
di: Violeta, Lester Phillip, et al.
Pubblicazione: (2025)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
Automatic Voice Classification Of Autistic Subjects
di: Vacca, Jessica, et al.
Pubblicazione: (2024)
di: Vacca, Jessica, et al.
Pubblicazione: (2024)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
di: Xu, Haiying, et al.
Pubblicazione: (2025)
di: Xu, Haiying, et al.
Pubblicazione: (2025)
EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences
di: Li, Baifeng, et al.
Pubblicazione: (2023)
di: Li, Baifeng, et al.
Pubblicazione: (2023)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
Human Voice is Unique
di: Singh, Rita, et al.
Pubblicazione: (2025)
di: Singh, Rita, et al.
Pubblicazione: (2025)
AdaProj: Adaptively Scaled Angular Margin Subspace Projections for Anomalous Sound Detection with Auxiliary Classification Tasks
di: Wilkinghoff, Kevin
Pubblicazione: (2024)
di: Wilkinghoff, Kevin
Pubblicazione: (2024)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
di: Seki, Kentaro, et al.
Pubblicazione: (2024)
di: Seki, Kentaro, et al.
Pubblicazione: (2024)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2025)
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2025)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2020)
di: Kameoka, Hirokazu, et al.
Pubblicazione: (2020)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
A Transversal Study of Fundamental Frequency Contours in Parkinsonian Voices
di: Rodriguez-Perez, Pablo, et al.
Pubblicazione: (2024)
di: Rodriguez-Perez, Pablo, et al.
Pubblicazione: (2024)
SingIt! Singer Voice Transformation
di: Eliav, Amit, et al.
Pubblicazione: (2024)
di: Eliav, Amit, et al.
Pubblicazione: (2024)
Controlling your Attributes in Voice
di: Li, Xuyuan, et al.
Pubblicazione: (2025)
di: Li, Xuyuan, et al.
Pubblicazione: (2025)
Objective Measurements of Voice Quality
di: Dhamyal, Hira, et al.
Pubblicazione: (2024)
di: Dhamyal, Hira, et al.
Pubblicazione: (2024)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2026)
di: Wang, Zhichao, et al.
Pubblicazione: (2026)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
di: Du, Zongyang, et al.
Pubblicazione: (2024)
di: Du, Zongyang, et al.
Pubblicazione: (2024)
TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
di: Farhadipour, Aref, et al.
Pubblicazione: (2026)
di: Farhadipour, Aref, et al.
Pubblicazione: (2026)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2024)
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2024)
Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
di: Ren, Zhao, et al.
Pubblicazione: (2025)
di: Ren, Zhao, et al.
Pubblicazione: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
di: Kim, Heeseung, et al.
Pubblicazione: (2025)
di: Kim, Heeseung, et al.
Pubblicazione: (2025)
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
di: Bhogale, Kaushal, et al.
Pubblicazione: (2026)
di: Bhogale, Kaushal, et al.
Pubblicazione: (2026)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
di: Sha, Binzhu, et al.
Pubblicazione: (2023)
di: Sha, Binzhu, et al.
Pubblicazione: (2023)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
di: Lin, Yueqian, et al.
Pubblicazione: (2025)
di: Lin, Yueqian, et al.
Pubblicazione: (2025)
Revolutionizing Personalized Voice Synthesis: The Journey towards Emotional and Individual Authenticity with DIVSE (Dynamic Individual Voice Synthesis Engine)
di: Shi, Fan
Pubblicazione: (2023)
di: Shi, Fan
Pubblicazione: (2023)
Generating Novel and Realistic Speakers for Voice Conversion
di: Chen, Meiying Melissa, et al.
Pubblicazione: (2025)
di: Chen, Meiying Melissa, et al.
Pubblicazione: (2025)
Robust Singing Voice Transcription Serves Synthesis
di: Li, Ruiqi, et al.
Pubblicazione: (2024)
di: Li, Ruiqi, et al.
Pubblicazione: (2024)
Implementation and Applications of WakeWords Integrated with Speaker Recognition: A Case Study
di: Filho, Alexandre Costa Ferro, et al.
Pubblicazione: (2024)
di: Filho, Alexandre Costa Ferro, et al.
Pubblicazione: (2024)
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
di: Ren, Zhao, et al.
Pubblicazione: (2025)
di: Ren, Zhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications
di: de Groot, Dimme, et al.
Pubblicazione: (2025) -
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
di: Togootogtokh, Enkhtogtokh, et al.
Pubblicazione: (2025) -
TidyVoice 2026 Challenge Evaluation Plan
di: Farhadipour, Aref, et al.
Pubblicazione: (2026) -
Descriptor:: Extended-Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD-SVDSR)
di: Vijaykumar, Rahul, et al.
Pubblicazione: (2025) -
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
di: Zheng, Zhisheng, et al.
Pubblicazione: (2025)