TidyVoice: A Curated Multilingual Dataset for Speaker Verification Derived from Common Voice
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Farhadipour, Aref, Marquenie, Jan, Madikeri, Srikanth, Chodroff, Eleanor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TidyVoice 2026 Challenge Evaluation Plan
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
von: Li, Ze, et al.
Veröffentlicht: (2026)
von: Li, Ze, et al.
Veröffentlicht: (2026)
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
Adaptive Multimodal Person Recognition: A Robust Framework for Handling Missing Modalities
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
Residual Speaker Representation for One-Shot Voice Conversion
von: Xu, Le, et al.
Veröffentlicht: (2023)
von: Xu, Le, et al.
Veröffentlicht: (2023)
Descriptor:: Extended-Length Audio Dataset for Synthetic Voice Detection and Speaker Recognition (ELAD-SVDSR)
von: Vijaykumar, Rahul, et al.
Veröffentlicht: (2025)
von: Vijaykumar, Rahul, et al.
Veröffentlicht: (2025)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
von: Tomashenko, Natalia, et al.
Veröffentlicht: (2024)
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
von: Lee, Jaejun, et al.
Veröffentlicht: (2025)
von: Lee, Jaejun, et al.
Veröffentlicht: (2025)
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yanzhe, et al.
Veröffentlicht: (2024)
Voices of Civilizations: A Multilingual QA Benchmark for Global Music Understanding
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
von: Wu, Shangda, et al.
Veröffentlicht: (2026)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
von: Yu, Fan, et al.
Veröffentlicht: (2025)
von: Yu, Fan, et al.
Veröffentlicht: (2025)
Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
von: Ferreira, Alef Iury Siqueira, et al.
Veröffentlicht: (2025)
Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
ReverbFX: A Dataset of Room Impulse Responses Derived from Reverb Effect Plugins for Singing Voice Dereverberation
von: Richter, Julius, et al.
Veröffentlicht: (2025)
von: Richter, Julius, et al.
Veröffentlicht: (2025)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
NeuralMultiling: A Novel Neural Architecture Search for Smartphone based Multilingual Speaker Verification
von: PN, Aravinda Reddy, et al.
Veröffentlicht: (2024)
von: PN, Aravinda Reddy, et al.
Veröffentlicht: (2024)
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
von: Xu, Haiying, et al.
Veröffentlicht: (2025)
AdvSV: An Over-the-Air Adversarial Attack Dataset for Speaker Verification
von: Wang, Li, et al.
Veröffentlicht: (2023)
von: Wang, Li, et al.
Veröffentlicht: (2023)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
von: Yeom, Jiheum, et al.
Veröffentlicht: (2024)
The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TidyVoice 2026 Challenge Evaluation Plan
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026) -
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
von: Zhang, Miao, et al.
Veröffentlicht: (2025) -
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
von: Li, Ze, et al.
Veröffentlicht: (2026) -
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026) -
CL-UZH submission to the NIST SRE 2024 Speaker Recognition Evaluation
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)