LibriVAD: A Scalable Open Dataset with Deep Learning Benchmarks for Voice Activity Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Stylianou, Ioannis, Sarkar, Achintya kr., Dawalatabad, Nauman, Glass, James, Tan, Zheng-Hua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Vocal Tract Length Warped Features for Spoken Keyword Spotting
di: Sarkar, Achintya kr., et al.
Pubblicazione: (2025)
di: Sarkar, Achintya kr., et al.
Pubblicazione: (2025)
SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization
di: Wang, Chien-Chun, et al.
Pubblicazione: (2025)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2025)
One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
di: Stylianou, Ioannis, et al.
Pubblicazione: (2026)
di: Stylianou, Ioannis, et al.
Pubblicazione: (2026)
Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
di: Wang, Liming, et al.
Pubblicazione: (2024)
di: Wang, Liming, et al.
Pubblicazione: (2024)
sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
di: Yang, Qu, et al.
Pubblicazione: (2024)
di: Yang, Qu, et al.
Pubblicazione: (2024)
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
di: Landau, Gilad, et al.
Pubblicazione: (2025)
di: Landau, Gilad, et al.
Pubblicazione: (2025)
LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
di: Ohmura, Junki, et al.
Pubblicazione: (2025)
di: Ohmura, Junki, et al.
Pubblicazione: (2025)
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
di: Bovbjerg, Holger Severin, et al.
Pubblicazione: (2023)
di: Bovbjerg, Holger Severin, et al.
Pubblicazione: (2023)
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
di: Bovbjerg, Holger Severin, et al.
Pubblicazione: (2025)
di: Bovbjerg, Holger Severin, et al.
Pubblicazione: (2025)
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
di: Liu, Yun, et al.
Pubblicazione: (2024)
di: Liu, Yun, et al.
Pubblicazione: (2024)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
di: Li, Zongyi, et al.
Pubblicazione: (2025)
di: Li, Zongyi, et al.
Pubblicazione: (2025)
Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
di: Wu, Weijie, et al.
Pubblicazione: (2025)
di: Wu, Weijie, et al.
Pubblicazione: (2025)
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
VoiceWukong: Benchmarking Deepfake Voice Detection
di: Yan, Ziwei, et al.
Pubblicazione: (2024)
di: Yan, Ziwei, et al.
Pubblicazione: (2024)
Voices of the Mountains: Deep Learning-Based Vocal Error Detection System for Kurdish Maqams
di: Khairaldeen, Darvan Shvan, et al.
Pubblicazione: (2026)
di: Khairaldeen, Darvan Shvan, et al.
Pubblicazione: (2026)
The First Voice Timbre Attribute Detection Challenge
di: Chen, Liping, et al.
Pubblicazione: (2025)
di: Chen, Liping, et al.
Pubblicazione: (2025)
Bird detection in audio: a survey and a challenge
di: Stowell, Dan, et al.
Pubblicazione: (2016)
di: Stowell, Dan, et al.
Pubblicazione: (2016)
SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
di: Tang, Yuxun, et al.
Pubblicazione: (2024)
di: Tang, Yuxun, et al.
Pubblicazione: (2024)
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
di: Wang, Dongmei, et al.
Pubblicazione: (2023)
di: Wang, Dongmei, et al.
Pubblicazione: (2023)
Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race
di: Mao, Xutao, et al.
Pubblicazione: (2025)
di: Mao, Xutao, et al.
Pubblicazione: (2025)
Content Leakage in LibriSpeech and Its Impact on the Privacy Evaluation of Speaker Anonymization
di: Franzreb, Carlos, et al.
Pubblicazione: (2026)
di: Franzreb, Carlos, et al.
Pubblicazione: (2026)
VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding
di: Ye, Jashin, et al.
Pubblicazione: (2026)
di: Ye, Jashin, et al.
Pubblicazione: (2026)
R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
di: Chang, Heng-Jui, et al.
Pubblicazione: (2023)
di: Chang, Heng-Jui, et al.
Pubblicazione: (2023)
VoiceBench: Benchmarking LLM-Based Voice Assistants
di: Chen, Yiming, et al.
Pubblicazione: (2024)
di: Chen, Yiming, et al.
Pubblicazione: (2024)
Automatic acoustic detection of birds through deep learning: the first Bird Audio Detection challenge
di: Stowell, Dan, et al.
Pubblicazione: (2018)
di: Stowell, Dan, et al.
Pubblicazione: (2018)
AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features
di: Xu, Weiye, et al.
Pubblicazione: (2025)
di: Xu, Weiye, et al.
Pubblicazione: (2025)
LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments
di: Opochinsky, Renana, et al.
Pubblicazione: (2024)
di: Opochinsky, Renana, et al.
Pubblicazione: (2024)
Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection
di: Mariotte, Théo, et al.
Pubblicazione: (2024)
di: Mariotte, Théo, et al.
Pubblicazione: (2024)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2024)
di: Thienpondt, Jenthe, et al.
Pubblicazione: (2024)
$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
di: Ray, Soham, et al.
Pubblicazione: (2026)
di: Ray, Soham, et al.
Pubblicazione: (2026)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
di: Wu, Zhiyu, et al.
Pubblicazione: (2025)
di: Wu, Zhiyu, et al.
Pubblicazione: (2025)
A Holistic Framework for Robust Bangla ASR and Speaker Diarization with Optimized VAD and CTC Alignment
di: Ishmam, Zarif, et al.
Pubblicazione: (2026)
di: Ishmam, Zarif, et al.
Pubblicazione: (2026)
OpenVoice: Versatile Instant Voice Cloning
di: Qin, Zengyi, et al.
Pubblicazione: (2023)
di: Qin, Zengyi, et al.
Pubblicazione: (2023)
Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
di: Zeng, Bang, et al.
Pubblicazione: (2025)
di: Zeng, Bang, et al.
Pubblicazione: (2025)
How Much Does Machine Identity Matter in Anomalous Sound Detection at Test Time?
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
di: Mahapatra, Ishaan, et al.
Pubblicazione: (2025)
di: Mahapatra, Ishaan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Vocal Tract Length Warped Features for Spoken Keyword Spotting
di: Sarkar, Achintya kr., et al.
Pubblicazione: (2025) -
SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization
di: Wang, Chien-Chun, et al.
Pubblicazione: (2025) -
One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
di: Stylianou, Ioannis, et al.
Pubblicazione: (2026) -
Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
di: Wang, Liming, et al.
Pubblicazione: (2024) -
sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
di: Yang, Qu, et al.
Pubblicazione: (2024)