Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bovbjerg, Holger Severin, Østergaard, Jan, Jensen, Jesper, Tan, Zheng-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023)
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025)
Noise-Robust Keyword Spotting through Self-supervised Pretraining
von: Mørk, Jacob, et al.
Veröffentlicht: (2024)
von: Mørk, Jacob, et al.
Veröffentlicht: (2024)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)
von: Li, Haowen, et al.
Veröffentlicht: (2025)
Audio-based Kinship Verification Using Age Domain Conversion
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
Deep low-latency joint speech transmission and enhancement over a gaussian channel
von: Bokaei, Mohammad, et al.
Veröffentlicht: (2024)
von: Bokaei, Mohammad, et al.
Veröffentlicht: (2024)
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
von: Cheripally, Sowmya
Veröffentlicht: (2024)
von: Cheripally, Sowmya
Veröffentlicht: (2024)
ChordSync: Conformer-Based Alignment of Chord Annotations to Music Audio
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
von: Poltronieri, Andrea, et al.
Veröffentlicht: (2024)
Symbolic Audio Classification via Modal Decision Tree Learning
von: Marzano, Enrico, et al.
Veröffentlicht: (2025)
von: Marzano, Enrico, et al.
Veröffentlicht: (2025)
Joint Minimum Processing Beamforming and Near-end Listening Enhancement
von: Fuglsig, Andreas J., et al.
Veröffentlicht: (2023)
von: Fuglsig, Andreas J., et al.
Veröffentlicht: (2023)
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0
von: Kang, Taein, et al.
Veröffentlicht: (2024)
von: Kang, Taein, et al.
Veröffentlicht: (2024)
Monaural Multi-Speaker Speech Separation Using Efficient Transformer Model
von: Rijal, S., et al.
Veröffentlicht: (2023)
von: Rijal, S., et al.
Veröffentlicht: (2023)
STAR: Speech-to-Audio Generation via Representation Learning
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
FakeSound: Deepfake General Audio Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Passive Underwater Acoustic Signal Separation based on Feature Decoupling Dual-path Network
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
von: Liu, Yucheng, et al.
Veröffentlicht: (2025)
Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
von: Pankov, Vikentii, et al.
Veröffentlicht: (2023)
von: Pankov, Vikentii, et al.
Veröffentlicht: (2023)
Investigating the Design Space of Diffusion Models for Speech Enhancement
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2023)
Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
von: Cao, Xinwei, et al.
Veröffentlicht: (2026)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
TuneGenie: Reasoning-based LLM agents for preferential music generation
von: Pandey, Amitesh, et al.
Veröffentlicht: (2025)
von: Pandey, Amitesh, et al.
Veröffentlicht: (2025)
KinSPEAK: Improving speech recognition for Kinyarwanda via semi-supervised learning methods
von: Nzeyimana, Antoine
Veröffentlicht: (2023)
von: Nzeyimana, Antoine
Veröffentlicht: (2023)
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2023) -
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
von: Bovbjerg, Holger Severin, et al.
Veröffentlicht: (2025) -
Noise-Robust Keyword Spotting through Self-supervised Pretraining
von: Mørk, Jacob, et al.
Veröffentlicht: (2024) -
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026) -
Joint Feature and Output Distillation for Low-complexity Acoustic Scene Classification
von: Li, Haowen, et al.
Veröffentlicht: (2025)