ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Yafeng, Zheng, Siqi, Wang, Hui, Cheng, Luyao, Chen, Qian, Zhang, Shiliang, Li, Junjie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2023)
por: Chen, Yafeng, et al.
Publicado: (2023)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
por: Cheng, Luyao, et al.
Publicado: (2023)
por: Cheng, Luyao, et al.
Publicado: (2023)
An Adaptive X-vector Model for Text-independent Speaker Verification
por: Gu, Bin, et al.
Publicado: (2020)
por: Gu, Bin, et al.
Publicado: (2020)
Advanced Signal Analysis in Detecting Replay Attacks for Automatic Speaker Verification Systems
por: Kuang, Lee Shih
Publicado: (2024)
por: Kuang, Lee Shih
Publicado: (2024)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
por: Sato, Hiroshi, et al.
Publicado: (2024)
por: Sato, Hiroshi, et al.
Publicado: (2024)
Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios
por: Fiorio, Luan Vinícius, et al.
Publicado: (2025)
por: Fiorio, Luan Vinícius, et al.
Publicado: (2025)
The Overview of Segmental Durations Modification Algorithms on Speech Signal Characteristics
por: Jang, Kyeomeun, et al.
Publicado: (2025)
por: Jang, Kyeomeun, et al.
Publicado: (2025)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
por: Iatariene, Taous, et al.
Publicado: (2025)
por: Iatariene, Taous, et al.
Publicado: (2025)
Speakers Localization Using Batch EM In Unfolding Neural Network
por: Veler, Rina, et al.
Publicado: (2026)
por: Veler, Rina, et al.
Publicado: (2026)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
por: Xie, Yuying, et al.
Publicado: (2024)
por: Xie, Yuying, et al.
Publicado: (2024)
Tracking of Intermittent and Moving Speakers : Dataset and Metrics
por: Iatariene, Taous, et al.
Publicado: (2025)
por: Iatariene, Taous, et al.
Publicado: (2025)
Binaural Selective Attention Model for Target Speaker Extraction
por: Meng, Hanyu, et al.
Publicado: (2024)
por: Meng, Hanyu, et al.
Publicado: (2024)
SyncNet: correlating objective for time delay estimation in audio signals
por: Raina, Akshay, et al.
Publicado: (2022)
por: Raina, Akshay, et al.
Publicado: (2022)
Prompt-driven Target Speech Diarization
por: Jiang, Yidi, et al.
Publicado: (2023)
por: Jiang, Yidi, et al.
Publicado: (2023)
Comparison of Frequency-Fusion Mechanisms for Binaural Direction-of-Arrival Estimation for Multiple Speakers
por: Fejgin, Daniel, et al.
Publicado: (2024)
por: Fejgin, Daniel, et al.
Publicado: (2024)
PLDNet: PLD-Guided Lightweight Deep Network Boosted by Efficient Attention for Handheld Dual-Microphone Speech Enhancement
por: Zhou, Nan, et al.
Publicado: (2024)
por: Zhou, Nan, et al.
Publicado: (2024)
Modeling and Link Budget Feasibility Analysis of Secure LoRa-Based Peer-to-Peer Communication for Short-Range Tactical Networks
por: Agrawal, Ayush Kumar, et al.
Publicado: (2026)
por: Agrawal, Ayush Kumar, et al.
Publicado: (2026)
Self-Boosted Weight-Constrained FxLMS: A Robustness Distributed Active Noise Control Algorithm Without Internode Communication
por: Ji, Junwei, et al.
Publicado: (2025)
por: Ji, Junwei, et al.
Publicado: (2025)
Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
por: Fejgin, Daniel, et al.
Publicado: (2023)
por: Fejgin, Daniel, et al.
Publicado: (2023)
Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
por: Fejgin, Daniel, et al.
Publicado: (2025)
por: Fejgin, Daniel, et al.
Publicado: (2025)
FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
por: Choi, Yuseon, et al.
Publicado: (2025)
por: Choi, Yuseon, et al.
Publicado: (2025)
Automotive sound field reproduction using deep optimization with spatial domain constraint
por: Qian, Yufan, et al.
Publicado: (2025)
por: Qian, Yufan, et al.
Publicado: (2025)
DAME: Duration-Aware Matryoshka Embedding for Duration-Robust Speaker Verification
por: Jung, Youngmoon, et al.
Publicado: (2026)
por: Jung, Youngmoon, et al.
Publicado: (2026)
Reduce Computational Complexity for Continuous Wavelet Transform in Acoustic Recognition Using Hop Size
por: Phan, Dang Thoai
Publicado: (2024)
por: Phan, Dang Thoai
Publicado: (2024)
Robustness of Speech Separation Models for Similar-pitch Speakers
por: Lay, Bunlong, et al.
Publicado: (2024)
por: Lay, Bunlong, et al.
Publicado: (2024)
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
por: Xu, Yanze, et al.
Publicado: (2026)
por: Xu, Yanze, et al.
Publicado: (2026)
S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
por: Wang, Ziqian, et al.
Publicado: (2026)
por: Wang, Ziqian, et al.
Publicado: (2026)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
por: Du, Hui-Peng, et al.
Publicado: (2024)
por: Du, Hui-Peng, et al.
Publicado: (2024)
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
por: Kim, Miseul, et al.
Publicado: (2025)
por: Kim, Miseul, et al.
Publicado: (2025)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
por: Yang, Shu-wen, et al.
Publicado: (2025)
por: Yang, Shu-wen, et al.
Publicado: (2025)
Auditory Attention Decoding from Ear-EEG Signals: A Dataset with Dynamic Attention Switching and Rigorous Cross-Validation
por: Zhang, Yuanming, et al.
Publicado: (2025)
por: Zhang, Yuanming, et al.
Publicado: (2025)
ReDimNet2: Scaling Speaker Verification via Time-Pooled Dimension Reshaping
por: Yakovlev, Ivan, et al.
Publicado: (2026)
por: Yakovlev, Ivan, et al.
Publicado: (2026)
Noise Suppression for Time Difference of Arrival: Performance Evaluation of a Generalized Cross-Correlation Method Using Mean Signal and Inverse Filter
por: Obo, Hirotaka, et al.
Publicado: (2025)
por: Obo, Hirotaka, et al.
Publicado: (2025)
SuperM2M: Supervised and Mixture-to-Mixture Co-Learning for Speech Enhancement and Noise-Robust ASR
por: Wang, Zhong-Qiu
Publicado: (2024)
por: Wang, Zhong-Qiu
Publicado: (2024)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
por: Yang, Yujie, et al.
Publicado: (2025)
por: Yang, Yujie, et al.
Publicado: (2025)
Unsupervised Face-Masked Speech Enhancement Using Generative Adversarial Networks With Human-in-the-Loop Assessment Metrics
por: Wang, Syu-Siang, et al.
Publicado: (2024)
por: Wang, Syu-Siang, et al.
Publicado: (2024)
Future Full-Ocean Deep SSPs Prediction based on Hierarchical Long Short-Term Memory Neural Networks
por: Lu, Jiajun, et al.
Publicado: (2023)
por: Lu, Jiajun, et al.
Publicado: (2023)
Ejemplares similares
-
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
por: Chen, Yafeng, et al.
Publicado: (2024) -
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2024) -
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2023) -
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
por: Cheng, Luyao, et al.
Publicado: (2023) -
An Adaptive X-vector Model for Text-independent Speaker Verification
por: Gu, Bin, et al.
Publicado: (2020)