Speaker Recognition -- Wavelet Packet Based Multiresolution Feature Extraction Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhardwaj, Saurabh, Srivastava, Smriti, Bhandari, Abhishek, Gupta, Krit, Bahl, Hitesh, Gupta, J. R. P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
von: Shore, Noah
Veröffentlicht: (2025)
von: Shore, Noah
Veröffentlicht: (2025)
Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
von: Goel, Arnav, et al.
Veröffentlicht: (2024)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025)
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
von: Agrawal, Saurabh, et al.
Veröffentlicht: (2025)
von: Agrawal, Saurabh, et al.
Veröffentlicht: (2025)
Brainprint-Modulated Target Speaker Extraction
von: Han, Qiushi, et al.
Veröffentlicht: (2025)
von: Han, Qiushi, et al.
Veröffentlicht: (2025)
Training-Free Multi-Step Inference for Target Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2026)
von: You, Zhenghai, et al.
Veröffentlicht: (2026)
Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
von: Dai, Yuhang, et al.
Veröffentlicht: (2026)
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
von: Wu, Shu, et al.
Veröffentlicht: (2025)
von: Wu, Shu, et al.
Veröffentlicht: (2025)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
Target Speaker Extraction with Curriculum Learning
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
USED: Universal Speaker Extraction and Diarization
von: Ao, Junyi, et al.
Veröffentlicht: (2023)
von: Ao, Junyi, et al.
Veröffentlicht: (2023)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction
von: Liu, Yun, et al.
Veröffentlicht: (2026)
von: Liu, Yun, et al.
Veröffentlicht: (2026)
Online Audio-Visual Autoregressive Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
U3-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty
von: Li, Junjie, et al.
Veröffentlicht: (2026)
von: Li, Junjie, et al.
Veröffentlicht: (2026)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
von: Yin, Han, et al.
Veröffentlicht: (2025)
von: Yin, Han, et al.
Veröffentlicht: (2025)
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2026)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Listen to Extract: Onset-Prompted Target Speaker Extraction
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
Binaural Target Speaker Extraction using Individualized HRTF
von: Ellinson, Yoav, et al.
Veröffentlicht: (2025)
von: Ellinson, Yoav, et al.
Veröffentlicht: (2025)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Self-Tuning Spectral Clustering for Speaker Diarization
von: Raghav, Nikhil, et al.
Veröffentlicht: (2024)
von: Raghav, Nikhil, et al.
Veröffentlicht: (2024)
RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
von: Matiyali, Neeraj, et al.
Veröffentlicht: (2025)
von: Matiyali, Neeraj, et al.
Veröffentlicht: (2025)
Multi-Target Backdoor Attacks Against Speaker Recognition
von: Fortier, Alexandrine, et al.
Veröffentlicht: (2025)
von: Fortier, Alexandrine, et al.
Veröffentlicht: (2025)
Beyond Speaker Identity: Text Guided Target Speech Extraction
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
von: Li, Xiao, et al.
Veröffentlicht: (2025)
von: Li, Xiao, et al.
Veröffentlicht: (2025)
Two-Stage Adaptation for Non-Normative Speech Recognition: Revisiting Speaker-Independent Initialization for Personalization
von: Jiang, Shan, et al.
Veröffentlicht: (2026)
von: Jiang, Shan, et al.
Veröffentlicht: (2026)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
End-to-End Multi-Microphone Speaker Extraction Using Relative Transfer Functions
von: Eisenberg, Aviad, et al.
Veröffentlicht: (2025)
von: Eisenberg, Aviad, et al.
Veröffentlicht: (2025)
oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
von: Dip, Muhammad Sudipto Siam, et al.
Veröffentlicht: (2024)
von: Dip, Muhammad Sudipto Siam, et al.
Veröffentlicht: (2024)
On the application of Visibility Graphs in the Spectral Domain for Speaker Recognition
von: Bocaccio, Hernan, et al.
Veröffentlicht: (2025)
von: Bocaccio, Hernan, et al.
Veröffentlicht: (2025)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025) -
Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
von: Shore, Noah
Veröffentlicht: (2025) -
Exploring Multilingual Unseen Speaker Emotion Recognition: Leveraging Co-Attention Cues in Multitask Learning
von: Goel, Arnav, et al.
Veröffentlicht: (2024) -
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
von: Nfissi, Alaa, et al.
Veröffentlicht: (2025) -
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025)