Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jin Sob, Park, Hyun Joon, Shin, Wooseok, Yun, Juan, Han, Sung Won |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
by: Park, Hyun Joon, et al.
Published: (2024)
by: Park, Hyun Joon, et al.
Published: (2024)
Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
by: Kim, Jin Sob, et al.
Published: (2025)
by: Kim, Jin Sob, et al.
Published: (2025)
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
by: Heo, Hyun-Jun, et al.
Published: (2023)
by: Heo, Hyun-Jun, et al.
Published: (2023)
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
by: Li, Ya, et al.
Published: (2025)
by: Li, Ya, et al.
Published: (2025)
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
by: Park, Hyun Joon, et al.
Published: (2025)
by: Park, Hyun Joon, et al.
Published: (2025)
Short-Segment Speaker Verification with Pre-trained Models and Multi-Resolution Encoder
by: Myoung, Jisoo, et al.
Published: (2025)
by: Myoung, Jisoo, et al.
Published: (2025)
Infant Cry Emotion Recognition Using Improved ECAPA-TDNN with Multiscale Feature Fusion and Attention Enhancement
by: Zhou, Junyu, et al.
Published: (2025)
by: Zhou, Junyu, et al.
Published: (2025)
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
by: Weng, Shilong, et al.
Published: (2025)
by: Weng, Shilong, et al.
Published: (2025)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
by: Kunešová, Marie, et al.
Published: (2025)
by: Kunešová, Marie, et al.
Published: (2025)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
by: Nam, KiHyun, et al.
Published: (2024)
by: Nam, KiHyun, et al.
Published: (2024)
LG Uplus System with Multi-Speaker IDs and Discriminator-based Sub-Judges for the WildSpoof Challenge
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
by: Han, Wooseok, et al.
Published: (2024)
by: Han, Wooseok, et al.
Published: (2024)
Deep Filter Estimation from Inter-Frame Correlations for Monaural Speech Dereverberation
by: Shin, Ui-Hyeop, et al.
Published: (2026)
by: Shin, Ui-Hyeop, et al.
Published: (2026)
LightCAM: A Fast and Light Implementation of Context-Aware Masking based D-TDNN for Speaker Verification
by: Cao, Di, et al.
Published: (2024)
by: Cao, Di, et al.
Published: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
by: Park, Nohil, et al.
Published: (2024)
by: Park, Nohil, et al.
Published: (2024)
Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTS
by: Li, Haoyu, et al.
Published: (2025)
by: Li, Haoyu, et al.
Published: (2025)
SEED: Speaker Embedding Enhancement Diffusion Model
by: Nam, KiHyun, et al.
Published: (2025)
by: Nam, KiHyun, et al.
Published: (2025)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
by: Shao, Yiwen, et al.
Published: (2024)
by: Shao, Yiwen, et al.
Published: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
by: Li, Guinan, et al.
Published: (2024)
by: Li, Guinan, et al.
Published: (2024)
Study on Inter and Intra Speaker Variability in Speaker Recognition
by: Okhotnikov, Anton, et al.
Published: (2024)
by: Okhotnikov, Anton, et al.
Published: (2024)
MR-RawNet: Speaker verification system with multiple temporal resolutions for variable duration utterances using raw waveforms
by: Kim, Seung-bin, et al.
Published: (2024)
by: Kim, Seung-bin, et al.
Published: (2024)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
by: Dixit, Satvik, et al.
Published: (2024)
by: Dixit, Satvik, et al.
Published: (2024)
Multi-View Based Audio Visual Target Speaker Extraction
by: Yang, Peijun, et al.
Published: (2026)
by: Yang, Peijun, et al.
Published: (2026)
SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification
by: Heo, Jungwoo, et al.
Published: (2025)
by: Heo, Jungwoo, et al.
Published: (2025)
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
by: Aronowitz, Hagai, et al.
Published: (2026)
by: Aronowitz, Hagai, et al.
Published: (2026)
Rhythm Features for Speaker Identification
by: Mehlman, Nick, et al.
Published: (2025)
by: Mehlman, Nick, et al.
Published: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
by: Chen, Zhengyang, et al.
Published: (2024)
by: Chen, Zhengyang, et al.
Published: (2024)
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning
by: Kang, Ju Yeon, et al.
Published: (2025)
by: Kang, Ju Yeon, et al.
Published: (2025)
Reshape Dimensions Network for Speaker Recognition
by: Yakovlev, Ivan, et al.
Published: (2024)
by: Yakovlev, Ivan, et al.
Published: (2024)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
by: Zhou, Zhenyu, et al.
Published: (2024)
by: Zhou, Zhenyu, et al.
Published: (2024)
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
by: Deegen, Marc, et al.
Published: (2026)
by: Deegen, Marc, et al.
Published: (2026)
DNN based HRIRs Identification with a Continuously Rotating Speaker Array
by: Ko, Byeong-Yun, et al.
Published: (2025)
by: Ko, Byeong-Yun, et al.
Published: (2025)
An Effective Energy Mask-based Adversarial Evasion Attacks against Misclassification in Speaker Recognition Systems
by: Park, Chanwoo, et al.
Published: (2026)
by: Park, Chanwoo, et al.
Published: (2026)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
by: Sakuma, Asahi, et al.
Published: (2025)
by: Sakuma, Asahi, et al.
Published: (2025)
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
by: Joshi, Sonal, et al.
Published: (2021)
by: Joshi, Sonal, et al.
Published: (2021)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
by: Jeon, Yejin, et al.
Published: (2024)
by: Jeon, Yejin, et al.
Published: (2024)
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
by: Lin, Yuke, et al.
Published: (2025)
by: Lin, Yuke, et al.
Published: (2025)
Multi-Label Training for Text-Independent Speaker Identification
by: Xue, Yuqi
Published: (2022)
by: Xue, Yuqi
Published: (2022)
Similar Items
-
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
by: Park, Hyun Joon, et al.
Published: (2024) -
Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
by: Kim, Jin Sob, et al.
Published: (2025) -
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
by: Heo, Hyun-Jun, et al.
Published: (2023) -
MGFF-TDNN: A Multi-Granularity Feature Fusion TDNN Model with Depth-Wise Separable Module for Speaker Verification
by: Li, Ya, et al.
Published: (2025) -
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
by: Park, Hyun Joon, et al.
Published: (2025)