Rhythm Features for Speaker Identification
Fuente:
arXiv
Guardado en:
| Autores principales: | Mehlman, Nick, Thebaud, Thomas, Byrd, Dani, Narayanan, Shri |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
por: He, Xiluo, et al.
Publicado: (2025)
por: He, Xiluo, et al.
Publicado: (2025)
On the Relationship between Accent Strength and Articulatory Features
por: Huang, Kevin, et al.
Publicado: (2025)
por: Huang, Kevin, et al.
Publicado: (2025)
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification
por: Joshi, Sonal, et al.
Publicado: (2024)
por: Joshi, Sonal, et al.
Publicado: (2024)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
por: Xu, Anfeng, et al.
Publicado: (2026)
por: Xu, Anfeng, et al.
Publicado: (2026)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
por: Xu, Anfeng, et al.
Publicado: (2024)
por: Xu, Anfeng, et al.
Publicado: (2024)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
por: Park, Jay, et al.
Publicado: (2025)
por: Park, Jay, et al.
Publicado: (2025)
Joint ASR and Speaker Role Tagging with Serialized Output Training
por: Xu, Anfeng, et al.
Publicado: (2025)
por: Xu, Anfeng, et al.
Publicado: (2025)
VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
por: Lin, Yuke, et al.
Publicado: (2024)
por: Lin, Yuke, et al.
Publicado: (2024)
Phone Duration Modeling for Speaker Age Estimation in Children
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2021)
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2021)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
por: Morrone, Giovanni, et al.
Publicado: (2024)
por: Morrone, Giovanni, et al.
Publicado: (2024)
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
por: Deegen, Marc, et al.
Publicado: (2026)
por: Deegen, Marc, et al.
Publicado: (2026)
Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
por: Ahmed, Sabbir, et al.
Publicado: (2025)
por: Ahmed, Sabbir, et al.
Publicado: (2025)
DNN based HRIRs Identification with a Continuously Rotating Speaker Array
por: Ko, Byeong-Yun, et al.
Publicado: (2025)
por: Ko, Byeong-Yun, et al.
Publicado: (2025)
SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion
por: Chen, Zhiyong, et al.
Publicado: (2026)
por: Chen, Zhiyong, et al.
Publicado: (2026)
Emotion Recognition in Multi-Speaker Conversations through Speaker Identification, Knowledge Distillation, and Hierarchical Fusion
por: Li, Xiao, et al.
Publicado: (2025)
por: Li, Xiao, et al.
Publicado: (2025)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
por: Xu, Anfeng, et al.
Publicado: (2026)
por: Xu, Anfeng, et al.
Publicado: (2026)
Multi-Label Training for Text-Independent Speaker Identification
por: Xue, Yuqi
Publicado: (2022)
por: Xue, Yuqi
Publicado: (2022)
Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample
por: Chen, Zhiyong, et al.
Publicado: (2024)
por: Chen, Zhiyong, et al.
Publicado: (2024)
openFEAT: Improving Speaker Identification by Open-set Few-shot Embedding Adaptation with Transformer
por: C, Kishan K, et al.
Publicado: (2022)
por: C, Kishan K, et al.
Publicado: (2022)
Magnitude and Phase-based Feature Fusion Using Co-attention Mechanism for Speaker recognition
por: Su, Rongfeng, et al.
Publicado: (2025)
por: Su, Rongfeng, et al.
Publicado: (2025)
Noise-robust Speech Separation with Fast Generative Correction
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification
por: McKnight, Simon W., et al.
Publicado: (2023)
por: McKnight, Simon W., et al.
Publicado: (2023)
Study on Inter and Intra Speaker Variability in Speaker Recognition
por: Okhotnikov, Anton, et al.
Publicado: (2024)
por: Okhotnikov, Anton, et al.
Publicado: (2024)
Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
por: Thebaud, Thomas, et al.
Publicado: (2025)
por: Thebaud, Thomas, et al.
Publicado: (2025)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
por: Zhao, Yiyang, et al.
Publicado: (2024)
por: Zhao, Yiyang, et al.
Publicado: (2024)
Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
por: Zhang, Jing-Xuan, et al.
Publicado: (2025)
Joint Optimization of Speaker and Spoof Detectors for Spoofing-Robust Automatic Speaker Verification
por: Kurnaz, Oğuzhan, et al.
Publicado: (2025)
por: Kurnaz, Oğuzhan, et al.
Publicado: (2025)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
por: Shao, Yiwen, et al.
Publicado: (2024)
por: Shao, Yiwen, et al.
Publicado: (2024)
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
por: Fan, Cunhang, et al.
Publicado: (2025)
por: Fan, Cunhang, et al.
Publicado: (2025)
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
por: Wang, Qing, et al.
Publicado: (2025)
por: Wang, Qing, et al.
Publicado: (2025)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
por: Zhang, Yanzhe, et al.
Publicado: (2024)
por: Zhang, Yanzhe, et al.
Publicado: (2024)
SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
por: Li, Yue, et al.
Publicado: (2024)
por: Li, Yue, et al.
Publicado: (2024)
Generating Rhythm Game Music with Jukebox
por: Yan, Nicholas
Publicado: (2023)
por: Yan, Nicholas
Publicado: (2023)
Frequency Tracking Features for Data-Efficient Deep Siren Identification
por: Damiano, Stefano, et al.
Publicado: (2024)
por: Damiano, Stefano, et al.
Publicado: (2024)
MC-LExt: Multi-Channel Target Speaker Extraction with Onset-Prompted Speaker Conditioning Mechanism
por: Ling, Tongtao, et al.
Publicado: (2025)
por: Ling, Tongtao, et al.
Publicado: (2025)
Ejemplares similares
-
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
por: Feng, Tiantian, et al.
Publicado: (2025) -
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
por: Feng, Tiantian, et al.
Publicado: (2025) -
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
por: He, Xiluo, et al.
Publicado: (2025) -
On the Relationship between Accent Strength and Articulatory Features
por: Huang, Kevin, et al.
Publicado: (2025) -
Unraveling Adversarial Examples against Speaker Identification -- Techniques for Attack Detection and Victim Model Classification
por: Joshi, Sonal, et al.
Publicado: (2024)