Phone Duration Modeling for Speaker Age Estimation in Children
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shivakumar, Prashanth Gurunath, Bishop, Somer, Lord, Catherine, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
von: Prescott, Jordan, et al.
Veröffentlicht: (2026)
von: Prescott, Jordan, et al.
Veröffentlicht: (2026)
Large Language Models based ASR Error Correction for Child Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Enhancing Age-Related Robustness in Children Speaker Verification
von: Shetty, Vishwas M., et al.
Veröffentlicht: (2025)
von: Shetty, Vishwas M., et al.
Veröffentlicht: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
Audio-visual child-adult speaker classification in dyadic interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
von: Xu, Anfeng, et al.
Veröffentlicht: (2023)
Emotion-Aligned Contrastive Learning Between Images and Music
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
Robust Target Speaker Direction of Arrival Estimation
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
Group Relative Policy Optimization for Speech Recognition
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2025)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2025)
Multi-Speaker DOA Estimation in Binaural Hearing Aids using Deep Learning and Speaker Count Fusion
von: Jazaeri, Farnaz, et al.
Veröffentlicht: (2025)
von: Jazaeri, Farnaz, et al.
Veröffentlicht: (2025)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
von: Wu, Shu, et al.
Veröffentlicht: (2025)
von: Wu, Shu, et al.
Veröffentlicht: (2025)
ChildAugment: Data Augmentation Methods for Zero-Resource Children's Speaker Verification
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2024)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2024)
Speaker Distance Estimation in Enclosures from Single-Channel Audio
von: Neri, Michael, et al.
Veröffentlicht: (2024)
von: Neri, Michael, et al.
Veröffentlicht: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
von: Du, Chenpeng, et al.
Veröffentlicht: (2021)
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
von: Boeddeker, Christoph, et al.
Veröffentlicht: (2023)
von: Boeddeker, Christoph, et al.
Veröffentlicht: (2023)
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
Speaker Contrastive Learning for Source Speaker Tracing
von: Wang, Qing, et al.
Veröffentlicht: (2024)
von: Wang, Qing, et al.
Veröffentlicht: (2024)
Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation
von: Lin, Jackie, et al.
Veröffentlicht: (2025)
von: Lin, Jackie, et al.
Veröffentlicht: (2025)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
A Comprehensive Investigation on Speaker Augmentation for Speaker Recognition
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
Mamba-based Segmentation Model for Speaker Diarization
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
von: Kolehmainen, Jari, et al.
Veröffentlicht: (2024)
von: Kolehmainen, Jari, et al.
Veröffentlicht: (2024)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026) -
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
von: Feng, Tiantian, et al.
Veröffentlicht: (2024) -
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024) -
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025) -
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)