Audio-visual child-adult speaker classification in dyadic interactions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Anfeng, Huang, Kevin, Feng, Tiantian, Tager-Flusberg, Helen, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Joint ASR and Speaker Role Tagging with Serialized Output Training
par: Xu, Anfeng, et autres
Publié: (2025)
par: Xu, Anfeng, et autres
Publié: (2025)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
par: Xu, Anfeng, et autres
Publié: (2024)
par: Xu, Anfeng, et autres
Publié: (2024)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
par: Feng, Tiantian, et autres
Publié: (2026)
par: Feng, Tiantian, et autres
Publié: (2026)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
par: Xu, Anfeng, et autres
Publié: (2026)
par: Xu, Anfeng, et autres
Publié: (2026)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
par: Xu, Anfeng, et autres
Publié: (2024)
par: Xu, Anfeng, et autres
Publié: (2024)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
par: Feng, Tiantian, et autres
Publié: (2023)
par: Feng, Tiantian, et autres
Publié: (2023)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
par: Feng, Tiantian, et autres
Publié: (2024)
par: Feng, Tiantian, et autres
Publié: (2024)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
par: Feng, Tiantian, et autres
Publié: (2024)
par: Feng, Tiantian, et autres
Publié: (2024)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
par: Feng, Tiantian, et autres
Publié: (2024)
par: Feng, Tiantian, et autres
Publié: (2024)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
par: Feng, Tiantian, et autres
Publié: (2025)
par: Feng, Tiantian, et autres
Publié: (2025)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
par: Xu, Anfeng, et autres
Publié: (2026)
par: Xu, Anfeng, et autres
Publié: (2026)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
par: Feng, Tiantian, et autres
Publié: (2025)
par: Feng, Tiantian, et autres
Publié: (2025)
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
par: Prescott, Jordan, et autres
Publié: (2026)
par: Prescott, Jordan, et autres
Publié: (2026)
Emotion-Aligned Contrastive Learning Between Images and Music
par: Stewart, Shanti, et autres
Publié: (2023)
par: Stewart, Shanti, et autres
Publié: (2023)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
par: Feng, Tiantian, et autres
Publié: (2024)
par: Feng, Tiantian, et autres
Publié: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
par: Feng, Tiantian, et autres
Publié: (2025)
par: Feng, Tiantian, et autres
Publié: (2025)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
par: Shi, Zhonghao, et autres
Publié: (2024)
par: Shi, Zhonghao, et autres
Publié: (2024)
Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech
par: Sun, Anchen, et autres
Publié: (2025)
par: Sun, Anchen, et autres
Publié: (2025)
Phone Duration Modeling for Speaker Age Estimation in Children
par: Shivakumar, Prashanth Gurunath, et autres
Publié: (2021)
par: Shivakumar, Prashanth Gurunath, et autres
Publié: (2021)
Hierarchical speaker representation for target speaker extraction
par: He, Shulin, et autres
Publié: (2022)
par: He, Shulin, et autres
Publié: (2022)
A framework of text-dependent speaker verification for chinese numerical string corpus
par: Zheng, Litong, et autres
Publié: (2024)
par: Zheng, Litong, et autres
Publié: (2024)
Text adaptation for speaker verification with speaker-text factorized embeddings
par: Yang, Yexin, et autres
Publié: (2025)
par: Yang, Yexin, et autres
Publié: (2025)
Improving curriculum learning for target speaker extraction with synthetic speakers
par: Liu, Yun, et autres
Publié: (2024)
par: Liu, Yun, et autres
Publié: (2024)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
par: Nguyen, Hong, et autres
Publié: (2024)
par: Nguyen, Hong, et autres
Publié: (2024)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
par: Yang, Dongchao, et autres
Publié: (2024)
par: Yang, Dongchao, et autres
Publié: (2024)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
par: Li, Guinan, et autres
Publié: (2024)
par: Li, Guinan, et autres
Publié: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
par: Lee, Jihwan, et autres
Publié: (2024)
par: Lee, Jihwan, et autres
Publié: (2024)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
par: Yang, Dongchao, et autres
Publié: (2023)
par: Yang, Dongchao, et autres
Publié: (2023)
The NeurIPS 2023 Machine Learning for Audio Workshop: Affective Audio Benchmarks and Novel Data
par: Baird, Alice, et autres
Publié: (2024)
par: Baird, Alice, et autres
Publié: (2024)
Effective Integration of KAN for Keyword Spotting
par: Xu, Anfeng, et autres
Publié: (2024)
par: Xu, Anfeng, et autres
Publié: (2024)
How phonemes contribute to deep speaker models?
par: Li, Pengqi, et autres
Publié: (2024)
par: Li, Pengqi, et autres
Publié: (2024)
The importance of spatial and spectral information in multiple speaker tracking
par: Beit-On, Hanan, et autres
Publié: (2024)
par: Beit-On, Hanan, et autres
Publié: (2024)
Spoken language change detection inspired by speaker change detection
par: Mishra, Jagabandhu, et autres
Publié: (2023)
par: Mishra, Jagabandhu, et autres
Publié: (2023)
On the influence of language similarity in non-target speaker verification trials
par: Reuter, Paul M., et autres
Publié: (2025)
par: Reuter, Paul M., et autres
Publié: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
par: Han, Runduo, et autres
Publié: (2024)
par: Han, Runduo, et autres
Publié: (2024)
AudioCIL: A Python Toolbox for Audio Class-Incremental Learning with Multiple Scenes
par: Xu, Qisheng, et autres
Publié: (2024)
par: Xu, Qisheng, et autres
Publié: (2024)
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
par: Ma, Yi, et autres
Publié: (2024)
par: Ma, Yi, et autres
Publié: (2024)
Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
par: Eisenberg, Aviad, et autres
Publié: (2025)
par: Eisenberg, Aviad, et autres
Publié: (2025)
Why disentanglement-based speaker anonymization systems fail at preserving emotions?
par: Gaznepoglu, Ünal Ege, et autres
Publié: (2025)
par: Gaznepoglu, Ünal Ege, et autres
Publié: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
par: Murata, Masato, et autres
Publié: (2025)
par: Murata, Masato, et autres
Publié: (2025)
Documents similaires
-
Joint ASR and Speaker Role Tagging with Serialized Output Training
par: Xu, Anfeng, et autres
Publié: (2025) -
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
par: Xu, Anfeng, et autres
Publié: (2024) -
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
par: Feng, Tiantian, et autres
Publié: (2026) -
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
par: Xu, Anfeng, et autres
Publié: (2026) -
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
par: Xu, Anfeng, et autres
Publié: (2024)