Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Jialu, Hasegawa-Johnson, Mark, Karahalios, Karrie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
por: Zheng, Xiuwen, et al.
Publicado: (2024)
por: Zheng, Xiuwen, et al.
Publicado: (2024)
Revealing the Hidden Temporal Structure of HubertSoft Embeddings based on the Russian Phonetic Corpus
por: Ananeva, Anastasia, et al.
Publicado: (2025)
por: Ananeva, Anastasia, et al.
Publicado: (2025)
Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis
por: Zhang, Yucong, et al.
Publicado: (2024)
por: Zhang, Yucong, et al.
Publicado: (2024)
Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
por: Li, Jialu, et al.
Publicado: (2025)
por: Li, Jialu, et al.
Publicado: (2025)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
por: Kang, Wonjune, et al.
Publicado: (2022)
por: Kang, Wonjune, et al.
Publicado: (2022)
PhiNet: Speaker Verification with Phonetic Interpretability
por: Ma, Yi, et al.
Publicado: (2026)
por: Ma, Yi, et al.
Publicado: (2026)
High-Fidelity Neural Phonetic Posteriorgrams
por: Churchwell, Cameron, et al.
Publicado: (2024)
por: Churchwell, Cameron, et al.
Publicado: (2024)
Phonetic Richness for Improved Automatic Speaker Verification
por: Klein, Nicholas, et al.
Publicado: (2024)
por: Klein, Nicholas, et al.
Publicado: (2024)
A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features For Classical Vocal Performance
por: Park, Jiyun, et al.
Publicado: (2024)
por: Park, Jiyun, et al.
Publicado: (2024)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
por: Zhou, Xuanru, et al.
Publicado: (2025)
por: Zhou, Xuanru, et al.
Publicado: (2025)
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
por: Wang, Ju-Chiang, et al.
Publicado: (2024)
por: Wang, Ju-Chiang, et al.
Publicado: (2024)
Computational Extraction of Intonation and Tuning Systems from Multiple Microtonal Monophonic Vocal Recordings with Diverse Modes
por: Shafiei, Sepideh, et al.
Publicado: (2025)
por: Shafiei, Sepideh, et al.
Publicado: (2025)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
por: Chodroff, Eleanor, et al.
Publicado: (2024)
por: Chodroff, Eleanor, et al.
Publicado: (2024)
Sound Tagging in Infant-centric Home Soundscapes
por: Khan, Mohammad Nur Hossain, et al.
Publicado: (2024)
por: Khan, Mohammad Nur Hossain, et al.
Publicado: (2024)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
por: Zheng, Xiuwen, et al.
Publicado: (2026)
por: Zheng, Xiuwen, et al.
Publicado: (2026)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
por: Zhang, Miao, et al.
Publicado: (2025)
por: Zhang, Miao, et al.
Publicado: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
por: Zhou, Kun, et al.
Publicado: (2024)
por: Zhou, Kun, et al.
Publicado: (2024)
CVSM: Contrastive Vocal Similarity Modeling
por: Garoufis, Christos, et al.
Publicado: (2025)
por: Garoufis, Christos, et al.
Publicado: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
por: Han, Seungu, et al.
Publicado: (2026)
por: Han, Seungu, et al.
Publicado: (2026)
Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning
por: Shi, Runwu, et al.
Publicado: (2024)
por: Shi, Runwu, et al.
Publicado: (2024)
Vision Transformer Segmentation for Visual Bird Sound Denoising
por: Kumar, Sahil, et al.
Publicado: (2024)
por: Kumar, Sahil, et al.
Publicado: (2024)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
por: Yamashita, Natsuo, et al.
Publicado: (2026)
por: Yamashita, Natsuo, et al.
Publicado: (2026)
Interpreting the Dimensions of Speaker Embedding Space
por: Huckvale, Mark
Publicado: (2025)
por: Huckvale, Mark
Publicado: (2025)
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
por: Pozorski, Paweł, et al.
Publicado: (2026)
por: Pozorski, Paweł, et al.
Publicado: (2026)
Feature Representations for Automatic Meerkat Vocalization Classification
por: Mahmoud, Imen Ben, et al.
Publicado: (2024)
por: Mahmoud, Imen Ben, et al.
Publicado: (2024)
Complex Image-Generative Diffusion Transformer for Audio Denoising
por: Li, Junhui, et al.
Publicado: (2024)
por: Li, Junhui, et al.
Publicado: (2024)
Auditory Representation Effective for Estimating Vocal Tract Information
por: Irino, Toshio, et al.
Publicado: (2023)
por: Irino, Toshio, et al.
Publicado: (2023)
Melodic and Metrical Elements of Expressiveness in Hindustani Vocal Music
por: Bhake, Yash, et al.
Publicado: (2025)
por: Bhake, Yash, et al.
Publicado: (2025)
A Reliable and Efficient Detection Pipeline for Rodent Ultrasonic Vocalizations
por: Anis, Sabah Shahnoor, et al.
Publicado: (2025)
por: Anis, Sabah Shahnoor, et al.
Publicado: (2025)
Biodenoising: Animal Vocalization Denoising without Access to Clean Data
por: Miron, Marius, et al.
Publicado: (2024)
por: Miron, Marius, et al.
Publicado: (2024)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
por: Nobukawa, Rinka, et al.
Publicado: (2025)
por: Nobukawa, Rinka, et al.
Publicado: (2025)
Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
por: Wada, Aogu, et al.
Publicado: (2025)
por: Wada, Aogu, et al.
Publicado: (2025)
Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
por: Wang, Zilai, et al.
Publicado: (2026)
por: Wang, Zilai, et al.
Publicado: (2026)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
por: Wei, Haojie, et al.
Publicado: (2023)
por: Wei, Haojie, et al.
Publicado: (2023)
voc2vec: A Foundation Model for Non-Verbal Vocalization
por: Koudounas, Alkis, et al.
Publicado: (2025)
por: Koudounas, Alkis, et al.
Publicado: (2025)
Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
por: Sorrenti, Adam
Publicado: (2024)
por: Sorrenti, Adam
Publicado: (2024)
Learning Vocal-Tract Area and Radiation with a Physics-Informed Webster Model
por: Lu, Minhui, et al.
Publicado: (2026)
por: Lu, Minhui, et al.
Publicado: (2026)
DiffVox: A Differentiable Model for Capturing and Analysing Vocal Effects Distributions
por: Yu, Chin-Yun, et al.
Publicado: (2025)
por: Yu, Chin-Yun, et al.
Publicado: (2025)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
por: Ashvin, Aditya, et al.
Publicado: (2024)
por: Ashvin, Aditya, et al.
Publicado: (2024)
Ejemplares similares
-
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
por: Li, Jialu, et al.
Publicado: (2024) -
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
por: Zheng, Xiuwen, et al.
Publicado: (2024) -
Revealing the Hidden Temporal Structure of HubertSoft Embeddings based on the Russian Phonetic Corpus
por: Ananeva, Anastasia, et al.
Publicado: (2025) -
Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis
por: Zhang, Yucong, et al.
Publicado: (2024) -
Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
por: Li, Jialu, et al.
Publicado: (2025)