Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Foley, Sean, Nguyen, Hong, Lee, Jihwan, Kadiri, Sudarsana Reddy, Byrd, Dani, Goldstein, Louis, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
by: Justus, Aju Ani, et al.
Published: (2026)
by: Justus, Aju Ani, et al.
Published: (2026)
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
by: Laukkanen, Anne-Maria, et al.
Published: (2024)
by: Laukkanen, Anne-Maria, et al.
Published: (2024)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
by: Sinha, Abhijit, et al.
Published: (2025)
by: Sinha, Abhijit, et al.
Published: (2025)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025)
by: Park, Jay, et al.
Published: (2025)
Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
On the Relationship between Accent Strength and Articulatory Features
by: Huang, Kevin, et al.
Published: (2025)
by: Huang, Kevin, et al.
Published: (2025)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
by: Ashvin, Aditya, et al.
Published: (2024)
by: Ashvin, Aditya, et al.
Published: (2024)
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
by: Sinha, Abhijit, et al.
Published: (2025)
by: Sinha, Abhijit, et al.
Published: (2025)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
by: Lee, Jihwan, et al.
Published: (2024)
by: Lee, Jihwan, et al.
Published: (2024)
Towards Child-Inclusive Clinical Video Understanding for Autism Spectrum Disorder
by: Kommineni, Aditya, et al.
Published: (2024)
by: Kommineni, Aditya, et al.
Published: (2024)
A long-form single-speaker real-time MRI speech dataset and benchmark
by: Foley, Sean, et al.
Published: (2025)
by: Foley, Sean, et al.
Published: (2025)
Learning-free L2-Accented Speech Generation using Phonological Rules
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
Neural Responses to Affective Sentences Reveal Signatures of Depression
by: Kommineni, Aditya, et al.
Published: (2025)
by: Kommineni, Aditya, et al.
Published: (2025)
An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
by: Lee, Jihwan, et al.
Published: (2026)
by: Lee, Jihwan, et al.
Published: (2026)
Deep Learning Characterizes Depression and Suicidal Ideation from Eye Movements
by: Avramidis, Kleanthis, et al.
Published: (2025)
by: Avramidis, Kleanthis, et al.
Published: (2025)
MMSD-Net: Towards Multi-modal Stuttering Detection
by: Nie, Liangyu, et al.
Published: (2024)
by: Nie, Liangyu, et al.
Published: (2024)
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
Neural Codecs as Biosignal Tokenizers
by: Avramidis, Kleanthis, et al.
Published: (2025)
by: Avramidis, Kleanthis, et al.
Published: (2025)
Time-Resolved EEG Decoding of Semantic Processing Reveals Altered Neural Dynamics in Depression and Suicidality
by: Jeong, Woojae, et al.
Published: (2025)
by: Jeong, Woojae, et al.
Published: (2025)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
Early Detection of Coffee Leaf Rust Through Convolutional Neural Networks Trained on Low-Resolution Images
by: Cabrera, Angelly, et al.
Published: (2024)
by: Cabrera, Angelly, et al.
Published: (2024)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
Knowledge-guided EEG Representation Learning
by: Kommineni, Aditya, et al.
Published: (2024)
by: Kommineni, Aditya, et al.
Published: (2024)
LLM Agents for Bargaining with Utility-based Feedback
by: Oh, Jihwan
Published: (2025)
by: Oh, Jihwan
Published: (2025)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
Towards Diverse Evaluation of Class Incremental Learning: A Representation Learning Perspective
by: Cha, Sungmin, et al.
Published: (2022)
by: Cha, Sungmin, et al.
Published: (2022)
Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
Goodness-of-pronunciation without phoneme time alignment
by: Wong, Jeremy H. M., et al.
Published: (2026)
by: Wong, Jeremy H. M., et al.
Published: (2026)
Informed Bootstrap Augmentation Improves EEG Decoding
by: Jeong, Woojae, et al.
Published: (2025)
by: Jeong, Woojae, et al.
Published: (2025)
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
by: Cuervo, Santiago, et al.
Published: (2021)
by: Cuervo, Santiago, et al.
Published: (2021)
IndiSeek learns information-guided disentangled representations
by: Gui, Yu, et al.
Published: (2025)
by: Gui, Yu, et al.
Published: (2025)
Superposition disentanglement of neural representations reveals hidden alignment
by: Longon, André, et al.
Published: (2025)
by: Longon, André, et al.
Published: (2025)
Partial Inverse Design of High-Performance Concrete Using Cooperative Neural Networks for Constraint-Aware Mix Generation
by: Nugraha, Agung, et al.
Published: (2025)
by: Nugraha, Agung, et al.
Published: (2025)
Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
by: Zhmoginov, Andrey, et al.
Published: (2025)
by: Zhmoginov, Andrey, et al.
Published: (2025)
Robust and Consistent Ski Rental with Distributional Advice
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Similar Items
-
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
by: Justus, Aju Ani, et al.
Published: (2026) -
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
by: Laukkanen, Anne-Maria, et al.
Published: (2024) -
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
by: Sinha, Abhijit, et al.
Published: (2025) -
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
by: Lee, Jihwan, et al.
Published: (2025) -
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025)