ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jihwan, Foley, Sean, Lertpetchpun, Thanathai, Huang, Kevin, Lee, Yoonjeong, Feng, Tiantian, Goldstein, Louis, Byrd, Dani, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Relationship between Accent Strength and Articulatory Features
by: Huang, Kevin, et al.
Published: (2025)
by: Huang, Kevin, et al.
Published: (2025)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025)
by: Park, Jay, et al.
Published: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
by: Prescott, Jordan, et al.
Published: (2026)
by: Prescott, Jordan, et al.
Published: (2026)
Articulatory Feature Prediction from Surface EMG during Speech Production
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
by: Lee, Jihwan, et al.
Published: (2026)
by: Lee, Jihwan, et al.
Published: (2026)
Learning-free L2-Accented Speech Generation using Phonological Rules
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
by: Trachu, Thanapat, et al.
Published: (2025)
by: Trachu, Thanapat, et al.
Published: (2025)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
by: Xu, Anfeng, et al.
Published: (2026)
by: Xu, Anfeng, et al.
Published: (2026)
Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
by: Lertpetchpun, Thanathai, et al.
Published: (2025)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
by: Feng, Tiantian, et al.
Published: (2026)
by: Feng, Tiantian, et al.
Published: (2026)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
by: Lee, Jihwan, et al.
Published: (2024)
by: Lee, Jihwan, et al.
Published: (2024)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
by: Lertpetchpun, Thanathai, et al.
Published: (2026)
Joint ASR and Speaker Role Tagging with Serialized Output Training
by: Xu, Anfeng, et al.
Published: (2025)
by: Xu, Anfeng, et al.
Published: (2025)
Rhythm Features for Speaker Identification
by: Mehlman, Nick, et al.
Published: (2025)
by: Mehlman, Nick, et al.
Published: (2025)
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
by: Lee, Junhyeok, et al.
Published: (2026)
by: Lee, Junhyeok, et al.
Published: (2026)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
by: Pistrosch, Simon, et al.
Published: (2026)
by: Pistrosch, Simon, et al.
Published: (2026)
Audio-visual child-adult speaker classification in dyadic interactions
by: Xu, Anfeng, et al.
Published: (2023)
by: Xu, Anfeng, et al.
Published: (2023)
Speech Codec Probing from Semantic and Phonetic Perspectives
by: Shi, Xuan, et al.
Published: (2026)
by: Shi, Xuan, et al.
Published: (2026)
Emotion-Aligned Contrastive Learning Between Images and Music
by: Stewart, Shanti, et al.
Published: (2023)
by: Stewart, Shanti, et al.
Published: (2023)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
by: Xu, Anfeng, et al.
Published: (2024)
by: Xu, Anfeng, et al.
Published: (2024)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
by: Xu, Anfeng, et al.
Published: (2024)
by: Xu, Anfeng, et al.
Published: (2024)
Deep Speech Synthesis from Multimodal Articulatory Representations
by: Wu, Peter, et al.
Published: (2024)
by: Wu, Peter, et al.
Published: (2024)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
by: Xu, Anfeng, et al.
Published: (2026)
by: Xu, Anfeng, et al.
Published: (2026)
Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal Insufficiency
by: Tabatabaee, Saba, et al.
Published: (2025)
by: Tabatabaee, Saba, et al.
Published: (2025)
Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating Nasality
by: Tabatabaee, Saba, et al.
Published: (2025)
by: Tabatabaee, Saba, et al.
Published: (2025)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
by: Lee, Chia-Yu, et al.
Published: (2026)
by: Lee, Chia-Yu, et al.
Published: (2026)
Perceptual Ratings Predict Speech Inversion Articulatory Kinematics in Childhood Speech Sound Disorders
by: Benway, Nina R., et al.
Published: (2025)
by: Benway, Nina R., et al.
Published: (2025)
Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model
by: Azzouz, Sofiane, et al.
Published: (2026)
by: Azzouz, Sofiane, et al.
Published: (2026)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
by: Feng, Tiantian, et al.
Published: (2024)
by: Feng, Tiantian, et al.
Published: (2024)
Towards a Quantitative Analysis of Coarticulation with a Phoneme-to-Articulatory Model
by: Fan, Chaofei, et al.
Published: (2024)
by: Fan, Chaofei, et al.
Published: (2024)
Similar Items
-
On the Relationship between Accent Strength and Articulatory Features
by: Huang, Kevin, et al.
Published: (2025) -
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
by: Feng, Tiantian, et al.
Published: (2025) -
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
by: Feng, Tiantian, et al.
Published: (2025) -
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025) -
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
by: Feng, Tiantian, et al.
Published: (2025)