A long-form single-speaker real-time MRI speech dataset and benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Foley, Sean, Lee, Jihwan, Huang, Kevin, Shi, Xuan, Lee, Yoonjeong, Goldstein, Louis, Narayanan, Shrikanth |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
por: Park, Jay, et al.
Publicado: (2025)
por: Park, Jay, et al.
Publicado: (2025)
On the Relationship between Accent Strength and Articulatory Features
por: Huang, Kevin, et al.
Publicado: (2025)
por: Huang, Kevin, et al.
Publicado: (2025)
Articulatory Feature Prediction from Surface EMG during Speech Production
por: Lee, Jihwan, et al.
Publicado: (2025)
por: Lee, Jihwan, et al.
Publicado: (2025)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
por: Nguyen, Hong, et al.
Publicado: (2024)
por: Nguyen, Hong, et al.
Publicado: (2024)
Audio-visual child-adult speaker classification in dyadic interactions
por: Xu, Anfeng, et al.
Publicado: (2023)
por: Xu, Anfeng, et al.
Publicado: (2023)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
por: Lee, Jihwan, et al.
Publicado: (2025)
por: Lee, Jihwan, et al.
Publicado: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
por: Feng, Tiantian, et al.
Publicado: (2024)
por: Feng, Tiantian, et al.
Publicado: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
por: Lee, Jihwan, et al.
Publicado: (2024)
por: Lee, Jihwan, et al.
Publicado: (2024)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
por: Feng, Tiantian, et al.
Publicado: (2026)
por: Feng, Tiantian, et al.
Publicado: (2026)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
por: Feng, Tiantian, et al.
Publicado: (2023)
por: Feng, Tiantian, et al.
Publicado: (2023)
Multi-channel multi-speaker transformer for speech recognition
por: Yifan, Guo, et al.
Publicado: (2026)
por: Yifan, Guo, et al.
Publicado: (2026)
Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
por: Wang, Shih-Heng, et al.
Publicado: (2026)
por: Wang, Shih-Heng, et al.
Publicado: (2026)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
por: Feng, Tiantian, et al.
Publicado: (2024)
por: Feng, Tiantian, et al.
Publicado: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
por: Kwak, Doyeop, et al.
Publicado: (2026)
por: Kwak, Doyeop, et al.
Publicado: (2026)
Multi-speaker Text-to-speech Training with Speaker Anonymized Data
por: Huang, Wen-Chin, et al.
Publicado: (2024)
por: Huang, Wen-Chin, et al.
Publicado: (2024)
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
por: Foley, Sean, et al.
Publicado: (2025)
por: Foley, Sean, et al.
Publicado: (2025)
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
por: Lee, Jaejun, et al.
Publicado: (2026)
por: Lee, Jaejun, et al.
Publicado: (2026)
Learning-free L2-Accented Speech Generation using Phonological Rules
por: Lertpetchpun, Thanathai, et al.
Publicado: (2026)
por: Lertpetchpun, Thanathai, et al.
Publicado: (2026)
Joint ASR and Speaker Role Tagging with Serialized Output Training
por: Xu, Anfeng, et al.
Publicado: (2025)
por: Xu, Anfeng, et al.
Publicado: (2025)
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
por: Prescott, Jordan, et al.
Publicado: (2026)
por: Prescott, Jordan, et al.
Publicado: (2026)
Quantifying the effect of speech pathology on automatic and human speaker verification
por: Halpern, Bence Mark, et al.
Publicado: (2024)
por: Halpern, Bence Mark, et al.
Publicado: (2024)
VorTEX: Various overlap ratio for Target speech EXtraction
por: Oh, Ro-hoon, et al.
Publicado: (2026)
por: Oh, Ro-hoon, et al.
Publicado: (2026)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
por: Feng, Tiantian, et al.
Publicado: (2024)
por: Feng, Tiantian, et al.
Publicado: (2024)
Online speaker diarization of meetings guided by speech separation
por: Gruttadauria, Elio, et al.
Publicado: (2024)
por: Gruttadauria, Elio, et al.
Publicado: (2024)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
por: Feng, Tiantian, et al.
Publicado: (2025)
por: Feng, Tiantian, et al.
Publicado: (2025)
End-to-end multi-channel speaker extraction and binaural speech synthesis
por: Chi, Cheng, et al.
Publicado: (2024)
por: Chi, Cheng, et al.
Publicado: (2024)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
por: Pistrosch, Simon, et al.
Publicado: (2026)
por: Pistrosch, Simon, et al.
Publicado: (2026)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
por: Zhang, Yiru, et al.
Publicado: (2025)
por: Zhang, Yiru, et al.
Publicado: (2025)
Emotion-Aligned Contrastive Learning Between Images and Music
por: Stewart, Shanti, et al.
Publicado: (2023)
por: Stewart, Shanti, et al.
Publicado: (2023)
HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
por: Kacprzak, Stanisław, et al.
Publicado: (2024)
por: Kacprzak, Stanisław, et al.
Publicado: (2024)
Phone Duration Modeling for Speaker Age Estimation in Children
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2021)
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2021)
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
por: Justus, Aju Ani, et al.
Publicado: (2026)
por: Justus, Aju Ani, et al.
Publicado: (2026)
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
por: Lertpetchpun, Thanathai, et al.
Publicado: (2026)
por: Lertpetchpun, Thanathai, et al.
Publicado: (2026)
Context-aware child-directed speech detection from long-form recordings
por: Charlot, Théo, et al.
Publicado: (2026)
por: Charlot, Théo, et al.
Publicado: (2026)
VoxCare: Studying Natural Communication Behaviors of Hospital Caregivers through Wearable Sensing of Egocentric Audio
por: Feng, Tiantian, et al.
Publicado: (2026)
por: Feng, Tiantian, et al.
Publicado: (2026)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
por: Shi, Zhonghao, et al.
Publicado: (2024)
por: Shi, Zhonghao, et al.
Publicado: (2024)
WhisperFlow: speech foundation models in real time
por: Wang, Rongxiang, et al.
Publicado: (2024)
por: Wang, Rongxiang, et al.
Publicado: (2024)
ModalityMirror: Improving Audio Classification in Modality Heterogeneity Federated Learning with Multimodal Distillation
por: Feng, Tiantian, et al.
Publicado: (2024)
por: Feng, Tiantian, et al.
Publicado: (2024)
Ejemplares similares
-
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
por: Park, Jay, et al.
Publicado: (2025) -
On the Relationship between Accent Strength and Articulatory Features
por: Huang, Kevin, et al.
Publicado: (2025) -
Articulatory Feature Prediction from Surface EMG during Speech Production
por: Lee, Jihwan, et al.
Publicado: (2025) -
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
por: Feng, Tiantian, et al.
Publicado: (2025) -
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
por: Nguyen, Hong, et al.
Publicado: (2024)