Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jihwan, Feng, Tiantian, Kommineni, Aditya, Kadiri, Sudarsana Reddy, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
von: Laukkanen, Anne-Maria, et al.
Veröffentlicht: (2024)
von: Laukkanen, Anne-Maria, et al.
Veröffentlicht: (2024)
An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
von: Lee, Jihwan, et al.
Veröffentlicht: (2026)
von: Lee, Jihwan, et al.
Veröffentlicht: (2026)
Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
von: Kutum, Subham, et al.
Veröffentlicht: (2025)
von: Kutum, Subham, et al.
Veröffentlicht: (2025)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
von: Park, Jay, et al.
Veröffentlicht: (2025)
von: Park, Jay, et al.
Veröffentlicht: (2025)
Independent Feature Enhanced Crossmodal Fusion for Match-Mismatch Classification of Speech Stimulus and EEG Response
von: Fan, Shitong, et al.
Veröffentlicht: (2024)
von: Fan, Shitong, et al.
Veröffentlicht: (2024)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Auditory Attention Decoding from Ear-EEG Signals: A Dataset with Dynamic Attention Switching and Rigorous Cross-Validation
von: Zhang, Yuanming, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanming, et al.
Veröffentlicht: (2025)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment
von: Zhu, Haolin, et al.
Veröffentlicht: (2024)
von: Zhu, Haolin, et al.
Veröffentlicht: (2024)
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble Diffusion Models
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
von: Kim, Soowon, et al.
Veröffentlicht: (2024)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
Speech Enhancement based on cascaded two flows
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
FlowSE: Flow Matching-based Speech Enhancement
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
MMSD-Net: Towards Multi-modal Stuttering Detection
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
von: Xie, Yuying, et al.
Veröffentlicht: (2024)
von: Xie, Yuying, et al.
Veröffentlicht: (2024)
SpeechMLC: Speech Multi-label Classification
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
von: Lenz, Isabella, et al.
Veröffentlicht: (2025)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
Binaural Localization Model for Speech in Noise
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
Speech-Based Prioritization for Schizophrenia Intervention
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
Prompt-driven Target Speech Diarization
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
von: Jiang, Yidi, et al.
Veröffentlicht: (2023)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024) -
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
von: Sinha, Abhijit, et al.
Veröffentlicht: (2025) -
Evaluation of Speech Foundation Models for ASR on Child-Adult Conversations in Autism Diagnostic Sessions
von: Ashvin, Aditya, et al.
Veröffentlicht: (2024) -
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023) -
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)