Brain-to-Speech: Prosody Feature Engineering and Transformer-Based Reconstruction
Fuente:
arXiv
Saved in:
| Main Authors: | Al-Radhi, Mohammed Salah, Németh, Géza, Tchechmedjiev, Andon, Xu, Binbin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
by: Al-Radhi, Mohammed Salah, et al.
Published: (2025)
by: Al-Radhi, Mohammed Salah, et al.
Published: (2025)
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
by: Al-Radhi, Mohammed Salah, et al.
Published: (2026)
by: Al-Radhi, Mohammed Salah, et al.
Published: (2026)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
by: Zaiem, Salah, et al.
Published: (2023)
by: Zaiem, Salah, et al.
Published: (2023)
Real-Time Streamable Generative Speech Restoration with Flow Matching
by: Welker, Simon, et al.
Published: (2025)
by: Welker, Simon, et al.
Published: (2025)
A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
by: Gupta, Kishan, et al.
Published: (2022)
by: Gupta, Kishan, et al.
Published: (2022)
Speech Watermarking with Discrete Intermediate Representations
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
by: Parker, Julian D, et al.
Published: (2024)
by: Parker, Julian D, et al.
Published: (2024)
Reconstruction of Sound Field through Diffusion Models
by: Miotello, Federico, et al.
Published: (2023)
by: Miotello, Federico, et al.
Published: (2023)
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
by: Mandal, Atanu, et al.
Published: (2024)
by: Mandal, Atanu, et al.
Published: (2024)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
by: Sinha, Abhijit, et al.
Published: (2025)
by: Sinha, Abhijit, et al.
Published: (2025)
A Convolutional Framework for Mapping Imagined Auditory MEG into Listened Brain Responses
by: Maghsoudi, Maryam, et al.
Published: (2025)
by: Maghsoudi, Maryam, et al.
Published: (2025)
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
by: Xu, Zhongweiyang, et al.
Published: (2025)
by: Xu, Zhongweiyang, et al.
Published: (2025)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
by: Baoueb, Teysir, et al.
Published: (2025)
by: Baoueb, Teysir, et al.
Published: (2025)
Resource-Efficient Separation Transformer
by: Della Libera, Luca, et al.
Published: (2022)
by: Della Libera, Luca, et al.
Published: (2022)
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
by: Baoueb, Teysir, et al.
Published: (2024)
by: Baoueb, Teysir, et al.
Published: (2024)
Lightweight DNN for Full-Band Speech Denoising on Mobile Devices: Exploiting Long and Short Temporal Patterns
by: Drossos, Konstantinos, et al.
Published: (2025)
by: Drossos, Konstantinos, et al.
Published: (2025)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
by: Kawamura, Masaya, et al.
Published: (2025)
by: Kawamura, Masaya, et al.
Published: (2025)
Room Transfer Function Reconstruction Using Complex-valued Neural Networks and Irregularly Distributed Microphones
by: Ronchini, Francesca, et al.
Published: (2024)
by: Ronchini, Francesca, et al.
Published: (2024)
MaskSR: Masked Language Model for Full-band Speech Restoration
by: Li, Xu, et al.
Published: (2024)
by: Li, Xu, et al.
Published: (2024)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
by: Ziogas, Ioannis, et al.
Published: (2024)
by: Ziogas, Ioannis, et al.
Published: (2024)
Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features
by: Meng, Hanyu, et al.
Published: (2024)
by: Meng, Hanyu, et al.
Published: (2024)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
by: Sankar, Ashwin, et al.
Published: (2024)
by: Sankar, Ashwin, et al.
Published: (2024)
Learnable Adaptive Time-Frequency Representation via Differentiable Short-Time Fourier Transform
by: Leiber, Maxime, et al.
Published: (2025)
by: Leiber, Maxime, et al.
Published: (2025)
Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels
by: Bedir, Oguz, et al.
Published: (2025)
by: Bedir, Oguz, et al.
Published: (2025)
Time Series Diffusion Method: A Denoising Diffusion Probabilistic Model for Vibration Signal Generation
by: Yi, Haiming, et al.
Published: (2023)
by: Yi, Haiming, et al.
Published: (2023)
EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG
by: Park, Hanbeot, et al.
Published: (2025)
by: Park, Hanbeot, et al.
Published: (2025)
Respiratory Disease Classification and Biometric Analysis Using Biosignals from Digital Stethoscopes
by: Casado, Constantino Álvarez, et al.
Published: (2023)
by: Casado, Constantino Álvarez, et al.
Published: (2023)
Scaling to Multimodal and Multichannel Heart Sound Classification with Synthetic and Augmented Biosignals
by: Marocchi, Milan, et al.
Published: (2025)
by: Marocchi, Milan, et al.
Published: (2025)
O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization
by: Gruttadauria, Elio, et al.
Published: (2025)
by: Gruttadauria, Elio, et al.
Published: (2025)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
by: Kwon, Younghoo, et al.
Published: (2024)
by: Kwon, Younghoo, et al.
Published: (2024)
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
Deep Active Speech Cancellation with Mamba-Masking Network
by: Mishaly, Yehuda, et al.
Published: (2025)
by: Mishaly, Yehuda, et al.
Published: (2025)
Joint Source-Environment Adaptation for Deep Learning-Based Underwater Acoustic Source Ranging
by: Kari, Dariush, et al.
Published: (2025)
by: Kari, Dariush, et al.
Published: (2025)
PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
by: Hono, Yukiya, et al.
Published: (2024)
by: Hono, Yukiya, et al.
Published: (2024)
A Physics-Informed Neural Network-Based Approach for the Spatial Upsampling of Spherical Microphone Arrays
by: Miotello, Federico, et al.
Published: (2024)
by: Miotello, Federico, et al.
Published: (2024)
Joint Source-Environment Adaptation of Data-Driven Underwater Acoustic Source Ranging Based on Model Uncertainty
by: Kari, Dariush, et al.
Published: (2025)
by: Kari, Dariush, et al.
Published: (2025)
Is Attention always needed? A Case Study on Language Identification from Speech
by: Mandal, Atanu, et al.
Published: (2021)
by: Mandal, Atanu, et al.
Published: (2021)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
by: Wang, Yuancheng, et al.
Published: (2025)
by: Wang, Yuancheng, et al.
Published: (2025)
Condition-Invariant fMRI Decoding of Speech Intelligibility with Deep State Space Model
by: Sung, Ching-Chih, et al.
Published: (2025)
by: Sung, Ching-Chih, et al.
Published: (2025)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
by: Qi, Tianhua, et al.
Published: (2024)
by: Qi, Tianhua, et al.
Published: (2024)
Similar Items
-
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
by: Al-Radhi, Mohammed Salah, et al.
Published: (2025) -
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
by: Al-Radhi, Mohammed Salah, et al.
Published: (2026) -
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
by: Zaiem, Salah, et al.
Published: (2023) -
Real-Time Streamable Generative Speech Restoration with Flow Matching
by: Welker, Simon, et al.
Published: (2025) -
A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain
by: Gupta, Kishan, et al.
Published: (2022)