Neural2Speech: A Transfer Learning Framework for Neural-Driven Speech Reconstruction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Jiawei, Guo, Chunxu, Fu, Li, Fan, Lu, Chang, Edward F., Li, Yuanning
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911768153948160
author Li, Jiawei
Guo, Chunxu
Fu, Li
Fan, Lu
Chang, Edward F.
Li, Yuanning
author_facet Li, Jiawei
Guo, Chunxu
Fu, Li
Fan, Lu
Chang, Edward F.
Li, Yuanning
contents Reconstructing natural speech from neural activity is vital for enabling direct communication via brain-computer interfaces. Previous efforts have explored the conversion of neural recordings into speech using complex deep neural network (DNN) models trained on extensive neural recording data, which is resource-intensive under regular clinical constraints. However, achieving satisfactory performance in reconstructing speech from limited-scale neural recordings has been challenging, mainly due to the complexity of speech representations and the neural data constraints. To overcome these challenges, we propose a novel transfer learning framework for neural-driven speech reconstruction, called Neural2Speech, which consists of two distinct training phases. First, a speech autoencoder is pre-trained on readily available speech corpora to decode speech waveforms from the encoded speech representations. Second, a lightweight adaptor is trained on the small-scale neural recordings to align the neural activity and the speech representation for decoding. Remarkably, our proposed Neural2Speech demonstrates the feasibility of neural-driven speech reconstruction even with only 20 minutes of intracranial data, which significantly outperforms existing baseline methods in terms of speech fidelity and intelligibility.
format Preprint
id arxiv_https___arxiv_org_abs_2310_04644
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Neural2Speech: A Transfer Learning Framework for Neural-Driven Speech Reconstruction
Li, Jiawei
Guo, Chunxu
Fu, Li
Fan, Lu
Chang, Edward F.
Li, Yuanning
Sound
Audio and Speech Processing
Neurons and Cognition
Reconstructing natural speech from neural activity is vital for enabling direct communication via brain-computer interfaces. Previous efforts have explored the conversion of neural recordings into speech using complex deep neural network (DNN) models trained on extensive neural recording data, which is resource-intensive under regular clinical constraints. However, achieving satisfactory performance in reconstructing speech from limited-scale neural recordings has been challenging, mainly due to the complexity of speech representations and the neural data constraints. To overcome these challenges, we propose a novel transfer learning framework for neural-driven speech reconstruction, called Neural2Speech, which consists of two distinct training phases. First, a speech autoencoder is pre-trained on readily available speech corpora to decode speech waveforms from the encoded speech representations. Second, a lightweight adaptor is trained on the small-scale neural recordings to align the neural activity and the speech representation for decoding. Remarkably, our proposed Neural2Speech demonstrates the feasibility of neural-driven speech reconstruction even with only 20 minutes of intracranial data, which significantly outperforms existing baseline methods in terms of speech fidelity and intelligibility.
title Neural2Speech: A Transfer Learning Framework for Neural-Driven Speech Reconstruction
topic Sound
Audio and Speech Processing
Neurons and Cognition
url https://arxiv.org/abs/2310.04644