NeuGPT: Unified multi-modal Neural GPT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yiqian, Duan, Yiqun, Jo, Hyejeong, Zhang, Qiang, Xu, Renjing, Jones, Oiwi Parker, Hu, Xuming, Lin, Chin-teng, Xiong, Hui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910673063116800
author Yang, Yiqian
Duan, Yiqun
Jo, Hyejeong
Zhang, Qiang
Xu, Renjing
Jones, Oiwi Parker
Hu, Xuming
Lin, Chin-teng
Xiong, Hui
author_facet Yang, Yiqian
Duan, Yiqun
Jo, Hyejeong
Zhang, Qiang
Xu, Renjing
Jones, Oiwi Parker
Hu, Xuming
Lin, Chin-teng
Xiong, Hui
contents This paper introduces NeuGPT, a groundbreaking multi-modal language generation model designed to harmonize the fragmented landscape of neural recording research. Traditionally, studies in the field have been compartmentalized by signal type, with EEG, MEG, ECoG, SEEG, fMRI, and fNIRS data being analyzed in isolation. Recognizing the untapped potential for cross-pollination and the adaptability of neural signals across varying experimental conditions, we set out to develop a unified model capable of interfacing with multiple modalities. Drawing inspiration from the success of pre-trained large models in NLP, computer vision, and speech processing, NeuGPT is architected to process a diverse array of neural recordings and interact with speech and text data. Our model mainly focus on brain-to-text decoding, improving SOTA from 6.94 to 12.92 on BLEU-1 and 6.93 to 13.06 on ROUGE-1F. It can also simulate brain signals, thereby serving as a novel neural interface. Code is available at \href{https://github.com/NeuSpeech/NeuGPT}{NeuSpeech/NeuGPT (https://github.com/NeuSpeech/NeuGPT) .}
format Preprint
id arxiv_https___arxiv_org_abs_2410_20916
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NeuGPT: Unified multi-modal Neural GPT
Yang, Yiqian
Duan, Yiqun
Jo, Hyejeong
Zhang, Qiang
Xu, Renjing
Jones, Oiwi Parker
Hu, Xuming
Lin, Chin-teng
Xiong, Hui
Computation and Language
This paper introduces NeuGPT, a groundbreaking multi-modal language generation model designed to harmonize the fragmented landscape of neural recording research. Traditionally, studies in the field have been compartmentalized by signal type, with EEG, MEG, ECoG, SEEG, fMRI, and fNIRS data being analyzed in isolation. Recognizing the untapped potential for cross-pollination and the adaptability of neural signals across varying experimental conditions, we set out to develop a unified model capable of interfacing with multiple modalities. Drawing inspiration from the success of pre-trained large models in NLP, computer vision, and speech processing, NeuGPT is architected to process a diverse array of neural recordings and interact with speech and text data. Our model mainly focus on brain-to-text decoding, improving SOTA from 6.94 to 12.92 on BLEU-1 and 6.93 to 13.06 on ROUGE-1F. It can also simulate brain signals, thereby serving as a novel neural interface. Code is available at \href{https://github.com/NeuSpeech/NeuGPT}{NeuSpeech/NeuGPT (https://github.com/NeuSpeech/NeuGPT) .}
title NeuGPT: Unified multi-modal Neural GPT
topic Computation and Language
url https://arxiv.org/abs/2410.20916