BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jilong, Song, Zhenxi, Wang, Jiaqi, Zhang, Meishan, Liu, Honghai, Zhang, Min, Zhang, Zhiguo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
ASM: Audio Spectrogram Mixer
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
FAST: Fast Audio Spectrogram Transformer
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
von: Naman, Anugunj, et al.
Veröffentlicht: (2025)
DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
A Practical Guide to Spectrogram Analysis for Audio Signal Processing
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2024)
von: Zhang, Li, et al.
Veröffentlicht: (2024)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
von: Fang, Yuan, et al.
Veröffentlicht: (2024)
SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
Sparse Direction of Arrival Estimation Method Based on Vector Signal Reconstruction with a Single Vector Sensor
von: Guo, Jiabin
Veröffentlicht: (2024)
von: Guo, Jiabin
Veröffentlicht: (2024)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Decoding Linguistic Representations of Human Brain
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
ECHO: Frequency-aware Hierarchical Encoding for Variable-length Signals
von: Zhang, Yucong, et al.
Veröffentlicht: (2025)
von: Zhang, Yucong, et al.
Veröffentlicht: (2025)
Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
von: Zuo, Keying, et al.
Veröffentlicht: (2024)
von: Zuo, Keying, et al.
Veröffentlicht: (2024)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
von: Fang, Zihao, et al.
Veröffentlicht: (2026)
von: Fang, Zihao, et al.
Veröffentlicht: (2026)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
A Study on Incorporating Whisper for Robust Speech Assessment
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
von: Fucci, Dennis, et al.
Veröffentlicht: (2024)
von: Fucci, Dennis, et al.
Veröffentlicht: (2024)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms
von: Wen, Penghui, et al.
Veröffentlicht: (2023)
von: Wen, Penghui, et al.
Veröffentlicht: (2023)
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task
von: Phan, Dang Thoai
Veröffentlicht: (2024)
von: Phan, Dang Thoai
Veröffentlicht: (2024)
Leveraging AM and FM Rhythm Spectrograms for Dementia Classification and Assessment
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
von: Selvaraj, Nithish Muthuchamy, et al.
Veröffentlicht: (2023)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
von: Chen, Yushen, et al.
Veröffentlicht: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Quantizing Whisper-small: How design choices affect ASR performance
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
von: Chaudhuri, Yashwardhan, et al.
Veröffentlicht: (2024)
Prompting Whisper for Joint Speech Transcription and Diarization
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026) -
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025) -
ASM: Audio Spectrogram Mixer
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024) -
FAST: Fast Audio Spectrogram Transformer
von: Naman, Anugunj, et al.
Veröffentlicht: (2025) -
DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)