NeuSpeech: Decode Neural signal as Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Yiqian, Duan, Yiqun, Zhang, Qiang, Jo, Hyejeong, Zhou, Jinni, Lee, Won Hee, Xu, Renjing, Xiong, Hui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MAD: Multi-Alignment MEG-to-Text Decoding
di: Yang, Yiqian, et al.
Pubblicazione: (2024)
di: Yang, Yiqian, et al.
Pubblicazione: (2024)
Are EEG-to-Text Models Working?
di: Jo, Hyejeong, et al.
Pubblicazione: (2024)
di: Jo, Hyejeong, et al.
Pubblicazione: (2024)
NeuGPT: Unified multi-modal Neural GPT
di: Yang, Yiqian, et al.
Pubblicazione: (2024)
di: Yang, Yiqian, et al.
Pubblicazione: (2024)
Prompting Multi-Modal Tokens to Enhance End-to-End Autonomous Driving Imitation Learning with LLMs
di: Duan, Yiqun, et al.
Pubblicazione: (2024)
di: Duan, Yiqun, et al.
Pubblicazione: (2024)
NeuGaze: Reshaping the future BCI
di: Yang, Yiqian
Pubblicazione: (2025)
di: Yang, Yiqian
Pubblicazione: (2025)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
di: Hyeon, Sieun, et al.
Pubblicazione: (2024)
di: Hyeon, Sieun, et al.
Pubblicazione: (2024)
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
di: Zhao, Xingjian, et al.
Pubblicazione: (2025)
di: Zhao, Xingjian, et al.
Pubblicazione: (2025)
Measuring the Redundancy of Decoder Layers in SpeechLLMs
di: Moumen, Adel, et al.
Pubblicazione: (2026)
di: Moumen, Adel, et al.
Pubblicazione: (2026)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
di: Zhang, Wei, et al.
Pubblicazione: (2025)
di: Zhang, Wei, et al.
Pubblicazione: (2025)
Decoding Emotion: Speech Perception Patterns in Individuals with Self-reported Depression
di: Vats, Guneesh, et al.
Pubblicazione: (2024)
di: Vats, Guneesh, et al.
Pubblicazione: (2024)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
di: Kim, Haechan, et al.
Pubblicazione: (2026)
di: Kim, Haechan, et al.
Pubblicazione: (2026)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
di: Lee, Jung-Sun, et al.
Pubblicazione: (2024)
di: Lee, Jung-Sun, et al.
Pubblicazione: (2024)
SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
di: Yang, Wanqi, et al.
Pubblicazione: (2025)
di: Yang, Wanqi, et al.
Pubblicazione: (2025)
UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions
di: Han, Wenkang, et al.
Pubblicazione: (2025)
di: Han, Wenkang, et al.
Pubblicazione: (2025)
Pretraining Large Brain Language Model for Active BCI: Silent Speech
di: Zhou, Jinzhao, et al.
Pubblicazione: (2025)
di: Zhou, Jinzhao, et al.
Pubblicazione: (2025)
R-BI: Regularized Batched Inputs enhance Incremental Decoding Framework for Low-Latency Simultaneous Speech Translation
di: Guo, Jiaxin, et al.
Pubblicazione: (2024)
di: Guo, Jiaxin, et al.
Pubblicazione: (2024)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
di: Lee, Wonjun, et al.
Pubblicazione: (2026)
di: Lee, Wonjun, et al.
Pubblicazione: (2026)
Raon-Speech Technical Report
di: Kim, Beomsoo, et al.
Pubblicazione: (2026)
di: Kim, Beomsoo, et al.
Pubblicazione: (2026)
Toward Robust EEG-based Intention Decoding during Misarticulated Speech in Dysarthria
di: Jo, Ha-Na, et al.
Pubblicazione: (2025)
di: Jo, Ha-Na, et al.
Pubblicazione: (2025)
Speech to Speech Translation with Translatotron: A State of the Art Review
di: Kala, Jules R., et al.
Pubblicazione: (2025)
di: Kala, Jules R., et al.
Pubblicazione: (2025)
Efficient Training for Cross-lingual Speech Language Models
di: Zhou, Yan, et al.
Pubblicazione: (2026)
di: Zhou, Yan, et al.
Pubblicazione: (2026)
NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering
di: Cao, Ruisheng, et al.
Pubblicazione: (2025)
di: Cao, Ruisheng, et al.
Pubblicazione: (2025)
Alternative Speech: Complementary Method to Counter-Narrative for Better Discourse
di: Lee, Seungyoon, et al.
Pubblicazione: (2024)
di: Lee, Seungyoon, et al.
Pubblicazione: (2024)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
di: Kim, Yumin, et al.
Pubblicazione: (2025)
di: Kim, Yumin, et al.
Pubblicazione: (2025)
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection
di: Sun, Runze, et al.
Pubblicazione: (2026)
di: Sun, Runze, et al.
Pubblicazione: (2026)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
di: Cettolo, Mauro, et al.
Pubblicazione: (2025)
di: Cettolo, Mauro, et al.
Pubblicazione: (2025)
StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation
di: Chen, Xi, et al.
Pubblicazione: (2025)
di: Chen, Xi, et al.
Pubblicazione: (2025)
From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation
di: Pan, Yu, et al.
Pubblicazione: (2026)
di: Pan, Yu, et al.
Pubblicazione: (2026)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models
di: Xiang, Bajian, et al.
Pubblicazione: (2026)
di: Xiang, Bajian, et al.
Pubblicazione: (2026)
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
di: Serai, Prashant, et al.
Pubblicazione: (2024)
di: Serai, Prashant, et al.
Pubblicazione: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
di: Zhang, Shaolei, et al.
Pubblicazione: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
di: Du, Yuhao, et al.
Pubblicazione: (2025)
di: Du, Yuhao, et al.
Pubblicazione: (2025)
DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
di: Yang, Jianing, et al.
Pubblicazione: (2026)
di: Yang, Jianing, et al.
Pubblicazione: (2026)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
FASST: Fast LLM-based Simultaneous Speech Translation
di: Ouyang, Siqi, et al.
Pubblicazione: (2024)
di: Ouyang, Siqi, et al.
Pubblicazione: (2024)
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
di: Chen, Sirry, et al.
Pubblicazione: (2026)
di: Chen, Sirry, et al.
Pubblicazione: (2026)
How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena
di: Gaido, Marco, et al.
Pubblicazione: (2024)
di: Gaido, Marco, et al.
Pubblicazione: (2024)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
di: Guo, Rongchen, et al.
Pubblicazione: (2025)
di: Guo, Rongchen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MAD: Multi-Alignment MEG-to-Text Decoding
di: Yang, Yiqian, et al.
Pubblicazione: (2024) -
Are EEG-to-Text Models Working?
di: Jo, Hyejeong, et al.
Pubblicazione: (2024) -
NeuGPT: Unified multi-modal Neural GPT
di: Yang, Yiqian, et al.
Pubblicazione: (2024) -
Prompting Multi-Modal Tokens to Enhance End-to-End Autonomous Driving Imitation Learning with LLMs
di: Duan, Yiqun, et al.
Pubblicazione: (2024) -
NeuGaze: Reshaping the future BCI
di: Yang, Yiqian
Pubblicazione: (2025)