Decoding EEG Speech Perception with Transformers and VAE-based Data Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Terrance Yu-Hao, Chen, Yulin, Soederhaell, Pontus, Agrawal, Sadrishya, Shapovalenko, Kateryna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
von: Hiraki, Hirotaka, et al.
Veröffentlicht: (2024)
von: Hiraki, Hirotaka, et al.
Veröffentlicht: (2024)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
von: Woodard, Brandon, et al.
Veröffentlicht: (2025)
von: Woodard, Brandon, et al.
Veröffentlicht: (2025)
Enhanced DareFightingICE Competitions: Sound Design and AI Competitions
von: Khan, Ibrahim, et al.
Veröffentlicht: (2024)
von: Khan, Ibrahim, et al.
Veröffentlicht: (2024)
A Penny for Your Thoughts: Decoding Speech from Inexpensive Brain Signals
von: Auster, Quentin, et al.
Veröffentlicht: (2025)
von: Auster, Quentin, et al.
Veröffentlicht: (2025)
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024)
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024)
Whisphone: Whispering Input Earbuds
von: Fukumoto, Masaaki
Veröffentlicht: (2025)
von: Fukumoto, Masaaki
Veröffentlicht: (2025)
MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface Vibrations for Real-time Noise-Robust Speech Input
von: Hiraki, Hirotaka, et al.
Veröffentlicht: (2025)
von: Hiraki, Hirotaka, et al.
Veröffentlicht: (2025)
Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis
von: Zhang, Pengfei, et al.
Veröffentlicht: (2026)
von: Zhang, Pengfei, et al.
Veröffentlicht: (2026)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
von: Mehta, Shivam, et al.
Veröffentlicht: (2024)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
Cross-Utterance Conditioned VAE for Speech Generation
von: Li, Yang, et al.
Veröffentlicht: (2023)
von: Li, Yang, et al.
Veröffentlicht: (2023)
Matcha-TTS: A fast TTS architecture with conditional flow matching
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
von: Mehta, Shivam, et al.
Veröffentlicht: (2023)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
A Framework for Multimodal Medical Image Interaction
von: Schütz, Laura, et al.
Veröffentlicht: (2024)
von: Schütz, Laura, et al.
Veröffentlicht: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
von: Mehta, Shivam, et al.
Veröffentlicht: (2025)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
Non-Invasive Suicide Risk Prediction Through Speech Analysis
von: Amiriparian, Shahin, et al.
Veröffentlicht: (2024)
von: Amiriparian, Shahin, et al.
Veröffentlicht: (2024)
Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception
von: Chen, Jiaxin, et al.
Veröffentlicht: (2025)
von: Chen, Jiaxin, et al.
Veröffentlicht: (2025)
REMAST: Real-time Emotion-based Music Arrangement with Soft Transition
von: Wang, Zihao, et al.
Veröffentlicht: (2023)
von: Wang, Zihao, et al.
Veröffentlicht: (2023)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
Everyday Speech in the Indian Subcontinent
von: P, Utkarsh
Veröffentlicht: (2024)
von: P, Utkarsh
Veröffentlicht: (2024)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2026)
Generation of Musical Timbres using a Text-Guided Diffusion Model
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
von: Mahfi, Muntahi Safwan, et al.
Veröffentlicht: (2025)
Measuring the Accuracy of Automatic Speech Recognition Solutions
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Sonify Anything: Towards Context-Aware Sonic Interactions in AR
von: Schütz, Laura, et al.
Veröffentlicht: (2025)
von: Schütz, Laura, et al.
Veröffentlicht: (2025)
Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
von: Chary, Luis Felipe, et al.
Veröffentlicht: (2025)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
von: McKinnon, Max, et al.
Veröffentlicht: (2026)
ToMoBrush: Exploring Dental Health Sensing using a Sonic Toothbrush
von: Yuan, Kuang, et al.
Veröffentlicht: (2024)
von: Yuan, Kuang, et al.
Veröffentlicht: (2024)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
Compositional Phoneme Approximation for L1-Grounded L2 Pronunciation Training
von: Park, Jisang, et al.
Veröffentlicht: (2024)
von: Park, Jisang, et al.
Veröffentlicht: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
von: Shi, Hao, et al.
Veröffentlicht: (2023)
von: Shi, Hao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
von: Hiraki, Hirotaka, et al.
Veröffentlicht: (2024) -
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
von: Woodard, Brandon, et al.
Veröffentlicht: (2025) -
Enhanced DareFightingICE Competitions: Sound Design and AI Competitions
von: Khan, Ibrahim, et al.
Veröffentlicht: (2024) -
A Penny for Your Thoughts: Decoding Speech from Inexpensive Brain Signals
von: Auster, Quentin, et al.
Veröffentlicht: (2025) -
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
von: Cho, Hyunsung, et al.
Veröffentlicht: (2024)