Decoding EEG Speech Perception with Transformers and VAE-based Data Augmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Terrance Yu-Hao, Chen, Yulin, Soederhaell, Pontus, Agrawal, Sadrishya, Shapovalenko, Kateryna |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
por: Hiraki, Hirotaka, et al.
Publicado: (2024)
por: Hiraki, Hirotaka, et al.
Publicado: (2024)
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
por: Woodard, Brandon, et al.
Publicado: (2025)
por: Woodard, Brandon, et al.
Publicado: (2025)
Enhanced DareFightingICE Competitions: Sound Design and AI Competitions
por: Khan, Ibrahim, et al.
Publicado: (2024)
por: Khan, Ibrahim, et al.
Publicado: (2024)
A Penny for Your Thoughts: Decoding Speech from Inexpensive Brain Signals
por: Auster, Quentin, et al.
Publicado: (2025)
por: Auster, Quentin, et al.
Publicado: (2025)
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
por: Cho, Hyunsung, et al.
Publicado: (2024)
por: Cho, Hyunsung, et al.
Publicado: (2024)
Whisphone: Whispering Input Earbuds
por: Fukumoto, Masaaki
Publicado: (2025)
por: Fukumoto, Masaaki
Publicado: (2025)
MaskClip: Detachable Clip-on Piezoelectric Sensing of Mask Surface Vibrations for Real-time Noise-Robust Speech Input
por: Hiraki, Hirotaka, et al.
Publicado: (2025)
por: Hiraki, Hirotaka, et al.
Publicado: (2025)
Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis
por: Zhang, Pengfei, et al.
Publicado: (2026)
por: Zhang, Pengfei, et al.
Publicado: (2026)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
por: Mehta, Shivam, et al.
Publicado: (2024)
por: Mehta, Shivam, et al.
Publicado: (2024)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
por: Viswanathan, Janaki, et al.
Publicado: (2025)
por: Viswanathan, Janaki, et al.
Publicado: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
por: Wang, Shaowen, et al.
Publicado: (2025)
por: Wang, Shaowen, et al.
Publicado: (2025)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
por: Soni, Aniket Abhishek
Publicado: (2025)
por: Soni, Aniket Abhishek
Publicado: (2025)
Cross-Utterance Conditioned VAE for Speech Generation
por: Li, Yang, et al.
Publicado: (2023)
por: Li, Yang, et al.
Publicado: (2023)
Matcha-TTS: A fast TTS architecture with conditional flow matching
por: Mehta, Shivam, et al.
Publicado: (2023)
por: Mehta, Shivam, et al.
Publicado: (2023)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
por: Li, Pengcheng, et al.
Publicado: (2024)
por: Li, Pengcheng, et al.
Publicado: (2024)
A Framework for Multimodal Medical Image Interaction
por: Schütz, Laura, et al.
Publicado: (2024)
por: Schütz, Laura, et al.
Publicado: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
por: Li, Bohan, et al.
Publicado: (2024)
por: Li, Bohan, et al.
Publicado: (2024)
Non-Invasive Suicide Risk Prediction Through Speech Analysis
por: Amiriparian, Shahin, et al.
Publicado: (2024)
por: Amiriparian, Shahin, et al.
Publicado: (2024)
Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception
por: Chen, Jiaxin, et al.
Publicado: (2025)
por: Chen, Jiaxin, et al.
Publicado: (2025)
REMAST: Real-time Emotion-based Music Arrangement with Soft Transition
por: Wang, Zihao, et al.
Publicado: (2023)
por: Wang, Zihao, et al.
Publicado: (2023)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
por: Wang, Haoyu, et al.
Publicado: (2025)
por: Wang, Haoyu, et al.
Publicado: (2025)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
por: Ahn, Taekyung, et al.
Publicado: (2024)
por: Ahn, Taekyung, et al.
Publicado: (2024)
Everyday Speech in the Indian Subcontinent
por: P, Utkarsh
Publicado: (2024)
por: P, Utkarsh
Publicado: (2024)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
por: Donepudi, Dharma Teja
Publicado: (2025)
por: Donepudi, Dharma Teja
Publicado: (2025)
Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
por: Niizumi, Daisuke, et al.
Publicado: (2026)
por: Niizumi, Daisuke, et al.
Publicado: (2026)
Generation of Musical Timbres using a Text-Guided Diffusion Model
por: Yuan, Weixuan, et al.
Publicado: (2025)
por: Yuan, Weixuan, et al.
Publicado: (2025)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
por: Mahfi, Muntahi Safwan, et al.
Publicado: (2025)
por: Mahfi, Muntahi Safwan, et al.
Publicado: (2025)
Measuring the Accuracy of Automatic Speech Recognition Solutions
por: Kuhn, Korbinian, et al.
Publicado: (2024)
por: Kuhn, Korbinian, et al.
Publicado: (2024)
Sonify Anything: Towards Context-Aware Sonic Interactions in AR
por: Schütz, Laura, et al.
Publicado: (2025)
por: Schütz, Laura, et al.
Publicado: (2025)
Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
por: Chary, Luis Felipe, et al.
Publicado: (2025)
por: Chary, Luis Felipe, et al.
Publicado: (2025)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
por: Bhadra, Dipayan, et al.
Publicado: (2025)
por: Bhadra, Dipayan, et al.
Publicado: (2025)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
por: McKinnon, Max, et al.
Publicado: (2026)
por: McKinnon, Max, et al.
Publicado: (2026)
ToMoBrush: Exploring Dental Health Sensing using a Sonic Toothbrush
por: Yuan, Kuang, et al.
Publicado: (2024)
por: Yuan, Kuang, et al.
Publicado: (2024)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
por: Robertson, Sean, et al.
Publicado: (2023)
por: Robertson, Sean, et al.
Publicado: (2023)
Compositional Phoneme Approximation for L1-Grounded L2 Pronunciation Training
por: Park, Jisang, et al.
Publicado: (2024)
por: Park, Jisang, et al.
Publicado: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
por: Deng, Keqi, et al.
Publicado: (2025)
por: Deng, Keqi, et al.
Publicado: (2025)
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
por: Shi, Hao, et al.
Publicado: (2023)
por: Shi, Hao, et al.
Publicado: (2023)
Ejemplares similares
-
WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
por: Hiraki, Hirotaka, et al.
Publicado: (2024) -
AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality
por: Woodard, Brandon, et al.
Publicado: (2025) -
Enhanced DareFightingICE Competitions: Sound Design and AI Competitions
por: Khan, Ibrahim, et al.
Publicado: (2024) -
A Penny for Your Thoughts: Decoding Speech from Inexpensive Brain Signals
por: Auster, Quentin, et al.
Publicado: (2025) -
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
por: Cho, Hyunsung, et al.
Publicado: (2024)