A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Benster, Tyler, Wilson, Guy, Elisha, Reshef, Willett, Francis R, Druckmann, Shaul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
por: Yu, Luca Jiang-Tao, et al.
Publicado: (2024)
por: Yu, Luca Jiang-Tao, et al.
Publicado: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
por: Chen, Youjun, et al.
Publicado: (2025)
por: Chen, Youjun, et al.
Publicado: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
por: Feng, Tiantian, et al.
Publicado: (2023)
por: Feng, Tiantian, et al.
Publicado: (2023)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
por: Nowrin, Sadia, et al.
Publicado: (2024)
por: Nowrin, Sadia, et al.
Publicado: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
por: Dutta, Satwik, et al.
Publicado: (2025)
por: Dutta, Satwik, et al.
Publicado: (2025)
DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality
por: Choi, Youngwon, et al.
Publicado: (2025)
por: Choi, Youngwon, et al.
Publicado: (2025)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
por: Cai, Zhuojiang, et al.
Publicado: (2024)
por: Cai, Zhuojiang, et al.
Publicado: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
por: Nishida, Naoto, et al.
Publicado: (2025)
por: Nishida, Naoto, et al.
Publicado: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
por: Mishra, Ruchik, et al.
Publicado: (2024)
por: Mishra, Ruchik, et al.
Publicado: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
por: Li, Yue, et al.
Publicado: (2024)
por: Li, Yue, et al.
Publicado: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
por: Park, Seohyun, et al.
Publicado: (2025)
por: Park, Seohyun, et al.
Publicado: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
por: Wang, Hongbin, et al.
Publicado: (2025)
por: Wang, Hongbin, et al.
Publicado: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
por: Mertes, Silvan, et al.
Publicado: (2024)
por: Mertes, Silvan, et al.
Publicado: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
por: Xiao, Yi, et al.
Publicado: (2022)
por: Xiao, Yi, et al.
Publicado: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
por: Khanday, Owais Mujtaba, et al.
Publicado: (2025)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
por: Zhao, Zhixian, et al.
Publicado: (2024)
por: Zhao, Zhixian, et al.
Publicado: (2024)
Human Feedback Driven Dynamic Speech Emotion Recognition
por: Fedorov, Ilya, et al.
Publicado: (2025)
por: Fedorov, Ilya, et al.
Publicado: (2025)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
por: Liu, Ailin, et al.
Publicado: (2024)
por: Liu, Ailin, et al.
Publicado: (2024)
Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
por: Hou, Yuanbo, et al.
Publicado: (2024)
por: Hou, Yuanbo, et al.
Publicado: (2024)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
por: Chang, Yi, et al.
Publicado: (2024)
por: Chang, Yi, et al.
Publicado: (2024)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
por: Sankey-Olsen, Cuno, et al.
Publicado: (2025)
Enhancing DMI Interactions by Integrating Haptic Feedback for Intricate Vibrato Technique
por: Piao, Ziyue, et al.
Publicado: (2024)
por: Piao, Ziyue, et al.
Publicado: (2024)
Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
por: Liu, Shuiyun, et al.
Publicado: (2024)
por: Liu, Shuiyun, et al.
Publicado: (2024)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
por: Ma, Yong, et al.
Publicado: (2025)
por: Ma, Yong, et al.
Publicado: (2025)
Silent Speech Sentence Recognition with Six-Axis Accelerometers using Conformer and CTC Algorithm
por: Xie, Yudong, et al.
Publicado: (2025)
por: Xie, Yudong, et al.
Publicado: (2025)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
por: Han, Zhichen, et al.
Publicado: (2024)
por: Han, Zhichen, et al.
Publicado: (2024)
Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
por: Tiwari, Upasana, et al.
Publicado: (2025)
por: Tiwari, Upasana, et al.
Publicado: (2025)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
por: Sharma, Roshan, et al.
Publicado: (2024)
por: Sharma, Roshan, et al.
Publicado: (2024)
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
por: Kuhn, Korbinian, et al.
Publicado: (2025)
por: Kuhn, Korbinian, et al.
Publicado: (2025)
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
por: Li, Yinan, et al.
Publicado: (2026)
por: Li, Yinan, et al.
Publicado: (2026)
A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
por: Zheng, Shuoyang, et al.
Publicado: (2024)
por: Zheng, Shuoyang, et al.
Publicado: (2024)
A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings
por: Han, Hyewon, et al.
Publicado: (2024)
por: Han, Hyewon, et al.
Publicado: (2024)
Interactive Sonification for Health and Energy using ChucK and Unity
por: Zhao, Yichun, et al.
Publicado: (2024)
por: Zhao, Yichun, et al.
Publicado: (2024)
Seeing Beyond Sound: Visualization and Abstraction in Audio Data Representation
por: Blum'e, Ashlae
Publicado: (2025)
por: Blum'e, Ashlae
Publicado: (2025)
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
por: Manukalpa, J. M. Chan Sri, et al.
Publicado: (2025)
por: Manukalpa, J. M. Chan Sri, et al.
Publicado: (2025)
Interfacing with history: Curating with audio augmented objects
por: Cliffe, Laurence
Publicado: (2024)
por: Cliffe, Laurence
Publicado: (2024)
Transhuman Ansambl - Voice Beyond Language
por: Ivsic, Lucija, et al.
Publicado: (2024)
por: Ivsic, Lucija, et al.
Publicado: (2024)
Ejemplares similares
-
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
por: Yu, Luca Jiang-Tao, et al.
Publicado: (2024) -
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
por: Chen, Youjun, et al.
Publicado: (2025) -
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
por: Zhou, Dongliang, et al.
Publicado: (2025) -
Directional Source Separation for Robust Speech Recognition on Smart Glasses
por: Feng, Tiantian, et al.
Publicado: (2023) -
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
por: Nowrin, Sadia, et al.
Publicado: (2024)