Robust Dual-Modal Speech Keyword Spotting for XR Headsets
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Zhuojiang, Ma, Yuhan, Lu, Feng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
di: Nishida, Naoto, et al.
Pubblicazione: (2025)
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
di: Hopkins, Torin, et al.
Pubblicazione: (2026)
di: Hopkins, Torin, et al.
Pubblicazione: (2026)
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024)
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024)
Creating Aesthetic Sonifications on the Web with SIREN
di: Peng, Tristan, et al.
Pubblicazione: (2024)
di: Peng, Tristan, et al.
Pubblicazione: (2024)
Assessing the Viability of Wave Field Synthesis in VR-Based Cognitive Research
di: Kahl, Benjamin
Pubblicazione: (2025)
di: Kahl, Benjamin
Pubblicazione: (2025)
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
di: Huh, Mina, et al.
Pubblicazione: (2026)
di: Huh, Mina, et al.
Pubblicazione: (2026)
NeoLightning: A Modern Reimagination of Gesture-Based Sound Design
di: Kim, Yonghyun, et al.
Pubblicazione: (2025)
di: Kim, Yonghyun, et al.
Pubblicazione: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
di: Yang, Sicheng, et al.
Pubblicazione: (2024)
Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
di: Enge, Kajetan, et al.
Pubblicazione: (2024)
Towards Reliable Large Audio Language Model
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
di: Liu, Haoxuan, et al.
Pubblicazione: (2024)
di: Liu, Haoxuan, et al.
Pubblicazione: (2024)
AI TrackMate: Finally, Someone Who Will Give Your Music More Than Just "Sounds Great!"
di: Jiang, Yi-Lin, et al.
Pubblicazione: (2024)
di: Jiang, Yi-Lin, et al.
Pubblicazione: (2024)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
di: Peng, Jing, et al.
Pubblicazione: (2026)
di: Peng, Jing, et al.
Pubblicazione: (2026)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
di: Zhang, Fan, et al.
Pubblicazione: (2024)
di: Zhang, Fan, et al.
Pubblicazione: (2024)
Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
di: Bryan-Kinns, Nick, et al.
Pubblicazione: (2024)
di: Bryan-Kinns, Nick, et al.
Pubblicazione: (2024)
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
di: Selvamani, Shaja Arul, et al.
Pubblicazione: (2025)
di: Selvamani, Shaja Arul, et al.
Pubblicazione: (2025)
Workflow-Based Evaluation of Music Generation Systems
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
di: Dadman, Shayan, et al.
Pubblicazione: (2025)
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
di: Chen, Maximillian, et al.
Pubblicazione: (2026)
di: Chen, Maximillian, et al.
Pubblicazione: (2026)
Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
di: Liang, Huihong, et al.
Pubblicazione: (2025)
di: Liang, Huihong, et al.
Pubblicazione: (2025)
Evaluating Spatialized Auditory Cues for Rapid Attention Capture in XR
di: Kim, Yoonsang, et al.
Pubblicazione: (2026)
di: Kim, Yoonsang, et al.
Pubblicazione: (2026)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
di: Wang, Haoxu, et al.
Pubblicazione: (2024)
di: Wang, Haoxu, et al.
Pubblicazione: (2024)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
Acoustic Wave Modeling Using 2D FDTD: Applications in Unreal Engine For Dynamic Sound Rendering
di: Samsurya, Bilkent
Pubblicazione: (2025)
di: Samsurya, Bilkent
Pubblicazione: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
di: Sun, Licai, et al.
Pubblicazione: (2024)
di: Sun, Licai, et al.
Pubblicazione: (2024)
Soundify: Matching Sound Effects to Video
di: Lin, David Chuan-En, et al.
Pubblicazione: (2021)
di: Lin, David Chuan-En, et al.
Pubblicazione: (2021)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
di: Su, Fei, et al.
Pubblicazione: (2026)
di: Su, Fei, et al.
Pubblicazione: (2026)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
di: Wang, Hongbin, et al.
Pubblicazione: (2025)
di: Wang, Hongbin, et al.
Pubblicazione: (2025)
Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
di: Liu, Shuiyun, et al.
Pubblicazione: (2024)
di: Liu, Shuiyun, et al.
Pubblicazione: (2024)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
Revival: Collaborative Artistic Creation through Human-AI Interactions in Musical Creativity
di: Lee, Keon Ju M., et al.
Pubblicazione: (2025)
di: Lee, Keon Ju M., et al.
Pubblicazione: (2025)
Low-latency Speech Enhancement via Speech Token Generation
di: Xue, Huaying, et al.
Pubblicazione: (2023)
di: Xue, Huaying, et al.
Pubblicazione: (2023)
RespEar: Earable-Based Robust Respiratory Rate Monitoring
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
di: Li, Yue, et al.
Pubblicazione: (2024)
di: Li, Yue, et al.
Pubblicazione: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
di: Park, Seohyun, et al.
Pubblicazione: (2025)
di: Park, Seohyun, et al.
Pubblicazione: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
di: Mertes, Silvan, et al.
Pubblicazione: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
di: Xiao, Yi, et al.
Pubblicazione: (2022)
di: Xiao, Yi, et al.
Pubblicazione: (2022)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
di: Liu, Qianhui, et al.
Pubblicazione: (2024)
di: Liu, Qianhui, et al.
Pubblicazione: (2024)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
di: Nowrin, Sadia, et al.
Pubblicazione: (2024)
di: Nowrin, Sadia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025) -
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
di: Nishida, Naoto, et al.
Pubblicazione: (2025) -
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
di: Hopkins, Torin, et al.
Pubblicazione: (2026) -
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
di: Hnatyshyn, Rostyslav, et al.
Pubblicazione: (2024) -
Creating Aesthetic Sonifications on the Web with SIREN
di: Peng, Tristan, et al.
Pubblicazione: (2024)