Saved in:
| Main Authors: | Dementyev, Artem, Kanevsky, Dimitri, Yang, Samuel J., Parvaix, Mathieu, Lai, Chiong, Olwal, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.08848 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
by: Hiraki, Hirotaka, et al.
Published: (2024)
by: Hiraki, Hirotaka, et al.
Published: (2024)
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
by: Jaiswal, Pranay, et al.
Published: (2024)
by: Jaiswal, Pranay, et al.
Published: (2024)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
Improved Dysarthric Speech to Text Conversion via TTS Personalization
by: Mihajlik, Péter, et al.
Published: (2025)
by: Mihajlik, Péter, et al.
Published: (2025)
LightBeam: An Accurate and Memory-Efficient CTC Decoder for Speech Neuroprostheses
by: Feghhi, Ebrahim, et al.
Published: (2026)
by: Feghhi, Ebrahim, et al.
Published: (2026)
EvolveCaptions: Empowering DHH Users Through Real-Time Collaborative Captioning
by: Wu, Liang-Yuan, et al.
Published: (2025)
by: Wu, Liang-Yuan, et al.
Published: (2025)
$R$-equivalence on Cubic Surfaces I: Existing Cases with Non-Trivial Universal Equivalence
by: Kanevsky, Dimitri, et al.
Published: (2026)
by: Kanevsky, Dimitri, et al.
Published: (2026)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
by: Yuan, Kuang, et al.
Published: (2025)
by: Yuan, Kuang, et al.
Published: (2025)
Exploring Gender Bias in Alzheimer's Disease Detection: Insights from Mandarin and Greek Speech Perception
by: He, Liu, et al.
Published: (2025)
by: He, Liu, et al.
Published: (2025)
EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition
by: Paul, Durjoy Chandra, et al.
Published: (2025)
by: Paul, Durjoy Chandra, et al.
Published: (2025)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
by: Park, Seohyun, et al.
Published: (2025)
by: Park, Seohyun, et al.
Published: (2025)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
by: Benster, Tyler, et al.
Published: (2024)
by: Benster, Tyler, et al.
Published: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
by: Wang, Hongbin, et al.
Published: (2025)
by: Wang, Hongbin, et al.
Published: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
by: Mertes, Silvan, et al.
Published: (2024)
by: Mertes, Silvan, et al.
Published: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
by: Xiao, Yi, et al.
Published: (2022)
by: Xiao, Yi, et al.
Published: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
by: Chen, Youjun, et al.
Published: (2025)
by: Chen, Youjun, et al.
Published: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
by: Nowrin, Sadia, et al.
Published: (2024)
by: Nowrin, Sadia, et al.
Published: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
by: Dutta, Satwik, et al.
Published: (2025)
by: Dutta, Satwik, et al.
Published: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
by: Liu, Ailin, et al.
Published: (2024)
by: Liu, Ailin, et al.
Published: (2024)
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
Bridging the Gap between Micro-scale Traffic Simulation and 4D Digital Cityscapes
by: Jiao, Longxiang, et al.
Published: (2026)
by: Jiao, Longxiang, et al.
Published: (2026)
Erie: A Declarative Grammar for Data Sonification
by: Kim, Hyeok, et al.
Published: (2024)
by: Kim, Hyeok, et al.
Published: (2024)
Effect of Avatar Head Movement on Communication Behaviour, Experience of Presence and Conversation Success in Triadic Conversations
by: Kothe, Angelika, et al.
Published: (2025)
by: Kothe, Angelika, et al.
Published: (2025)
Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
by: Xiong, Xiangrui, et al.
Published: (2025)
by: Xiong, Xiangrui, et al.
Published: (2025)
A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
by: Li, Kai, et al.
Published: (2026)
by: Li, Kai, et al.
Published: (2026)
Accessible Fine-grained Data Representation via Spatial Audio
by: Liu, Can, et al.
Published: (2026)
by: Liu, Can, et al.
Published: (2026)
FlueBricks: A Construction Kit of Flute-like Instruments for Acoustic Reasoning
by: Chen, Bo-Yu, et al.
Published: (2026)
by: Chen, Bo-Yu, et al.
Published: (2026)
EarResp-ANS : Audio-Based On-Device Respiration Rate Estimation on Earphones with Adaptive Noise Suppression
by: Küttner, Michael, et al.
Published: (2026)
by: Küttner, Michael, et al.
Published: (2026)
AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features
by: Xu, Weiye, et al.
Published: (2025)
by: Xu, Weiye, et al.
Published: (2025)
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
by: Martin, Charles Patrick
Published: (2026)
by: Martin, Charles Patrick
Published: (2026)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
by: Sharma, Roshan, et al.
Published: (2024)
by: Sharma, Roshan, et al.
Published: (2024)
Between the AI and Me: Analysing Listeners' Perspectives on AI- and Human-Composed Progressive Metal Music
by: Sarmento, Pedro, et al.
Published: (2024)
by: Sarmento, Pedro, et al.
Published: (2024)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024)
by: Cai, Zhuojiang, et al.
Published: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
Similar Items
-
WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
by: Hiraki, Hirotaka, et al.
Published: (2024) -
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
by: Jaiswal, Pranay, et al.
Published: (2024) -
Directional Source Separation for Robust Speech Recognition on Smart Glasses
by: Feng, Tiantian, et al.
Published: (2023) -
Improved Dysarthric Speech to Text Conversion via TTS Personalization
by: Mihajlik, Péter, et al.
Published: (2025) -
LightBeam: An Accurate and Memory-Efficient CTC Decoder for Speech Neuroprostheses
by: Feghhi, Ebrahim, et al.
Published: (2026)