Saved in:
| Main Authors: | Jiao, Longxiang, Hofmann, Lukas, Yang, Yiru, Wu, Zhanyi, Egeler, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.08497 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
by: Xiong, Xiangrui, et al.
Published: (2025)
by: Xiong, Xiangrui, et al.
Published: (2025)
SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
by: Dementyev, Artem, et al.
Published: (2025)
by: Dementyev, Artem, et al.
Published: (2025)
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
by: Chen, Junjie, et al.
Published: (2025)
by: Chen, Junjie, et al.
Published: (2025)
MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
by: Li, Lu, et al.
Published: (2025)
by: Li, Lu, et al.
Published: (2025)
Improved Dysarthric Speech to Text Conversion via TTS Personalization
by: Mihajlik, Péter, et al.
Published: (2025)
by: Mihajlik, Péter, et al.
Published: (2025)
Erie: A Declarative Grammar for Data Sonification
by: Kim, Hyeok, et al.
Published: (2024)
by: Kim, Hyeok, et al.
Published: (2024)
Effect of Avatar Head Movement on Communication Behaviour, Experience of Presence and Conversation Success in Triadic Conversations
by: Kothe, Angelika, et al.
Published: (2025)
by: Kothe, Angelika, et al.
Published: (2025)
A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
by: Li, Kai, et al.
Published: (2026)
by: Li, Kai, et al.
Published: (2026)
Accessible Fine-grained Data Representation via Spatial Audio
by: Liu, Can, et al.
Published: (2026)
by: Liu, Can, et al.
Published: (2026)
LightBeam: An Accurate and Memory-Efficient CTC Decoder for Speech Neuroprostheses
by: Feghhi, Ebrahim, et al.
Published: (2026)
by: Feghhi, Ebrahim, et al.
Published: (2026)
FlueBricks: A Construction Kit of Flute-like Instruments for Acoustic Reasoning
by: Chen, Bo-Yu, et al.
Published: (2026)
by: Chen, Bo-Yu, et al.
Published: (2026)
EarResp-ANS : Audio-Based On-Device Respiration Rate Estimation on Earphones with Adaptive Noise Suppression
by: Küttner, Michael, et al.
Published: (2026)
by: Küttner, Michael, et al.
Published: (2026)
AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features
by: Xu, Weiye, et al.
Published: (2025)
by: Xu, Weiye, et al.
Published: (2025)
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
by: Martin, Charles Patrick
Published: (2026)
by: Martin, Charles Patrick
Published: (2026)
Exploring Gender Bias in Alzheimer's Disease Detection: Insights from Mandarin and Greek Speech Perception
by: He, Liu, et al.
Published: (2025)
by: He, Liu, et al.
Published: (2025)
An Intelligent AI glasses System with Multi-Agent Architecture for Real-Time Voice Processing and Task Execution
by: Chen, Sheng-Kai, et al.
Published: (2026)
by: Chen, Sheng-Kai, et al.
Published: (2026)
Sound Clouds: Exploring ambient intelligence in public spaces to elicit deep human experience of awe, wonder, and beauty
by: Zhang, Chengzhi, et al.
Published: (2025)
by: Zhang, Chengzhi, et al.
Published: (2025)
Semi-Automatic Flute Robot and Its Acoustic Sensing
by: Kuriyama, Hikari, et al.
Published: (2026)
by: Kuriyama, Hikari, et al.
Published: (2026)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
by: Yang, Tianle, et al.
Published: (2026)
by: Yang, Tianle, et al.
Published: (2026)
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
by: Hopkins, Torin, et al.
Published: (2026)
by: Hopkins, Torin, et al.
Published: (2026)
Multilingual and Continuous Backchannel Prediction: A Cross-lingual Study
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Beyond-Voice: Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
by: Li, Yin, et al.
Published: (2023)
by: Li, Yin, et al.
Published: (2023)
Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
by: Xue, Liumeng, et al.
Published: (2024)
by: Xue, Liumeng, et al.
Published: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
Echoes of Humanity: Exploring the Perceived Humanness of AI Music
by: Figueiredo, Flavio, et al.
Published: (2025)
by: Figueiredo, Flavio, et al.
Published: (2025)
EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition
by: Paul, Durjoy Chandra, et al.
Published: (2025)
by: Paul, Durjoy Chandra, et al.
Published: (2025)
Same Words, Different Judgments: How Preferences Vary Across Modalities
by: Broukhim, Aaron, et al.
Published: (2026)
by: Broukhim, Aaron, et al.
Published: (2026)
Live Music Models
by: Lyria Team, et al.
Published: (2025)
by: Lyria Team, et al.
Published: (2025)
Sona: Real-Time Multi-Target Sound Attenuation for Noise Sensitivity
by: Huang, Jeremy Zhengqi, et al.
Published: (2026)
by: Huang, Jeremy Zhengqi, et al.
Published: (2026)
BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation
by: Wang, Yunzhe, et al.
Published: (2025)
by: Wang, Yunzhe, et al.
Published: (2025)
LoopLens: Supporting Search as Creation in Loop-Based Music Composition
by: Long, Sheng, et al.
Published: (2026)
by: Long, Sheng, et al.
Published: (2026)
Opening Musical Creativity? Embedded Ideologies in Generative-AI Music Systems
by: Pram, Liam, et al.
Published: (2025)
by: Pram, Liam, et al.
Published: (2025)
The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity
by: Bradshaw, Louis, et al.
Published: (2025)
by: Bradshaw, Louis, et al.
Published: (2025)
Apollo: An Interactive Environment for Generating Symbolic Musical Phrases using Corpus-based Style Imitation
by: Tchemeube, Renaud Bougueng, et al.
Published: (2025)
by: Tchemeube, Renaud Bougueng, et al.
Published: (2025)
Calliope: An Online Generative Music System for Symbolic Multi-Track Composition
by: Tchemeube, Renaud Bougueng, et al.
Published: (2025)
by: Tchemeube, Renaud Bougueng, et al.
Published: (2025)
Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
by: Morreale, Fabio, et al.
Published: (2025)
by: Morreale, Fabio, et al.
Published: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
by: Wang, Hongbin, et al.
Published: (2025)
by: Wang, Hongbin, et al.
Published: (2025)
WSCoach: Wearable Real-time Auditory Feedback for Reducing Unwanted Words in Daily Communication
by: Youpeng, Zhang, et al.
Published: (2025)
by: Youpeng, Zhang, et al.
Published: (2025)
Similar Items
-
Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
by: Xiong, Xiangrui, et al.
Published: (2025) -
SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
by: Dementyev, Artem, et al.
Published: (2025) -
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
by: Chen, Junjie, et al.
Published: (2025) -
MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
by: Li, Lu, et al.
Published: (2025) -
Improved Dysarthric Speech to Text Conversion via TTS Personalization
by: Mihajlik, Péter, et al.
Published: (2025)