MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Maximillian, Zhang, Xuanming, Peng, Michael, Yu, Zhou, Papangelis, Alexandros, Jo, Yohan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024)
by: Cai, Zhuojiang, et al.
Published: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
Creating Aesthetic Sonifications on the Web with SIREN
by: Peng, Tristan, et al.
Published: (2024)
by: Peng, Tristan, et al.
Published: (2024)
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
by: Hnatyshyn, Rostyslav, et al.
Published: (2024)
by: Hnatyshyn, Rostyslav, et al.
Published: (2024)
Assessing the Viability of Wave Field Synthesis in VR-Based Cognitive Research
by: Kahl, Benjamin
Published: (2025)
by: Kahl, Benjamin
Published: (2025)
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
by: Hopkins, Torin, et al.
Published: (2026)
by: Hopkins, Torin, et al.
Published: (2026)
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
by: Huh, Mina, et al.
Published: (2026)
by: Huh, Mina, et al.
Published: (2026)
NeoLightning: A Modern Reimagination of Gesture-Based Sound Design
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
by: Yang, Sicheng, et al.
Published: (2024)
by: Yang, Sicheng, et al.
Published: (2024)
Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
by: Enge, Kajetan, et al.
Published: (2024)
by: Enge, Kajetan, et al.
Published: (2024)
Towards Reliable Large Audio Language Model
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
by: Peng, Jing, et al.
Published: (2026)
by: Peng, Jing, et al.
Published: (2026)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
by: Liu, Haoxuan, et al.
Published: (2024)
by: Liu, Haoxuan, et al.
Published: (2024)
AI TrackMate: Finally, Someone Who Will Give Your Music More Than Just "Sounds Great!"
by: Jiang, Yi-Lin, et al.
Published: (2024)
by: Jiang, Yi-Lin, et al.
Published: (2024)
Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
by: Selvamani, Shaja Arul, et al.
Published: (2025)
by: Selvamani, Shaja Arul, et al.
Published: (2025)
Workflow-Based Evaluation of Music Generation Systems
by: Dadman, Shayan, et al.
Published: (2025)
by: Dadman, Shayan, et al.
Published: (2025)
Beyond-Voice: Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
by: Li, Yin, et al.
Published: (2023)
by: Li, Yin, et al.
Published: (2023)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality
by: Choi, Youngwon, et al.
Published: (2025)
by: Choi, Youngwon, et al.
Published: (2025)
Revival: Collaborative Artistic Creation through Human-AI Interactions in Musical Creativity
by: Lee, Keon Ju M., et al.
Published: (2025)
by: Lee, Keon Ju M., et al.
Published: (2025)
Acoustic Wave Modeling Using 2D FDTD: Applications in Unreal Engine For Dynamic Sound Rendering
by: Samsurya, Bilkent
Published: (2025)
by: Samsurya, Bilkent
Published: (2025)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
by: Sun, Licai, et al.
Published: (2024)
by: Sun, Licai, et al.
Published: (2024)
Soundify: Matching Sound Effects to Video
by: Lin, David Chuan-En, et al.
Published: (2021)
by: Lin, David Chuan-En, et al.
Published: (2021)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
by: Ma, Yong, et al.
Published: (2025)
by: Ma, Yong, et al.
Published: (2025)
A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language
by: Kraack, Kris
Published: (2024)
by: Kraack, Kris
Published: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
BioSonix: Can Physics-Based Sonification Perceptualize Tissue Deformations From Tool Interactions?
by: Ruozzi, Veronica, et al.
Published: (2025)
by: Ruozzi, Veronica, et al.
Published: (2025)
LLAMAPIE: Proactive In-Ear Conversation Assistants
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
by: Park, Seohyun, et al.
Published: (2025)
by: Park, Seohyun, et al.
Published: (2025)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
by: Mishra, Ruchik, et al.
Published: (2024)
by: Mishra, Ruchik, et al.
Published: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
by: Wang, Hongbin, et al.
Published: (2025)
by: Wang, Hongbin, et al.
Published: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
by: Mertes, Silvan, et al.
Published: (2024)
by: Mertes, Silvan, et al.
Published: (2024)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
by: Xue, Liumeng, et al.
Published: (2024)
by: Xue, Liumeng, et al.
Published: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
by: Xiao, Yi, et al.
Published: (2022)
by: Xiao, Yi, et al.
Published: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
by: Chen, Youjun, et al.
Published: (2025)
by: Chen, Youjun, et al.
Published: (2025)
Similar Items
-
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025) -
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024) -
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
by: Nishida, Naoto, et al.
Published: (2025) -
Creating Aesthetic Sonifications on the Web with SIREN
by: Peng, Tristan, et al.
Published: (2024) -
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
by: Hnatyshyn, Rostyslav, et al.
Published: (2024)