TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Weiyan, Upadhyay, Sunaya, Quek, Geraldine, Choo, Kenny Tsu Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
More Than 1v1: Human-AI Alignment in Early Developmental Communities with Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
Sound Clouds: Exploring ambient intelligence in public spaces to elicit deep human experience of awe, wonder, and beauty
by: Zhang, Chengzhi, et al.
Published: (2025)
by: Zhang, Chengzhi, et al.
Published: (2025)
Towards Multimodal Large-Language Models for Parent-Child Interaction: A Focus on Joint Attention
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
Proceedings of The third international workshop on eXplainable AI for the Arts (XAIxArts)
by: Ford, Corey, et al.
Published: (2025)
by: Ford, Corey, et al.
Published: (2025)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024)
by: Cai, Zhuojiang, et al.
Published: (2024)
Music of Changing Lines: Toward a Culturally Situated Approach to the I-Ching
by: Qi, Ling, et al.
Published: (2026)
by: Qi, Ling, et al.
Published: (2026)
MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments
by: Tong, Yuqi, et al.
Published: (2024)
by: Tong, Yuqi, et al.
Published: (2024)
VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails
by: Huh, Mina, et al.
Published: (2026)
by: Huh, Mina, et al.
Published: (2026)
Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching
by: Lin, David Chuan-En, et al.
Published: (2025)
by: Lin, David Chuan-En, et al.
Published: (2025)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
by: Yang, Sicheng, et al.
Published: (2024)
by: Yang, Sicheng, et al.
Published: (2024)
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
by: Chen, Maximillian, et al.
Published: (2026)
by: Chen, Maximillian, et al.
Published: (2026)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Capturing Cancer as Music: Cancer Mechanisms Expressed through Musification
by: Hnatyshyn, Rostyslav, et al.
Published: (2024)
by: Hnatyshyn, Rostyslav, et al.
Published: (2024)
Assessing the Viability of Wave Field Synthesis in VR-Based Cognitive Research
by: Kahl, Benjamin
Published: (2025)
by: Kahl, Benjamin
Published: (2025)
Creating Aesthetic Sonifications on the Web with SIREN
by: Peng, Tristan, et al.
Published: (2024)
by: Peng, Tristan, et al.
Published: (2024)
MR-DAW: Towards Collaborative Digital Audio Workstations in Mixed Reality
by: Hopkins, Torin, et al.
Published: (2026)
by: Hopkins, Torin, et al.
Published: (2026)
NeoLightning: A Modern Reimagination of Gesture-Based Sound Design
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Proceedings of The second international workshop on eXplainable AI for the Arts (XAIxArts)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
by: Bryan-Kinns, Nick, et al.
Published: (2024)
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
by: Selvamani, Shaja Arul, et al.
Published: (2025)
by: Selvamani, Shaja Arul, et al.
Published: (2025)
AI TrackMate: Finally, Someone Who Will Give Your Music More Than Just "Sounds Great!"
by: Jiang, Yi-Lin, et al.
Published: (2024)
by: Jiang, Yi-Lin, et al.
Published: (2024)
Workflow-Based Evaluation of Music Generation Systems
by: Dadman, Shayan, et al.
Published: (2025)
by: Dadman, Shayan, et al.
Published: (2025)
Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
by: Enge, Kajetan, et al.
Published: (2024)
by: Enge, Kajetan, et al.
Published: (2024)
ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
by: Luo, Cheng, et al.
Published: (2026)
by: Luo, Cheng, et al.
Published: (2026)
Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft
by: Sun, Xin, et al.
Published: (2025)
by: Sun, Xin, et al.
Published: (2025)
Towards Reliable Large Audio Language Model
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Language-Guided Multimodal Texture Authoring via Generative Models
by: Qian, Wanli, et al.
Published: (2026)
by: Qian, Wanli, et al.
Published: (2026)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024)
by: Fu, Chencan, et al.
Published: (2024)
Multimodal Infusion Tuning for Large Models
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
by: Wan, Ninghao, et al.
Published: (2026)
by: Wan, Ninghao, et al.
Published: (2026)
Exploring Gaze Pattern Differences Between Autistic and Neurotypical Children: Clustering, Visualisation, and Prediction
by: Shi, Weiyan, et al.
Published: (2024)
by: Shi, Weiyan, et al.
Published: (2024)
MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
by: Liu, Haoxuan, et al.
Published: (2024)
by: Liu, Haoxuan, et al.
Published: (2024)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
by: Peng, Jing, et al.
Published: (2026)
by: Peng, Jing, et al.
Published: (2026)
Soundify: Matching Sound Effects to Video
by: Lin, David Chuan-En, et al.
Published: (2021)
by: Lin, David Chuan-En, et al.
Published: (2021)
Similar Items
-
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026) -
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025) -
A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation
by: Shi, Weiyan, et al.
Published: (2026) -
Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions
by: Shi, Weiyan, et al.
Published: (2025) -
More Than 1v1: Human-AI Alignment in Early Developmental Communities with Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)