When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Weiyan, Herremans, Dorien, Choo, Kenny Tsu Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
More Than 1v1: Human-AI Alignment in Early Developmental Communities with Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
Towards Multimodal Large-Language Models for Parent-Child Interaction: A Focus on Joint Attention
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Exploring Gaze Pattern Differences Between Autistic and Neurotypical Children: Clustering, Visualisation, and Prediction
by: Shi, Weiyan, et al.
Published: (2024)
by: Shi, Weiyan, et al.
Published: (2024)
Multimodal Infusion Tuning for Large Models
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
MetaDragonBoat: Exploring Paddling Techniques of Virtual Dragon Boating in a Metaverse Campus
by: He, Wei, et al.
Published: (2024)
by: He, Wei, et al.
Published: (2024)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024)
by: Fu, Chencan, et al.
Published: (2024)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
Language-Guided Multimodal Texture Authoring via Generative Models
by: Qian, Wanli, et al.
Published: (2026)
by: Qian, Wanli, et al.
Published: (2026)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
From Multimodal Signals to Adaptive XR Experiences for De-escalation Training
by: Nierula, Birgit, et al.
Published: (2026)
by: Nierula, Birgit, et al.
Published: (2026)
MULTI-CASE: A Transformer-based Ethics-aware Multimodal Investigative Intelligence Framework
by: Fischer, Maximilian T., et al.
Published: (2024)
by: Fischer, Maximilian T., et al.
Published: (2024)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
by: Wan, Ninghao, et al.
Published: (2026)
by: Wan, Ninghao, et al.
Published: (2026)
Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
by: Ghosh, Indrajeet, et al.
Published: (2025)
by: Ghosh, Indrajeet, et al.
Published: (2025)
MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments
by: Tong, Yuqi, et al.
Published: (2024)
by: Tong, Yuqi, et al.
Published: (2024)
Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions
by: Liu, Qixuan, et al.
Published: (2025)
by: Liu, Qixuan, et al.
Published: (2025)
Multimodal Digital Sensing of Early-Life Laying Hens: A Pilot Study Integrating Thermal, Acoustic, Optical-Flow and Environmental Data
by: Dhaliwal, Yashan, et al.
Published: (2026)
by: Dhaliwal, Yashan, et al.
Published: (2026)
IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models
by: Rahimi, Hamed, et al.
Published: (2026)
by: Rahimi, Hamed, et al.
Published: (2026)
Foreign Domestic Workers' Perspectives on an LLM-Based Emotional Support tool for Caregiving Burden
by: Teng, Shin Shoon Nicholas, et al.
Published: (2026)
by: Teng, Shin Shoon Nicholas, et al.
Published: (2026)
User-Generated Content and Editors in Games: A Comprehensive Survey
by: Liu, Yuyue, et al.
Published: (2024)
by: Liu, Yuyue, et al.
Published: (2024)
MV-Crafter: An Intelligent System for Music-guided Video Generation
by: Chen, Chuer, et al.
Published: (2025)
by: Chen, Chuer, et al.
Published: (2025)
Crafting Dynamic Virtual Activities with Advanced Multimodal Models
by: Li, Changyang, et al.
Published: (2024)
by: Li, Changyang, et al.
Published: (2024)
Evaluating the Usability of Microgestures for Text Editing Tasks in Virtual Reality
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
by: Wang, Yingna, et al.
Published: (2025)
by: Wang, Yingna, et al.
Published: (2025)
Designing Effective AI Explanations for Misinformation Detection: A Comparative Study of Content, Social, and Combined Explanations
by: Gong, Yeaeun, et al.
Published: (2025)
by: Gong, Yeaeun, et al.
Published: (2025)
Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML
by: Wang, Bokang, et al.
Published: (2026)
by: Wang, Bokang, et al.
Published: (2026)
Sound Clouds: Exploring ambient intelligence in public spaces to elicit deep human experience of awe, wonder, and beauty
by: Zhang, Chengzhi, et al.
Published: (2025)
by: Zhang, Chengzhi, et al.
Published: (2025)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
Examining Augmented Virtuality Impairment Simulation for Mobile App Accessibility Design
by: Choo, Kenny Tsu Wei, et al.
Published: (2025)
by: Choo, Kenny Tsu Wei, et al.
Published: (2025)
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
by: Chen, Xiaolin, et al.
Published: (2022)
by: Chen, Xiaolin, et al.
Published: (2022)
Privileged Contrastive Pretraining for Multimodal Affect Modelling
by: Pinitas, Kosmas, et al.
Published: (2025)
by: Pinitas, Kosmas, et al.
Published: (2025)
MRATTS: An MR-Based Acupoint Therapy Training System with Real-Time Acupoint Detection and Evaluation Standards
by: Liu, Jiacheng, et al.
Published: (2026)
by: Liu, Jiacheng, et al.
Published: (2026)
Fostering Emotional Perspective-Taking: An Exploration of Affective Face-Tracking Interactions in the VR Narrative Rekindle
by: Fan, Hector, et al.
Published: (2026)
by: Fan, Hector, et al.
Published: (2026)
Whispering Water: Materializing Human-AI Dialogue as Interactive Ripples
by: Wang, Ruipeng, et al.
Published: (2026)
by: Wang, Ruipeng, et al.
Published: (2026)
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
AffectMachine-Pop: A controllable expert system for real-time pop music generation
by: Agres, Kat R., et al.
Published: (2025)
by: Agres, Kat R., et al.
Published: (2025)
Similar Items
-
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
by: Shi, Weiyan, et al.
Published: (2025) -
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025) -
Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions
by: Shi, Weiyan, et al.
Published: (2025) -
More Than 1v1: Human-AI Alignment in Early Developmental Communities with Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026) -
A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation
by: Shi, Weiyan, et al.
Published: (2026)