Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Weiyan, Choo, Kenny Tsu Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
Towards Multimodal Large-Language Models for Parent-Child Interaction: A Focus on Joint Attention
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
von: Shi, Weiyan, et al.
Veröffentlicht: (2025)
More Than 1v1: Human-AI Alignment in Early Developmental Communities with Multimodal LLMs
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)
Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML
von: Wang, Bokang, et al.
Veröffentlicht: (2026)
von: Wang, Bokang, et al.
Veröffentlicht: (2026)
Applying LLM-Powered Virtual Humans to Child Interviews in Child-Centered Design
von: Li, Linshi, et al.
Veröffentlicht: (2025)
von: Li, Linshi, et al.
Veröffentlicht: (2025)
Whispering Water: Materializing Human-AI Dialogue as Interactive Ripples
von: Wang, Ruipeng, et al.
Veröffentlicht: (2026)
von: Wang, Ruipeng, et al.
Veröffentlicht: (2026)
SentiAvatar: Towards Expressive and Interactive Digital Humans
von: Jin, Chuhao, et al.
Veröffentlicht: (2026)
von: Jin, Chuhao, et al.
Veröffentlicht: (2026)
Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions
von: Liu, Qixuan, et al.
Veröffentlicht: (2025)
von: Liu, Qixuan, et al.
Veröffentlicht: (2025)
Multimodal Infusion Tuning for Large Models
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
CvhSlicer 2.0: Immersive and Interactive Visualization of Chinese Visible Human Data in XR Environments
von: Qiu, Yue, et al.
Veröffentlicht: (2025)
von: Qiu, Yue, et al.
Veröffentlicht: (2025)
Exploring Gaze Pattern Differences Between Autistic and Neurotypical Children: Clustering, Visualisation, and Prediction
von: Shi, Weiyan, et al.
Veröffentlicht: (2024)
von: Shi, Weiyan, et al.
Veröffentlicht: (2024)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2025)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2025)
Explainable Multimodal Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
MULTI-CASE: A Transformer-based Ethics-aware Multimodal Investigative Intelligence Framework
von: Fischer, Maximilian T., et al.
Veröffentlicht: (2024)
von: Fischer, Maximilian T., et al.
Veröffentlicht: (2024)
Ichiyo: Fragile and Transient Interaction in Neighborhood
von: Shibata, Hirofumi, et al.
Veröffentlicht: (2025)
von: Shibata, Hirofumi, et al.
Veröffentlicht: (2025)
Language-Guided Multimodal Texture Authoring via Generative Models
von: Qian, Wanli, et al.
Veröffentlicht: (2026)
von: Qian, Wanli, et al.
Veröffentlicht: (2026)
THE WASTIVE: An Interactive Ebb and Flow of Digital Fabrication Waste
von: Shan, Yifan, et al.
Veröffentlicht: (2025)
von: Shan, Yifan, et al.
Veröffentlicht: (2025)
From Multimodal Signals to Adaptive XR Experiences for De-escalation Training
von: Nierula, Birgit, et al.
Veröffentlicht: (2026)
von: Nierula, Birgit, et al.
Veröffentlicht: (2026)
Winds Through Time: Interactive Data Visualization and Physicalization for Paleoclimate Communication
von: Hunter, David, et al.
Veröffentlicht: (2025)
von: Hunter, David, et al.
Veröffentlicht: (2025)
ICE: Interactive 3D Game Character Editing via Dialogue
von: Wu, Haoqian, et al.
Veröffentlicht: (2024)
von: Wu, Haoqian, et al.
Veröffentlicht: (2024)
Multimodal Digital Sensing of Early-Life Laying Hens: A Pilot Study Integrating Thermal, Acoustic, Optical-Flow and Environmental Data
von: Dhaliwal, Yashan, et al.
Veröffentlicht: (2026)
von: Dhaliwal, Yashan, et al.
Veröffentlicht: (2026)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
von: Wan, Ninghao, et al.
Veröffentlicht: (2026)
von: Wan, Ninghao, et al.
Veröffentlicht: (2026)
INDCOR white paper 4: Evaluation of Interactive Narrative Design For Complexity Representations
von: Roth, Christian, et al.
Veröffentlicht: (2023)
von: Roth, Christian, et al.
Veröffentlicht: (2023)
Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
von: Ghosh, Indrajeet, et al.
Veröffentlicht: (2025)
von: Ghosh, Indrajeet, et al.
Veröffentlicht: (2025)
The Rhythm of Tai Chi: Revitalizing Cultural Heritage in Virtual Reality through Interactive Visuals
von: Wang, Xianghan
Veröffentlicht: (2025)
von: Wang, Xianghan
Veröffentlicht: (2025)
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
GustosonicSense: Towards understanding the design of playful gustosonic eating experiences
von: Wang, Yan, et al.
Veröffentlicht: (2024)
von: Wang, Yan, et al.
Veröffentlicht: (2024)
Fostering Emotional Perspective-Taking: An Exploration of Affective Face-Tracking Interactions in the VR Narrative Rekindle
von: Fan, Hector, et al.
Veröffentlicht: (2026)
von: Fan, Hector, et al.
Veröffentlicht: (2026)
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
von: Song, Jiarun, et al.
Veröffentlicht: (2026)
von: Song, Jiarun, et al.
Veröffentlicht: (2026)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
Physical-aware Cross-modal Adversarial Network for Wearable Sensor-based Human Action Recognition
von: Ni, Jianyuan, et al.
Veröffentlicht: (2023)
von: Ni, Jianyuan, et al.
Veröffentlicht: (2023)
MuDoC: An Interactive Multimodal Document-grounded Conversational AI System
von: Taneja, Karan, et al.
Veröffentlicht: (2025)
von: Taneja, Karan, et al.
Veröffentlicht: (2025)
User-Generated Content and Editors in Games: A Comprehensive Survey
von: Liu, Yuyue, et al.
Veröffentlicht: (2024)
von: Liu, Yuyue, et al.
Veröffentlicht: (2024)
"Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported Interactions
von: Sim, Kellie Yu Hui, et al.
Veröffentlicht: (2025)
von: Sim, Kellie Yu Hui, et al.
Veröffentlicht: (2025)
Human-Machine Collaboration-Guided Space Design: Combination of Machine Learning Models and Humanistic Design Concepts
von: Yang, Yuxuan
Veröffentlicht: (2025)
von: Yang, Yuxuan
Veröffentlicht: (2025)
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
von: Wang, Yingna, et al.
Veröffentlicht: (2025)
von: Wang, Yingna, et al.
Veröffentlicht: (2025)
Designing Effective AI Explanations for Misinformation Detection: A Comparative Study of Content, Social, and Combined Explanations
von: Gong, Yeaeun, et al.
Veröffentlicht: (2025)
von: Gong, Yeaeun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
von: Shi, Weiyan, et al.
Veröffentlicht: (2026) -
Towards Multimodal Large-Language Models for Parent-Child Interaction: A Focus on Joint Attention
von: Shi, Weiyan, et al.
Veröffentlicht: (2025) -
Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions
von: Shi, Weiyan, et al.
Veröffentlicht: (2025) -
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
von: Shi, Weiyan, et al.
Veröffentlicht: (2025) -
More Than 1v1: Human-AI Alignment in Early Developmental Communities with Multimodal LLMs
von: Shi, Weiyan, et al.
Veröffentlicht: (2026)