TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Weiyan, Upadhyay, Sunaya, Quek, Geraldine, Choo, Kenny Tsu Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914162776473600
author Shi, Weiyan
Upadhyay, Sunaya
Quek, Geraldine
Choo, Kenny Tsu Wei
author_facet Shi, Weiyan
Upadhyay, Sunaya
Quek, Geraldine
Choo, Kenny Tsu Wei
contents Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05817
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
Shi, Weiyan
Upadhyay, Sunaya
Quek, Geraldine
Choo, Kenny Tsu Wei
Human-Computer Interaction
Multimedia
Sound
Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output.
title TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
topic Human-Computer Interaction
Multimedia
Sound
url https://arxiv.org/abs/2511.05817