TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914162776473600 |
|---|---|
| author | Shi, Weiyan Upadhyay, Sunaya Quek, Geraldine Choo, Kenny Tsu Wei |
| author_facet | Shi, Weiyan Upadhyay, Sunaya Quek, Geraldine Choo, Kenny Tsu Wei |
| contents | Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_05817 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech Shi, Weiyan Upadhyay, Sunaya Quek, Geraldine Choo, Kenny Tsu Wei Human-Computer Interaction Multimedia Sound Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output. |
| title | TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech |
| topic | Human-Computer Interaction Multimedia Sound |
| url | https://arxiv.org/abs/2511.05817 |