Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917900308185088 |
|---|---|
| author | Chung, John Joon Young Roemmele, Melissa Kreminski, Max |
| author_facet | Chung, John Joon Young Roemmele, Melissa Kreminski, Max |
| contents | We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_13284 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols Chung, John Joon Young Roemmele, Melissa Kreminski, Max Human-Computer Interaction Artificial Intelligence Computation and Language We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction. |
| title | Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols |
| topic | Human-Computer Interaction Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2501.13284 |