Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chung, John Joon Young, Roemmele, Melissa, Kreminski, Max
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917900308185088
author Chung, John Joon Young
Roemmele, Melissa
Kreminski, Max
author_facet Chung, John Joon Young
Roemmele, Melissa
Kreminski, Max
contents We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
Chung, John Joon Young
Roemmele, Melissa
Kreminski, Max
Human-Computer Interaction
Artificial Intelligence
Computation and Language
We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction.
title Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2501.13284