Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Kexue, Huang, Jingfei, Ling, Long, Hong, Sumin, Zuo, Yihang, LC, Ray, Li, Toby Jia-jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908815094448128
author Fu, Kexue
Huang, Jingfei
Ling, Long
Hong, Sumin
Zuo, Yihang
LC, Ray
Li, Toby Jia-jun
author_facet Fu, Kexue
Huang, Jingfei
Ling, Long
Hong, Sumin
Zuo, Yihang
LC, Ray
Li, Toby Jia-jun
contents Humans think visually-we remember in images, dream in pictures, and use visual metaphors to communicate. Yet, most creative writing tools remain text-centric, limiting how authors plan and translate ideas. We present Vistoria, a system for synchronized text-image co-editing in fictional story writing that treats visuals and text as coequal narrative materials. A formative Wizard-of-Oz co-design study with 10 story writers revealed how sketches, images, and annotations serve as essential instruments for ideation and organization. Drawing on theories of Instrumental Interaction and Structural Mapping, Vistoria introduces multimodal operations-lasso, collage, filters, and perspective shifts that enable seamless narrative exploration across modalities. A controlled study with 12 participants shows that co-editing enhances expressiveness, immersion, and collaboration, enabling writers to explore divergent directions, embrace serendipitous randomness, and trace evolving storylines. While multimodality increased cognitive demand, participants reported stronger senses of authorship and agency. These findings demonstrate how multimodal co-editing expands creative potential by balancing abstraction and concreteness in narrative development.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing
Fu, Kexue
Huang, Jingfei
Ling, Long
Hong, Sumin
Zuo, Yihang
LC, Ray
Li, Toby Jia-jun
Human-Computer Interaction
Humans think visually-we remember in images, dream in pictures, and use visual metaphors to communicate. Yet, most creative writing tools remain text-centric, limiting how authors plan and translate ideas. We present Vistoria, a system for synchronized text-image co-editing in fictional story writing that treats visuals and text as coequal narrative materials. A formative Wizard-of-Oz co-design study with 10 story writers revealed how sketches, images, and annotations serve as essential instruments for ideation and organization. Drawing on theories of Instrumental Interaction and Structural Mapping, Vistoria introduces multimodal operations-lasso, collage, filters, and perspective shifts that enable seamless narrative exploration across modalities. A controlled study with 12 participants shows that co-editing enhances expressiveness, immersion, and collaboration, enabling writers to explore divergent directions, embrace serendipitous randomness, and trace evolving storylines. While multimodality increased cognitive demand, participants reported stronger senses of authorship and agency. These findings demonstrate how multimodal co-editing expands creative potential by balancing abstraction and concreteness in narrative development.
title Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Text-Image Co-Editing
topic Human-Computer Interaction
url https://arxiv.org/abs/2509.13646