Generating Visual Stories with Grounded and Coreferent Characters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Danyang, Lapata, Mirella, Keller, Frank
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915177987833856
author Liu, Danyang
Lapata, Mirella
Keller, Frank
author_facet Liu, Danyang
Lapata, Mirella
Keller, Frank
contents Characters are important in narratives. They move the plot forward, create emotional connections, and embody the story's themes. Visual storytelling methods focus more on the plot and events relating to it, without building the narrative around specific characters. As a result, the generated stories feel generic, with character mentions being absent, vague, or incorrect. To mitigate these issues, we introduce the new task of character-centric story generation and present the first model capable of predicting visual stories with consistently grounded and coreferent character mentions. Our model is finetuned on a new dataset which we build on top of the widely used VIST benchmark. Specifically, we develop an automated pipeline to enrich VIST with visual and textual character coreference chains. We also propose new evaluation metrics to measure the richness of characters and coreference in stories. Experimental results show that our model generates stories with recurring characters which are consistent and coreferent to larger extent compared to baselines and state-of-the-art systems.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13555
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generating Visual Stories with Grounded and Coreferent Characters
Liu, Danyang
Lapata, Mirella
Keller, Frank
Computation and Language
Artificial Intelligence
Characters are important in narratives. They move the plot forward, create emotional connections, and embody the story's themes. Visual storytelling methods focus more on the plot and events relating to it, without building the narrative around specific characters. As a result, the generated stories feel generic, with character mentions being absent, vague, or incorrect. To mitigate these issues, we introduce the new task of character-centric story generation and present the first model capable of predicting visual stories with consistently grounded and coreferent character mentions. Our model is finetuned on a new dataset which we build on top of the widely used VIST benchmark. Specifically, we develop an automated pipeline to enrich VIST with visual and textual character coreference chains. We also propose new evaluation metrics to measure the richness of characters and coreference in stories. Experimental results show that our model generates stories with recurring characters which are consistent and coreferent to larger extent compared to baselines and state-of-the-art systems.
title Generating Visual Stories with Grounded and Coreferent Characters
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2409.13555