StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oliveira, Daniel A. P., de Matos, David Martins
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915472052584448
author Oliveira, Daniel A. P.
de Matos, David Martins
author_facet Oliveira, Daniel A. P.
de Matos, David Martins
contents Visual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations. These issues can be addressed through grounding of characters, objects, and other entities on the visual elements. We propose StoryReasoning, a dataset containing 4,178 stories derived from 52,016 movie images, with both structured scene analyses and grounded stories. Each story maintains character and object consistency across frames while explicitly modeling multi-frame relationships through structured tabular representations. Our approach features cross-frame object re-identification using visual similarity and face recognition, chain-of-thought reasoning for explicit narrative modeling, and a grounding scheme that links textual elements to visual entities across multiple frames. We establish baseline performance by fine-tuning Qwen2.5-VL 7B, creating Qwen Storyteller, which performs end-to-end object detection, re-identification, and landmark detection while maintaining consistent object references throughout the story. Evaluation demonstrates a reduction from 4.06 to 3.56 (-12.3%) hallucinations on average per story and an improvement in creativity from 2.58 to 3.38 (+31.0%) when compared to a non-fine-tuned model.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10292
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
Oliveira, Daniel A. P.
de Matos, David Martins
Computer Vision and Pattern Recognition
Computation and Language
I.2.10; I.2.7
Visual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations. These issues can be addressed through grounding of characters, objects, and other entities on the visual elements. We propose StoryReasoning, a dataset containing 4,178 stories derived from 52,016 movie images, with both structured scene analyses and grounded stories. Each story maintains character and object consistency across frames while explicitly modeling multi-frame relationships through structured tabular representations. Our approach features cross-frame object re-identification using visual similarity and face recognition, chain-of-thought reasoning for explicit narrative modeling, and a grounding scheme that links textual elements to visual entities across multiple frames. We establish baseline performance by fine-tuning Qwen2.5-VL 7B, creating Qwen Storyteller, which performs end-to-end object detection, re-identification, and landmark detection while maintaining consistent object references throughout the story. Evaluation demonstrates a reduction from 4.06 to 3.56 (-12.3%) hallucinations on average per story and an improvement in creativity from 2.58 to 3.38 (+31.0%) when compared to a non-fine-tuned model.
title StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
topic Computer Vision and Pattern Recognition
Computation and Language
I.2.10; I.2.7
url https://arxiv.org/abs/2505.10292