What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Dingyi, Jin, Qin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
di: Yang, Dingyi, et al.
Pubblicazione: (2024)
StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment
di: Cui, Hanwen, et al.
Pubblicazione: (2026)
di: Cui, Hanwen, et al.
Pubblicazione: (2026)
LongStory: Coherent, Complete and Length Controlled Long story Generation
di: Park, Kyeongman, et al.
Pubblicazione: (2023)
di: Park, Kyeongman, et al.
Pubblicazione: (2023)
StoryAlign: Evaluating and Training Reward Models for Story Generation
di: Xia, Haotian, et al.
Pubblicazione: (2026)
di: Xia, Haotian, et al.
Pubblicazione: (2026)
STORYSUMM: Evaluating Faithfulness in Story Summarization
di: Subbiah, Melanie, et al.
Pubblicazione: (2024)
di: Subbiah, Melanie, et al.
Pubblicazione: (2024)
StoryWriter: A Multi-Agent Framework for Long Story Generation
di: Xia, Haotian, et al.
Pubblicazione: (2025)
di: Xia, Haotian, et al.
Pubblicazione: (2025)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
di: Li, Junjie, et al.
Pubblicazione: (2026)
di: Li, Junjie, et al.
Pubblicazione: (2026)
Evaluation Framework for AI Creativity: A Case Study Based on Story Generation
di: Sathya, Pharath, et al.
Pubblicazione: (2026)
di: Sathya, Pharath, et al.
Pubblicazione: (2026)
Text2Stories: Evaluating the Alignment Between Stakeholder Interviews and Generated User Stories
di: Dente, Francesco, et al.
Pubblicazione: (2025)
di: Dente, Francesco, et al.
Pubblicazione: (2025)
Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation
di: Chhun, Cyril, et al.
Pubblicazione: (2024)
di: Chhun, Cyril, et al.
Pubblicazione: (2024)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
di: Wan, Luanbo, et al.
Pubblicazione: (2025)
di: Wan, Luanbo, et al.
Pubblicazione: (2025)
Learning to Reason for Long-Form Story Generation
di: Gurung, Alexander, et al.
Pubblicazione: (2025)
di: Gurung, Alexander, et al.
Pubblicazione: (2025)
BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation
di: Ran, Yiting, et al.
Pubblicazione: (2025)
di: Ran, Yiting, et al.
Pubblicazione: (2025)
Previously on the Stories: Recap Snippet Identification for Story Reading
di: Li, Jiangnan, et al.
Pubblicazione: (2024)
di: Li, Jiangnan, et al.
Pubblicazione: (2024)
Event Causality Is Key to Computational Story Understanding
di: Sun, Yidan, et al.
Pubblicazione: (2023)
di: Sun, Yidan, et al.
Pubblicazione: (2023)
Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles
di: Wigler, Ben, et al.
Pubblicazione: (2026)
di: Wigler, Ben, et al.
Pubblicazione: (2026)
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
di: Wang, Xinda, et al.
Pubblicazione: (2025)
di: Wang, Xinda, et al.
Pubblicazione: (2025)
Extending Automatic Machine Translation Evaluation to Book-Length Documents
di: Wang, Kuang-Da, et al.
Pubblicazione: (2025)
di: Wang, Kuang-Da, et al.
Pubblicazione: (2025)
Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers
di: Subbiah, Melanie, et al.
Pubblicazione: (2024)
di: Subbiah, Melanie, et al.
Pubblicazione: (2024)
Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
di: Hida, Rem, et al.
Pubblicazione: (2024)
di: Hida, Rem, et al.
Pubblicazione: (2024)
AmharicStoryQA: A Multicultural Story Question Answering Benchmark in Amharic
di: Azime, Israel Abebe, et al.
Pubblicazione: (2026)
di: Azime, Israel Abebe, et al.
Pubblicazione: (2026)
Long Story Short: Story-level Video Understanding from 20K Short Films
di: Ghermi, Ridouane, et al.
Pubblicazione: (2024)
di: Ghermi, Ridouane, et al.
Pubblicazione: (2024)
Making a Long Story Short in Conversation Modeling
di: Tao, Yufei, et al.
Pubblicazione: (2024)
di: Tao, Yufei, et al.
Pubblicazione: (2024)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
di: Bai, Xiaoyan, et al.
Pubblicazione: (2026)
di: Bai, Xiaoyan, et al.
Pubblicazione: (2026)
Evaluating Creative Short Story Generation in Humans and Large Language Models
di: Ismayilzada, Mete, et al.
Pubblicazione: (2024)
di: Ismayilzada, Mete, et al.
Pubblicazione: (2024)
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
di: Rashkin, Hannah, et al.
Pubblicazione: (2025)
di: Rashkin, Hannah, et al.
Pubblicazione: (2025)
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization
di: Joshi, Brihi, et al.
Pubblicazione: (2025)
di: Joshi, Brihi, et al.
Pubblicazione: (2025)
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
di: Monazzah, Erfan Moosavi, et al.
Pubblicazione: (2025)
di: Monazzah, Erfan Moosavi, et al.
Pubblicazione: (2025)
BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training
di: Theodoropoulos, Nikitas, et al.
Pubblicazione: (2024)
di: Theodoropoulos, Nikitas, et al.
Pubblicazione: (2024)
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
di: Argese, Alex, et al.
Pubblicazione: (2026)
di: Argese, Alex, et al.
Pubblicazione: (2026)
MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation
di: Ma, Yan, et al.
Pubblicazione: (2024)
di: Ma, Yan, et al.
Pubblicazione: (2024)
Where Do People Tell Stories Online? Story Detection Across Online Communities
di: Antoniak, Maria, et al.
Pubblicazione: (2023)
di: Antoniak, Maria, et al.
Pubblicazione: (2023)
What the Heck is a MOO? And What's the Story with All Those Cows?
di: Falsetti, Julie
Pubblicazione: (1995)
di: Falsetti, Julie
Pubblicazione: (1995)
Subjective Evaluation Profile Analysis of Science Fiction Short Stories and its Critical-Theoretical Significance
di: Otsuka, Kazuyoshi
Pubblicazione: (2025)
di: Otsuka, Kazuyoshi
Pubblicazione: (2025)
CollabStory: Multi-LLM Collaborative Story Generation and Authorship Analysis
di: Venkatraman, Saranya, et al.
Pubblicazione: (2024)
di: Venkatraman, Saranya, et al.
Pubblicazione: (2024)
Long Story Generation via Knowledge Graph and Literary Theory
di: Shi, Ge, et al.
Pubblicazione: (2025)
di: Shi, Ge, et al.
Pubblicazione: (2025)
StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models
di: Chen, Zehao, et al.
Pubblicazione: (2025)
di: Chen, Zehao, et al.
Pubblicazione: (2025)
Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement
di: Wang, Qianyue, et al.
Pubblicazione: (2024)
di: Wang, Qianyue, et al.
Pubblicazione: (2024)
Multilingual TinyStories: A Synthetic Combinatorial Corpus of Indic Children's Stories for Training Small Language Models
di: Halder, Deepon, et al.
Pubblicazione: (2026)
di: Halder, Deepon, et al.
Pubblicazione: (2026)
Language Models Might Not Understand You: Evaluating Theory of Mind via Story Prompting
di: Getachew, Nathaniel, et al.
Pubblicazione: (2025)
di: Getachew, Nathaniel, et al.
Pubblicazione: (2025)
Documenti analoghi
-
What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation
di: Yang, Dingyi, et al.
Pubblicazione: (2024) -
StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment
di: Cui, Hanwen, et al.
Pubblicazione: (2026) -
LongStory: Coherent, Complete and Length Controlled Long story Generation
di: Park, Kyeongman, et al.
Pubblicazione: (2023) -
StoryAlign: Evaluating and Training Reward Models for Story Generation
di: Xia, Haotian, et al.
Pubblicazione: (2026) -
STORYSUMM: Evaluating Faithfulness in Story Summarization
di: Subbiah, Melanie, et al.
Pubblicazione: (2024)