LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Zikai, Huang, Fei, Tu, Jianhong, Wei, Jianhui, Ma, Wen, Zhou, Yuxuan, Wu, Jian, Yu, Bowen, Liu, Zuozhu, Lin, Junyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915582775918592
author Xiao, Zikai
Huang, Fei
Tu, Jianhong
Wei, Jianhui
Ma, Wen
Zhou, Yuxuan
Wu, Jian
Yu, Bowen
Liu, Zuozhu
Lin, Junyang
author_facet Xiao, Zikai
Huang, Fei
Tu, Jianhong
Wei, Jianhui
Ma, Wen
Zhou, Yuxuan
Wu, Jian
Yu, Bowen
Liu, Zuozhu
Lin, Junyang
contents Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies. In this paper, we introduce \textbf{LongWeave}, which balances real-world and verifiable assessment with Constraint-Verifier Evaluation (CoV-Eval). CoV-Eval constructs tasks by first defining verifiable targets within real-world scenarios, then systematically generating corresponding queries, textual materials, and constraints based on these targets. This ensures that tasks are both realistic and objectively assessable, enabling rigorous assessment of model capabilities in meeting complex real-world constraints. LongWeave supports customizable input/output lengths (up to 64K/8K tokens) across seven distinct tasks. Evaluation on 23 LLMs shows that even state-of-the-art models encounter significant challenges in long-form generation as real-world complexity and output length increase.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24345
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
Xiao, Zikai
Huang, Fei
Tu, Jianhong
Wei, Jianhui
Ma, Wen
Zhou, Yuxuan
Wu, Jian
Yu, Bowen
Liu, Zuozhu
Lin, Junyang
Computation and Language
Artificial Intelligence
Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies. In this paper, we introduce \textbf{LongWeave}, which balances real-world and verifiable assessment with Constraint-Verifier Evaluation (CoV-Eval). CoV-Eval constructs tasks by first defining verifiable targets within real-world scenarios, then systematically generating corresponding queries, textual materials, and constraints based on these targets. This ensures that tasks are both realistic and objectively assessable, enabling rigorous assessment of model capabilities in meeting complex real-world constraints. LongWeave supports customizable input/output lengths (up to 64K/8K tokens) across seven distinct tasks. Evaluation on 23 LLMs shows that even state-of-the-art models encounter significant challenges in long-form generation as real-world complexity and output length increase.
title LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.24345