LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915582775918592 |
|---|---|
| author | Xiao, Zikai Huang, Fei Tu, Jianhong Wei, Jianhui Ma, Wen Zhou, Yuxuan Wu, Jian Yu, Bowen Liu, Zuozhu Lin, Junyang |
| author_facet | Xiao, Zikai Huang, Fei Tu, Jianhong Wei, Jianhui Ma, Wen Zhou, Yuxuan Wu, Jian Yu, Bowen Liu, Zuozhu Lin, Junyang |
| contents | Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies. In this paper, we introduce \textbf{LongWeave}, which balances real-world and verifiable assessment with Constraint-Verifier Evaluation (CoV-Eval). CoV-Eval constructs tasks by first defining verifiable targets within real-world scenarios, then systematically generating corresponding queries, textual materials, and constraints based on these targets. This ensures that tasks are both realistic and objectively assessable, enabling rigorous assessment of model capabilities in meeting complex real-world constraints. LongWeave supports customizable input/output lengths (up to 64K/8K tokens) across seven distinct tasks. Evaluation on 23 LLMs shows that even state-of-the-art models encounter significant challenges in long-form generation as real-world complexity and output length increase. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_24345 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability Xiao, Zikai Huang, Fei Tu, Jianhong Wei, Jianhui Ma, Wen Zhou, Yuxuan Wu, Jian Yu, Bowen Liu, Zuozhu Lin, Junyang Computation and Language Artificial Intelligence Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies. In this paper, we introduce \textbf{LongWeave}, which balances real-world and verifiable assessment with Constraint-Verifier Evaluation (CoV-Eval). CoV-Eval constructs tasks by first defining verifiable targets within real-world scenarios, then systematically generating corresponding queries, textual materials, and constraints based on these targets. This ensures that tasks are both realistic and objectively assessable, enabling rigorous assessment of model capabilities in meeting complex real-world constraints. LongWeave supports customizable input/output lengths (up to 64K/8K tokens) across seven distinct tasks. Evaluation on 23 LLMs shows that even state-of-the-art models encounter significant challenges in long-form generation as real-world complexity and output length increase. |
| title | LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.24345 |