SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Su, Weihang, Xie, Anzhe, Ai, Qingyao, Long, Jianming, Chen, Xuanyi, Mao, Jiaxin, Ye, Ziyi, Liu, Yiqun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911639128768512
author Su, Weihang
Xie, Anzhe
Ai, Qingyao
Long, Jianming
Chen, Xuanyi
Mao, Jiaxin
Ye, Ziyi
Liu, Yiqun
author_facet Su, Weihang
Xie, Anzhe
Ai, Qingyao
Long, Jianming
Chen, Xuanyi
Mao, Jiaxin
Ye, Ziyi
Liu, Yiqun
contents The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show promise for automating this process, progress in this area is hindered by the absence of standardized benchmarks and evaluation protocols. To bridge this critical gap, we introduce SurGE (Survey Generation Evaluation), a new benchmark for scientific survey generation in computer science. SurGE consists of (1) a collection of test instances, each including a topic description, an expert-written survey, and its full set of cited references, and (2) a large-scale academic corpus of over one million papers. In addition, we propose an automated evaluation framework that measures the quality of generated surveys across four dimensions: comprehensiveness, citation accuracy, structural organization, and content quality. Our evaluation of diverse LLM-based methods demonstrates a significant performance gap, revealing that even advanced agentic frameworks struggle with the complexities of survey generation and highlighting the need for future research in this area. We have open-sourced all the code, data, and models at: https://github.com/oneal2000/SurGE
format Preprint
id arxiv_https___arxiv_org_abs_2508_15658
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
Su, Weihang
Xie, Anzhe
Ai, Qingyao
Long, Jianming
Chen, Xuanyi
Mao, Jiaxin
Ye, Ziyi
Liu, Yiqun
Computation and Language
Artificial Intelligence
Information Retrieval
The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show promise for automating this process, progress in this area is hindered by the absence of standardized benchmarks and evaluation protocols. To bridge this critical gap, we introduce SurGE (Survey Generation Evaluation), a new benchmark for scientific survey generation in computer science. SurGE consists of (1) a collection of test instances, each including a topic description, an expert-written survey, and its full set of cited references, and (2) a large-scale academic corpus of over one million papers. In addition, we propose an automated evaluation framework that measures the quality of generated surveys across four dimensions: comprehensiveness, citation accuracy, structural organization, and content quality. Our evaluation of diverse LLM-based methods demonstrates a significant performance gap, revealing that even advanced agentic frameworks struggle with the complexities of survey generation and highlighting the need for future research in this area. We have open-sourced all the code, data, and models at: https://github.com/oneal2000/SurGE
title SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2508.15658