SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Jiahao, Zhang, Shuaixing, Xu, Nan, Wang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Meow: End-to-End Outline Writing for Automatic Academic Survey
by: Ma, Zhaoyu, et al.
Published: (2025)
by: Ma, Zhaoyu, et al.
Published: (2025)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)
by: Zhang, Guo-Biao, et al.
Published: (2026)
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
by: Zhu, Jiachen, et al.
Published: (2025)
by: Zhu, Jiachen, et al.
Published: (2025)
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
by: Wu, Wanxing, et al.
Published: (2026)
by: Wu, Wanxing, et al.
Published: (2026)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
by: Guo, Beichen, et al.
Published: (2025)
by: Guo, Beichen, et al.
Published: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT
by: Tao, Zhen, et al.
Published: (2024)
by: Tao, Zhen, et al.
Published: (2024)
AcademicEval: Live Long-Context LLM Benchmark
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
A Comprehensive Survey on Relation Extraction: Recent Advances and New Frontiers
by: Zhao, Xiaoyan, et al.
Published: (2023)
by: Zhao, Xiaoyan, et al.
Published: (2023)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
by: Zhao, Jingqian, et al.
Published: (2025)
by: Zhao, Jingqian, et al.
Published: (2025)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
by: Tan, Haoran, et al.
Published: (2025)
by: Tan, Haoran, et al.
Published: (2025)
Evaluation of Retrieval-Augmented Generation: A Survey
by: Yu, Hao, et al.
Published: (2024)
by: Yu, Hao, et al.
Published: (2024)
A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
by: Lin, Minhua, et al.
Published: (2025)
by: Lin, Minhua, et al.
Published: (2025)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
by: Wang, Yiheng, et al.
Published: (2025)
by: Wang, Yiheng, et al.
Published: (2025)
Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey
by: Ni, Bo, et al.
Published: (2025)
by: Ni, Bo, et al.
Published: (2025)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
by: Liu, Songyang, et al.
Published: (2025)
by: Liu, Songyang, et al.
Published: (2025)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
by: Li, Minghao, et al.
Published: (2025)
by: Li, Minghao, et al.
Published: (2025)
A Survey of Pun Generation: Datasets, Evaluations and Methodologies
by: Su, Yuchen, et al.
Published: (2025)
by: Su, Yuchen, et al.
Published: (2025)
Large Language Models for Planning: A Comprehensive and Systematic Survey
by: Cao, Pengfei, et al.
Published: (2025)
by: Cao, Pengfei, et al.
Published: (2025)
A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
by: Zhong, Jialun, et al.
Published: (2025)
by: Zhong, Jialun, et al.
Published: (2025)
Vision Mamba: A Comprehensive Survey and Taxonomy
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
A Survey on Evaluation of Large Language Models
by: Chang, Yupeng, et al.
Published: (2023)
by: Chang, Yupeng, et al.
Published: (2023)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
by: Khoshnoodi, Mahsa, et al.
Published: (2024)
by: Khoshnoodi, Mahsa, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
Empowering LLMs with Logical Reasoning: A Comprehensive Survey
by: Cheng, Fengxiang, et al.
Published: (2025)
by: Cheng, Fengxiang, et al.
Published: (2025)
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
by: Gupta, Prannaya, et al.
Published: (2024)
by: Gupta, Prannaya, et al.
Published: (2024)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)
by: Zhang, Mengyuan, et al.
Published: (2024)
From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems
by: Zhou, Zekun, et al.
Published: (2025)
by: Zhou, Zekun, et al.
Published: (2025)
Medical Dialogue: A Survey of Categories, Methods, Evaluation and Challenges
by: Shi, Xiaoming, et al.
Published: (2024)
by: Shi, Xiaoming, et al.
Published: (2024)
LLMs for Explainable AI: A Comprehensive Survey
by: Bilal, Ahsan, et al.
Published: (2025)
by: Bilal, Ahsan, et al.
Published: (2025)
Domain Specialization as the Key to Make Large Language Models Disruptive: A Comprehensive Survey
by: Ling, Chen, et al.
Published: (2023)
by: Ling, Chen, et al.
Published: (2023)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
by: Khan, Haidar, et al.
Published: (2025)
by: Khan, Haidar, et al.
Published: (2025)
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
by: Yang, Chih-Kai, et al.
Published: (2025)
by: Yang, Chih-Kai, et al.
Published: (2025)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
by: Garg, Madhav Krishan, et al.
Published: (2025)
by: Garg, Madhav Krishan, et al.
Published: (2025)
Similar Items
-
Meow: End-to-End Outline Writing for Automatic Academic Survey
by: Ma, Zhaoyu, et al.
Published: (2025) -
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026) -
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
by: Zhu, Jiachen, et al.
Published: (2025) -
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems
by: Wu, Wanxing, et al.
Published: (2026) -
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
by: Fu, Chaoyou, et al.
Published: (2024)