RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Tianjun, Wen, Wei, Qiao, Ruizhi, Sun, Xing, Ma, Jianghong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
von: Lu, Wensheng, et al.
Veröffentlicht: (2025)
von: Lu, Wensheng, et al.
Veröffentlicht: (2025)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
von: Shen, Chengyu, et al.
Veröffentlicht: (2026)
von: Shen, Chengyu, et al.
Veröffentlicht: (2026)
Check-Eval: A Checklist-based Approach for Evaluating Text Quality
von: Pereira, Jayr, et al.
Veröffentlicht: (2024)
von: Pereira, Jayr, et al.
Veröffentlicht: (2024)
EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization
von: Wang, Yaoning, et al.
Veröffentlicht: (2025)
von: Wang, Yaoning, et al.
Veröffentlicht: (2025)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
von: Liang, Chumeng, et al.
Veröffentlicht: (2025)
von: Liang, Chumeng, et al.
Veröffentlicht: (2025)
ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs
von: Chen, Keyu, et al.
Veröffentlicht: (2025)
von: Chen, Keyu, et al.
Veröffentlicht: (2025)
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
von: Dou, Yao, et al.
Veröffentlicht: (2026)
von: Dou, Yao, et al.
Veröffentlicht: (2026)
LLM-based Automated Grading with Human-in-the-Loop
von: Chu, Yucheng, et al.
Veröffentlicht: (2025)
von: Chu, Yucheng, et al.
Veröffentlicht: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
A2Eval: Agentic and Automated Evaluation for Embodied Brain
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
von: Lior, Gili, et al.
Veröffentlicht: (2025)
von: Lior, Gili, et al.
Veröffentlicht: (2025)
RepEval: Effective Text Evaluation with LLM Representation
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
von: Sheng, Shuqian, et al.
Veröffentlicht: (2024)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoou, et al.
Veröffentlicht: (2025)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
von: Zhou, Karen, et al.
Veröffentlicht: (2026)
von: Zhou, Karen, et al.
Veröffentlicht: (2026)
Optimizing In-Context Demonstrations for LLM-based Automated Grading
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
von: Behnam, Payman, et al.
Veröffentlicht: (2025)
RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Lingfeng, et al.
Veröffentlicht: (2025)
AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
von: Liang, Yuanzhi, et al.
Veröffentlicht: (2024)
von: Liang, Yuanzhi, et al.
Veröffentlicht: (2024)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
von: Dejl, Adam, et al.
Veröffentlicht: (2026)
von: Dejl, Adam, et al.
Veröffentlicht: (2026)
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
von: Chu, Seongyeub, et al.
Veröffentlicht: (2026)
von: Chu, Seongyeub, et al.
Veröffentlicht: (2026)
CreativEval: Evaluating Creativity of LLM-Based Hardware Code Generation
von: DeLorenzo, Matthew, et al.
Veröffentlicht: (2024)
von: DeLorenzo, Matthew, et al.
Veröffentlicht: (2024)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
von: Zhou, Yixi, et al.
Veröffentlicht: (2026)
von: Zhou, Yixi, et al.
Veröffentlicht: (2026)
Checklist Engineering Empowers Multilingual LLM Judges
von: Mohammadkhani, Mohammad Ghiasvand, et al.
Veröffentlicht: (2025)
von: Mohammadkhani, Mohammad Ghiasvand, et al.
Veröffentlicht: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
von: Yuan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaohan, et al.
Veröffentlicht: (2024)
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
von: Yang, Haote, et al.
Veröffentlicht: (2025)
von: Yang, Haote, et al.
Veröffentlicht: (2025)
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2024)
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
von: Zhu, Zhiying, et al.
Veröffentlicht: (2024)
von: Zhu, Zhiying, et al.
Veröffentlicht: (2024)
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
MARIO Eval: Evaluate Your Math LLM with your Math LLM--A mathematical dataset evaluation toolkit
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
von: Zhang, Boning, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
von: Lu, Wensheng, et al.
Veröffentlicht: (2025) -
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
von: Shen, Chengyu, et al.
Veröffentlicht: (2026) -
Check-Eval: A Checklist-based Approach for Evaluating Text Quality
von: Pereira, Jayr, et al.
Veröffentlicht: (2024) -
EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization
von: Wang, Yaoning, et al.
Veröffentlicht: (2025) -
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
von: Liang, Chumeng, et al.
Veröffentlicht: (2025)