AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Karen, Tan, Chenhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes
by: Zhou, Karen, et al.
Published: (2025)
by: Zhou, Karen, et al.
Published: (2025)
Checklist Engineering Empowers Multilingual LLM Judges
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2025)
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2025)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024)
by: Cook, Jonathan, et al.
Published: (2024)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
by: Wei, Tianjun, et al.
Published: (2025)
by: Wei, Tianjun, et al.
Published: (2025)
Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists
by: Cui, Yue, et al.
Published: (2025)
by: Cui, Yue, et al.
Published: (2025)
Data Checklist: On Unit-Testing Datasets with Usable Information
by: Zhang, Heidi C., et al.
Published: (2024)
by: Zhang, Heidi C., et al.
Published: (2024)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation
by: Garbacea, Cristina, et al.
Published: (2025)
by: Garbacea, Cristina, et al.
Published: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist
by: Seo, Kwangwook, et al.
Published: (2026)
by: Seo, Kwangwook, et al.
Published: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Checklists Are Better Than Reward Models For Aligning Language Models
by: Viswanathan, Vijay, et al.
Published: (2025)
by: Viswanathan, Vijay, et al.
Published: (2025)
Commitment Checklist: Auditing Author Commitments in Peer Review
by: Chen, Chung-Chi, et al.
Published: (2026)
by: Chen, Chung-Chi, et al.
Published: (2026)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
by: Lee, Young-Jun, et al.
Published: (2025)
by: Lee, Young-Jun, et al.
Published: (2025)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
by: Zhou, Zihao, et al.
Published: (2024)
by: Zhou, Zihao, et al.
Published: (2024)
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
by: Dou, Yao, et al.
Published: (2026)
by: Dou, Yao, et al.
Published: (2026)
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
by: Ruan, Jie, et al.
Published: (2025)
by: Ruan, Jie, et al.
Published: (2025)
Check-Eval: A Checklist-based Approach for Evaluating Text Quality
by: Pereira, Jayr, et al.
Published: (2024)
by: Pereira, Jayr, et al.
Published: (2024)
The Minimum Information about CLinical Artificial Intelligence Checklist for Generative Modeling Research (MI-CLAIM-GEN)
by: Miao, Brenda Y., et al.
Published: (2024)
by: Miao, Brenda Y., et al.
Published: (2024)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
by: Huang, Hui, et al.
Published: (2024)
by: Huang, Hui, et al.
Published: (2024)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
by: Fujinuma, Yoshinari
Published: (2025)
by: Fujinuma, Yoshinari
Published: (2025)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
by: Hu, Renjun, et al.
Published: (2025)
by: Hu, Renjun, et al.
Published: (2025)
AutoJudge: Judge Decoding Without Manual Annotation
by: Garipov, Roman, et al.
Published: (2025)
by: Garipov, Roman, et al.
Published: (2025)
Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment
by: Goldberg, Alexander, et al.
Published: (2024)
by: Goldberg, Alexander, et al.
Published: (2024)
Composable Cross-prompt Essay Scoring by Merging Models
by: Lee, Sanwoo, et al.
Published: (2025)
by: Lee, Sanwoo, et al.
Published: (2025)
ConfReady: A RAG based Assistant and Dataset for Conference Checklist Responses
by: Galarnyk, Michael, et al.
Published: (2024)
by: Galarnyk, Michael, et al.
Published: (2024)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
xList-Hate: A Checklist-Based Framework for Interpretable and Generalizable Hate Speech Detection
by: Girón, Adrián, et al.
Published: (2026)
by: Girón, Adrián, et al.
Published: (2026)
Same Input, Different Scores: A Multi Model Study on the Inconsistency of LLM Judge
by: Lau, Fiona
Published: (2026)
by: Lau, Fiona
Published: (2026)
Better Generalizing to Unseen Concepts: An Evaluation Framework and An LLM-Based Auto-Labeled Pipeline for Biomedical Concept Recognition
by: Liu, Shanshan, et al.
Published: (2026)
by: Liu, Shanshan, et al.
Published: (2026)
Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport
by: Morandi, Andrea
Published: (2026)
by: Morandi, Andrea
Published: (2026)
Improve LLM-as-a-Judge Ability as a General Ability
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
by: Wu, Junjie, et al.
Published: (2026)
by: Wu, Junjie, et al.
Published: (2026)
Similar Items
-
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes
by: Zhou, Karen, et al.
Published: (2025) -
Checklist Engineering Empowers Multilingual LLM Judges
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2025) -
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
by: Cook, Jonathan, et al.
Published: (2024) -
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
by: Furuhashi, Momoka, et al.
Published: (2025) -
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
by: Wei, Tianjun, et al.
Published: (2025)