Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Furuhashi, Momoka, Nakayama, Kouta, Kodama, Takashi, Sugawara, Saku |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2026)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2026)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
von: Emura, Rei, et al.
Veröffentlicht: (2026)
von: Emura, Rei, et al.
Veröffentlicht: (2026)
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
What Makes Language Models Good-enough?
von: Asami, Daiki, et al.
Veröffentlicht: (2024)
von: Asami, Daiki, et al.
Veröffentlicht: (2024)
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
von: Suzuki, Hisami, et al.
Veröffentlicht: (2025)
von: Suzuki, Hisami, et al.
Veröffentlicht: (2025)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
llm-jp-modernbert: A ModernBERT Model Trained on a Large-Scale Japanese Corpus with Long Context Length
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
MoreHopQA: More Than Multi-hop Reasoning
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
Exclusive Unlearning
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2026)
von: Sasaki, Mutsumi, et al.
Veröffentlicht: (2026)
Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
von: Oba, Miyu, et al.
Veröffentlicht: (2024)
von: Oba, Miyu, et al.
Veröffentlicht: (2024)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
von: Sun, Xin, et al.
Veröffentlicht: (2026)
von: Sun, Xin, et al.
Veröffentlicht: (2026)
Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment
von: Goldberg, Alexander, et al.
Veröffentlicht: (2024)
von: Goldberg, Alexander, et al.
Veröffentlicht: (2024)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
von: Zhou, Karen, et al.
Veröffentlicht: (2026)
von: Zhou, Karen, et al.
Veröffentlicht: (2026)
A Survey of Useful LLM Evaluation
von: Peng, Ji-Lun, et al.
Veröffentlicht: (2024)
von: Peng, Ji-Lun, et al.
Veröffentlicht: (2024)
Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
von: Simonotti, Martina, et al.
Veröffentlicht: (2026)
von: Simonotti, Martina, et al.
Veröffentlicht: (2026)
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
von: Ruan, Jie, et al.
Veröffentlicht: (2025)
von: Ruan, Jie, et al.
Veröffentlicht: (2025)
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes
von: Zhou, Karen, et al.
Veröffentlicht: (2025)
von: Zhou, Karen, et al.
Veröffentlicht: (2025)
Finding Blind Spots in Evaluator LLMs with Interpretable Checklists
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
Evaluating Diversity in Automatic Poetry Generation
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
Automatic Answerability Evaluation for Question Generation
von: Wang, Zifan, et al.
Veröffentlicht: (2023)
von: Wang, Zifan, et al.
Veröffentlicht: (2023)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2024)
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2024)
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
von: Chu, KuanChao, et al.
Veröffentlicht: (2024)
von: Chu, KuanChao, et al.
Veröffentlicht: (2024)
Comparing Two Model Designs for Clinical Note Generation; Is an LLM a Useful Evaluator of Consistency?
von: Brake, Nathan, et al.
Veröffentlicht: (2024)
von: Brake, Nathan, et al.
Veröffentlicht: (2024)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
RecMind: Japanese Movie Recommendation Dialogue with Seeker's Internal State
von: Kodama, Takashi, et al.
Veröffentlicht: (2024)
von: Kodama, Takashi, et al.
Veröffentlicht: (2024)
An Automatic Prompt Generation System for Tabular Data Tasks
von: Akella, Ashlesha, et al.
Veröffentlicht: (2024)
von: Akella, Ashlesha, et al.
Veröffentlicht: (2024)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
von: Wei, Tianjun, et al.
Veröffentlicht: (2025)
von: Wei, Tianjun, et al.
Veröffentlicht: (2025)
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2025)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructions
von: Geiger, Luisa, et al.
Veröffentlicht: (2025)
von: Geiger, Luisa, et al.
Veröffentlicht: (2025)
Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision
von: Kadasi, Pritam, et al.
Veröffentlicht: (2026)
von: Kadasi, Pritam, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2026) -
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025) -
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
von: Emura, Rei, et al.
Veröffentlicht: (2026) -
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
von: Kayano, Yoko, et al.
Veröffentlicht: (2025) -
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)