TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cook, Jonathan, Rocktäschel, Tim, Foerster, Jakob, Aumiller, Dennis, Wang, Alex |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024)
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024)
LLM Attributor: Interactive Visual Attribution for LLM Generation
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
von: Fu, Xiao, et al.
Veröffentlicht: (2025)
von: Fu, Xiao, et al.
Veröffentlicht: (2025)
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
von: Sun, Yuxuan, et al.
Veröffentlicht: (2025)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2025)
Properties and Challenges of LLM-Generated Explanations
von: Kunz, Jenny, et al.
Veröffentlicht: (2024)
von: Kunz, Jenny, et al.
Veröffentlicht: (2024)
Large Language Models for Cancer Communication: Evaluating Linguistic Quality, Safety, and Accessibility in Generative AI
von: Saha, Agnik, et al.
Veröffentlicht: (2025)
von: Saha, Agnik, et al.
Veröffentlicht: (2025)
Generative UI: LLMs are Effective UI Generators
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2026)
von: Leviathan, Yaniv, et al.
Veröffentlicht: (2026)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
von: Rajagopal, Shreya, et al.
Veröffentlicht: (2024)
von: Rajagopal, Shreya, et al.
Veröffentlicht: (2024)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
von: Kahng, Minsuk, et al.
Veröffentlicht: (2024)
von: Kahng, Minsuk, et al.
Veröffentlicht: (2024)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
von: Park, Jung In, et al.
Veröffentlicht: (2024)
von: Park, Jung In, et al.
Veröffentlicht: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
von: Baidya, Avinash, et al.
Veröffentlicht: (2025)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
von: Kim, Jane Paik
Veröffentlicht: (2026)
von: Kim, Jane Paik
Veröffentlicht: (2026)
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
von: Phutane, Mahika, et al.
Veröffentlicht: (2025)
von: Phutane, Mahika, et al.
Veröffentlicht: (2025)
Can LLM-Generated Misinformation Be Detected?
von: Chen, Canyu, et al.
Veröffentlicht: (2023)
von: Chen, Canyu, et al.
Veröffentlicht: (2023)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
von: Cook, Jonathan, et al.
Veröffentlicht: (2025)
von: Cook, Jonathan, et al.
Veröffentlicht: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
Transformer Explainer: Interactive Learning of Text-Generative Models
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
von: Cho, Aeree, et al.
Veröffentlicht: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
Survey of User Interface Design and Interaction Techniques in Generative AI Applications
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
von: Liu, Jonathan, et al.
Veröffentlicht: (2025)
von: Liu, Jonathan, et al.
Veröffentlicht: (2025)
Agent Laboratory: Using LLM Agents as Research Assistants
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
UniAutoML: A Human-Centered Framework for Unified Discriminative and Generative AutoML with Large Language Models
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
BADGE: BADminton report Generation and Evaluation with LLM
von: Chiang, Shang-Hsuan, et al.
Veröffentlicht: (2024)
von: Chiang, Shang-Hsuan, et al.
Veröffentlicht: (2024)
Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
von: Lam, Michelle S., et al.
Veröffentlicht: (2024)
von: Lam, Michelle S., et al.
Veröffentlicht: (2024)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
von: Zhang, Lechen, et al.
Veröffentlicht: (2024)
von: Zhang, Lechen, et al.
Veröffentlicht: (2024)
Multimodal Behavioral Patterns Analysis with Eye-Tracking and LLM-Based Reasoning
von: Guo, Dongyang, et al.
Veröffentlicht: (2025)
von: Guo, Dongyang, et al.
Veröffentlicht: (2025)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
von: Wu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyuan, et al.
Veröffentlicht: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
von: Swain, Sankalp Tattwadarshi, et al.
Veröffentlicht: (2025)
von: Swain, Sankalp Tattwadarshi, et al.
Veröffentlicht: (2025)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
von: Zhang, Lechen, et al.
Veröffentlicht: (2025)
von: Zhang, Lechen, et al.
Veröffentlicht: (2025)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
von: Lee, Dong Won, et al.
Veröffentlicht: (2024)
von: Lee, Dong Won, et al.
Veröffentlicht: (2024)
Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Heterogeneous Value Alignment Evaluation for Large Language Models
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2023)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2023)
Language Models as Zero-Shot Trajectory Generators
von: Kwon, Teyun, et al.
Veröffentlicht: (2023)
von: Kwon, Teyun, et al.
Veröffentlicht: (2023)
PRECISE Framework: GPT-based Text For Improved Readability, Reliability, and Understandability of Radiology Reports For Patient-Centered Care
von: Tripathi, Satvik, et al.
Veröffentlicht: (2024)
von: Tripathi, Satvik, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Health-related Queries with Presuppositions
von: Kaur, Navreet, et al.
Veröffentlicht: (2023)
von: Kaur, Navreet, et al.
Veröffentlicht: (2023)
Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
von: Thakkar, Nitya, et al.
Veröffentlicht: (2025)
von: Thakkar, Nitya, et al.
Veröffentlicht: (2025)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
von: Zeng, Qiuhai, et al.
Veröffentlicht: (2025)
von: Zeng, Qiuhai, et al.
Veröffentlicht: (2025)
Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
von: Ben-Zion, Ziv, et al.
Veröffentlicht: (2025)
von: Ben-Zion, Ziv, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Creative Beam Search: LLM-as-a-Judge For Improving Response Generation
von: Franceschelli, Giorgio, et al.
Veröffentlicht: (2024) -
LLM Attributor: Interactive Visual Attribution for LLM Generation
von: Lee, Seongmin, et al.
Veröffentlicht: (2024) -
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
von: Fu, Xiao, et al.
Veröffentlicht: (2025) -
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
von: Sun, Yuxuan, et al.
Veröffentlicht: (2025) -
Properties and Challenges of LLM-Generated Explanations
von: Kunz, Jenny, et al.
Veröffentlicht: (2024)