AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
Fuente:
arXiv
Saved in:
| Main Authors: | Ryan, Michael J., Zhang, Yanzhe, Salunkhe, Amol, Chu, Yi, Xu, Di, Yang, Diyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping
by: Li, Ryan, et al.
Published: (2024)
by: Li, Ryan, et al.
Published: (2024)
Searching for Privacy Risks in LLM Agents via Simulation
by: Zhang, Yanzhe, et al.
Published: (2025)
by: Zhang, Yanzhe, et al.
Published: (2025)
Distilling an End-to-End Voice Assistant Without Instruction Training Data
by: Held, William, et al.
Published: (2024)
by: Held, William, et al.
Published: (2024)
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
by: Zhu, Hao, et al.
Published: (2025)
by: Zhu, Hao, et al.
Published: (2025)
Generative Interfaces for Language Models
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
by: Liu, Zijun, et al.
Published: (2023)
by: Liu, Zijun, et al.
Published: (2023)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
Improved LLM Agents for Financial Document Question Answering
by: Tan, Nelvin, et al.
Published: (2025)
by: Tan, Nelvin, et al.
Published: (2025)
Towards Automatic Evaluation of Task-Oriented Dialogue Flows
by: Mirtaheri, Mehrnoosh, et al.
Published: (2024)
by: Mirtaheri, Mehrnoosh, et al.
Published: (2024)
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
by: Tan, Nelvin, et al.
Published: (2025)
by: Tan, Nelvin, et al.
Published: (2025)
Auditing Gender Presentation Differences in Text-to-Image Models
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
by: Gan, Woody Haosheng, et al.
Published: (2026)
by: Gan, Woody Haosheng, et al.
Published: (2026)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
Automatic Information Extraction From Employment Tribunal Judgements Using Large Language Models
by: de Faria, Joana Ribeiro, et al.
Published: (2024)
by: de Faria, Joana Ribeiro, et al.
Published: (2024)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
by: GUO, Jiaxin, et al.
Published: (2025)
by: GUO, Jiaxin, et al.
Published: (2025)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
by: Fu, Jia, et al.
Published: (2024)
by: Fu, Jia, et al.
Published: (2024)
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
by: Martynov, Nikita, et al.
Published: (2025)
by: Martynov, Nikita, et al.
Published: (2025)
AutoPatent: A Multi-Agent Framework for Automatic Patent Generation
by: Wang, Qiyao, et al.
Published: (2024)
by: Wang, Qiyao, et al.
Published: (2024)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
by: Kodali, Prashant, et al.
Published: (2024)
by: Kodali, Prashant, et al.
Published: (2024)
AutoMix: Automatically Mixing Language Models
by: Aggarwal, Pranjal, et al.
Published: (2023)
by: Aggarwal, Pranjal, et al.
Published: (2023)
SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
by: Ryan, Michael J, et al.
Published: (2025)
by: Ryan, Michael J, et al.
Published: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
by: Barnes, Jeremy, et al.
Published: (2025)
by: Barnes, Jeremy, et al.
Published: (2025)
Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
Auto FAQ Generation
by: Kalvakolanu, Anjaneya Teja, et al.
Published: (2024)
by: Kalvakolanu, Anjaneya Teja, et al.
Published: (2024)
A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics
by: Soh, Yun Joon, et al.
Published: (2024)
by: Soh, Yun Joon, et al.
Published: (2024)
AutoReason: Automatic Few-Shot Reasoning Decomposition
by: Sevinc, Arda, et al.
Published: (2024)
by: Sevinc, Arda, et al.
Published: (2024)
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
by: Chu, Yucheng, et al.
Published: (2024)
by: Chu, Yucheng, et al.
Published: (2024)
EgoNormia: Benchmarking Physical Social Norm Understanding
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)
by: Zhang, Guo-Biao, et al.
Published: (2026)
Are Large Language Models Consistent over Value-laden Questions?
by: Moore, Jared, et al.
Published: (2024)
by: Moore, Jared, et al.
Published: (2024)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
AutoLife: Automatic Life Journaling with Smartphones and LLMs
by: Xu, Huatao, et al.
Published: (2024)
by: Xu, Huatao, et al.
Published: (2024)
Improving Legal Judgement Prediction in Romanian with Long Text Encoders
by: Masala, Mihai, et al.
Published: (2024)
by: Masala, Mihai, et al.
Published: (2024)
Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation
by: Jiang, Zhaokun, et al.
Published: (2024)
by: Jiang, Zhaokun, et al.
Published: (2024)
Similar Items
-
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping
by: Li, Ryan, et al.
Published: (2024) -
Searching for Privacy Risks in LLM Agents via Simulation
by: Zhang, Yanzhe, et al.
Published: (2025) -
Distilling an End-to-End Voice Assistant Without Instruction Training Data
by: Held, William, et al.
Published: (2024) -
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
by: Zhu, Hao, et al.
Published: (2025) -
Generative Interfaces for Language Models
by: Chen, Jiaqi, et al.
Published: (2025)