Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Fuente:
arXiv
Saved in:
| Main Authors: | Dubois, Yann, Galambosi, Balázs, Liang, Percy, Hashimoto, Tatsunori B. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
by: Dubois, Yann, et al.
Published: (2023)
by: Dubois, Yann, et al.
Published: (2023)
Evaluating Self-Supervised Learning via Risk Decomposition
by: Dubois, Yann, et al.
Published: (2023)
by: Dubois, Yann, et al.
Published: (2023)
s1: Simple test-time scaling
by: Muennighoff, Niklas, et al.
Published: (2025)
by: Muennighoff, Niklas, et al.
Published: (2025)
Language Models with Conformal Factuality Guarantees
by: Mohri, Christopher, et al.
Published: (2024)
by: Mohri, Christopher, et al.
Published: (2024)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
Eliciting Language Model Behaviors with Investigator Agents
by: Li, Xiang Lisa, et al.
Published: (2025)
by: Li, Xiang Lisa, et al.
Published: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)
by: Band, Neil, et al.
Published: (2024)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Reasoning to Learn from Latent Thoughts
by: Ruan, Yangjun, et al.
Published: (2025)
by: Ruan, Yangjun, et al.
Published: (2025)
Synthetic continued pretraining
by: Yang, Zitong, et al.
Published: (2024)
by: Yang, Zitong, et al.
Published: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
by: Grari, Vincent, et al.
Published: (2026)
by: Grari, Vincent, et al.
Published: (2026)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
by: Jiang, Mingjian, et al.
Published: (2024)
by: Jiang, Mingjian, et al.
Published: (2024)
Prompting Fairness: Integrating Causality to Debias Large Language Models
by: Li, Jingling, et al.
Published: (2024)
by: Li, Jingling, et al.
Published: (2024)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Towards Execution-Grounded Automated AI Research
by: Si, Chenglei, et al.
Published: (2026)
by: Si, Chenglei, et al.
Published: (2026)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)
by: Yu, Xiaodong, et al.
Published: (2023)
QualEval: Qualitative Evaluation for Model Improvement
by: Murahari, Vishvak, et al.
Published: (2023)
by: Murahari, Vishvak, et al.
Published: (2023)
Robust Distortion-free Watermarks for Language Models
by: Kuditipudi, Rohith, et al.
Published: (2023)
by: Kuditipudi, Rohith, et al.
Published: (2023)
AIMA at SemEval-2024 Task 3: Simple Yet Powerful Emotion Cause Pair Analysis
by: Kure, Alireza Ghahramani, et al.
Published: (2025)
by: Kure, Alireza Ghahramani, et al.
Published: (2025)
ScholarEval: Research Idea Evaluation Grounded in Literature
by: Moussa, Hanane Nour, et al.
Published: (2025)
by: Moussa, Hanane Nour, et al.
Published: (2025)
On the Entropy Calibration of Language Models
by: Cao, Steven, et al.
Published: (2025)
by: Cao, Steven, et al.
Published: (2025)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
by: Li, Mingxuan, et al.
Published: (2025)
by: Li, Mingxuan, et al.
Published: (2025)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
by: Wang, Ganghua, et al.
Published: (2025)
by: Wang, Ganghua, et al.
Published: (2025)
Synthetic Data for any Differentiable Target
by: Thrush, Tristan, et al.
Published: (2026)
by: Thrush, Tristan, et al.
Published: (2026)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
On the Learnability of Watermarks for Language Models
by: Gu, Chenchen, et al.
Published: (2023)
by: Gu, Chenchen, et al.
Published: (2023)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
by: Cheng, Jiali, et al.
Published: (2025)
by: Cheng, Jiali, et al.
Published: (2025)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
by: Sansford, Hannah, et al.
Published: (2024)
by: Sansford, Hannah, et al.
Published: (2024)
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
by: Liu, Minqian, et al.
Published: (2023)
by: Liu, Minqian, et al.
Published: (2023)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
by: Cao, Boxi, et al.
Published: (2024)
by: Cao, Boxi, et al.
Published: (2024)
Extracting books from production language models
by: Ahmed, Ahmed, et al.
Published: (2026)
by: Ahmed, Ahmed, et al.
Published: (2026)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
by: Chung, Tsz Ting, et al.
Published: (2025)
by: Chung, Tsz Ting, et al.
Published: (2025)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
by: He, Yuanqin, et al.
Published: (2024)
by: He, Yuanqin, et al.
Published: (2024)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024)
by: Boyeau, Pierre, et al.
Published: (2024)
Measuring all the noises of LLM Evals
by: Wang, Sida
Published: (2025)
by: Wang, Sida
Published: (2025)
AttributionBench: How Hard is Automatic Attribution Evaluation?
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Similar Items
-
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
by: Dubois, Yann, et al.
Published: (2023) -
Evaluating Self-Supervised Learning via Risk Decomposition
by: Dubois, Yann, et al.
Published: (2023) -
s1: Simple test-time scaling
by: Muennighoff, Niklas, et al.
Published: (2025) -
Language Models with Conformal Factuality Guarantees
by: Mohri, Christopher, et al.
Published: (2024) -
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)