Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dubois, Yann, Galambosi, Balázs, Liang, Percy, Hashimoto, Tatsunori B. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
von: Dubois, Yann, et al.
Veröffentlicht: (2023)
von: Dubois, Yann, et al.
Veröffentlicht: (2023)
Evaluating Self-Supervised Learning via Risk Decomposition
von: Dubois, Yann, et al.
Veröffentlicht: (2023)
von: Dubois, Yann, et al.
Veröffentlicht: (2023)
s1: Simple test-time scaling
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2025)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2025)
Language Models with Conformal Factuality Guarantees
von: Mohri, Christopher, et al.
Veröffentlicht: (2024)
von: Mohri, Christopher, et al.
Veröffentlicht: (2024)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
von: Ruan, Yangjun, et al.
Veröffentlicht: (2023)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2023)
Eliciting Language Model Behaviors with Investigator Agents
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2024)
Linguistic Calibration of Long-Form Generations
von: Band, Neil, et al.
Veröffentlicht: (2024)
von: Band, Neil, et al.
Veröffentlicht: (2024)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
Reasoning to Learn from Latent Thoughts
von: Ruan, Yangjun, et al.
Veröffentlicht: (2025)
von: Ruan, Yangjun, et al.
Veröffentlicht: (2025)
Synthetic continued pretraining
von: Yang, Zitong, et al.
Veröffentlicht: (2024)
von: Yang, Zitong, et al.
Veröffentlicht: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
von: Grari, Vincent, et al.
Veröffentlicht: (2026)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
von: Jiang, Mingjian, et al.
Veröffentlicht: (2024)
von: Jiang, Mingjian, et al.
Veröffentlicht: (2024)
Prompting Fairness: Integrating Causality to Debias Large Language Models
von: Li, Jingling, et al.
Veröffentlicht: (2024)
von: Li, Jingling, et al.
Veröffentlicht: (2024)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
Towards Execution-Grounded Automated AI Research
von: Si, Chenglei, et al.
Veröffentlicht: (2026)
von: Si, Chenglei, et al.
Veröffentlicht: (2026)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
von: Yu, Xiaodong, et al.
Veröffentlicht: (2023)
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
Robust Distortion-free Watermarks for Language Models
von: Kuditipudi, Rohith, et al.
Veröffentlicht: (2023)
von: Kuditipudi, Rohith, et al.
Veröffentlicht: (2023)
AIMA at SemEval-2024 Task 3: Simple Yet Powerful Emotion Cause Pair Analysis
von: Kure, Alireza Ghahramani, et al.
Veröffentlicht: (2025)
von: Kure, Alireza Ghahramani, et al.
Veröffentlicht: (2025)
ScholarEval: Research Idea Evaluation Grounded in Literature
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2025)
von: Moussa, Hanane Nour, et al.
Veröffentlicht: (2025)
On the Entropy Calibration of Language Models
von: Cao, Steven, et al.
Veröffentlicht: (2025)
von: Cao, Steven, et al.
Veröffentlicht: (2025)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
von: Li, Mingxuan, et al.
Veröffentlicht: (2025)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
Synthetic Data for any Differentiable Target
von: Thrush, Tristan, et al.
Veröffentlicht: (2026)
von: Thrush, Tristan, et al.
Veröffentlicht: (2026)
Reliable and Efficient Amortized Model-based Evaluation
von: Truong, Sang, et al.
Veröffentlicht: (2025)
von: Truong, Sang, et al.
Veröffentlicht: (2025)
On the Learnability of Watermarks for Language Models
von: Gu, Chenchen, et al.
Veröffentlicht: (2023)
von: Gu, Chenchen, et al.
Veröffentlicht: (2023)
Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
von: Liu, Minqian, et al.
Veröffentlicht: (2023)
von: Liu, Minqian, et al.
Veröffentlicht: (2023)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Extracting books from production language models
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
von: Ahmed, Ahmed, et al.
Veröffentlicht: (2026)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
von: Boyeau, Pierre, et al.
Veröffentlicht: (2024)
von: Boyeau, Pierre, et al.
Veröffentlicht: (2024)
Measuring all the noises of LLM Evals
von: Wang, Sida
Veröffentlicht: (2025)
von: Wang, Sida
Veröffentlicht: (2025)
AttributionBench: How Hard is Automatic Attribution Evaluation?
von: Li, Yifei, et al.
Veröffentlicht: (2024)
von: Li, Yifei, et al.
Veröffentlicht: (2024)
Less is More for Improving Automatic Evaluation of Factual Consistency
von: Wang, Tong, et al.
Veröffentlicht: (2024)
von: Wang, Tong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
von: Dubois, Yann, et al.
Veröffentlicht: (2023) -
Evaluating Self-Supervised Learning via Risk Decomposition
von: Dubois, Yann, et al.
Veröffentlicht: (2023) -
s1: Simple test-time scaling
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2025) -
Language Models with Conformal Factuality Guarantees
von: Mohri, Christopher, et al.
Veröffentlicht: (2024) -
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
von: Ruan, Yangjun, et al.
Veröffentlicht: (2023)