CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Badshah, Sher, Moustafa, Moamen, Sajjad, Hassan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909896706883584
author Badshah, Sher
Moustafa, Moamen
Sajjad, Hassan
author_facet Badshah, Sher
Moustafa, Moamen
Sajjad, Hassan
contents Evaluating free-form Question Answering (QA) remains a challenge due to its diverse and open-ended nature. Traditional automatic metrics fail to capture semantic equivalence or accommodate the variability of open-ended responses. Leveraging Large Language Models (LLMs) as evaluators offers a promising alternative due to their strong language understanding and instruction-following capabilities. We propose Consensus via Lightweight Efficient Voting (CLEV), which employs two primary LLMs as judges and invokes a third judge only in cases of disagreement. This approach prioritizes evaluation reliability while reducing unnecessary computational demands. Through experiments, including human evaluation, we demonstrate CLEV's ability to provide consistent, scalable, and resource-efficient assessments, establishing it as a robust framework for evaluating LLMs on free-form QA.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08542
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
Badshah, Sher
Moustafa, Moamen
Sajjad, Hassan
Computation and Language
Artificial Intelligence
Primary 68T50, Secondary 68T45
I.2.7; I.2.6; I.2.3; I.2.0
Evaluating free-form Question Answering (QA) remains a challenge due to its diverse and open-ended nature. Traditional automatic metrics fail to capture semantic equivalence or accommodate the variability of open-ended responses. Leveraging Large Language Models (LLMs) as evaluators offers a promising alternative due to their strong language understanding and instruction-following capabilities. We propose Consensus via Lightweight Efficient Voting (CLEV), which employs two primary LLMs as judges and invokes a third judge only in cases of disagreement. This approach prioritizes evaluation reliability while reducing unnecessary computational demands. Through experiments, including human evaluation, we demonstrate CLEV's ability to provide consistent, scalable, and resource-efficient assessments, establishing it as a robust framework for evaluating LLMs on free-form QA.
title CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
topic Computation and Language
Artificial Intelligence
Primary 68T50, Secondary 68T45
I.2.7; I.2.6; I.2.3; I.2.0
url https://arxiv.org/abs/2503.08542