Efficient Evaluation of Large Language Models via Collaborative Filtering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Xu-Xiang, Yi, Chao, Ye, Han-Jia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909576130985984
author Zhong, Xu-Xiang
Yi, Chao
Ye, Han-Jia
author_facet Zhong, Xu-Xiang
Yi, Chao
Ye, Han-Jia
contents With the development of Large Language Models (LLMs), numerous benchmarks have been proposed to measure and compare the capabilities of different LLMs. However, evaluating LLMs is costly due to the large number of test instances and their slow inference speed. In this paper, we aim to explore how to efficiently estimate a model's real performance on a given benchmark based on its evaluation results on a small number of instances sampled from the benchmark. Inspired by Collaborative Filtering (CF) in Recommendation Systems (RS), we treat LLMs as users and test instances as items and propose a two-stage method. In the first stage, we treat instance selection as recommending products to users to choose instances that can easily distinguish model performance. In the second stage, we see performance prediction as rating prediction problem in RS to predict the target LLM's behavior on unselected instances. Experiments on multiple LLMs and datasets imply that our method can accurately estimate the target model's performance while largely reducing its inference overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Evaluation of Large Language Models via Collaborative Filtering
Zhong, Xu-Xiang
Yi, Chao
Ye, Han-Jia
Computation and Language
Artificial Intelligence
Information Retrieval
With the development of Large Language Models (LLMs), numerous benchmarks have been proposed to measure and compare the capabilities of different LLMs. However, evaluating LLMs is costly due to the large number of test instances and their slow inference speed. In this paper, we aim to explore how to efficiently estimate a model's real performance on a given benchmark based on its evaluation results on a small number of instances sampled from the benchmark. Inspired by Collaborative Filtering (CF) in Recommendation Systems (RS), we treat LLMs as users and test instances as items and propose a two-stage method. In the first stage, we treat instance selection as recommending products to users to choose instances that can easily distinguish model performance. In the second stage, we see performance prediction as rating prediction problem in RS to predict the target LLM's behavior on unselected instances. Experiments on multiple LLMs and datasets imply that our method can accurately estimate the target model's performance while largely reducing its inference overhead.
title Efficient Evaluation of Large Language Models via Collaborative Filtering
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2504.08781