RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Lin, Gu, Zhouhong, Shi, Xiaoran, Feng, Hongwei, Xiao, Yanghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910900895612928
author Zhang, Lin
Gu, Zhouhong
Shi, Xiaoran
Feng, Hongwei
Xiao, Yanghua
author_facet Zhang, Lin
Gu, Zhouhong
Shi, Xiaoran
Feng, Hongwei
Xiao, Yanghua
contents As large language models (LLMs) advance, efficient knowledge evaluation becomes crucial to verifying their capabilities. Traditional methods, relying on benchmarks, face limitations such as high resource costs and information loss. We propose the Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model (RECKON), which directly uses reference data to evaluate models. RECKON organizes unstructured data into manageable units and generates targeted questions for each cluster, improving evaluation accuracy and efficiency. Experimental results show that RECKON reduces resource consumption by 56.5% compared to traditional methods while achieving over 97% accuracy across various domains, including world knowledge, code, legal, and biomedical datasets. Code is available at https://github.com/MikeGu721/reckon
format Preprint
id arxiv_https___arxiv_org_abs_2504_00756
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model
Zhang, Lin
Gu, Zhouhong
Shi, Xiaoran
Feng, Hongwei
Xiao, Yanghua
Computation and Language
As large language models (LLMs) advance, efficient knowledge evaluation becomes crucial to verifying their capabilities. Traditional methods, relying on benchmarks, face limitations such as high resource costs and information loss. We propose the Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model (RECKON), which directly uses reference data to evaluate models. RECKON organizes unstructured data into manageable units and generates targeted questions for each cluster, improving evaluation accuracy and efficiency. Experimental results show that RECKON reduces resource consumption by 56.5% compared to traditional methods while achieving over 97% accuracy across various domains, including world knowledge, code, legal, and biomedical datasets. Code is available at https://github.com/MikeGu721/reckon
title RECKON: Large-scale Reference-based Efficient Knowledge Evaluation for Large Language Model
topic Computation and Language
url https://arxiv.org/abs/2504.00756