Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jun, Gu, Ninglun, Zhang, Kailai, Zhang, Zijiao, Bao, Yelun, Yang, Jin, Yin, Xu, Liu, Liwei, Liu, Yihuan, Li, Pengyong, Yen, Gary G., Yan, Junchi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914161636671488
author Wang, Jun
Gu, Ninglun
Zhang, Kailai
Zhang, Zijiao
Bao, Yelun
Yang, Jin
Yin, Xu
Liu, Liwei
Liu, Yihuan
Li, Pengyong
Yen, Gary G.
Yan, Junchi
author_facet Wang, Jun
Gu, Ninglun
Zhang, Kailai
Zhang, Zijiao
Bao, Yelun
Yang, Jin
Yin, Xu
Liu, Liwei
Liu, Yihuan
Li, Pengyong
Yen, Gary G.
Yan, Junchi
contents For Large Language Models (LLMs), a disconnect persists between benchmark performance and real-world utility. Current evaluation frameworks remain fragmented, prioritizing technical metrics while neglecting holistic assessment for deployment. This survey introduces an anthropomorphic evaluation paradigm through the lens of human intelligence, proposing a novel three-dimensional taxonomy: Intelligence Quotient (IQ)-General Intelligence for foundational capacity, Emotional Quotient (EQ)-Alignment Ability for value-based interactions, and Professional Quotient (PQ)-Professional Expertise for specialized proficiency. For practical value, we pioneer a Value-oriented Evaluation (VQ) framework assessing economic viability, social impact, ethical alignment, and environmental sustainability. Our modular architecture integrates six components with an implementation roadmap. Through analysis of 200+ benchmarks, we identify key challenges including dynamic assessment needs and interpretability gaps. It provides actionable guidance for developing LLMs that are technically proficient, contextually relevant, and ethically sound. We maintain a curated repository of open-source evaluation resources at: https://github.com/onejune2018/Awesome-LLM-Eval.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18646
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
Wang, Jun
Gu, Ninglun
Zhang, Kailai
Zhang, Zijiao
Bao, Yelun
Yang, Jin
Yin, Xu
Liu, Liwei
Liu, Yihuan
Li, Pengyong
Yen, Gary G.
Yan, Junchi
Artificial Intelligence
Computation and Language
For Large Language Models (LLMs), a disconnect persists between benchmark performance and real-world utility. Current evaluation frameworks remain fragmented, prioritizing technical metrics while neglecting holistic assessment for deployment. This survey introduces an anthropomorphic evaluation paradigm through the lens of human intelligence, proposing a novel three-dimensional taxonomy: Intelligence Quotient (IQ)-General Intelligence for foundational capacity, Emotional Quotient (EQ)-Alignment Ability for value-based interactions, and Professional Quotient (PQ)-Professional Expertise for specialized proficiency. For practical value, we pioneer a Value-oriented Evaluation (VQ) framework assessing economic viability, social impact, ethical alignment, and environmental sustainability. Our modular architecture integrates six components with an implementation roadmap. Through analysis of 200+ benchmarks, we identify key challenges including dynamic assessment needs and interpretability gaps. It provides actionable guidance for developing LLMs that are technically proficient, contextually relevant, and ethically sound. We maintain a curated repository of open-source evaluation resources at: https://github.com/onejune2018/Awesome-LLM-Eval.
title Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.18646