Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914161636671488 |
|---|---|
| author | Wang, Jun Gu, Ninglun Zhang, Kailai Zhang, Zijiao Bao, Yelun Yang, Jin Yin, Xu Liu, Liwei Liu, Yihuan Li, Pengyong Yen, Gary G. Yan, Junchi |
| author_facet | Wang, Jun Gu, Ninglun Zhang, Kailai Zhang, Zijiao Bao, Yelun Yang, Jin Yin, Xu Liu, Liwei Liu, Yihuan Li, Pengyong Yen, Gary G. Yan, Junchi |
| contents | For Large Language Models (LLMs), a disconnect persists between benchmark performance and real-world utility. Current evaluation frameworks remain fragmented, prioritizing technical metrics while neglecting holistic assessment for deployment. This survey introduces an anthropomorphic evaluation paradigm through the lens of human intelligence, proposing a novel three-dimensional taxonomy: Intelligence Quotient (IQ)-General Intelligence for foundational capacity, Emotional Quotient (EQ)-Alignment Ability for value-based interactions, and Professional Quotient (PQ)-Professional Expertise for specialized proficiency. For practical value, we pioneer a Value-oriented Evaluation (VQ) framework assessing economic viability, social impact, ethical alignment, and environmental sustainability. Our modular architecture integrates six components with an implementation roadmap. Through analysis of 200+ benchmarks, we identify key challenges including dynamic assessment needs and interpretability gaps. It provides actionable guidance for developing LLMs that are technically proficient, contextually relevant, and ethically sound. We maintain a curated repository of open-source evaluation resources at: https://github.com/onejune2018/Awesome-LLM-Eval. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_18646 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap Wang, Jun Gu, Ninglun Zhang, Kailai Zhang, Zijiao Bao, Yelun Yang, Jin Yin, Xu Liu, Liwei Liu, Yihuan Li, Pengyong Yen, Gary G. Yan, Junchi Artificial Intelligence Computation and Language For Large Language Models (LLMs), a disconnect persists between benchmark performance and real-world utility. Current evaluation frameworks remain fragmented, prioritizing technical metrics while neglecting holistic assessment for deployment. This survey introduces an anthropomorphic evaluation paradigm through the lens of human intelligence, proposing a novel three-dimensional taxonomy: Intelligence Quotient (IQ)-General Intelligence for foundational capacity, Emotional Quotient (EQ)-Alignment Ability for value-based interactions, and Professional Quotient (PQ)-Professional Expertise for specialized proficiency. For practical value, we pioneer a Value-oriented Evaluation (VQ) framework assessing economic viability, social impact, ethical alignment, and environmental sustainability. Our modular architecture integrates six components with an implementation roadmap. Through analysis of 200+ benchmarks, we identify key challenges including dynamic assessment needs and interpretability gaps. It provides actionable guidance for developing LLMs that are technically proficient, contextually relevant, and ethically sound. We maintain a curated repository of open-source evaluation resources at: https://github.com/onejune2018/Awesome-LLM-Eval. |
| title | Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2508.18646 |