Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jun, Gu, Ninglun, Zhang, Kailai, Zhang, Zijiao, Bao, Yelun, Yang, Jin, Yin, Xu, Liu, Liwei, Liu, Yihuan, Li, Pengyong, Yen, Gary G., Yan, Junchi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
by: Wu, Xiaojun, et al.
Published: (2025)
by: Wu, Xiaojun, et al.
Published: (2025)
TN-AutoRCA: Benchmark Construction and Agentic Framework for Self-Improving Alarm-Based Root Cause Analysis in Telecommunication Networks
by: Wu, Keyu, et al.
Published: (2025)
by: Wu, Keyu, et al.
Published: (2025)
Beyond Anthropomorphism: a Spectrum of Interface Metaphors for LLMs
by: So, Jianna, et al.
Published: (2026)
by: So, Jianna, et al.
Published: (2026)
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
by: Liu, Zhiwei, et al.
Published: (2025)
by: Liu, Zhiwei, et al.
Published: (2025)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
by: Zhang, Hongbo, et al.
Published: (2025)
by: Zhang, Hongbo, et al.
Published: (2025)
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
by: Zhang, Yazhou, et al.
Published: (2025)
by: Zhang, Yazhou, et al.
Published: (2025)
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Biological Effects of Dietary Restriction on Alzheimer's Disease: Experimental and Clinical Investigations
by: Zijiao Liu, et al.
Published: (2025)
by: Zijiao Liu, et al.
Published: (2025)
Multi‐class financial distress prediction based on stacking ensemble method
by: Xiaofang Chen, et al.
Published: (2024)
by: Xiaofang Chen, et al.
Published: (2024)
Stereodivergent Construction of 1,5/1,7‐Nonadjacent Tetrasubstituted Stereocenters Enabled by Pd/Cu‐Cocatalyzed Asymmetric Heck Cascade Reaction
by: Panpan Li, et al.
Published: (2024)
by: Panpan Li, et al.
Published: (2024)
EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
A Benchmark of Dexterity for Anthropomorphic Robotic Hands
by: Liconti, Davide, et al.
Published: (2026)
by: Liconti, Davide, et al.
Published: (2026)
POLIS-Bench: Towards Multi-Dimensional Evaluation of LLMs for Bilingual Policy Tasks in Governmental Scenarios
by: Yang, Tingyue, et al.
Published: (2025)
by: Yang, Tingyue, et al.
Published: (2025)
Entropy-Tree: Tree-Based Decoding with Entropy-Guided Exploration
by: Wei, Longxuan, et al.
Published: (2026)
by: Wei, Longxuan, et al.
Published: (2026)
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
by: Zhong, Yiwu, et al.
Published: (2024)
by: Zhong, Yiwu, et al.
Published: (2024)
Beyond Benchmarking: A New Paradigm for Evaluation and Assessment of Large Language Models
by: Liu, Jin, et al.
Published: (2024)
by: Liu, Jin, et al.
Published: (2024)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
by: Li, Haoming, et al.
Published: (2025)
by: Li, Haoming, et al.
Published: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
by: Liu, Zhiqiang, et al.
Published: (2025)
by: Liu, Zhiqiang, et al.
Published: (2025)
Beyond Personhood: Agency, Accountability, and the Limits of Anthropomorphic Ethical Analysis
by: Dai, Jessica
Published: (2024)
by: Dai, Jessica
Published: (2024)
Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
by: Sophia, Franny Philos
Published: (2026)
by: Sophia, Franny Philos
Published: (2026)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
by: Wang, Xintao, et al.
Published: (2026)
by: Wang, Xintao, et al.
Published: (2026)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Software Security Analysis in 2030 and Beyond: A Research Roadmap
by: Böhme, Marcel, et al.
Published: (2024)
by: Böhme, Marcel, et al.
Published: (2024)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Flames: Benchmarking Value Alignment of LLMs in Chinese
by: Huang, Kexin, et al.
Published: (2023)
by: Huang, Kexin, et al.
Published: (2023)
Robo-Advisors Beyond Automation: Principles and Roadmap for AI-Driven Financial Planning
by: Feng, Runhuan, et al.
Published: (2025)
by: Feng, Runhuan, et al.
Published: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
From Pixels to Personas: Investigating and Modeling Self-Anthropomorphism in Human-Robot Dialogues
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
The Roadmap to 6G -- AI Empowered Wireless Networks
by: Letaief, Khaled B., et al.
Published: (2019)
by: Letaief, Khaled B., et al.
Published: (2019)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
by: Zhu, Hongda, et al.
Published: (2025)
by: Zhu, Hongda, et al.
Published: (2025)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
by: Yin, Zhiyu, et al.
Published: (2026)
by: Yin, Zhiyu, et al.
Published: (2026)
UCRBench: Benchmarking LLMs on Use Case Recovery
by: Xiao, Shuyuan, et al.
Published: (2025)
by: Xiao, Shuyuan, et al.
Published: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
by: Costa, Mariana Lins
Published: (2026)
by: Costa, Mariana Lins
Published: (2026)
Visual Anthropomorphism Shifts Evaluations of Gendered AI Managers
by: Han, Ruiqing, et al.
Published: (2026)
by: Han, Ruiqing, et al.
Published: (2026)
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
Similar Items
-
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
by: Wu, Xiaojun, et al.
Published: (2025) -
TN-AutoRCA: Benchmark Construction and Agentic Framework for Self-Improving Alarm-Based Root Cause Analysis in Telecommunication Networks
by: Wu, Keyu, et al.
Published: (2025) -
Beyond Anthropomorphism: a Spectrum of Interface Metaphors for LLMs
by: So, Jianna, et al.
Published: (2026) -
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
by: Liu, Zhiwei, et al.
Published: (2025) -
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)