Position: AI Evaluation Should Learn from How We Test Humans
Fuente:
arXiv
Saved in:
| Main Authors: | Zhuang, Yan, Liu, Qi, Pardos, Zachary A., Kyllonen, Patrick C., Zu, Jiyun, Huang, Zhenya, Wang, Shijin, Chen, Enhong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
by: Zhuang, Yan, et al.
Published: (2024)
by: Zhuang, Yan, et al.
Published: (2024)
Explainable Automatic Grading with Neural Additive Models
by: Condor, Aubrey, et al.
Published: (2024)
by: Condor, Aubrey, et al.
Published: (2024)
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
by: Zhang, Zhexin, et al.
Published: (2025)
by: Zhang, Zhexin, et al.
Published: (2025)
What Makes In-context Learning Effective for Mathematical Reasoning: A Theoretical Analysis
by: Liu, Jiayu, et al.
Published: (2024)
by: Liu, Jiayu, et al.
Published: (2024)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)
by: Liu, Yunting, et al.
Published: (2024)
Learning to Solve Geometry Problems via Simulating Human Dual-Reasoning Process
by: Xiao, Tong, et al.
Published: (2024)
by: Xiao, Tong, et al.
Published: (2024)
TestAgent: An Adaptive and Intelligent Expert for Human Assessment
by: Yu, Junhao, et al.
Published: (2025)
by: Yu, Junhao, et al.
Published: (2025)
How Should We Model the Probability of a Language?
by: Dent, Rasul, et al.
Published: (2026)
by: Dent, Rasul, et al.
Published: (2026)
We Should Evaluate Real-World Impact
by: Reiter, Ehud
Published: (2025)
by: Reiter, Ehud
Published: (2025)
Towards Personalized Evaluation of Large Language Models with An Anonymous Crowd-Sourcing Platform
by: Cheng, Mingyue, et al.
Published: (2024)
by: Cheng, Mingyue, et al.
Published: (2024)
Automated Coding of Communication Data Using ChatGPT: Consistency Across Subgroups
by: Hao, Jiangang, et al.
Published: (2025)
by: Hao, Jiangang, et al.
Published: (2025)
A Knowledge-Injected Curriculum Pretraining Framework for Question Answering
by: Lin, Xin, et al.
Published: (2024)
by: Lin, Xin, et al.
Published: (2024)
Refining Sentence Embedding Model through Ranking Sentences Generation with Large Language Models
by: He, Liyang, et al.
Published: (2025)
by: He, Liyang, et al.
Published: (2025)
pyBKT: An Accessible Python Library of Bayesian Knowledge Tracing Models
by: Badrinath, Anirudhan, et al.
Published: (2021)
by: Badrinath, Anirudhan, et al.
Published: (2021)
How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling
by: Liu, Yuhang, et al.
Published: (2026)
by: Liu, Yuhang, et al.
Published: (2026)
Unified Uncertainty Estimation for Cognitive Diagnosis Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Human-AI Collaboration Increases Skill Tagging Speed but Degrades Accuracy
by: Ren, Cheng, et al.
Published: (2024)
by: Ren, Cheng, et al.
Published: (2024)
From Objectives to Questions: A Planning-based Framework for Educational Mathematical Question Generation
by: Cheng, Cheng, et al.
Published: (2025)
by: Cheng, Cheng, et al.
Published: (2025)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
by: Zhao, Yuze, et al.
Published: (2026)
by: Zhao, Yuze, et al.
Published: (2026)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
End-to-End Graph Flattening Method for Large Language Models
by: Hong, Bin, et al.
Published: (2024)
by: Hong, Bin, et al.
Published: (2024)
Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
by: Gan, Aoran, et al.
Published: (2025)
by: Gan, Aoran, et al.
Published: (2025)
Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
by: Alghisi, Simone, et al.
Published: (2024)
by: Alghisi, Simone, et al.
Published: (2024)
Automated Coding of Communications in Collaborative Problem-solving Tasks Using ChatGPT
by: Hao, Jiangang, et al.
Published: (2024)
by: Hao, Jiangang, et al.
Published: (2024)
EduNLP: Towards a Unified and Modularized Library for Educational Resources
by: Huang, Zhenya, et al.
Published: (2024)
by: Huang, Zhenya, et al.
Published: (2024)
Just Because We Camp, Doesn't Mean We Should: The Ethics of Modelling Queer Voices
by: Sigurgeirsson, Atli, et al.
Published: (2024)
by: Sigurgeirsson, Atli, et al.
Published: (2024)
Should We Still Pretrain Encoders with Masked Language Modeling?
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2025)
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2025)
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
by: Mousavi, Pooneh, et al.
Published: (2024)
by: Mousavi, Pooneh, et al.
Published: (2024)
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
Bit-mask Robust Contrastive Knowledge Distillation for Unsupervised Semantic Hashing
by: He, Liyang, et al.
Published: (2024)
by: He, Liyang, et al.
Published: (2024)
LLMs as Data Annotators: How Close Are We to Human Performance
by: Haq, Muhammad Uzair Ul, et al.
Published: (2025)
by: Haq, Muhammad Uzair Ul, et al.
Published: (2025)
Auditing an Automatic Grading Model with deep Reinforcement Learning
by: Condor, Aubrey, et al.
Published: (2024)
by: Condor, Aubrey, et al.
Published: (2024)
Dense X Retrieval: What Retrieval Granularity Should We Use?
by: Chen, Tong, et al.
Published: (2023)
by: Chen, Tong, et al.
Published: (2023)
LLMs Should Incorporate Explicit Mechanisms for Human Empathy
by: You, Xiaoxing, et al.
Published: (2026)
by: You, Xiaoxing, et al.
Published: (2026)
Should We be Pedantic About Reasoning Errors in Machine Translation?
by: Bao, Calvin, et al.
Published: (2026)
by: Bao, Calvin, et al.
Published: (2026)
Are Word Embedding Methods Stable and Should We Care About It?
by: Borah, Angana, et al.
Published: (2021)
by: Borah, Angana, et al.
Published: (2021)
The RIGID Framework: Research-Integrated, Generative AI-Mediated Instructional Design
by: Kwak, Yerin, et al.
Published: (2026)
by: Kwak, Yerin, et al.
Published: (2026)
Position: Towards Bidirectional Human-AI Alignment
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
We Should Chart an Atlas of All the World's Models
by: Horwitz, Eliahu, et al.
Published: (2025)
by: Horwitz, Eliahu, et al.
Published: (2025)
Environment-Aware Code Generation: How far are We?
by: Wu, Tongtong, et al.
Published: (2026)
by: Wu, Tongtong, et al.
Published: (2026)
Similar Items
-
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
by: Zhuang, Yan, et al.
Published: (2024) -
Explainable Automatic Grading with Neural Additive Models
by: Condor, Aubrey, et al.
Published: (2024) -
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
by: Zhang, Zhexin, et al.
Published: (2025) -
What Makes In-context Learning Effective for Mathematical Reasoning: A Theoretical Analysis
by: Liu, Jiayu, et al.
Published: (2024) -
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)