Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Peiyu, Tang, Xiuxiu, Chen, Si, Cheng, Ying, Metoyer, Ronald, Hua, Ting, Chawla, Nitesh V. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
by: Chen, Si, et al.
Published: (2026)
by: Chen, Si, et al.
Published: (2026)
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
by: Molnar, Isabel, et al.
Published: (2026)
by: Molnar, Isabel, et al.
Published: (2026)
Building Scaffolding Dialogue Data with LLM-Simulated Novices
by: Chen, Si, et al.
Published: (2025)
by: Chen, Si, et al.
Published: (2025)
Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
by: Chen, Si, et al.
Published: (2025)
by: Chen, Si, et al.
Published: (2025)
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
by: Li, Peiyu, et al.
Published: (2025)
by: Li, Peiyu, et al.
Published: (2025)
Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams
by: Tang, Xiuxiu, et al.
Published: (2026)
by: Tang, Xiuxiu, et al.
Published: (2026)
Understanding Student Attitudes and Acceptability of GenAI Tools in Higher Ed: Scale Development and Evaluation
by: Tang, Xiuxiu, et al.
Published: (2025)
by: Tang, Xiuxiu, et al.
Published: (2025)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
by: Guo, Taicheng, et al.
Published: (2026)
by: Guo, Taicheng, et al.
Published: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
AI Academy: Building Generative AI Literacy in Higher Ed Instructors
by: Chen, Si, et al.
Published: (2025)
by: Chen, Si, et al.
Published: (2025)
FLAME: Towards Federated Fine-Tuning Large Language Models Through Adaptive SMoE
by: Le, Khiem, et al.
Published: (2025)
by: Le, Khiem, et al.
Published: (2025)
Breaking Language Barriers: Equitable Performance in Multilingual Language Models
by: Nagar, Tanay, et al.
Published: (2025)
by: Nagar, Tanay, et al.
Published: (2025)
Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQL
by: Wang, Yihan, et al.
Published: (2026)
by: Wang, Yihan, et al.
Published: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
by: Shi, Zhichao, et al.
Published: (2025)
by: Shi, Zhichao, et al.
Published: (2025)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
by: Tian, Yijun, et al.
Published: (2024)
by: Tian, Yijun, et al.
Published: (2024)
AgentDrug: Utilizing Large Language Models in An Agentic Workflow for Zero-Shot Molecular Editing
by: Le, Khiem, et al.
Published: (2024)
by: Le, Khiem, et al.
Published: (2024)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
by: Ye, Jiayi, et al.
Published: (2024)
by: Ye, Jiayi, et al.
Published: (2024)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
Evaluating Implicit Bias in Large Language Models by Attacking From a Psychometric Perspective
by: Wen, Yuchen, et al.
Published: (2024)
by: Wen, Yuchen, et al.
Published: (2024)
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
by: Karr Jr., Jonathan A., et al.
Published: (2025)
by: Karr Jr., Jonathan A., et al.
Published: (2025)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)
by: Liu, Yunting, et al.
Published: (2024)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
by: Yoa, Seungdong, et al.
Published: (2026)
by: Yoa, Seungdong, et al.
Published: (2026)
Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking
by: Zheng, Tianpeng, et al.
Published: (2026)
by: Zheng, Tianpeng, et al.
Published: (2026)
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
by: Tang, Wenjie, et al.
Published: (2025)
by: Tang, Wenjie, et al.
Published: (2025)
Human Psychometric Questionnaires Mischaracterize LLM Behavior
by: Song, Woojung, et al.
Published: (2025)
by: Song, Woojung, et al.
Published: (2025)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
by: Wei, Tianxin, et al.
Published: (2025)
by: Wei, Tianxin, et al.
Published: (2025)
Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions
by: Baidya, Madhav S., et al.
Published: (2026)
by: Baidya, Madhav S., et al.
Published: (2026)
Context Attribution with Multi-Armed Bandit Optimization
by: Pan, Deng, et al.
Published: (2025)
by: Pan, Deng, et al.
Published: (2025)
Improving LLM Leaderboards with Psychometrical Methodology
by: Federiakin, Denis
Published: (2025)
by: Federiakin, Denis
Published: (2025)
T-RAG: Lessons from the LLM Trenches
by: Fatehkia, Masoomali, et al.
Published: (2024)
by: Fatehkia, Masoomali, et al.
Published: (2024)
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
by: Lee, Huije, et al.
Published: (2026)
by: Lee, Huije, et al.
Published: (2026)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
GenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing
by: Wang, Ming, et al.
Published: (2026)
by: Wang, Ming, et al.
Published: (2026)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
by: Su, Zhaochen, et al.
Published: (2024)
by: Su, Zhaochen, et al.
Published: (2024)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
Large Language Model based Multi-Agents: A Survey of Progress and Challenges
by: Guo, Taicheng, et al.
Published: (2024)
by: Guo, Taicheng, et al.
Published: (2024)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
by: Sun, Shengyin, et al.
Published: (2025)
by: Sun, Shengyin, et al.
Published: (2025)
Recent Advances in Large Langauge Model Benchmarks against Data Contamination: From Static to Dynamic Evaluation
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
by: Song, Yiliang, et al.
Published: (2026)
by: Song, Yiliang, et al.
Published: (2026)
Similar Items
-
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
by: Chen, Si, et al.
Published: (2026) -
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
by: Molnar, Isabel, et al.
Published: (2026) -
Building Scaffolding Dialogue Data with LLM-Simulated Novices
by: Chen, Si, et al.
Published: (2025) -
Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
by: Chen, Si, et al.
Published: (2025) -
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
by: Li, Peiyu, et al.
Published: (2025)