Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Hongli, Huang, Hui, Zhao, Ziqing, Han, Lvyuan, Wang, Huicheng, Chen, Kehai, Yang, Muyun, Bao, Wei, Dong, Jian, Xu, Bing, Zhu, Conghui, Cao, Hailong, Zhao, Tiejun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mitigating the Bias of Large Language Model Evaluation
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
di: Song, Yuchen, et al.
Pubblicazione: (2025)
di: Song, Yuchen, et al.
Pubblicazione: (2025)
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models
di: Xu, Mufan, et al.
Pubblicazione: (2026)
di: Xu, Mufan, et al.
Pubblicazione: (2026)
LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
di: Xu, Mufan, et al.
Pubblicazione: (2024)
di: Xu, Mufan, et al.
Pubblicazione: (2024)
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
di: Chen, Andong, et al.
Pubblicazione: (2025)
di: Chen, Andong, et al.
Pubblicazione: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
DUAL-REFLECT: Enhancing Large Language Models for Reflective Translation through Dual Learning Feedback Mechanisms
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
Self-Evaluation of Large Language Model based on Glass-box Features
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
Auditing LLM Benchmarks with Item Response Theory
di: Land, Sander, et al.
Pubblicazione: (2026)
di: Land, Sander, et al.
Pubblicazione: (2026)
Enhancing Large Language Models'Machine Translation via Dynamic Focus Anchoring
di: Ding, Qiuyu, et al.
Pubblicazione: (2025)
di: Ding, Qiuyu, et al.
Pubblicazione: (2025)
Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
di: Tang, Yihong, et al.
Pubblicazione: (2025)
di: Tang, Yihong, et al.
Pubblicazione: (2025)
Dynamic Planning for LLM-based Graphical User Interface Automation
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
User-Aware Active Knowledge Acquisition for Emotional Support Dialogue
di: Xu, Mufan, et al.
Pubblicazione: (2026)
di: Xu, Mufan, et al.
Pubblicazione: (2026)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
A Survey on Human Preference Learning for Large Language Models
di: Jiang, Ruili, et al.
Pubblicazione: (2024)
di: Jiang, Ruili, et al.
Pubblicazione: (2024)
Memory-augmented Query Reconstruction for LLM-based Knowledge Graph Reasoning
di: Xu, Mufan, et al.
Pubblicazione: (2025)
di: Xu, Mufan, et al.
Pubblicazione: (2025)
LLM-based Translation Inference with Iterative Bilingual Understanding
di: Chen, Andong, et al.
Pubblicazione: (2024)
di: Chen, Andong, et al.
Pubblicazione: (2024)
Cross-Domain Bilingual Lexicon Induction via Pretrained Language Models
di: Ding, Qiuyu, et al.
Pubblicazione: (2025)
di: Ding, Qiuyu, et al.
Pubblicazione: (2025)
Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs
di: Zhu, Longyuan, et al.
Pubblicazione: (2026)
di: Zhu, Longyuan, et al.
Pubblicazione: (2026)
Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling
di: Chen, Andong, et al.
Pubblicazione: (2026)
di: Chen, Andong, et al.
Pubblicazione: (2026)
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment
di: Liu, Dongxu, et al.
Pubblicazione: (2024)
di: Liu, Dongxu, et al.
Pubblicazione: (2024)
DesignProbe: A Graphic Design Benchmark for Multimodal Large Language Models
di: Lin, Jieru, et al.
Pubblicazione: (2024)
di: Lin, Jieru, et al.
Pubblicazione: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
di: Zhu, Yingjie, et al.
Pubblicazione: (2024)
di: Zhu, Yingjie, et al.
Pubblicazione: (2024)
Rethinking RGB-D Salient Object Detection: Models, Data Sets, and Large-Scale Benchmarks
di: Fan, Deng-Ping, et al.
Pubblicazione: (2019)
di: Fan, Deng-Ping, et al.
Pubblicazione: (2019)
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
di: Li, Xinyuan, et al.
Pubblicazione: (2025)
di: Li, Xinyuan, et al.
Pubblicazione: (2025)
Exploring the Role of Response Time in Item Response Theory: Rethinking the PISA 2022 Creative Thinking Assessment
di: Lihong Xie, et al.
Pubblicazione: (2025)
di: Lihong Xie, et al.
Pubblicazione: (2025)
Beyond Rigid: Benchmarking Non-Rigid Video Editing
di: Qu, Bingzheng, et al.
Pubblicazione: (2026)
di: Qu, Bingzheng, et al.
Pubblicazione: (2026)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
di: Plaza, Irene, et al.
Pubblicazione: (2024)
di: Plaza, Irene, et al.
Pubblicazione: (2024)
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
di: Uebayashi, Shunki, et al.
Pubblicazione: (2026)
di: Uebayashi, Shunki, et al.
Pubblicazione: (2026)
Dual Instruction Tuning with Large Language Models for Mathematical Reasoning
di: Zhou, Yongwei, et al.
Pubblicazione: (2024)
di: Zhou, Yongwei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Mitigating the Bias of Large Language Model Evaluation
di: Zhou, Hongli, et al.
Pubblicazione: (2024) -
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
di: Zhou, Hongli, et al.
Pubblicazione: (2026) -
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
di: Huang, Hui, et al.
Pubblicazione: (2026) -
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
di: Zhou, Hongli, et al.
Pubblicazione: (2026) -
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
di: Song, Yuchen, et al.
Pubblicazione: (2025)