Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Myrzakhan, Aidar, Bsharat, Sondos Mahmoud, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
Sink-Aware Pruning for Diffusion Language Models
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
von: Rao, Varun, et al.
Veröffentlicht: (2025)
von: Rao, Varun, et al.
Veröffentlicht: (2025)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
von: Kim, Minsang, et al.
Veröffentlicht: (2024)
von: Kim, Minsang, et al.
Veröffentlicht: (2024)
Exploring the Latest LLMs for Leaderboard Extraction
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
von: Li, Yanhong, et al.
Veröffentlicht: (2025)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
Improving LLM Leaderboards with Psychometrical Methodology
von: Federiakin, Denis
Veröffentlicht: (2025)
von: Federiakin, Denis
Veröffentlicht: (2025)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
von: Tran, Khanh-Tung, et al.
Veröffentlicht: (2025)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
von: Cheng, Aileen, et al.
Veröffentlicht: (2025)
von: Cheng, Aileen, et al.
Veröffentlicht: (2025)
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
von: Luo, Yaxin, et al.
Veröffentlicht: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
von: Fu, Weiping, et al.
Veröffentlicht: (2024)
von: Fu, Weiping, et al.
Veröffentlicht: (2024)
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
MIRROR: A Novel Approach for the Automated Evaluation of Open-Ended Question Generation
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
von: Deroy, Aniket, et al.
Veröffentlicht: (2024)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering
von: Keluskar, Aryan, et al.
Veröffentlicht: (2024)
von: Keluskar, Aryan, et al.
Veröffentlicht: (2024)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
von: Veuthey, Jaime Raldua, et al.
Veröffentlicht: (2025)
Open Domain Question Answering with Conflicting Contexts
von: Liu, Siyi, et al.
Veröffentlicht: (2024)
von: Liu, Siyi, et al.
Veröffentlicht: (2024)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
von: Ji, Jiabao, et al.
Veröffentlicht: (2025)
League: Leaderboard Generation on Demand
von: Wu, Jian, et al.
Veröffentlicht: (2025)
von: Wu, Jian, et al.
Veröffentlicht: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
von: Warner, Benjamin, et al.
Veröffentlicht: (2026)
von: Warner, Benjamin, et al.
Veröffentlicht: (2026)
Closing the Data Loop: Using OpenDataArena to Engineer Superior Training Datasets
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
MFORT-QA: Multi-hop Few-shot Open Rich Table Question Answering
von: Guan, Che, et al.
Veröffentlicht: (2024)
von: Guan, Che, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2023) -
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025) -
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025) -
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025) -
Sink-Aware Pruning for Diffusion Language Models
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2026)