ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Alyahya, Hisham A., Khan, Haidar, Alnumay, Yazeed, Bari, M Saiful, Yener, Bülent |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
von: Khan, Haidar, et al.
Veröffentlicht: (2025)
von: Khan, Haidar, et al.
Veröffentlicht: (2025)
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
von: Alzahrani, Norah, et al.
Veröffentlicht: (2024)
von: Alzahrani, Norah, et al.
Veröffentlicht: (2024)
ALLaM: Large Language Models for Arabic and English
von: Bari, M Saiful, et al.
Veröffentlicht: (2024)
von: Bari, M Saiful, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
von: Lee, Yuho, et al.
Veröffentlicht: (2024)
von: Lee, Yuho, et al.
Veröffentlicht: (2024)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
von: Garg, Madhav Krishan, et al.
Veröffentlicht: (2025)
von: Garg, Madhav Krishan, et al.
Veröffentlicht: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
von: Sansford, Hannah, et al.
Veröffentlicht: (2024)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2025)
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2025)
Competition-Level Problems are Effective LLM Evaluators
von: Huang, Yiming, et al.
Veröffentlicht: (2023)
von: Huang, Yiming, et al.
Veröffentlicht: (2023)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
IndicEval: A Bilingual Indian Educational Evaluation Framework for Large Language Models
von: Bharti, Saurabh, et al.
Veröffentlicht: (2026)
von: Bharti, Saurabh, et al.
Veröffentlicht: (2026)
Zero-Shot Commonsense Validation and Reasoning with Large Language Models: An Evaluation on SemEval-2020 Task 4 Dataset
von: Alfugaha, Rawand, et al.
Veröffentlicht: (2025)
von: Alfugaha, Rawand, et al.
Veröffentlicht: (2025)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
von: Mohammadi, Hadi, et al.
Veröffentlicht: (2025)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
CriticEval: Evaluating Large Language Model as Critic
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
SHROOM-INDElab at SemEval-2024 Task 6: Zero- and Few-Shot LLM-Based Classification for Hallucination Detection
von: Allen, Bradley P., et al.
Veröffentlicht: (2024)
von: Allen, Bradley P., et al.
Veröffentlicht: (2024)
EvalCards: A Framework for Standardized Evaluation Reporting
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
von: Dhar, Ruchira, et al.
Veröffentlicht: (2025)
Structure-BiEval: A Self-Supervised, Dual-Track Framework for Decoupling Structure and Content in LLM Evaluation for Web Information Systems
von: Zhao, Boxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Boxiang, et al.
Veröffentlicht: (2026)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
von: Pokharel, Rhitabrat, et al.
Veröffentlicht: (2025)
von: Pokharel, Rhitabrat, et al.
Veröffentlicht: (2025)
LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
von: Diao, Shizhe, et al.
Veröffentlicht: (2023)
von: Diao, Shizhe, et al.
Veröffentlicht: (2023)
MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
von: Song, Seyoung, et al.
Veröffentlicht: (2025)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
von: Wang, Ganghua, et al.
Veröffentlicht: (2025)
VHDL-Eval: A Framework for Evaluating Large Language Models in VHDL Code Generation
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2024)
von: Vijayaraghavan, Prashanth, et al.
Veröffentlicht: (2024)
ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature
von: Sinha, Aarush, et al.
Veröffentlicht: (2025)
von: Sinha, Aarush, et al.
Veröffentlicht: (2025)
QualEval: Qualitative Evaluation for Model Improvement
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
von: Murahari, Vishvak, et al.
Veröffentlicht: (2023)
Evaluating Zero-Shot Long-Context LLM Compression
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023)
von: Shu, Lei, et al.
Veröffentlicht: (2023)
Large EEG-U-Transformer for Time-Step Level Detection Without Pre-Training
von: Wu, Kerui, et al.
Veröffentlicht: (2025)
von: Wu, Kerui, et al.
Veröffentlicht: (2025)
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges
von: Abdellatif, Mohamed Hisham
Veröffentlicht: (2025)
von: Abdellatif, Mohamed Hisham
Veröffentlicht: (2025)
Ähnliche Einträge
-
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
von: Khan, Haidar, et al.
Veröffentlicht: (2025) -
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
von: Alzahrani, Norah, et al.
Veröffentlicht: (2024) -
ALLaM: Large Language Models for Arabic and English
von: Bari, M Saiful, et al.
Veröffentlicht: (2024) -
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024) -
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
von: Alyahya, Ghadi, et al.
Veröffentlicht: (2024)