Saved in:
| Main Authors: | Khan, Haidar, Alyahya, Hisham A., Alnumay, Yazeed, Bari, M Saiful, Yener, Bülent |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.12562 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
by: Alyahya, Hisham A., et al.
Published: (2025)
by: Alyahya, Hisham A., et al.
Published: (2025)
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
by: Alzahrani, Norah, et al.
Published: (2024)
by: Alzahrani, Norah, et al.
Published: (2024)
ALLaM: Large Language Models for Arabic and English
by: Bari, M Saiful, et al.
Published: (2024)
by: Bari, M Saiful, et al.
Published: (2024)
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024)
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
by: Alyahya, Ghadi, et al.
Published: (2024)
by: Alyahya, Ghadi, et al.
Published: (2024)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
by: Nayeem, Mir Tafseer, et al.
Published: (2025)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
by: Alqahtani, Sawsan, et al.
Published: (2026)
by: Alqahtani, Sawsan, et al.
Published: (2026)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024)
by: Lee, Yuho, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
Competition-Level Problems are Effective LLM Evaluators
by: Huang, Yiming, et al.
Published: (2023)
by: Huang, Yiming, et al.
Published: (2023)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
Zero-Shot Commonsense Validation and Reasoning with Large Language Models: An Evaluation on SemEval-2020 Task 4 Dataset
by: Alfugaha, Rawand, et al.
Published: (2025)
by: Alfugaha, Rawand, et al.
Published: (2025)
Large EEG-U-Transformer for Time-Step Level Detection Without Pre-Training
by: Wu, Kerui, et al.
Published: (2025)
by: Wu, Kerui, et al.
Published: (2025)
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
by: Prato, Gabriele, et al.
Published: (2023)
by: Prato, Gabriele, et al.
Published: (2023)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
SHROOM-INDElab at SemEval-2024 Task 6: Zero- and Few-Shot LLM-Based Classification for Hallucination Detection
by: Allen, Bradley P., et al.
Published: (2024)
by: Allen, Bradley P., et al.
Published: (2024)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2025)
by: Pokharel, Rhitabrat, et al.
Published: (2025)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
by: He, Yuanqin, et al.
Published: (2024)
by: He, Yuanqin, et al.
Published: (2024)
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges
by: Abdellatif, Mohamed Hisham
Published: (2025)
by: Abdellatif, Mohamed Hisham
Published: (2025)
QualEval: Qualitative Evaluation for Model Improvement
by: Murahari, Vishvak, et al.
Published: (2023)
by: Murahari, Vishvak, et al.
Published: (2023)
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs
by: Haq, Saiful, et al.
Published: (2025)
by: Haq, Saiful, et al.
Published: (2025)
ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature
by: Sinha, Aarush, et al.
Published: (2025)
by: Sinha, Aarush, et al.
Published: (2025)
AraSpot: Arabic Spoken Command Spotting
by: Salhab, Mahmoud, et al.
Published: (2023)
by: Salhab, Mahmoud, et al.
Published: (2023)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
by: Sansford, Hannah, et al.
Published: (2024)
by: Sansford, Hannah, et al.
Published: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
Evaluating Zero-Shot Long-Context LLM Compression
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Measuring all the noises of LLM Evals
by: Wang, Sida
Published: (2025)
by: Wang, Sida
Published: (2025)
LuxVeri at GenAI Detection Task 3: Cross-Domain Detection of AI-Generated Text Using Inverse Perplexity-Weighted Ensemble of Fine-Tuned Transformer Models
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
by: Mobin, Md Kamrujjaman, et al.
Published: (2025)
TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy
by: McCammon, James
Published: (2025)
by: McCammon, James
Published: (2025)
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
by: Gupta, Prannaya, et al.
Published: (2024)
by: Gupta, Prannaya, et al.
Published: (2024)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
by: Yu, Linhao, et al.
Published: (2024)
by: Yu, Linhao, et al.
Published: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025)
by: Ren, Huimin, et al.
Published: (2025)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
by: Zhao, Jingqian, et al.
Published: (2025)
by: Zhao, Jingqian, et al.
Published: (2025)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
by: Garg, Madhav Krishan, et al.
Published: (2025)
by: Garg, Madhav Krishan, et al.
Published: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
IndicEval: A Bilingual Indian Educational Evaluation Framework for Large Language Models
by: Bharti, Saurabh, et al.
Published: (2026)
by: Bharti, Saurabh, et al.
Published: (2026)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Similar Items
-
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
by: Alyahya, Hisham A., et al.
Published: (2025) -
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
by: Alzahrani, Norah, et al.
Published: (2024) -
ALLaM: Large Language Models for Arabic and English
by: Bari, M Saiful, et al.
Published: (2024) -
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
by: Laskar, Md Tahmid Rahman, et al.
Published: (2024) -
Hatred Stems from Ignorance! Distillation of the Persuasion Modes in Countering Conversational Hate Speech
by: Alyahya, Ghadi, et al.
Published: (2024)