Reactor Mk.1 performances: MMLU, HumanEval and BBH test results
Fuente:
arXiv
Saved in:
| Main Authors: | Dunham, TJ, Syahputra, Henry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HumanEval on Latest GPT Models -- 2024
by: Li, Daniel, et al.
Published: (2024)
by: Li, Daniel, et al.
Published: (2024)
Are We Done with MMLU?
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
by: Plaza, Irene, et al.
Published: (2024)
by: Plaza, Irene, et al.
Published: (2024)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
On the Foundations of Trustworthy Artificial Intelligence
by: Dunham, TJ
Published: (2026)
by: Dunham, TJ
Published: (2026)
HumanEval-V: Benchmarking High-Level Visual Reasoning with Complex Diagrams in Coding Tasks
by: Zhang, Fengji, et al.
Published: (2024)
by: Zhang, Fengji, et al.
Published: (2024)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
by: Ghahroodi, Omid, et al.
Published: (2024)
by: Ghahroodi, Omid, et al.
Published: (2024)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
Qiskit HumanEval: An Evaluation Benchmark For Quantum Code Generative Models
by: Vishwakarma, Sanjay, et al.
Published: (2024)
by: Vishwakarma, Sanjay, et al.
Published: (2024)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
by: KJ, Sankalp, et al.
Published: (2025)
by: KJ, Sankalp, et al.
Published: (2025)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
by: Wang, Wentian, et al.
Published: (2024)
by: Wang, Wentian, et al.
Published: (2024)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
by: Zhao, Qihao, et al.
Published: (2024)
by: Zhao, Qihao, et al.
Published: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
by: Peng, Qiwei, et al.
Published: (2024)
by: Peng, Qiwei, et al.
Published: (2024)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
by: Bradbury, Jeremy S., et al.
Published: (2024)
by: Bradbury, Jeremy S., et al.
Published: (2024)
TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy
by: McCammon, James
Published: (2025)
by: McCammon, James
Published: (2025)
PatentEval: Understanding Errors in Patent Generation
by: Zuo, You, et al.
Published: (2024)
by: Zuo, You, et al.
Published: (2024)
WangchanLion and WangchanX MRC Eval
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2024)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
COGNAC at SemEval-2026 Task 5: LLM Ensembles for Human-Level Word Sense Plausibility Rating in Challenging Narratives
by: Islam, Azwad Anjum, et al.
Published: (2026)
by: Islam, Azwad Anjum, et al.
Published: (2026)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
by: Zhang, Shudan, et al.
Published: (2024)
by: Zhang, Shudan, et al.
Published: (2024)
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
by: Garg, Madhav Krishan, et al.
Published: (2025)
by: Garg, Madhav Krishan, et al.
Published: (2025)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
by: Sun, Chongyan, et al.
Published: (2024)
by: Sun, Chongyan, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
CausalEval: Towards Better Causal Reasoning in Language Models
by: Yu, Longxuan, et al.
Published: (2024)
by: Yu, Longxuan, et al.
Published: (2024)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
by: Wang, Clinton J., et al.
Published: (2025)
by: Wang, Clinton J., et al.
Published: (2025)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
by: Pokharel, Rhitabrat, et al.
Published: (2025)
by: Pokharel, Rhitabrat, et al.
Published: (2025)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
by: Prato, Gabriele, et al.
Published: (2023)
by: Prato, Gabriele, et al.
Published: (2023)
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
by: Ma, Zi-Ao, et al.
Published: (2025)
by: Ma, Zi-Ao, et al.
Published: (2025)
CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks
by: Qian, Zhaozhi, et al.
Published: (2024)
by: Qian, Zhaozhi, et al.
Published: (2024)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)
by: Zhang, Mengyuan, et al.
Published: (2024)
SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition
by: Osman, Mohamed, et al.
Published: (2024)
by: Osman, Mohamed, et al.
Published: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
Check-Eval: A Checklist-based Approach for Evaluating Text Quality
by: Pereira, Jayr, et al.
Published: (2024)
by: Pereira, Jayr, et al.
Published: (2024)
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
by: Ye, Fangda, et al.
Published: (2026)
by: Ye, Fangda, et al.
Published: (2026)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
A Single Character can Make or Break Your LLM Evals
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
by: Yin, Joy Lim Jia, et al.
Published: (2025)
by: Yin, Joy Lim Jia, et al.
Published: (2025)
Similar Items
-
HumanEval on Latest GPT Models -- 2024
by: Li, Daniel, et al.
Published: (2024) -
Are We Done with MMLU?
by: Gema, Aryo Pradipta, et al.
Published: (2024) -
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
by: Plaza, Irene, et al.
Published: (2024) -
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025) -
On the Foundations of Trustworthy Artificial Intelligence
by: Dunham, TJ
Published: (2026)