MastermindEval: A Simple But Scalable Reasoning Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Golde, Jonas, Haller, Patrick, Barth, Fabio, Akbik, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2025)
by: Golde, Jonas, et al.
Published: (2025)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
by: Golde, Jonas, et al.
Published: (2026)
by: Golde, Jonas, et al.
Published: (2026)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
by: Golde, Jonas, et al.
Published: (2024)
by: Golde, Jonas, et al.
Published: (2024)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
by: Golde, Jonas, et al.
Published: (2023)
by: Golde, Jonas, et al.
Published: (2023)
Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2024)
by: Golde, Jonas, et al.
Published: (2024)
PECC: Problem Extraction and Coding Challenges
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
Question Decomposition for Retrieval-Augmented Generation
by: Ammann, Paul J. L., et al.
Published: (2025)
by: Ammann, Paul J. L., et al.
Published: (2025)
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
by: Aynetdinov, Ansar, et al.
Published: (2026)
by: Aynetdinov, Ansar, et al.
Published: (2026)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
by: Christoph, Daniel, et al.
Published: (2025)
by: Christoph, Daniel, et al.
Published: (2025)
Fundus: A Simple-to-Use News Scraper Optimized for High Quality Extractions
by: Dallabetta, Max, et al.
Published: (2024)
by: Dallabetta, Max, et al.
Published: (2024)
Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation
by: Rücker, Susanna, et al.
Published: (2025)
by: Rücker, Susanna, et al.
Published: (2025)
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
by: Aynetdinov, Ansar, et al.
Published: (2024)
by: Aynetdinov, Ansar, et al.
Published: (2024)
LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
by: Zhang, Yadong, et al.
Published: (2024)
by: Zhang, Yadong, et al.
Published: (2024)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
by: Merdjanovska, Elena, et al.
Published: (2024)
by: Merdjanovska, Elena, et al.
Published: (2024)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025)
by: Aynetdinov, Ansar, et al.
Published: (2025)
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
by: Wiland, Jacek, et al.
Published: (2024)
by: Wiland, Jacek, et al.
Published: (2024)
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
by: Garbas, Lukas, et al.
Published: (2024)
by: Garbas, Lukas, et al.
Published: (2024)
Multilingual European Language Models: Benchmarking Approaches and Challenges
by: Barth, Fabio, et al.
Published: (2025)
by: Barth, Fabio, et al.
Published: (2025)
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
by: Toporkov, Olia, et al.
Published: (2025)
by: Toporkov, Olia, et al.
Published: (2025)
Towards a Principled Evaluation of Knowledge Editors
by: Pohl, Sebastian, et al.
Published: (2025)
by: Pohl, Sebastian, et al.
Published: (2025)
Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration
by: Kissling, Christopher, et al.
Published: (2026)
by: Kissling, Christopher, et al.
Published: (2026)
Hierarchical Text Classification with LLM-Refined Taxonomies
by: Golde, Jonas, et al.
Published: (2026)
by: Golde, Jonas, et al.
Published: (2026)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
by: Schulte, David, et al.
Published: (2024)
by: Schulte, David, et al.
Published: (2024)
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
by: Williams, Tristan, et al.
Published: (2026)
by: Williams, Tristan, et al.
Published: (2026)
LM-PUB-QUIZ: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models
by: Ploner, Max, et al.
Published: (2024)
by: Ploner, Max, et al.
Published: (2024)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
by: Wang, Clinton J., et al.
Published: (2025)
by: Wang, Clinton J., et al.
Published: (2025)
MLissard: Multilingual Long and Simple Sequential Reasoning Benchmarks
by: Bueno, Mirelle, et al.
Published: (2024)
by: Bueno, Mirelle, et al.
Published: (2024)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
by: Chen, Yongrui, et al.
Published: (2025)
by: Chen, Yongrui, et al.
Published: (2025)
mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models
by: Mukherjee, Arka, et al.
Published: (2025)
by: Mukherjee, Arka, et al.
Published: (2025)
StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models
by: Chen, Yongrui, et al.
Published: (2026)
by: Chen, Yongrui, et al.
Published: (2026)
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
by: Ko, Donghyeon, et al.
Published: (2025)
by: Ko, Donghyeon, et al.
Published: (2025)
CRCL at SemEval-2024 Task 2: Simple prompt optimizations
by: Brutti-Mairesse, Clément, et al.
Published: (2024)
by: Brutti-Mairesse, Clément, et al.
Published: (2024)
MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs
by: Qu, Zhan, et al.
Published: (2025)
by: Qu, Zhan, et al.
Published: (2025)
ClonEval: An Open Voice Cloning Benchmark
by: Christop, Iwona, et al.
Published: (2025)
by: Christop, Iwona, et al.
Published: (2025)
Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy
by: McCammon, James
Published: (2025)
by: McCammon, James
Published: (2025)
Similar Items
-
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2025) -
What Matters When Building Universal Multilingual Named Entity Recognition Models?
by: Golde, Jonas, et al.
Published: (2026) -
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
by: Haller, Patrick, et al.
Published: (2024) -
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
by: Haller, Patrick, et al.
Published: (2025) -
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025)