Do You Get the Hint? Benchmarking LLMs on the Board Game Concept
Fuente:
arXiv
Saved in:
| Main Authors: | Gevers, Ine, Daelemans, Walter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bag of Lies: Robustness in Continuous Pre-training BERT
by: Gevers, Ine, et al.
Published: (2024)
by: Gevers, Ine, et al.
Published: (2024)
WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization
by: Gevers, Ine, et al.
Published: (2025)
by: Gevers, Ine, et al.
Published: (2025)
BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language
by: Banar, Nikolay, et al.
Published: (2024)
by: Banar, Nikolay, et al.
Published: (2024)
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
by: Banar, Nikolay, et al.
Published: (2025)
by: Banar, Nikolay, et al.
Published: (2025)
FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
by: Hu, Xuesi, et al.
Published: (2026)
by: Hu, Xuesi, et al.
Published: (2026)
PersonalityChat: Conversation Distillation for Personalized Dialog Modeling with Facts and Traits
by: Lotfi, Ehsan, et al.
Published: (2024)
by: Lotfi, Ehsan, et al.
Published: (2024)
Bilingual BSARD: Extending Statutory Article Retrieval to Dutch
by: Lotfi, Ehsan, et al.
Published: (2024)
by: Lotfi, Ehsan, et al.
Published: (2024)
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification
by: Van Nooten, Jens, et al.
Published: (2025)
by: Van Nooten, Jens, et al.
Published: (2025)
An Agentic AI Framework for Training General Practitioner Student Skills
by: De Marez, Victor, et al.
Published: (2025)
by: De Marez, Victor, et al.
Published: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
NavHint: Vision and Language Navigation Agent with a Hint Generator
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge
by: Fu, Jinlan, et al.
Published: (2024)
by: Fu, Jinlan, et al.
Published: (2024)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
Simultaneous Reward Distillation and Preference Learning: Get You a Language Model Who Can Do Both
by: Nath, Abhijnan, et al.
Published: (2024)
by: Nath, Abhijnan, et al.
Published: (2024)
HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions
by: Mozafari, Jamshid, et al.
Published: (2025)
by: Mozafari, Jamshid, et al.
Published: (2025)
WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation
by: Mozafari, Jamshid, et al.
Published: (2024)
by: Mozafari, Jamshid, et al.
Published: (2024)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
by: Wang, Jin, et al.
Published: (2026)
by: Wang, Jin, et al.
Published: (2026)
Reasoning Gets Harder for LLMs Inside A Dialogue
by: Kartáč, Ivan, et al.
Published: (2026)
by: Kartáč, Ivan, et al.
Published: (2026)
LLMs Get Lost In Multi-Turn Conversation
by: Laban, Philippe, et al.
Published: (2025)
by: Laban, Philippe, et al.
Published: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
by: Cai, Will, et al.
Published: (2025)
by: Cai, Will, et al.
Published: (2025)
LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
by: Chen, Junhao, et al.
Published: (2025)
by: Chen, Junhao, et al.
Published: (2025)
Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
Do as We Do, Not as You Think: the Conformity of Large Language Models
by: Weng, Zhiyuan, et al.
Published: (2025)
by: Weng, Zhiyuan, et al.
Published: (2025)
ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
by: Tang, Siao, et al.
Published: (2025)
by: Tang, Siao, et al.
Published: (2025)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
by: Zhang, Kaiyi, et al.
Published: (2025)
by: Zhang, Kaiyi, et al.
Published: (2025)
Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
by: Altinok, Duygu
Published: (2025)
by: Altinok, Duygu
Published: (2025)
Hint Tuning: Less Data Makes Better Reasoners
by: Fan, Siqi, et al.
Published: (2026)
by: Fan, Siqi, et al.
Published: (2026)
Enhancing Financial Sentiment Analysis with Expert-Designed Hint
by: Chen, Chung-Chi, et al.
Published: (2024)
by: Chen, Chung-Chi, et al.
Published: (2024)
MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
by: Singh, Jyotika, et al.
Published: (2026)
by: Singh, Jyotika, et al.
Published: (2026)
Easy Problems That LLMs Get Wrong
by: Williams, Sean, et al.
Published: (2024)
by: Williams, Sean, et al.
Published: (2024)
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models
by: Srivastava, Gaurav, et al.
Published: (2025)
by: Srivastava, Gaurav, et al.
Published: (2025)
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks
by: Chang, Ting-Yun, et al.
Published: (2023)
by: Chang, Ting-Yun, et al.
Published: (2023)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
by: Jiang, Jiyue, et al.
Published: (2024)
by: Jiang, Jiyue, et al.
Published: (2024)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
by: Toker, Gilat, et al.
Published: (2026)
by: Toker, Gilat, et al.
Published: (2026)
Concept Space Alignment in Multilingual LLMs
by: Peng, Qiwei, et al.
Published: (2024)
by: Peng, Qiwei, et al.
Published: (2024)
Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
by: Hryszko, Jaroslaw
Published: (2026)
by: Hryszko, Jaroslaw
Published: (2026)
Learning to Hint for Reinforcement Learning
by: Xia, Yu, et al.
Published: (2026)
by: Xia, Yu, et al.
Published: (2026)
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
by: Jangra, Anubhav, et al.
Published: (2025)
by: Jangra, Anubhav, et al.
Published: (2025)
Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviews
by: Zou, Chenye, et al.
Published: (2025)
by: Zou, Chenye, et al.
Published: (2025)
Similar Items
-
Bag of Lies: Robustness in Continuous Pre-training BERT
by: Gevers, Ine, et al.
Published: (2024) -
WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization
by: Gevers, Ine, et al.
Published: (2025) -
BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language
by: Banar, Nikolay, et al.
Published: (2024) -
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
by: Banar, Nikolay, et al.
Published: (2025) -
FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
by: Hu, Xuesi, et al.
Published: (2026)