LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Choukrani, Omar, Malek, Idriss, Orel, Daniil, Xie, Zhuohan, Iklassov, Zangir, Takáč, Martin, Lahlou, Salem |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
The AI Data Scientist
by: Akimov, Farkhad, et al.
Published: (2025)
by: Akimov, Farkhad, et al.
Published: (2025)
Self-Guiding Exploration for Combinatorial Problems
by: Iklassov, Zangir, et al.
Published: (2024)
by: Iklassov, Zangir, et al.
Published: (2024)
Reinforcement Learning for Solving Stochastic Vehicle Routing Problem with Time Windows
by: Iklassov, Zangir, et al.
Published: (2024)
by: Iklassov, Zangir, et al.
Published: (2024)
Measuring AI Reasoning: A Guide for Researchers
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
by: Nwadike, Munachiso, et al.
Published: (2025)
by: Nwadike, Munachiso, et al.
Published: (2025)
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
by: Mishra, Kshitij, et al.
Published: (2026)
by: Mishra, Kshitij, et al.
Published: (2026)
CORE: Collaborative Reasoning via Cross Teaching
by: Mishra, Kshitij, et al.
Published: (2026)
by: Mishra, Kshitij, et al.
Published: (2026)
Mitigating Societal Cognitive Overload in the Age of AI: Challenges and Directions
by: Lahlou, Salem
Published: (2025)
by: Lahlou, Salem
Published: (2025)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
by: Xie, Zhuohan, et al.
Published: (2025)
by: Xie, Zhuohan, et al.
Published: (2025)
Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
by: Li, Yanda, et al.
Published: (2026)
by: Li, Yanda, et al.
Published: (2026)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
by: Gu, Zhuohan, et al.
Published: (2026)
by: Gu, Zhuohan, et al.
Published: (2026)
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
Zero-Shot Off-Policy Learning
by: Asadulaev, Arip, et al.
Published: (2026)
by: Asadulaev, Arip, et al.
Published: (2026)
TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages
by: Ofer, Moshe, et al.
Published: (2025)
by: Ofer, Moshe, et al.
Published: (2025)
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
by: Richardson, Andrew Keenan, et al.
Published: (2025)
by: Richardson, Andrew Keenan, et al.
Published: (2025)
Breaking the Martingale Curse: Multi-Agent Debate via Asymmetric Cognitive Potential Energy
by: Liu, Yuhan, et al.
Published: (2026)
by: Liu, Yuhan, et al.
Published: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
by: Asadulaev, Arip, et al.
Published: (2025)
by: Asadulaev, Arip, et al.
Published: (2025)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
by: Cui, Hao, et al.
Published: (2025)
by: Cui, Hao, et al.
Published: (2025)
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
by: Jiayang, Cheng, et al.
Published: (2025)
by: Jiayang, Cheng, et al.
Published: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
by: Jung, Dongwon, et al.
Published: (2025)
by: Jung, Dongwon, et al.
Published: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025)
by: Malek, Alan, et al.
Published: (2025)
SIFT: Grounding LLM Reasoning in Contexts via Stickers
by: Zeng, Zihao, et al.
Published: (2025)
by: Zeng, Zihao, et al.
Published: (2025)
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
by: Wei, Hui, et al.
Published: (2025)
by: Wei, Hui, et al.
Published: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
ArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform
by: Lahlou, Salem
Published: (2026)
by: Lahlou, Salem
Published: (2026)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
by: Amirizaniani, Maryam, et al.
Published: (2024)
by: Amirizaniani, Maryam, et al.
Published: (2024)
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
by: Dou, Zhihao, et al.
Published: (2025)
by: Dou, Zhihao, et al.
Published: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
Can LLMs Automate Fact-Checking Article Writing?
by: Sahnan, Dhruv, et al.
Published: (2025)
by: Sahnan, Dhruv, et al.
Published: (2025)
GameArena: Evaluating LLM Reasoning through Live Computer Games
by: Hu, Lanxiang, et al.
Published: (2024)
by: Hu, Lanxiang, et al.
Published: (2024)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
by: Tran, Khanh-Tung, et al.
Published: (2025)
by: Tran, Khanh-Tung, et al.
Published: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
by: Gurgurov, Daniil, et al.
Published: (2024)
by: Gurgurov, Daniil, et al.
Published: (2024)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
by: Chen, Zichen, et al.
Published: (2023)
by: Chen, Zichen, et al.
Published: (2023)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
by: Hao, Shibo, et al.
Published: (2024)
by: Hao, Shibo, et al.
Published: (2024)
Similar Items
-
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
by: Heakl, Ahmed, et al.
Published: (2025) -
The AI Data Scientist
by: Akimov, Farkhad, et al.
Published: (2025) -
Self-Guiding Exploration for Combinatorial Problems
by: Iklassov, Zangir, et al.
Published: (2024) -
Reinforcement Learning for Solving Stochastic Vehicle Routing Problem with Time Windows
by: Iklassov, Zangir, et al.
Published: (2024) -
Measuring AI Reasoning: A Guide for Researchers
by: Nwadike, Munachiso Samuel, et al.
Published: (2026)