LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Choukrani, Omar, Malek, Idriss, Orel, Daniil, Xie, Zhuohan, Iklassov, Zangir, Takáč, Martin, Lahlou, Salem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
The AI Data Scientist
di: Akimov, Farkhad, et al.
Pubblicazione: (2025)
di: Akimov, Farkhad, et al.
Pubblicazione: (2025)
Self-Guiding Exploration for Combinatorial Problems
di: Iklassov, Zangir, et al.
Pubblicazione: (2024)
di: Iklassov, Zangir, et al.
Pubblicazione: (2024)
Reinforcement Learning for Solving Stochastic Vehicle Routing Problem with Time Windows
di: Iklassov, Zangir, et al.
Pubblicazione: (2024)
di: Iklassov, Zangir, et al.
Pubblicazione: (2024)
Measuring AI Reasoning: A Guide for Researchers
di: Nwadike, Munachiso Samuel, et al.
Pubblicazione: (2026)
di: Nwadike, Munachiso Samuel, et al.
Pubblicazione: (2026)
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
di: Nwadike, Munachiso, et al.
Pubblicazione: (2025)
di: Nwadike, Munachiso, et al.
Pubblicazione: (2025)
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
di: Mishra, Kshitij, et al.
Pubblicazione: (2026)
di: Mishra, Kshitij, et al.
Pubblicazione: (2026)
CORE: Collaborative Reasoning via Cross Teaching
di: Mishra, Kshitij, et al.
Pubblicazione: (2026)
di: Mishra, Kshitij, et al.
Pubblicazione: (2026)
Mitigating Societal Cognitive Overload in the Age of AI: Challenges and Directions
di: Lahlou, Salem
Pubblicazione: (2025)
di: Lahlou, Salem
Pubblicazione: (2025)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
di: Xie, Zhuohan, et al.
Pubblicazione: (2025)
di: Xie, Zhuohan, et al.
Pubblicazione: (2025)
Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
di: Li, Yanda, et al.
Pubblicazione: (2026)
di: Li, Yanda, et al.
Pubblicazione: (2026)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
di: Gu, Zhuohan, et al.
Pubblicazione: (2026)
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
di: Elbadry, Rania, et al.
Pubblicazione: (2026)
di: Elbadry, Rania, et al.
Pubblicazione: (2026)
Zero-Shot Off-Policy Learning
di: Asadulaev, Arip, et al.
Pubblicazione: (2026)
di: Asadulaev, Arip, et al.
Pubblicazione: (2026)
TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages
di: Ofer, Moshe, et al.
Pubblicazione: (2025)
di: Ofer, Moshe, et al.
Pubblicazione: (2025)
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
di: Richardson, Andrew Keenan, et al.
Pubblicazione: (2025)
di: Richardson, Andrew Keenan, et al.
Pubblicazione: (2025)
Breaking the Martingale Curse: Multi-Agent Debate via Asymmetric Cognitive Potential Energy
di: Liu, Yuhan, et al.
Pubblicazione: (2026)
di: Liu, Yuhan, et al.
Pubblicazione: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
di: Asadulaev, Arip, et al.
Pubblicazione: (2025)
di: Asadulaev, Arip, et al.
Pubblicazione: (2025)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
di: Cui, Hao, et al.
Pubblicazione: (2025)
di: Cui, Hao, et al.
Pubblicazione: (2025)
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
di: Wang, Siyuan, et al.
Pubblicazione: (2024)
di: Wang, Siyuan, et al.
Pubblicazione: (2024)
InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding
di: Jiayang, Cheng, et al.
Pubblicazione: (2025)
di: Jiayang, Cheng, et al.
Pubblicazione: (2025)
Code Execution as Grounded Supervision for LLM Reasoning
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
di: Jung, Dongwon, et al.
Pubblicazione: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
di: Malek, Alan, et al.
Pubblicazione: (2025)
di: Malek, Alan, et al.
Pubblicazione: (2025)
SIFT: Grounding LLM Reasoning in Contexts via Stickers
di: Zeng, Zihao, et al.
Pubblicazione: (2025)
di: Zeng, Zihao, et al.
Pubblicazione: (2025)
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
di: Wei, Hui, et al.
Pubblicazione: (2025)
di: Wei, Hui, et al.
Pubblicazione: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
di: Huan, Maggie, et al.
Pubblicazione: (2025)
di: Huan, Maggie, et al.
Pubblicazione: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
di: Badanin, Ilia, et al.
Pubblicazione: (2026)
di: Badanin, Ilia, et al.
Pubblicazione: (2026)
ArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform
di: Lahlou, Salem
Pubblicazione: (2026)
di: Lahlou, Salem
Pubblicazione: (2026)
Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
di: Amirizaniani, Maryam, et al.
Pubblicazione: (2024)
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
di: Dou, Zhihao, et al.
Pubblicazione: (2025)
di: Dou, Zhihao, et al.
Pubblicazione: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
di: Jiang, Botian, et al.
Pubblicazione: (2024)
di: Jiang, Botian, et al.
Pubblicazione: (2024)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
di: Jin, Keyan, et al.
Pubblicazione: (2025)
di: Jin, Keyan, et al.
Pubblicazione: (2025)
Can LLMs Automate Fact-Checking Article Writing?
di: Sahnan, Dhruv, et al.
Pubblicazione: (2025)
di: Sahnan, Dhruv, et al.
Pubblicazione: (2025)
GameArena: Evaluating LLM Reasoning through Live Computer Games
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
di: Hu, Lanxiang, et al.
Pubblicazione: (2024)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
di: Tran, Khanh-Tung, et al.
Pubblicazione: (2025)
di: Tran, Khanh-Tung, et al.
Pubblicazione: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
di: Yang, Siwei, et al.
Pubblicazione: (2024)
di: Yang, Siwei, et al.
Pubblicazione: (2024)
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
di: Gurgurov, Daniil, et al.
Pubblicazione: (2024)
di: Gurgurov, Daniil, et al.
Pubblicazione: (2024)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
di: Chen, Zichen, et al.
Pubblicazione: (2023)
di: Chen, Zichen, et al.
Pubblicazione: (2023)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
di: Hao, Shibo, et al.
Pubblicazione: (2024)
di: Hao, Shibo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
di: Heakl, Ahmed, et al.
Pubblicazione: (2025) -
The AI Data Scientist
di: Akimov, Farkhad, et al.
Pubblicazione: (2025) -
Self-Guiding Exploration for Combinatorial Problems
di: Iklassov, Zangir, et al.
Pubblicazione: (2024) -
Reinforcement Learning for Solving Stochastic Vehicle Routing Problem with Time Windows
di: Iklassov, Zangir, et al.
Pubblicazione: (2024) -
Measuring AI Reasoning: A Guide for Researchers
di: Nwadike, Munachiso Samuel, et al.
Pubblicazione: (2026)