ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Handa, Divij, Dolin, Pavel, Kumbhar, Shrinidhi, Son, Tran Cao, Baral, Chitta |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024)
by: Handa, Divij, et al.
Published: (2024)
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
by: Kumbhar, Shrinidhi, et al.
Published: (2025)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
by: Luo, Man, et al.
Published: (2023)
by: Luo, Man, et al.
Published: (2023)
Epistemic Skills: Reasoning about Knowledge and Oblivion
by: Liang, Xiaolong, et al.
Published: (2025)
by: Liang, Xiaolong, et al.
Published: (2025)
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)
by: RRV, Aswin, et al.
Published: (2025)
CSPs with Few Alien Constraints
by: Jonsson, Peter, et al.
Published: (2024)
by: Jonsson, Peter, et al.
Published: (2024)
FormulaOne: Measuring the Depth of Algorithmic Reasoning Beyond Competitive Programming
by: Beniamini, Gal, et al.
Published: (2025)
by: Beniamini, Gal, et al.
Published: (2025)
Limitations on Accurate, Trusted, Human-level Reasoning
by: Panigrahy, Rina, et al.
Published: (2025)
by: Panigrahy, Rina, et al.
Published: (2025)
Fast Approximation Algorithm for Non-Monotone DR-submodular Maximization under Size Constraint
by: Tran, Tan D., et al.
Published: (2025)
by: Tran, Tan D., et al.
Published: (2025)
Reasoning About Knowledge on Regular Expressions is 2EXPTIME-complete
by: Ghosh, Avijeet, et al.
Published: (2025)
by: Ghosh, Avijeet, et al.
Published: (2025)
Minimal Model Reasoning in Description Logics: Don't Try This at Home!
by: Di Stefano, Federica, et al.
Published: (2025)
by: Di Stefano, Federica, et al.
Published: (2025)
On Probabilistic and Causal Reasoning with Summation Operators
by: Ibeling, Duligur, et al.
Published: (2024)
by: Ibeling, Duligur, et al.
Published: (2024)
Exact Algorithms for Multiagent Path Finding with Communication Constraints on Tree-Like Structures
by: Fioravantes, Foivos, et al.
Published: (2024)
by: Fioravantes, Foivos, et al.
Published: (2024)
VERIFY-RL: Verifiable Recursive Decomposition for Reinforcement Learning in Mathematical Reasoning
by: Qasim, Kaleem Ullah, et al.
Published: (2026)
by: Qasim, Kaleem Ullah, et al.
Published: (2026)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
by: Huang, Zekai, et al.
Published: (2025)
by: Huang, Zekai, et al.
Published: (2025)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
by: RRV, Aswin, et al.
Published: (2026)
by: RRV, Aswin, et al.
Published: (2026)
The Scaling Properties of Implicit Deductive Reasoning in Transformers
by: Vompa, Enrico, et al.
Published: (2026)
by: Vompa, Enrico, et al.
Published: (2026)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
by: Chambon, Pierre, et al.
Published: (2025)
by: Chambon, Pierre, et al.
Published: (2025)
A Complexity Map of Probabilistic Reasoning for Neurosymbolic Classification Techniques
by: Ledaguenel, Arthur, et al.
Published: (2024)
by: Ledaguenel, Arthur, et al.
Published: (2024)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
by: Hazra, Rishi, et al.
Published: (2025)
by: Hazra, Rishi, et al.
Published: (2025)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
A Measure-Theoretic Analysis of Reasoning: Structural Generalization and Approximation Limits
by: Zhang, Yuyang, et al.
Published: (2026)
by: Zhang, Yuyang, et al.
Published: (2026)
On the Intrinsic Limits of Transformer Image Embeddings in Non-Solvable Spatial Reasoning
by: Lyu, Siyi, et al.
Published: (2026)
by: Lyu, Siyi, et al.
Published: (2026)
Partial Evidence Bench: Benchmarking Authorization-Limited Evidence in Agentic Systems
by: Tallam, Krti
Published: (2026)
by: Tallam, Krti
Published: (2026)
Transformer Encoder Satisfiability: Complexity and Impact on Formal Reasoning
by: Sälzer, Marco, et al.
Published: (2024)
by: Sälzer, Marco, et al.
Published: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
by: Fan, Lizhou, et al.
Published: (2023)
by: Fan, Lizhou, et al.
Published: (2023)
Higher-Order Responsibility
by: Jiang, Junli, et al.
Published: (2025)
by: Jiang, Junli, et al.
Published: (2025)
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
by: Varshney, Neeraj, et al.
Published: (2023)
by: Varshney, Neeraj, et al.
Published: (2023)
LLM+AL: Bridging Large Language Models and Action Languages for Complex Reasoning about Actions
by: Ishay, Adam, et al.
Published: (2025)
by: Ishay, Adam, et al.
Published: (2025)
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
by: Boche, Holger, et al.
Published: (2024)
by: Boche, Holger, et al.
Published: (2024)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
by: Kumbhar, Shrinidhi, et al.
Published: (2026)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
by: Dudeja, Divij, et al.
Published: (2025)
by: Dudeja, Divij, et al.
Published: (2025)
Solving Multiagent Path Finding on Highly Centralized Networks
by: Fioravantes, Foivos, et al.
Published: (2024)
by: Fioravantes, Foivos, et al.
Published: (2024)
Over the Edge of Chaos? Excess Complexity as a Roadblock to Artificial General Intelligence
by: Susnjak, Teo, et al.
Published: (2024)
by: Susnjak, Teo, et al.
Published: (2024)
Probabilistic Generating Circuits -- Demystified
by: Agarwal, Sanyam, et al.
Published: (2024)
by: Agarwal, Sanyam, et al.
Published: (2024)
From Probability to Counterfactuals: the Increasing Complexity of Satisfiability in Pearl's Causal Hierarchy
by: Dörfler, Julian, et al.
Published: (2024)
by: Dörfler, Julian, et al.
Published: (2024)
A Structural Complexity Analysis of Hierarchical Task Network Planning
by: Brand, Cornelius, et al.
Published: (2024)
by: Brand, Cornelius, et al.
Published: (2024)
Similar Items
-
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
by: Handa, Divij, et al.
Published: (2024) -
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
by: Kumbhar, Shrinidhi, et al.
Published: (2025) -
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
by: Luo, Man, et al.
Published: (2023) -
Epistemic Skills: Reasoning about Knowledge and Oblivion
by: Liang, Xiaolong, et al.
Published: (2025) -
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)