The Illusion of Procedural Reasoning: Measuring Long-Horizon FSM Execution in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Samiei, Mahdi, Mansouri, Mahdi, Baghshah, Mahdieh Soleymani |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
by: Dijujin, Negin Hashemi, et al.
Published: (2025)
by: Dijujin, Negin Hashemi, et al.
Published: (2025)
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
by: Sinha, Akshit, et al.
Published: (2025)
by: Sinha, Akshit, et al.
Published: (2025)
SUSD: Structured Unsupervised Skill Discovery through State Factorization
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Efficient Adversarial Attacks on High-dimensional Offline Bandits
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
by: Mehri, Faridoun, et al.
Published: (2024)
by: Mehri, Faridoun, et al.
Published: (2024)
LLM-Agent-Controller: A Universal Multi-Agent Large Language Model System as a Control Engineer
by: Zahedifar, Rasoul, et al.
Published: (2025)
by: Zahedifar, Rasoul, et al.
Published: (2025)
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
by: Saghafian, Armin, et al.
Published: (2024)
by: Saghafian, Armin, et al.
Published: (2024)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
by: Mozafari, Mohammad, et al.
Published: (2024)
by: Mozafari, Mohammad, et al.
Published: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Trained Models Tell Us How to Make Them Robust to Spurious Correlation without Group Annotation
by: Ghaznavi, Mahdi, et al.
Published: (2024)
by: Ghaznavi, Mahdi, et al.
Published: (2024)
CER: Confidence Enhanced Reasoning in LLMs
by: Razghandi, Ali, et al.
Published: (2025)
by: Razghandi, Ali, et al.
Published: (2025)
Intrinsic Stability Limits of Autoregressive Reasoning: Structural Consequences for Long-Horizon Execution
by: Liao, Hsien-Jyh
Published: (2026)
by: Liao, Hsien-Jyh
Published: (2026)
Procedural Content Generation in Games: A Survey with Insights on Emerging LLM Integration
by: Maleki, Mahdi Farrokhi, et al.
Published: (2024)
by: Maleki, Mahdi Farrokhi, et al.
Published: (2024)
How Well Do LLMs Understand Tunisian Arabic?
by: Mahdi, Mohamed
Published: (2025)
by: Mahdi, Mohamed
Published: (2025)
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
by: Abdollahi, Ali, et al.
Published: (2024)
by: Abdollahi, Ali, et al.
Published: (2024)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
by: Hasani, Hosein, et al.
Published: (2025)
by: Hasani, Hosein, et al.
Published: (2025)
Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
by: Hasani, Hosein, et al.
Published: (2025)
by: Hasani, Hosein, et al.
Published: (2025)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
by: Hu, Xavier, et al.
Published: (2026)
by: Hu, Xavier, et al.
Published: (2026)
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
by: Ghahroodi, Omid, et al.
Published: (2024)
by: Ghahroodi, Omid, et al.
Published: (2024)
ExGRG: Explicitly-Generated Relation Graph for Self-Supervised Representation Learning
by: Naseri, Mahdi, et al.
Published: (2024)
by: Naseri, Mahdi, et al.
Published: (2024)
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
by: Wang, Tianle, et al.
Published: (2026)
by: Wang, Tianle, et al.
Published: (2026)
CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
by: Vaghasiya, Jay, et al.
Published: (2025)
by: Vaghasiya, Jay, et al.
Published: (2025)
Dilated Balanced Cross Entropy Loss for Medical Image Segmentation
by: Hosseini, Seyed Mohsen, et al.
Published: (2024)
by: Hosseini, Seyed Mohsen, et al.
Published: (2024)
PARC: An Autonomous Self-Reflective Coding Agent for Robust Execution of Long-Horizon Tasks
by: Orimo, Yuki, et al.
Published: (2025)
by: Orimo, Yuki, et al.
Published: (2025)
SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
by: Li, Jialiang, et al.
Published: (2025)
by: Li, Jialiang, et al.
Published: (2025)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025)
by: Liang, Shanchao, et al.
Published: (2025)
LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning
by: Pushkin, Denys, et al.
Published: (2026)
by: Pushkin, Denys, et al.
Published: (2026)
The Illusion of Insight in Reasoning Models
by: d'Aliberti, Liv G., et al.
Published: (2026)
by: d'Aliberti, Liv G., et al.
Published: (2026)
Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
by: Forouzandeh, Saman, et al.
Published: (2025)
by: Forouzandeh, Saman, et al.
Published: (2025)
Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs
by: Salimi, Moein, et al.
Published: (2026)
by: Salimi, Moein, et al.
Published: (2026)
Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation
by: Deng, Zehao, et al.
Published: (2025)
by: Deng, Zehao, et al.
Published: (2025)
When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution
by: Zhu, Zilin, et al.
Published: (2026)
by: Zhu, Zilin, et al.
Published: (2026)
Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning
by: Hwang, Jaebak, et al.
Published: (2025)
by: Hwang, Jaebak, et al.
Published: (2025)
PADME: Procedure Aware DynaMic Execution
by: Garg, Deepeka, et al.
Published: (2025)
by: Garg, Deepeka, et al.
Published: (2025)
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
by: Lan, Yifan, et al.
Published: (2026)
by: Lan, Yifan, et al.
Published: (2026)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
Similar Items
-
Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025) -
VQEL: Enabling Self-Play in Emergent Language Games via Agent-Internal Vector Quantization
by: Paqaleh, Mohammad Mahdi Samiei, et al.
Published: (2025) -
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
by: Dijujin, Negin Hashemi, et al.
Published: (2025) -
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
by: Sinha, Akshit, et al.
Published: (2025) -
SUSD: Structured Unsupervised Skill Discovery through State Factorization
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2026)