Can LLMs Reason in the Wild with Programs?
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yuan, Xiong, Siheng, Payani, Ali, Shareghi, Ehsan, Fekri, Faramarz |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation
by: Xiong, Siheng, et al.
Published: (2025)
by: Xiong, Siheng, et al.
Published: (2025)
Large Language Models Can Learn Temporal Reasoning
by: Xiong, Siheng, et al.
Published: (2024)
by: Xiong, Siheng, et al.
Published: (2024)
The Compressor-Retriever Architecture for Language Model OS
by: Yang, Yuan, et al.
Published: (2024)
by: Yang, Yuan, et al.
Published: (2024)
Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
by: Xiong, Siheng, et al.
Published: (2024)
by: Xiong, Siheng, et al.
Published: (2024)
TEILP: Time Prediction over Knowledge Graphs via Logical Reasoning
by: Xiong, Siheng, et al.
Published: (2023)
by: Xiong, Siheng, et al.
Published: (2023)
PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs
by: Gungordu, Oguzhan, et al.
Published: (2026)
by: Gungordu, Oguzhan, et al.
Published: (2026)
Temporal Inductive Logic Reasoning over Hypergraphs
by: Yang, Yuan, et al.
Published: (2022)
by: Yang, Yuan, et al.
Published: (2022)
TILP: Differentiable Learning of Temporal Logical Rules on Knowledge Graphs
by: Xiong, Siheng, et al.
Published: (2024)
by: Xiong, Siheng, et al.
Published: (2024)
Scaling Search-Augmented LLM Reasoning via Adaptive Information Control
by: Xiong, Siheng, et al.
Published: (2026)
by: Xiong, Siheng, et al.
Published: (2026)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
A Closer Look at Logical Reasoning with LLMs: The Choice of Tool Matters
by: Lam, Long Hei Matthew, et al.
Published: (2024)
by: Lam, Long Hei Matthew, et al.
Published: (2024)
ReasonGraph: Visualisation of Reasoning Paths
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance
by: Theuma, Adrian, et al.
Published: (2024)
by: Theuma, Adrian, et al.
Published: (2024)
VerifiAgent: a Unified Verification Agent in Language Model Reasoning
by: Han, Jiuzhou, et al.
Published: (2025)
by: Han, Jiuzhou, et al.
Published: (2025)
Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning
by: Han, Jiuzhou, et al.
Published: (2025)
by: Han, Jiuzhou, et al.
Published: (2025)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
PiVe: Prompting with Iterative Verification Improving Graph-based Generative Capability of LLMs
by: Han, Jiuzhou, et al.
Published: (2023)
by: Han, Jiuzhou, et al.
Published: (2023)
GLIDR: Graph-Like Inductive Logic Programming with Differentiable Reasoning
by: Johnson, Blair, et al.
Published: (2025)
by: Johnson, Blair, et al.
Published: (2025)
Towards Inference-time Scaling for Continuous Space Reasoning
by: Wang, Minghan, et al.
Published: (2025)
by: Wang, Minghan, et al.
Published: (2025)
On The Theory of Semantic Information and Communication for Logical Inference
by: Saz, Ahmet Faruk, et al.
Published: (2024)
by: Saz, Ahmet Faruk, et al.
Published: (2024)
Lossy Semantic Communication for the Logical Deduction of the State of the World
by: Saz, Ahmet Faruk, et al.
Published: (2024)
by: Saz, Ahmet Faruk, et al.
Published: (2024)
Analysis of Semantic Communication for Logic-based Hypothesis Deduction
by: Saz, Ahmet Faruk, et al.
Published: (2025)
by: Saz, Ahmet Faruk, et al.
Published: (2025)
DISCD: Distributed Lossy Semantic Communication for Logical Deduction of Hypothesis
by: Saz, Ahmet Faruk, et al.
Published: (2025)
by: Saz, Ahmet Faruk, et al.
Published: (2025)
One STEP at a time: Language Agents are Stepwise Planners
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
Strategies for Improving NL-to-FOL Translation with LLMs: Data Generation, Incremental Fine-Tuning, and Verification
by: Thatikonda, Ramya Keerthy, et al.
Published: (2024)
by: Thatikonda, Ramya Keerthy, et al.
Published: (2024)
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
by: Zhang, Caiqi, et al.
Published: (2025)
by: Zhang, Caiqi, et al.
Published: (2025)
Towards Uncertainty-Aware Language Agent
by: Han, Jiuzhou, et al.
Published: (2024)
by: Han, Jiuzhou, et al.
Published: (2024)
Reward Engineering for Generating Semi-structured Explanation
by: Han, Jiuzhou, et al.
Published: (2023)
by: Han, Jiuzhou, et al.
Published: (2023)
Improving Symbolic Translation of Language Models for Logical Reasoning
by: Thatikonda, Ramya Keerthy, et al.
Published: (2026)
by: Thatikonda, Ramya Keerthy, et al.
Published: (2026)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Assessing the Sensitivity and Alignment of FOL Closeness Metrics
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
by: Thatikonda, Ramya Keerthy, et al.
Published: (2025)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
by: Mishra, Venkatesh, et al.
Published: (2025)
by: Mishra, Venkatesh, et al.
Published: (2025)
Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study
by: Han, Jiuzhou, et al.
Published: (2024)
by: Han, Jiuzhou, et al.
Published: (2024)
Can Knowledge Editing Really Correct Hallucinations?
by: Huang, Baixiang, et al.
Published: (2024)
by: Huang, Baixiang, et al.
Published: (2024)
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
by: Wang, Minghan, et al.
Published: (2026)
by: Wang, Minghan, et al.
Published: (2026)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Similar Items
-
Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation
by: Xiong, Siheng, et al.
Published: (2025) -
Large Language Models Can Learn Temporal Reasoning
by: Xiong, Siheng, et al.
Published: (2024) -
The Compressor-Retriever Architecture for Language Model OS
by: Yang, Yuan, et al.
Published: (2024) -
Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
by: Xiong, Siheng, et al.
Published: (2024) -
TEILP: Time Prediction over Knowledge Graphs via Logical Reasoning
by: Xiong, Siheng, et al.
Published: (2023)