Playing Psychic: Using Thought Trees to Predict Reasoning Models Accuracy on Coding Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Jiaxin, He, Runyuan, Bhatia, Sahil, Gajare, Neel, Cheung, Alvin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2025)
by: Hong, Charles, et al.
Published: (2025)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
by: Gong, Linyuan, et al.
Published: (2024)
by: Gong, Linyuan, et al.
Published: (2024)
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
by: Macar, Uzay, et al.
Published: (2025)
by: Macar, Uzay, et al.
Published: (2025)
Thought Anchors: Which LLM Reasoning Steps Matter?
by: Bogdan, Paul C., et al.
Published: (2025)
by: Bogdan, Paul C., et al.
Published: (2025)
Structure-Aware Fill-in-the-Middle Pretraining for Code
by: Gong, Linyuan, et al.
Published: (2025)
by: Gong, Linyuan, et al.
Published: (2025)
Can Large Reasoning Models Improve Accuracy on Mathematical Tasks Using Flawed Thinking?
by: Amjith, Saraswathy, et al.
Published: (2025)
by: Amjith, Saraswathy, et al.
Published: (2025)
Reasoning Models Struggle to Control their Chains of Thought
by: Yueh-Han, Chen, et al.
Published: (2026)
by: Yueh-Han, Chen, et al.
Published: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
by: Arcuschin, Iván, et al.
Published: (2025)
by: Arcuschin, Iván, et al.
Published: (2025)
Automatic Question Generation for Intuitive Learning Utilizing Causal Graph Guided Chain of Thought Reasoning
by: Wang, Nicholas X., et al.
Published: (2026)
by: Wang, Nicholas X., et al.
Published: (2026)
Domain-Specialized Tree of Thought through Plug-and-Play Predictors
by: Gao, Xuanqi, et al.
Published: (2026)
by: Gao, Xuanqi, et al.
Published: (2026)
Reasoning Topology Matters: Network-of-Thought for Complex Reasoning Tasks
by: Huang, Fan
Published: (2026)
by: Huang, Fan
Published: (2026)
Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks
by: Dunivin, Zackary Okun
Published: (2024)
by: Dunivin, Zackary Okun
Published: (2024)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation
by: Jiang, Zhaoyang, et al.
Published: (2026)
by: Jiang, Zhaoyang, et al.
Published: (2026)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
by: Jiao, Rui, et al.
Published: (2025)
by: Jiao, Rui, et al.
Published: (2025)
Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
by: Liu, Yule, et al.
Published: (2025)
by: Liu, Yule, et al.
Published: (2025)
Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
Code World Models for General Game Playing
by: Lehrach, Wolfgang, et al.
Published: (2025)
by: Lehrach, Wolfgang, et al.
Published: (2025)
From Code to Play: Benchmarking Program Search for Games Using Large Language Models
by: Eberhardinger, Manuel, et al.
Published: (2024)
by: Eberhardinger, Manuel, et al.
Published: (2024)
Subliminal Learning Is Steering Vector Distillation
by: Blank, Camila, et al.
Published: (2026)
by: Blank, Camila, et al.
Published: (2026)
Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection
by: Park, Jongseok, et al.
Published: (2026)
by: Park, Jongseok, et al.
Published: (2026)
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
by: Bie, Fuqing, et al.
Published: (2025)
by: Bie, Fuqing, et al.
Published: (2025)
Modeling and Discovering Direct Causes for Predictive Models
by: Chen, Yizuo, et al.
Published: (2024)
by: Chen, Yizuo, et al.
Published: (2024)
BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem Solving
by: Wang, Teng, et al.
Published: (2024)
by: Wang, Teng, et al.
Published: (2024)
Diagnosing Pathological Chain-of-Thought in Reasoning Models
by: Liu, Manqing, et al.
Published: (2026)
by: Liu, Manqing, et al.
Published: (2026)
A Text-Based Knowledge-Embedded Soft Sensing Modeling Approach for General Industrial Process Tasks Based on Large Language Model
by: Tong, Shuo, et al.
Published: (2025)
by: Tong, Shuo, et al.
Published: (2025)
Thought Propagation: An Analogical Approach to Complex Reasoning with Large Language Models
by: Yu, Junchi, et al.
Published: (2023)
by: Yu, Junchi, et al.
Published: (2023)
CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution
by: Xu, Ruiyang, et al.
Published: (2024)
by: Xu, Ruiyang, et al.
Published: (2024)
Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought
by: Guo, Yuchen, et al.
Published: (2026)
by: Guo, Yuchen, et al.
Published: (2026)
SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks
by: Wan, Wentao, et al.
Published: (2025)
by: Wan, Wentao, et al.
Published: (2025)
UTMath: Math Evaluation with Unit Test via Reasoning-to-Coding Thoughts
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
by: Guo, Shiyuan, et al.
Published: (2025)
by: Guo, Shiyuan, et al.
Published: (2025)
Reasoning-Based Approach with Chain-of-Thought for Alzheimer's Detection Using Speech and Large Language Models
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
Predicting Empirical AI Research Outcomes with Language Models
by: Wen, Jiaxin, et al.
Published: (2025)
by: Wen, Jiaxin, et al.
Published: (2025)
Analyzing Chain of Thought (CoT) Approaches in Control Flow Code Deobfuscation Tasks
by: Mohseni, Seyedreza, et al.
Published: (2026)
by: Mohseni, Seyedreza, et al.
Published: (2026)
Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
Reason-to-Recommend: Using Interaction-of-Thought Reasoning to Enhance LLM Recommendation
by: Zhao, Keyu, et al.
Published: (2025)
by: Zhao, Keyu, et al.
Published: (2025)
Verified Code Transpilation with LLMs
by: Bhatia, Sahil, et al.
Published: (2024)
by: Bhatia, Sahil, et al.
Published: (2024)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
by: Park, Jongseok, et al.
Published: (2026)
by: Park, Jongseok, et al.
Published: (2026)
Similar Items
-
Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
by: Hong, Charles, et al.
Published: (2025) -
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
by: Gong, Linyuan, et al.
Published: (2024) -
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025) -
Thought Branches: Interpreting LLM Reasoning Requires Resampling
by: Macar, Uzay, et al.
Published: (2025) -
Thought Anchors: Which LLM Reasoning Steps Matter?
by: Bogdan, Paul C., et al.
Published: (2025)