Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Boshi, Yue, Xiang, Su, Yu, Sun, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
by: Wang, Boshi, et al.
Published: (2025)
by: Wang, Boshi, et al.
Published: (2025)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
by: Kohli, Harsh, et al.
Published: (2026)
by: Kohli, Harsh, et al.
Published: (2026)
Mechanistic Insights into Grokking from the Embedding Layer
by: AlquBoj, H. V., et al.
Published: (2025)
by: AlquBoj, H. V., et al.
Published: (2025)
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
by: Mo, Lingbo, et al.
Published: (2023)
by: Mo, Lingbo, et al.
Published: (2023)
Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
by: He, Kaiyu, et al.
Published: (2026)
by: He, Kaiyu, et al.
Published: (2026)
Implicit Reasoning in Transformers is Reasoning through Shortcuts
by: Lin, Tianhe, et al.
Published: (2025)
by: Lin, Tianhe, et al.
Published: (2025)
Not Just the Destination, But the Journey: Reasoning Traces Causally Shape Generalization Behaviors
by: Wen, Pengcheng, et al.
Published: (2026)
by: Wen, Pengcheng, et al.
Published: (2026)
A Retrieve-and-Read Framework for Knowledge Graph Link Prediction
by: Pahuja, Vardaan, et al.
Published: (2022)
by: Pahuja, Vardaan, et al.
Published: (2022)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
by: Sun, Zhongxiang, et al.
Published: (2025)
by: Sun, Zhongxiang, et al.
Published: (2025)
Improving Code Localization with Repository Memory
by: Wang, Boshi, et al.
Published: (2025)
by: Wang, Boshi, et al.
Published: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning
by: Zhu, Rongzhi, et al.
Published: (2025)
by: Zhu, Rongzhi, et al.
Published: (2025)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
by: Wang, Boshi, et al.
Published: (2024)
by: Wang, Boshi, et al.
Published: (2024)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
by: Abramov, Roman, et al.
Published: (2025)
by: Abramov, Roman, et al.
Published: (2025)
Open-domain Implicit Format Control for Large Language Model Generation
by: Yao, Yiqun, et al.
Published: (2024)
by: Yao, Yiqun, et al.
Published: (2024)
GroundCocoa: A Benchmark for Evaluating Compositional & Conditional Reasoning in Language Models
by: Kohli, Harsh, et al.
Published: (2024)
by: Kohli, Harsh, et al.
Published: (2024)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
by: Wang, Atticus, et al.
Published: (2025)
by: Wang, Atticus, et al.
Published: (2025)
Do LLMs Really Think Step-by-step In Implicit Reasoning?
by: Yu, Yijiong
Published: (2024)
by: Yu, Yijiong
Published: (2024)
Reasoning with Graphs: Structuring Implicit Knowledge to Enhance LLMs Reasoning
by: Han, Haoyu, et al.
Published: (2025)
by: Han, Haoyu, et al.
Published: (2025)
Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning
by: Ye, Donald, et al.
Published: (2026)
by: Ye, Donald, et al.
Published: (2026)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
DeReason: A Difficulty-Aware Curriculum Improves Decoupled SFT-then-RL Training for General Reasoning
by: Hu, Hanxu, et al.
Published: (2026)
by: Hu, Hanxu, et al.
Published: (2026)
AuroraEdge-V-2B: A Faster And Stronger Edge Visual Large Language Model
by: Chen, Xiang
Published: (2026)
by: Chen, Xiang
Published: (2026)
TableLlama: Towards Open Large Generalist Models for Tables
by: Zhang, Tianshu, et al.
Published: (2023)
by: Zhang, Tianshu, et al.
Published: (2023)
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
by: He, Yinhan, et al.
Published: (2025)
by: He, Yinhan, et al.
Published: (2025)
Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
by: Liu, Kejin, et al.
Published: (2025)
by: Liu, Kejin, et al.
Published: (2025)
O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
by: Huang, Zhongzhen, et al.
Published: (2025)
by: Huang, Zhongzhen, et al.
Published: (2025)
Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Mechanistic Interpretability of Binary and Ternary Transformers
by: Li, Jason
Published: (2024)
by: Li, Jason
Published: (2024)
When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability
by: Sun, Zhongxiang, et al.
Published: (2024)
by: Sun, Zhongxiang, et al.
Published: (2024)
Application of Multiple Chain-of-Thought in Contrastive Reasoning for Implicit Sentiment Analysis
by: Yang, Liwei, et al.
Published: (2025)
by: Yang, Liwei, et al.
Published: (2025)
Fragile Reasoning: A Mechanistic Analysis of LLM Sensitivity to Meaning-Preserving Perturbations
by: Han, Shou-Tzu, et al.
Published: (2026)
by: Han, Shou-Tzu, et al.
Published: (2026)
Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures
by: Hu, Yi, et al.
Published: (2026)
by: Hu, Yi, et al.
Published: (2026)
Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations
by: Li, Shanghao, et al.
Published: (2026)
by: Li, Shanghao, et al.
Published: (2026)
AttributionBench: How Hard is Automatic Attribution Evaluation?
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
RexDrug: Reliable Multi-Drug Combination Extraction through Reasoning-Enhanced LLMs
by: Wang, Zhijun, et al.
Published: (2026)
by: Wang, Zhijun, et al.
Published: (2026)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
by: Zhang, Charlie, et al.
Published: (2025)
by: Zhang, Charlie, et al.
Published: (2025)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
by: Xiong, Xuan, et al.
Published: (2026)
by: Xiong, Xuan, et al.
Published: (2026)
Similar Items
-
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization Failure
by: Wang, Boshi, et al.
Published: (2025) -
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
by: Kohli, Harsh, et al.
Published: (2026) -
Mechanistic Insights into Grokking from the Embedding Layer
by: AlquBoj, H. V., et al.
Published: (2025) -
How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities
by: Mo, Lingbo, et al.
Published: (2023) -
Is Grokking Worthwhile? Functional Analysis and Transferability of Generalization Circuits in Transformers
by: He, Kaiyu, et al.
Published: (2026)