Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Chen, Hung-Hsuan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
par: Yang, Wang, et autres
Publié: (2025)
par: Yang, Wang, et autres
Publié: (2025)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
par: Kohli, Harsh, et autres
Publié: (2026)
par: Kohli, Harsh, et autres
Publié: (2026)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
par: McLeish, Sean, et autres
Publié: (2025)
par: McLeish, Sean, et autres
Publié: (2025)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
par: Lu, Wenquan, et autres
Publié: (2025)
par: Lu, Wenquan, et autres
Publié: (2025)
CeRA: Overcoming the Linear Ceiling of Low-Rank Adaptation via Capacity Expansion
par: Chen, Hung-Hsuan
Publié: (2026)
par: Chen, Hung-Hsuan
Publié: (2026)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
par: Barron, Joshua, et autres
Publié: (2025)
par: Barron, Joshua, et autres
Publié: (2025)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
par: Knupp, Jonas, et autres
Publié: (2026)
par: Knupp, Jonas, et autres
Publié: (2026)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
par: Lu, Ximing, et autres
Publié: (2025)
par: Lu, Ximing, et autres
Publié: (2025)
Fast Think-on-Graph: Wider, Deeper and Faster Reasoning of Large Language Model on Knowledge Graph
par: Liang, Xujian, et autres
Publié: (2025)
par: Liang, Xujian, et autres
Publié: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
par: Moosa, Ibraheem Muhammad, et autres
Publié: (2026)
par: Moosa, Ibraheem Muhammad, et autres
Publié: (2026)
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
par: Botev, Aleksandar, et autres
Publié: (2024)
par: Botev, Aleksandar, et autres
Publié: (2024)
Think Before You Act: Decision Transformers with Working Memory
par: Kang, Jikun, et autres
Publié: (2023)
par: Kang, Jikun, et autres
Publié: (2023)
SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
par: Zhang, Xichen, et autres
Publié: (2025)
par: Zhang, Xichen, et autres
Publié: (2025)
The Depth Delusion: Why Transformers Should Be Wider, Not Deeper
par: Fahim, Md Muhtasim Munif, et autres
Publié: (2026)
par: Fahim, Md Muhtasim Munif, et autres
Publié: (2026)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
par: Chen, Yilong, et autres
Publié: (2026)
par: Chen, Yilong, et autres
Publié: (2026)
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
par: Yehudai, Gilad, et autres
Publié: (2025)
par: Yehudai, Gilad, et autres
Publié: (2025)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
par: Shopkhoev, Dmitriy, et autres
Publié: (2025)
par: Shopkhoev, Dmitriy, et autres
Publié: (2025)
SQLong: Enhanced NL2SQL for Longer Contexts with LLMs
par: Nguyen, Dai Quoc, et autres
Publié: (2025)
par: Nguyen, Dai Quoc, et autres
Publié: (2025)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
par: Yang, Senqiao, et autres
Publié: (2024)
par: Yang, Senqiao, et autres
Publié: (2024)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
par: Wang, Shouren, et autres
Publié: (2025)
par: Wang, Shouren, et autres
Publié: (2025)
MeTHanol: Modularized Thinking Language Models with Intermediate Layer Thinking, Decoding and Bootstrapping Reasoning
par: Xi, Ningyuan, et autres
Publié: (2024)
par: Xi, Ningyuan, et autres
Publié: (2024)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
par: Xu, Xin, et autres
Publié: (2026)
par: Xu, Xin, et autres
Publié: (2026)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
par: Yehudai, Gilad, et autres
Publié: (2025)
par: Yehudai, Gilad, et autres
Publié: (2025)
From Faithfulness to Correctness: Generative Reward Models that Think Critically
par: Ma, Qiyao, et autres
Publié: (2025)
par: Ma, Qiyao, et autres
Publié: (2025)
Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
par: Singh, Joykirat, et autres
Publié: (2025)
par: Singh, Joykirat, et autres
Publié: (2025)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
par: Kapl, Ferdinand, et autres
Publié: (2025)
par: Kapl, Ferdinand, et autres
Publié: (2025)
Efficient Reasoning with Balanced Thinking
par: Li, Yulin, et autres
Publié: (2026)
par: Li, Yulin, et autres
Publié: (2026)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
par: Bergner, Benjamin, et autres
Publié: (2024)
par: Bergner, Benjamin, et autres
Publié: (2024)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
par: Zhu, Zihao, et autres
Publié: (2025)
par: Zhu, Zihao, et autres
Publié: (2025)
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
par: Wan, Xu, et autres
Publié: (2025)
par: Wan, Xu, et autres
Publié: (2025)
AdaptThink: Reasoning Models Can Learn When to Think
par: Zhang, Jiajie, et autres
Publié: (2025)
par: Zhang, Jiajie, et autres
Publié: (2025)
Associative Recurrent Memory Transformer
par: Rodkin, Ivan, et autres
Publié: (2024)
par: Rodkin, Ivan, et autres
Publié: (2024)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
par: Kothapalli, Vignesh, et autres
Publié: (2025)
par: Kothapalli, Vignesh, et autres
Publié: (2025)
Reverse Thinking Makes LLMs Stronger Reasoners
par: Chen, Justin Chih-Yao, et autres
Publié: (2024)
par: Chen, Justin Chih-Yao, et autres
Publié: (2024)
SCI-Verifier: Scientific Verifier with Thinking
par: Zheng, Shenghe, et autres
Publié: (2025)
par: Zheng, Shenghe, et autres
Publié: (2025)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
par: Guo, Zhenyuan, et autres
Publié: (2026)
par: Guo, Zhenyuan, et autres
Publié: (2026)
CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model
par: Chiang, Shang-Hsuan, et autres
Publié: (2024)
par: Chiang, Shang-Hsuan, et autres
Publié: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
par: Zhou, Yongchao, et autres
Publié: (2024)
par: Zhou, Yongchao, et autres
Publié: (2024)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
par: Yang, Diji, et autres
Publié: (2024)
par: Yang, Diji, et autres
Publié: (2024)
Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
par: Jha, Basab, et autres
Publié: (2025)
par: Jha, Basab, et autres
Publié: (2025)
Documents similaires
-
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
par: Yang, Wang, et autres
Publié: (2025) -
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
par: Kohli, Harsh, et autres
Publié: (2026) -
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
par: McLeish, Sean, et autres
Publié: (2025) -
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
par: Lu, Wenquan, et autres
Publié: (2025) -
CeRA: Overcoming the Linear Ceiling of Low-Rank Adaptation via Capacity Expansion
par: Chen, Hung-Hsuan
Publié: (2026)