What Makes Looped Transformers Perform Better Than Non-Recursive Ones
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gong, Zixuan, Liu, Yong, Teng, Jiaye |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
von: Gong, Zixuan, et al.
Veröffentlicht: (2025)
von: Gong, Zixuan, et al.
Veröffentlicht: (2025)
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
von: Yu, Hantao, et al.
Veröffentlicht: (2025)
von: Yu, Hantao, et al.
Veröffentlicht: (2025)
LoopQ: Quantization for Recursive Transformers
von: Fang, Rui, et al.
Veröffentlicht: (2026)
von: Fang, Rui, et al.
Veröffentlicht: (2026)
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
von: Xu, Ruihan, et al.
Veröffentlicht: (2026)
von: Xu, Ruihan, et al.
Veröffentlicht: (2026)
Predictive Inference With Fast Feature Conformal Prediction
von: Tang, Zihao, et al.
Veröffentlicht: (2024)
von: Tang, Zihao, et al.
Veröffentlicht: (2024)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
von: Xu, Huangyu, et al.
Veröffentlicht: (2026)
von: Xu, Huangyu, et al.
Veröffentlicht: (2026)
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
von: Sun, Yifan, et al.
Veröffentlicht: (2025)
Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection
von: Ryser, Alain, et al.
Veröffentlicht: (2024)
von: Ryser, Alain, et al.
Veröffentlicht: (2024)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
von: Mofakhami, Mehrnaz, et al.
Veröffentlicht: (2024)
von: Mofakhami, Mehrnaz, et al.
Veröffentlicht: (2024)
Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
von: Liu, Yong, et al.
Veröffentlicht: (2024)
von: Liu, Yong, et al.
Veröffentlicht: (2024)
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models
von: Cameron, Chris, et al.
Veröffentlicht: (2026)
von: Cameron, Chris, et al.
Veröffentlicht: (2026)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
FERRET: Private Deep Learning Faster And Better Than DPSGD
von: Zagardo, David
Veröffentlicht: (2025)
von: Zagardo, David
Veröffentlicht: (2025)
First Hallucination Tokens Are Different from Conditional Ones
von: Snel, Jakob, et al.
Veröffentlicht: (2025)
von: Snel, Jakob, et al.
Veröffentlicht: (2025)
MeSH: Memory-as-State-Highways for Recursive Transformers
von: Yu, Chengting, et al.
Veröffentlicht: (2025)
von: Yu, Chengting, et al.
Veröffentlicht: (2025)
On the Design Space Between Transformers and Recursive Neural Nets
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2024)
von: Chowdhury, Jishnu Ray, et al.
Veröffentlicht: (2024)
Stability and Generalization in Looped Transformers
von: Labovich, Asher
Veröffentlicht: (2026)
von: Labovich, Asher
Veröffentlicht: (2026)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
von: Razin, Noam, et al.
Veröffentlicht: (2025)
von: Razin, Noam, et al.
Veröffentlicht: (2025)
Better World Models Can Lead to Better Post-Training Performance
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
What Makes a Good Diffusion Planner for Decision Making?
von: Lu, Haofei, et al.
Veröffentlicht: (2025)
von: Lu, Haofei, et al.
Veröffentlicht: (2025)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction Correction
von: Loconte, Lorenzo, et al.
Veröffentlicht: (2026)
von: Loconte, Lorenzo, et al.
Veröffentlicht: (2026)
Balance-aware Sequence Sampling Makes Multi-modal Learning Better
von: Guan, Zhi-Hao
Veröffentlicht: (2025)
von: Guan, Zhi-Hao
Veröffentlicht: (2025)
Mixtraining: A Better Trade-Off Between Compute and Performance
von: Li, Zexin, et al.
Veröffentlicht: (2025)
von: Li, Zexin, et al.
Veröffentlicht: (2025)
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions
von: Nair, Inderjeet, et al.
Veröffentlicht: (2024)
von: Nair, Inderjeet, et al.
Veröffentlicht: (2024)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
von: Xu, Kevin, et al.
Veröffentlicht: (2025)
von: Xu, Kevin, et al.
Veröffentlicht: (2025)
Effective Frontiers: A Unification of Neural Scaling Laws
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
Making Large Language Models Better Knowledge Miners for Online Marketing with Progressive Prompting Augmentation
von: Gan, Chunjing, et al.
Veröffentlicht: (2023)
von: Gan, Chunjing, et al.
Veröffentlicht: (2023)
Unlocking Out-of-Distribution Generalization in Transformers via Recursive Latent Space Reasoning
von: Altabaa, Awni, et al.
Veröffentlicht: (2025)
von: Altabaa, Awni, et al.
Veröffentlicht: (2025)
Data or Language Supervision: What Makes CLIP Better than DINO?
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?
von: Omori, Yumi, et al.
Veröffentlicht: (2025)
von: Omori, Yumi, et al.
Veröffentlicht: (2025)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
von: Tian, Xiao, et al.
Veröffentlicht: (2026)
Offline Inverse Constrained Reinforcement Learning for Safe-Critical Decision Making in Healthcare
von: Fang, Nan, et al.
Veröffentlicht: (2024)
von: Fang, Nan, et al.
Veröffentlicht: (2024)
Comparative Explanations: Explanation Guided Decision Making for Human-in-the-Loop Preference Selection
von: Chakraborty, Tanmay, et al.
Veröffentlicht: (2025)
von: Chakraborty, Tanmay, et al.
Veröffentlicht: (2025)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
Relational Preference Encoding in Looped Transformer Internal States
von: Kirin, Jan
Veröffentlicht: (2026)
von: Kirin, Jan
Veröffentlicht: (2026)
Modality-Decoupled Online Recursive Editing
von: Li, Siyuan, et al.
Veröffentlicht: (2026)
von: Li, Siyuan, et al.
Veröffentlicht: (2026)
Do Transformer World Models Give Better Policy Gradients?
von: Ma, Michel, et al.
Veröffentlicht: (2024)
von: Ma, Michel, et al.
Veröffentlicht: (2024)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
von: Huang, Jianhao, et al.
Veröffentlicht: (2025)
von: Huang, Jianhao, et al.
Veröffentlicht: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
von: Gong, Zixuan, et al.
Veröffentlicht: (2025) -
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
von: Yu, Hantao, et al.
Veröffentlicht: (2025) -
LoopQ: Quantization for Recursive Transformers
von: Fang, Rui, et al.
Veröffentlicht: (2026) -
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
von: Xu, Ruihan, et al.
Veröffentlicht: (2026) -
Predictive Inference With Fast Feature Conformal Prediction
von: Tang, Zihao, et al.
Veröffentlicht: (2024)