What Makes Looped Transformers Perform Better Than Non-Recursive Ones
Fuente:
arXiv
Salvato in:
| Autori principali: | Gong, Zixuan, Liu, Yong, Teng, Jiaye |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
di: Yu, Hantao, et al.
Pubblicazione: (2025)
di: Yu, Hantao, et al.
Pubblicazione: (2025)
LoopQ: Quantization for Recursive Transformers
di: Fang, Rui, et al.
Pubblicazione: (2026)
di: Fang, Rui, et al.
Pubblicazione: (2026)
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
di: Xu, Ruihan, et al.
Pubblicazione: (2026)
di: Xu, Ruihan, et al.
Pubblicazione: (2026)
Predictive Inference With Fast Feature Conformal Prediction
di: Tang, Zihao, et al.
Pubblicazione: (2024)
di: Tang, Zihao, et al.
Pubblicazione: (2024)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
di: Xu, Huangyu, et al.
Pubblicazione: (2026)
di: Xu, Huangyu, et al.
Pubblicazione: (2026)
Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
di: Sun, Yifan, et al.
Pubblicazione: (2025)
di: Sun, Yifan, et al.
Pubblicazione: (2025)
Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection
di: Ryser, Alain, et al.
Pubblicazione: (2024)
di: Ryser, Alain, et al.
Pubblicazione: (2024)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2024)
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2024)
Sparse MeZO: Less Parameters for Better Performance in Zeroth-Order LLM Fine-Tuning
di: Liu, Yong, et al.
Pubblicazione: (2024)
di: Liu, Yong, et al.
Pubblicazione: (2024)
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models
di: Cameron, Chris, et al.
Pubblicazione: (2026)
di: Cameron, Chris, et al.
Pubblicazione: (2026)
What Makes Quantization for Large Language Models Hard? An Empirical Study from the Lens of Perturbation
di: Gong, Zhuocheng, et al.
Pubblicazione: (2024)
di: Gong, Zhuocheng, et al.
Pubblicazione: (2024)
FERRET: Private Deep Learning Faster And Better Than DPSGD
di: Zagardo, David
Pubblicazione: (2025)
di: Zagardo, David
Pubblicazione: (2025)
First Hallucination Tokens Are Different from Conditional Ones
di: Snel, Jakob, et al.
Pubblicazione: (2025)
di: Snel, Jakob, et al.
Pubblicazione: (2025)
MeSH: Memory-as-State-Highways for Recursive Transformers
di: Yu, Chengting, et al.
Pubblicazione: (2025)
di: Yu, Chengting, et al.
Pubblicazione: (2025)
On the Design Space Between Transformers and Recursive Neural Nets
di: Chowdhury, Jishnu Ray, et al.
Pubblicazione: (2024)
di: Chowdhury, Jishnu Ray, et al.
Pubblicazione: (2024)
Stability and Generalization in Looped Transformers
di: Labovich, Asher
Pubblicazione: (2026)
di: Labovich, Asher
Pubblicazione: (2026)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
di: Razin, Noam, et al.
Pubblicazione: (2025)
di: Razin, Noam, et al.
Pubblicazione: (2025)
Better World Models Can Lead to Better Post-Training Performance
di: Gupta, Prakhar, et al.
Pubblicazione: (2025)
di: Gupta, Prakhar, et al.
Pubblicazione: (2025)
What Makes a Good Diffusion Planner for Decision Making?
di: Lu, Haofei, et al.
Pubblicazione: (2025)
di: Lu, Haofei, et al.
Pubblicazione: (2025)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
di: Zhang, Jiazheng, et al.
Pubblicazione: (2025)
Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction Correction
di: Loconte, Lorenzo, et al.
Pubblicazione: (2026)
di: Loconte, Lorenzo, et al.
Pubblicazione: (2026)
Balance-aware Sequence Sampling Makes Multi-modal Learning Better
di: Guan, Zhi-Hao
Pubblicazione: (2025)
di: Guan, Zhi-Hao
Pubblicazione: (2025)
Mixtraining: A Better Trade-Off Between Compute and Performance
di: Li, Zexin, et al.
Pubblicazione: (2025)
di: Li, Zexin, et al.
Pubblicazione: (2025)
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions
di: Nair, Inderjeet, et al.
Pubblicazione: (2024)
di: Nair, Inderjeet, et al.
Pubblicazione: (2024)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
di: Xu, Kevin, et al.
Pubblicazione: (2025)
di: Xu, Kevin, et al.
Pubblicazione: (2025)
Effective Frontiers: A Unification of Neural Scaling Laws
di: Zou, Jiaxuan, et al.
Pubblicazione: (2026)
di: Zou, Jiaxuan, et al.
Pubblicazione: (2026)
Making Large Language Models Better Knowledge Miners for Online Marketing with Progressive Prompting Augmentation
di: Gan, Chunjing, et al.
Pubblicazione: (2023)
di: Gan, Chunjing, et al.
Pubblicazione: (2023)
Unlocking Out-of-Distribution Generalization in Transformers via Recursive Latent Space Reasoning
di: Altabaa, Awni, et al.
Pubblicazione: (2025)
di: Altabaa, Awni, et al.
Pubblicazione: (2025)
Data or Language Supervision: What Makes CLIP Better than DINO?
di: Liu, Yiming, et al.
Pubblicazione: (2025)
di: Liu, Yiming, et al.
Pubblicazione: (2025)
Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?
di: Omori, Yumi, et al.
Pubblicazione: (2025)
di: Omori, Yumi, et al.
Pubblicazione: (2025)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
di: Tian, Xiao, et al.
Pubblicazione: (2026)
di: Tian, Xiao, et al.
Pubblicazione: (2026)
Offline Inverse Constrained Reinforcement Learning for Safe-Critical Decision Making in Healthcare
di: Fang, Nan, et al.
Pubblicazione: (2024)
di: Fang, Nan, et al.
Pubblicazione: (2024)
Comparative Explanations: Explanation Guided Decision Making for Human-in-the-Loop Preference Selection
di: Chakraborty, Tanmay, et al.
Pubblicazione: (2025)
di: Chakraborty, Tanmay, et al.
Pubblicazione: (2025)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
di: Pecher, Branislav, et al.
Pubblicazione: (2026)
di: Pecher, Branislav, et al.
Pubblicazione: (2026)
Relational Preference Encoding in Looped Transformer Internal States
di: Kirin, Jan
Pubblicazione: (2026)
di: Kirin, Jan
Pubblicazione: (2026)
Modality-Decoupled Online Recursive Editing
di: Li, Siyuan, et al.
Pubblicazione: (2026)
di: Li, Siyuan, et al.
Pubblicazione: (2026)
Do Transformer World Models Give Better Policy Gradients?
di: Ma, Michel, et al.
Pubblicazione: (2024)
di: Ma, Michel, et al.
Pubblicazione: (2024)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
di: Huang, Jianhao, et al.
Pubblicazione: (2025)
di: Huang, Jianhao, et al.
Pubblicazione: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
di: Gong, Zixuan, et al.
Pubblicazione: (2025) -
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
di: Yu, Hantao, et al.
Pubblicazione: (2025) -
LoopQ: Quantization for Recursive Transformers
di: Fang, Rui, et al.
Pubblicazione: (2026) -
Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
di: Xu, Ruihan, et al.
Pubblicazione: (2026) -
Predictive Inference With Fast Feature Conformal Prediction
di: Tang, Zihao, et al.
Pubblicazione: (2024)