One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shen, Jucheng, Su, Barbara, Kyrillidis, Anastasios |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
par: Vandchali, Mahtab Alizadeh, et autres
Publié: (2025)
par: Vandchali, Mahtab Alizadeh, et autres
Publié: (2025)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
par: Liao, Fangshuo, et autres
Publié: (2026)
par: Liao, Fangshuo, et autres
Publié: (2026)
Unveiling Hidden Pivotal Players with GoalNet: A GNN-Based Soccer Player Evaluation System
par: Jiang, Jacky Hao, et autres
Publié: (2025)
par: Jiang, Jacky Hao, et autres
Publié: (2025)
Convergence Analysis of Two-Layer Neural Networks under Gaussian Input Masking
par: Kolomvaki, Afroditi, et autres
Publié: (2026)
par: Kolomvaki, Afroditi, et autres
Publié: (2026)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
par: Farhat, Yehya, et autres
Publié: (2023)
par: Farhat, Yehya, et autres
Publié: (2023)
TwIST: Rigging the Lottery in Transformers with Independent Subnetwork Training
par: Menezes, Michael, et autres
Publié: (2025)
par: Menezes, Michael, et autres
Publié: (2025)
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
par: Liao, Fangshuo, et autres
Publié: (2025)
par: Liao, Fangshuo, et autres
Publié: (2025)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
par: Liao, Fangshuo, et autres
Publié: (2023)
par: Liao, Fangshuo, et autres
Publié: (2023)
Thinking Out of the Box: Hybrid SAT Solving by Unconstrained Continuous Optimization
par: Zhang, Zhiwei, et autres
Publié: (2025)
par: Zhang, Zhiwei, et autres
Publié: (2025)
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
par: Liao, Fangshuo, et autres
Publié: (2025)
par: Liao, Fangshuo, et autres
Publié: (2025)
GHOST: Unmasking Phantom States in Mamba2 via Grouped Hidden-state Output-aware Selection & Truncation
par: Menezes, Michael, et autres
Publié: (2026)
par: Menezes, Michael, et autres
Publié: (2026)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
par: Li, Zihao, et autres
Publié: (2024)
par: Li, Zihao, et autres
Publié: (2024)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
par: Kim, Junhyung Lyle, et autres
Publié: (2022)
par: Kim, Junhyung Lyle, et autres
Publié: (2022)
On the Error-Propagation of Inexact Hotelling's Deflation for Principal Component Analysis
par: Liao, Fangshuo, et autres
Publié: (2023)
par: Liao, Fangshuo, et autres
Publié: (2023)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
par: Su, Weijie
Publié: (2025)
par: Su, Weijie
Publié: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
par: Zucchet, Nicolas, et autres
Publié: (2024)
par: Zucchet, Nicolas, et autres
Publié: (2024)
Exploiting Low-Rank Structure in Max-K-Cut Problems
par: Stevens, Ria, et autres
Publié: (2026)
par: Stevens, Ria, et autres
Publié: (2026)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
par: Wang, Jinbo, et autres
Publié: (2025)
par: Wang, Jinbo, et autres
Publié: (2025)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
par: Du, Zhehang, et autres
Publié: (2026)
par: Du, Zhehang, et autres
Publié: (2026)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
par: Sheen, Heejune, et autres
Publié: (2024)
par: Sheen, Heejune, et autres
Publié: (2024)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
par: Molodtsov, Gleb, et autres
Publié: (2026)
par: Molodtsov, Gleb, et autres
Publié: (2026)
Training Infinitely Deep and Wide Transformers
par: Barboni, Raphaël, et autres
Publié: (2026)
par: Barboni, Raphaël, et autres
Publié: (2026)
Stability of Transformers under Layer Normalization
par: Kan, Kelvin, et autres
Publié: (2025)
par: Kan, Kelvin, et autres
Publié: (2025)
Spatial Transformers for Radio Map Estimation
par: Viet, Pham Q., et autres
Publié: (2024)
par: Viet, Pham Q., et autres
Publié: (2024)
Deep Learning for Two-Stage Robust Integer Optimization
par: Dumouchelle, Justin, et autres
Publié: (2023)
par: Dumouchelle, Justin, et autres
Publié: (2023)
The Newton-Muon Optimizer
par: Du, Zhehang, et autres
Publié: (2026)
par: Du, Zhehang, et autres
Publié: (2026)
A multilevel approach to accelerate the training of Transformers
par: Lauga, Guillaume, et autres
Publié: (2025)
par: Lauga, Guillaume, et autres
Publié: (2025)
Quantum EigenGame for excited state calculation
par: Quiroga, David, et autres
Publié: (2025)
par: Quiroga, David, et autres
Publié: (2025)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
par: Zimin, Aleksandr, et autres
Publié: (2026)
par: Zimin, Aleksandr, et autres
Publié: (2026)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
par: Arda, Enes, et autres
Publié: (2026)
par: Arda, Enes, et autres
Publié: (2026)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
par: Yang, Tong, et autres
Publié: (2026)
par: Yang, Tong, et autres
Publié: (2026)
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
par: Kan, Kelvin, et autres
Publié: (2025)
par: Kan, Kelvin, et autres
Publié: (2025)
How Well Can Transformers Emulate In-context Newton's Method?
par: Giannou, Angeliki, et autres
Publié: (2024)
par: Giannou, Angeliki, et autres
Publié: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
par: Huang, Yu, et autres
Publié: (2025)
par: Huang, Yu, et autres
Publié: (2025)
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
par: Hrycej, Tomas, et autres
Publié: (2025)
par: Hrycej, Tomas, et autres
Publié: (2025)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
par: Lau, Tim Tsz-Kit, et autres
Publié: (2026)
par: Lau, Tim Tsz-Kit, et autres
Publié: (2026)
Revisiting Zeroth-Order Optimization: Minimum-Variance Two-Point Estimators and Directionally Aligned Perturbations
par: Ma, Shaocong, et autres
Publié: (2025)
par: Ma, Shaocong, et autres
Publié: (2025)
From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
par: Ameli, Mostafa, et autres
Publié: (2025)
par: Ameli, Mostafa, et autres
Publié: (2025)
Double-Bounded Optimal Transport for Advanced Clustering and Classification
par: Shi, Liangliang, et autres
Publié: (2024)
par: Shi, Liangliang, et autres
Publié: (2024)
On the Hopf-Cole Transform for Control-affine Schrödinger Bridge
par: Teter, Alexis, et autres
Publié: (2025)
par: Teter, Alexis, et autres
Publié: (2025)
Documents similaires
-
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
par: Vandchali, Mahtab Alizadeh, et autres
Publié: (2025) -
SGD at the Edge of Stability: The Stochastic Sharpness Gap
par: Liao, Fangshuo, et autres
Publié: (2026) -
Unveiling Hidden Pivotal Players with GoalNet: A GNN-Based Soccer Player Evaluation System
par: Jiang, Jacky Hao, et autres
Publié: (2025) -
Convergence Analysis of Two-Layer Neural Networks under Gaussian Input Masking
par: Kolomvaki, Afroditi, et autres
Publié: (2026) -
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
par: Farhat, Yehya, et autres
Publié: (2023)