Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zeyi, He, Xuehai, Ren, LiLiang, Wang, Yiping, Peng, Baolin, Cheng, Hao, Wang, Shuohang, He, Pengcheng, Gao, Jianfeng, Lee, Yong Jae, Shen, Yelong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
by: Wang, Yiping, et al.
Published: (2025)
by: Wang, Yiping, et al.
Published: (2025)
ThetaEvolve: Test-time Learning on Open Problems
by: Wang, Yiping, et al.
Published: (2025)
by: Wang, Yiping, et al.
Published: (2025)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
Mojito: Motion Trajectory and Intensity Control for Video Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
by: Ren, Liliang, et al.
Published: (2025)
by: Ren, Liliang, et al.
Published: (2025)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
by: Zhan, Zheng, et al.
Published: (2025)
by: Zhan, Zheng, et al.
Published: (2025)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
by: Zhang, Rongzhi, et al.
Published: (2024)
by: Zhang, Rongzhi, et al.
Published: (2024)
Reinforcement World Model Learning for LLM-based Agents
by: Yu, Xiao, et al.
Published: (2026)
by: Yu, Xiao, et al.
Published: (2026)
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
by: Ouyang, Siru, et al.
Published: (2024)
by: Ouyang, Siru, et al.
Published: (2024)
Synthetic Computers at Scale for Long-Horizon Productivity Simulation
by: Ge, Tao, et al.
Published: (2026)
by: Ge, Tao, et al.
Published: (2026)
Model-Driven Subspaces for Large-Scale Optimization with Local Approximation Strategy
by: He, Yitong, et al.
Published: (2025)
by: He, Yitong, et al.
Published: (2025)
Bridging the Gap Between Multimodal Foundation Models and World Models
by: He, Xuehai
Published: (2025)
by: He, Xuehai
Published: (2025)
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
by: Ren, Liliang, et al.
Published: (2026)
by: Ren, Liliang, et al.
Published: (2026)
Orchard: An Open-Source Agentic Modeling Framework
by: Peng, Baolin, et al.
Published: (2026)
by: Peng, Baolin, et al.
Published: (2026)
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction
by: He, Renjie
Published: (2026)
by: He, Renjie
Published: (2026)
Latent Manifold Reconstruction and Representation with Topological and Geometrical Regularization
by: Wang, Ren, et al.
Published: (2025)
by: Wang, Ren, et al.
Published: (2025)
Neural Field Transformations for Hybrid Monte Carlo: Architectural Design and Scaling
by: He, Jinchen, et al.
Published: (2025)
by: He, Jinchen, et al.
Published: (2025)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
by: Pappone, Francesco, et al.
Published: (2025)
by: Pappone, Francesco, et al.
Published: (2025)
Multi-LoRA Composition for Image Generation
by: Zhong, Ming, et al.
Published: (2024)
by: Zhong, Ming, et al.
Published: (2024)
DrivAer Transformer: A high-precision and fast prediction method for vehicle aerodynamic drag coefficient based on the DrivAerNet++ dataset
by: He, Jiaqi, et al.
Published: (2025)
by: He, Jiaqi, et al.
Published: (2025)
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
by: Chen, Yifang, et al.
Published: (2024)
by: Chen, Yifang, et al.
Published: (2024)
MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens
by: Zheng, Kaizhi, et al.
Published: (2023)
by: Zheng, Kaizhi, et al.
Published: (2023)
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
by: Chen, Zhengyu, et al.
Published: (2025)
by: Chen, Zhengyu, et al.
Published: (2025)
Early Dexmedetomidine in Ventilated Acute Myocardial Infarction: A Target Trial Emulation Using MIMIC‐IV
by: Shengkai Wang, et al.
Published: (2026)
by: Shengkai Wang, et al.
Published: (2026)
AutoSurfer -- Teaching Web Agents through Comprehensive Surfing, Learning, and Modeling
by: Faisal, Fazle Elahi, et al.
Published: (2026)
by: Faisal, Fazle Elahi, et al.
Published: (2026)
CAT: Contrastive Adversarial Training for Evaluating the Robustness of Protective Perturbations in Latent Diffusion Models
by: Peng, Sen, et al.
Published: (2025)
by: Peng, Sen, et al.
Published: (2025)
Global Convergence in Training Large-Scale Transformers
by: Gao, Cheng, et al.
Published: (2024)
by: Gao, Cheng, et al.
Published: (2024)
Feedback-induced interactive dynamics: unitary but dissipative evolution
by: Wu, Shuohang, et al.
Published: (2022)
by: Wu, Shuohang, et al.
Published: (2022)
SafeLink: Safety-Critical Control Under Dynamic and Irregular Unsafe Regions
by: Hu, Songqiao, et al.
Published: (2025)
by: Hu, Songqiao, et al.
Published: (2025)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
by: Wang, Yaoxiang, et al.
Published: (2025)
by: Wang, Yaoxiang, et al.
Published: (2025)
Breaking the Trade‐Off: An All‐in‐One Strategy for Nacre‐Inspired, Damage‐Tolerant SiC Ceramics With Shape Preservation and Pseudo‐Ductility
by: Chuming Ye, et al.
Published: (2026)
by: Chuming Ye, et al.
Published: (2026)
Multi-View Spectrogram Transformer for Respiratory Sound Classification
by: He, Wentao, et al.
Published: (2023)
by: He, Wentao, et al.
Published: (2023)
Teaching Language Models to Self-Improve through Interactive Demonstrations
by: Yu, Xiao, et al.
Published: (2023)
by: Yu, Xiao, et al.
Published: (2023)
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
by: Li, Miaoran, et al.
Published: (2023)
by: Li, Miaoran, et al.
Published: (2023)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
by: Lu, Wenquan, et al.
Published: (2025)
by: Lu, Wenquan, et al.
Published: (2025)
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
by: Aviss, Thea
Published: (2026)
by: Aviss, Thea
Published: (2026)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
by: Bhattamishra, Satwik, et al.
Published: (2024)
by: Bhattamishra, Satwik, et al.
Published: (2024)
Similar Items
-
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
by: Wang, Yiping, et al.
Published: (2025) -
ThetaEvolve: Test-time Learning on Open Problems
by: Wang, Yiping, et al.
Published: (2025) -
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024) -
Mojito: Motion Trajectory and Intensity Control for Video Generation
by: He, Xuehai, et al.
Published: (2024) -
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
by: Zhang, Zhen, et al.
Published: (2025)