SCORE: Replacing Layer Stacking with Contractive Recurrent Depth
Fuente:
arXiv
Saved in:
| Main Author: | Godin, Guillaume |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
All You Need Is Synthetic Task Augmentation
by: Godin, Guillaume
Published: (2025)
by: Godin, Guillaume
Published: (2025)
Z-Error Loss for Training Neural Networks
by: Godin, Guillaume
Published: (2025)
by: Godin, Guillaume
Published: (2025)
Bond-Centered Molecular Fingerprint Derivatives: A BBBP Dataset Study
by: Godin, Guillaume
Published: (2025)
by: Godin, Guillaume
Published: (2025)
Fast Leave-One-Out Approximation from Fragment-Target Prevalence Vectors (molFTP) : From Dummy Masking to Key-LOO for Leakage-Free Feature Construction
by: Godin, Guillaume
Published: (2025)
by: Godin, Guillaume
Published: (2025)
Be aware of overfitting by hyperparameter optimization!
by: Tetko, Igor V., et al.
Published: (2024)
by: Tetko, Igor V., et al.
Published: (2024)
Looped SSMs: Depth-Recurrence and Input Reshaping for Time Series Classification
by: Farsang, Mónika, et al.
Published: (2026)
by: Farsang, Mónika, et al.
Published: (2026)
Depth-Structured Music Recurrence: Budgeted Recurrent Attention for Full-Piece Symbolic Music Modeling
by: Yi, Yungang, et al.
Published: (2026)
by: Yi, Yungang, et al.
Published: (2026)
Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs
by: Yang, Zhichao, et al.
Published: (2026)
by: Yang, Zhichao, et al.
Published: (2026)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
by: Rodkin, Ivan, et al.
Published: (2025)
by: Rodkin, Ivan, et al.
Published: (2025)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
by: Lu, Wenquan, et al.
Published: (2025)
by: Lu, Wenquan, et al.
Published: (2025)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
by: Shopkhoev, Dmitriy, et al.
Published: (2025)
by: Shopkhoev, Dmitriy, et al.
Published: (2025)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
by: Chen, Hung-Hsuan
Published: (2026)
by: Chen, Hung-Hsuan
Published: (2026)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
by: Kohli, Harsh, et al.
Published: (2026)
by: Kohli, Harsh, et al.
Published: (2026)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
by: Knupp, Jonas, et al.
Published: (2026)
by: Knupp, Jonas, et al.
Published: (2026)
Dense Depth from Event Focal Stack
by: Horikawa, Kenta, et al.
Published: (2024)
by: Horikawa, Kenta, et al.
Published: (2024)
SCORE: Syntactic Code Representations for Static Script Malware Detection
by: Erdemir, Ecenaz, et al.
Published: (2024)
by: Erdemir, Ecenaz, et al.
Published: (2024)
Inverse Depth Scaling From Most Layers Being Similar
by: Liu, Yizhou, et al.
Published: (2026)
by: Liu, Yizhou, et al.
Published: (2026)
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities
by: Wang, Kevin, et al.
Published: (2025)
by: Wang, Kevin, et al.
Published: (2025)
Contraction Actor-Critic: Contraction Metric-Guided Reinforcement Learning for Robust Path Tracking
by: Cho, Minjae, et al.
Published: (2025)
by: Cho, Minjae, et al.
Published: (2025)
StackLiverNet: A Novel Stacked Ensemble Model for Accurate and Interpretable Liver Disease Detection
by: Haque, Md. Ehsanul, et al.
Published: (2025)
by: Haque, Md. Ehsanul, et al.
Published: (2025)
Improving the Diversity of Bootstrapped DQN by Replacing Priors With Noise
by: Meng, Li, et al.
Published: (2022)
by: Meng, Li, et al.
Published: (2022)
GRADE: Replacing Policy Gradients with Backpropagation for LLM Alignment
by: Nel, Lukas Abrie
Published: (2025)
by: Nel, Lukas Abrie
Published: (2025)
Denoising-based Contractive Imitation Learning
by: Shen, Macheng, et al.
Published: (2025)
by: Shen, Macheng, et al.
Published: (2025)
Stacking for Probabilistic Short-term Load Forecasting
by: Dudek, Grzegorz
Published: (2024)
by: Dudek, Grzegorz
Published: (2024)
Recurrent Reinforcement Learning with Memoroids
by: Morad, Steven, et al.
Published: (2024)
by: Morad, Steven, et al.
Published: (2024)
Recurrent Action Transformer with Memory
by: Cherepanov, Egor, et al.
Published: (2023)
by: Cherepanov, Egor, et al.
Published: (2023)
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
by: Xia, Xiaojie, et al.
Published: (2026)
by: Xia, Xiaojie, et al.
Published: (2026)
Replacing thinking with tool usage enables reasoning in small language models
by: Rainone, Corrado, et al.
Published: (2025)
by: Rainone, Corrado, et al.
Published: (2025)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
Symbol-Equivariant Recurrent Reasoning Models
by: Freinschlag, Richard, et al.
Published: (2026)
by: Freinschlag, Richard, et al.
Published: (2026)
Neural Contractive Dynamical Systems
by: Beik-Mohammadi, Hadi, et al.
Published: (2024)
by: Beik-Mohammadi, Hadi, et al.
Published: (2024)
Recurrent Aggregators in Neural Algorithmic Reasoning
by: Xu, Kaijia, et al.
Published: (2024)
by: Xu, Kaijia, et al.
Published: (2024)
READ: Recurrent Adaptation of Large Transformers
by: Nguyen, John, et al.
Published: (2023)
by: Nguyen, John, et al.
Published: (2023)
Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning
by: Dong, Jinzong, et al.
Published: (2026)
by: Dong, Jinzong, et al.
Published: (2026)
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
by: Hao, Jiawei, et al.
Published: (2026)
by: Hao, Jiawei, et al.
Published: (2026)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
by: Jeon, Wonseok, et al.
Published: (2024)
by: Jeon, Wonseok, et al.
Published: (2024)
Stacked Universal Successor Feature Approximators for Safety in Reinforcement Learning
by: Cannon, Ian, et al.
Published: (2024)
by: Cannon, Ian, et al.
Published: (2024)
Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?
by: Jin, Mingyu, et al.
Published: (2024)
by: Jin, Mingyu, et al.
Published: (2024)
Discovering Data Manifold Geometry via Non-Contracting Flows
by: Vigouroux, David, et al.
Published: (2026)
by: Vigouroux, David, et al.
Published: (2026)
Similar Items
-
All You Need Is Synthetic Task Augmentation
by: Godin, Guillaume
Published: (2025) -
Z-Error Loss for Training Neural Networks
by: Godin, Guillaume
Published: (2025) -
Bond-Centered Molecular Fingerprint Derivatives: A BBBP Dataset Study
by: Godin, Guillaume
Published: (2025) -
Fast Leave-One-Out Approximation from Fragment-Target Prevalence Vectors (molFTP) : From Dummy Masking to Key-LOO for Leakage-Free Feature Construction
by: Godin, Guillaume
Published: (2025) -
Be aware of overfitting by hyperparameter optimization!
by: Tetko, Igor V., et al.
Published: (2024)