Gradient Residual Connections
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Yangchen, Ying, Qizhen, Torr, Philip, Liu, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An MRP Formulation for Supervised Learning: Generalized Temporal Difference Learning Models
by: Pan, Yangchen, et al.
Published: (2024)
by: Pan, Yangchen, et al.
Published: (2024)
Measures of Variability for Risk-averse Policy Gradient
by: Luo, Yudong, et al.
Published: (2025)
by: Luo, Yudong, et al.
Published: (2025)
Reinforcement Learning in Dynamic Treatment Regimes Needs Critical Reexamination
by: Luo, Zhiyao, et al.
Published: (2024)
by: Luo, Zhiyao, et al.
Published: (2024)
ResidualDroppath: Enhancing Feature Reuse over Residual Connections
by: Park, Sejik
Published: (2024)
by: Park, Sejik
Published: (2024)
Real-Fake: Effective Training Data Synthesis Through Distribution Matching
by: Yuan, Jianhao, et al.
Published: (2023)
by: Yuan, Jianhao, et al.
Published: (2023)
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
by: Yang, Puning, et al.
Published: (2026)
by: Yang, Puning, et al.
Published: (2026)
DTR-Bench: An in silico Environment and Benchmark Platform for Reinforcement Learning Based Dynamic Treatment Regime
by: Luo, Zhiyao, et al.
Published: (2024)
by: Luo, Zhiyao, et al.
Published: (2024)
Mitigating Gradient Overlap in Deep Residual Networks with Gradient Normalization for Improved Non-Convex Optimization
by: Yun, Juyoung
Published: (2024)
by: Yun, Juyoung
Published: (2024)
Ablate and Rescue: A Causal Analysis of Residual Stream Hyper-Connections
by: Peng, William, et al.
Published: (2026)
by: Peng, William, et al.
Published: (2026)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Hierarchical Reinforcement Learning for Swarm Confrontation with High Uncertainty
by: Wu, Qizhen, et al.
Published: (2024)
by: Wu, Qizhen, et al.
Published: (2024)
Understanding Reasoning in Thinking Language Models via Steering Vectors
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
by: Aljaafari, Tala, et al.
Published: (2025)
by: Aljaafari, Tala, et al.
Published: (2025)
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Prompting a Pretrained Transformer Can Be a Universal Approximator
by: Petrov, Aleksandar, et al.
Published: (2024)
by: Petrov, Aleksandar, et al.
Published: (2024)
Gradient Regularized Natural Gradients
by: Dash, Satya Prakash, et al.
Published: (2026)
by: Dash, Satya Prakash, et al.
Published: (2026)
Support Vector Boosting Machine (SVBM): Enhancing Classification Performance with AdaBoost and Residual Connections
by: Lian, Junbo Jacob
Published: (2024)
by: Lian, Junbo Jacob
Published: (2024)
Conflict-Averse Gradient Descent for Multi-task Learning
by: Liu, Bo, et al.
Published: (2021)
by: Liu, Bo, et al.
Published: (2021)
Dynamic Context Adaptation and Information Flow Control in Transformers: Introducing the Evaluator Adjuster Unit and Gated Residual Connections
by: Dhayalkar, Sahil Rajesh
Published: (2024)
by: Dhayalkar, Sahil Rajesh
Published: (2024)
TabGen-ICL: Residual-Aware In-Context Example Selection for Tabular Data Generation
by: Fang, Liancheng, et al.
Published: (2025)
by: Fang, Liancheng, et al.
Published: (2025)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
by: Zhang, Wenxuan, et al.
Published: (2024)
by: Zhang, Wenxuan, et al.
Published: (2024)
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
by: Xiang, Maoyang, et al.
Published: (2026)
by: Xiang, Maoyang, et al.
Published: (2026)
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
by: Xiao, Da, et al.
Published: (2025)
by: Xiao, Da, et al.
Published: (2025)
Attention Sinks and Outliers in Attention Residuals
by: Luo, Haozheng, et al.
Published: (2026)
by: Luo, Haozheng, et al.
Published: (2026)
Fast Explanations via Policy Gradient-Optimized Explainer
by: Pan, Deng, et al.
Published: (2024)
by: Pan, Deng, et al.
Published: (2024)
Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
by: Chhabra, Anshuman, et al.
Published: (2024)
by: Chhabra, Anshuman, et al.
Published: (2024)
Select to Perfect: Imitating desired behavior from large multi-agent data
by: Franzmeyer, Tim, et al.
Published: (2024)
by: Franzmeyer, Tim, et al.
Published: (2024)
GradientStabilizer:Fix the Norm, Not the Gradient
by: Huang, Tianjin, et al.
Published: (2025)
by: Huang, Tianjin, et al.
Published: (2025)
Tackling the Non-IID Issue in Heterogeneous Federated Learning by Gradient Harmonization
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
Residual Stream Analysis of Overfitting And Structural Disruptions
by: Liu, Quan, et al.
Published: (2026)
by: Liu, Quan, et al.
Published: (2026)
A Simple Mixture Policy Parameterization for Improving Sample Efficiency of CVaR Optimization
by: Luo, Yudong, et al.
Published: (2024)
by: Luo, Yudong, et al.
Published: (2024)
A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning
by: Feng, Xidong, et al.
Published: (2021)
by: Feng, Xidong, et al.
Published: (2021)
Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
by: Abdi, Hossein, et al.
Published: (2025)
by: Abdi, Hossein, et al.
Published: (2025)
SAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation
by: Xie, Yu, et al.
Published: (2026)
by: Xie, Yu, et al.
Published: (2026)
Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation
by: Dai, Juntao, et al.
Published: (2024)
by: Dai, Juntao, et al.
Published: (2024)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Information Geometry
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Can Past Experience Accelerate LLM Reasoning?
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
Specify and Edit: Overcoming Ambiguity in Text-Based Image Editing
by: Iakovleva, Ekaterina, et al.
Published: (2024)
by: Iakovleva, Ekaterina, et al.
Published: (2024)
Similar Items
-
An MRP Formulation for Supervised Learning: Generalized Temporal Difference Learning Models
by: Pan, Yangchen, et al.
Published: (2024) -
Measures of Variability for Risk-averse Policy Gradient
by: Luo, Yudong, et al.
Published: (2025) -
Reinforcement Learning in Dynamic Treatment Regimes Needs Critical Reexamination
by: Luo, Zhiyao, et al.
Published: (2024) -
ResidualDroppath: Enhancing Feature Reuse over Residual Connections
by: Park, Sejik
Published: (2024) -
Real-Fake: Effective Training Data Synthesis Through Distribution Matching
by: Yuan, Jianhao, et al.
Published: (2023)