Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
Fuente:
arXiv
Saved in:
| Main Authors: | Grant, Satchel, Gillioz, Victor, Ward, Jake, McGrath, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Alignment Search
by: Grant, Satchel
Published: (2025)
by: Grant, Satchel
Published: (2025)
Emergent Symbol-like Number Variables in Artificial Neural Networks
by: Grant, Satchel, et al.
Published: (2025)
by: Grant, Satchel, et al.
Published: (2025)
Addressing divergent representations from causal interventions on neural networks
by: Grant, Satchel, et al.
Published: (2025)
by: Grant, Satchel, et al.
Published: (2025)
In-Training Defenses against Emergent Misalignment in Language Models
by: Kaczér, David, et al.
Published: (2025)
by: Kaczér, David, et al.
Published: (2025)
Training Deliberative Monitors for Black-Box Scheming Detection
by: Sinha, Aditya, et al.
Published: (2026)
by: Sinha, Aditya, et al.
Published: (2026)
From Frege to chatGPT: Compositionality in language, cognition, and deep neural networks
by: Russin, Jacob, et al.
Published: (2024)
by: Russin, Jacob, et al.
Published: (2024)
Rank-1 LoRAs Encode Interpretable Reasoning Signals
by: Ward, Jake, et al.
Published: (2025)
by: Ward, Jake, et al.
Published: (2025)
What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
by: Chang, Jiamin, et al.
Published: (2025)
by: Chang, Jiamin, et al.
Published: (2025)
Uncovering Gradient Inversion Risks in Practical Language Model Training
by: Feng, Xinguo, et al.
Published: (2025)
by: Feng, Xinguo, et al.
Published: (2025)
Learning-based estimation of cattle weight gain and its influencing factors
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
Mob-based cattle weight gain forecasting using ML models
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
by: Hossain, Muhammad Riaz Hasib, et al.
Published: (2025)
Training Large Language Models to Reason via EM Policy Gradient
by: Xu, Tianbing
Published: (2025)
by: Xu, Tianbing
Published: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
Understanding Gradient Boosting Classifier: Training, Prediction, and the Role of $γ_j$
by: Chen, Hung-Hsuan
Published: (2024)
by: Chen, Hung-Hsuan
Published: (2024)
Understanding Post-Training Structural Changes in Large Language Models
by: He, Xinyu, et al.
Published: (2025)
by: He, Xinyu, et al.
Published: (2025)
A Voter-Based Stochastic Rejection-Method Framework for Asymptotically Safe Language Model Outputs
by: Watts, Jake R., et al.
Published: (2024)
by: Watts, Jake R., et al.
Published: (2024)
How to Protect Models against Adversarial Unlearning?
by: Jasiorski, Patryk, et al.
Published: (2025)
by: Jasiorski, Patryk, et al.
Published: (2025)
Recontextualization Mitigates Specification Gaming without Modifying the Specification
by: Azarbal, Ariana, et al.
Published: (2025)
by: Azarbal, Ariana, et al.
Published: (2025)
Backward-Friendly Optimization: Training Large Language Models with Approximate Gradients under Memory Constraints
by: Yang, Jing, et al.
Published: (2025)
by: Yang, Jing, et al.
Published: (2025)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Optimal Defenses Against Gradient Reconstruction Attacks
by: Chen, Yuxiao, et al.
Published: (2024)
by: Chen, Yuxiao, et al.
Published: (2024)
Understanding Asynchronous Inference Methods for Vision-Language-Action Models
by: Agouzoul, Ayoub
Published: (2026)
by: Agouzoul, Ayoub
Published: (2026)
Understanding and Accelerating the Training of Masked Diffusion Language Models
by: Hong, Chunsan, et al.
Published: (2026)
by: Hong, Chunsan, et al.
Published: (2026)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
by: Bigelow, Eric, et al.
Published: (2026)
by: Bigelow, Eric, et al.
Published: (2026)
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
by: Ma, Yuhan, et al.
Published: (2024)
by: Ma, Yuhan, et al.
Published: (2024)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
by: Mousavi-Hosseini, Alireza, et al.
Published: (2026)
by: Mousavi-Hosseini, Alireza, et al.
Published: (2026)
DeepDefense: Layer-Wise Gradient-Feature Alignment for Building Robust Neural Networks
by: Lin, Ci, et al.
Published: (2025)
by: Lin, Ci, et al.
Published: (2025)
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Policy-Gradient Training of Language Models for Ranking
by: Gao, Ge, et al.
Published: (2023)
by: Gao, Ge, et al.
Published: (2023)
Zero-Sacrifice Persistent-Robustness Adversarial Defense for Pre-Trained Encoders
by: Lei, Zhuxin, et al.
Published: (2026)
by: Lei, Zhuxin, et al.
Published: (2026)
How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern Comprehension
by: Dai, Xinnan, et al.
Published: (2024)
by: Dai, Xinnan, et al.
Published: (2024)
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
by: Kong, Cong, et al.
Published: (2024)
by: Kong, Cong, et al.
Published: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
by: Das, Rocktim Jyoti, et al.
Published: (2023)
by: Das, Rocktim Jyoti, et al.
Published: (2023)
Towards Understanding Link Predictor Generalizability Under Distribution Shifts
by: Revolinsky, Jay, et al.
Published: (2024)
by: Revolinsky, Jay, et al.
Published: (2024)
Black-box Gradient Attack on Graph Neural Networks: Deeper Insights in Graph-based Attack and Defense
by: Zhan, Haoxi, et al.
Published: (2021)
by: Zhan, Haoxi, et al.
Published: (2021)
Combining Stochastic Defenses to Resist Gradient Inversion: An Ablation Study
by: Scheliga, Daniel, et al.
Published: (2022)
by: Scheliga, Daniel, et al.
Published: (2022)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Similar Items
-
Model Alignment Search
by: Grant, Satchel
Published: (2025) -
Emergent Symbol-like Number Variables in Artificial Neural Networks
by: Grant, Satchel, et al.
Published: (2025) -
Addressing divergent representations from causal interventions on neural networks
by: Grant, Satchel, et al.
Published: (2025) -
In-Training Defenses against Emergent Misalignment in Language Models
by: Kaczér, David, et al.
Published: (2025) -
Training Deliberative Monitors for Black-Box Scheming Detection
by: Sinha, Aditya, et al.
Published: (2026)