When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongkang, Zhang, Yihua, Zhang, Shuai, Wang, Meng, Liu, Sijia, Chen, Pin-Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
How does promoting the minority fraction affect generalization? A theoretical study of the one-hidden-layer neural network on group imbalance
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Visual prompting reimagined: The power of the Activation Prompts
by: Zhang, Yihua, et al.
Published: (2026)
by: Zhang, Yihua, et al.
Published: (2026)
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
Decomposing Task Vectors for Refined Model Editing
by: Damirchi, Hamed, et al.
Published: (2025)
by: Damirchi, Hamed, et al.
Published: (2025)
WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models
by: Jia, Jinghan, et al.
Published: (2024)
by: Jia, Jinghan, et al.
Published: (2024)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
by: Chowdhury, Mohammed Nowaz Rabbani, et al.
Published: (2024)
by: Chowdhury, Mohammed Nowaz Rabbani, et al.
Published: (2024)
Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing
by: Wang, Hanhui, et al.
Published: (2024)
by: Wang, Hanhui, et al.
Published: (2024)
The Power of Few: Accelerating and Enhancing Data Reweighting with Coreset Selection
by: Jafari, Mohammad, et al.
Published: (2024)
by: Jafari, Mohammad, et al.
Published: (2024)
Provable In-Context Learning of Nonlinear Regression with Transformers
by: Li, Hongbo, et al.
Published: (2025)
by: Li, Hongbo, et al.
Published: (2025)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
by: Zhang, Mohan, et al.
Published: (2025)
by: Zhang, Mohan, et al.
Published: (2025)
Forget Vectors at Play: Universal Input Perturbations Driving Machine Unlearning in Image Classification
by: Sun, Changchang, et al.
Published: (2024)
by: Sun, Changchang, et al.
Published: (2024)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
Theoretical Learning Performance of Graph Neural Networks: The Impact of Jumping Connections and Layer-wise Sparsification
by: Sun, Jiawei, et al.
Published: (2025)
by: Sun, Jiawei, et al.
Published: (2025)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
by: Shang, Bingqi, et al.
Published: (2025)
by: Shang, Bingqi, et al.
Published: (2025)
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
by: Fan, Chongyu, et al.
Published: (2023)
by: Fan, Chongyu, et al.
Published: (2023)
Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
by: Li, Hongkang, et al.
Published: (2026)
by: Li, Hongkang, et al.
Published: (2026)
Understand the Effectiveness of Shortcuts through the Lens of DCA
by: Sun, Youran, et al.
Published: (2024)
by: Sun, Youran, et al.
Published: (2024)
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
by: Lang, Yicheng, et al.
Published: (2025)
by: Lang, Yicheng, et al.
Published: (2025)
Beyond Task Diversity: Provable Representation Transfer for Sequential Multi-Task Linear Bandits
by: Duong, Thang, et al.
Published: (2025)
by: Duong, Thang, et al.
Published: (2025)
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
by: Chen, Yiwei, et al.
Published: (2025)
by: Chen, Yiwei, et al.
Published: (2025)
Generalized Radius and Integrated Codebook Transforms for Differentiable Vector Quantization
by: You, Haochen, et al.
Published: (2026)
by: You, Haochen, et al.
Published: (2026)
Powering Up Zeroth-Order Training via Subspace Gradient Orthogonalization
by: Lang, Yicheng, et al.
Published: (2026)
by: Lang, Yicheng, et al.
Published: (2026)
Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
by: Zhang, Bingjie, et al.
Published: (2025)
by: Zhang, Bingjie, et al.
Published: (2025)
DeepZero: Scaling up Zeroth-Order Optimization for Deep Model Training
by: Chen, Aochuan, et al.
Published: (2023)
by: Chen, Aochuan, et al.
Published: (2023)
Visual Prompting Upgrades Neural Network Sparsification: A Data-Model Perspective
by: Jin, Can, et al.
Published: (2023)
by: Jin, Can, et al.
Published: (2023)
Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark
by: Zhang, Yihua, et al.
Published: (2024)
by: Zhang, Yihua, et al.
Published: (2024)
Provable Training for Graph Contrastive Learning
by: Yu, Yue, et al.
Published: (2023)
by: Yu, Yue, et al.
Published: (2023)
Task-Specific Efficiency Analysis: When Small Language Models Outperform Large Language Models
by: Cao, Jinghan, et al.
Published: (2026)
by: Cao, Jinghan, et al.
Published: (2026)
Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation
by: Xu, Tian, et al.
Published: (2024)
by: Xu, Tian, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Similar Items
-
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024) -
How does promoting the minority fraction affect generalization? A theoretical study of the one-hidden-layer neural network on group imbalance
by: Li, Hongkang, et al.
Published: (2024) -
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024) -
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024) -
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
by: Li, Hongkang, et al.
Published: (2024)