MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Chenlu, Wu, Jiancan, Sheng, Leheng, Zhang, Fan, Yuan, Yancheng, Wang, Xiang, He, Xiangnan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911415673028608
author Ding, Chenlu
Wu, Jiancan
Sheng, Leheng
Zhang, Fan
Yuan, Yancheng
Wang, Xiang
He, Xiangnan
author_facet Ding, Chenlu
Wu, Jiancan
Sheng, Leheng
Zhang, Fan
Yuan, Yancheng
Wang, Xiang
He, Xiangnan
contents Multimodal large language models (MLLMs) have demonstrated remarkable capabilities across vision-language tasks, yet their large-scale deployment raises pressing concerns about memorized private data, outdated knowledge, and harmful content. Existing unlearning approaches for MLLMs typically adapt training-based strategies such as gradient ascent or preference optimization, but these methods are computationally expensive, irreversible, and often distort retained knowledge. In this work, we propose MLLMEraser, an input-aware, training-free framework for test-time unlearning. Our approach leverages activation steering to enable dynamic knowledge erasure without parameter updates. Specifically, we construct a multimodal erasure direction by contrasting adversarially perturbed, knowledge-recall image-text pairs with knowledge-erasure counterparts, capturing both textual and visual discrepancies. To prevent unnecessary interference, we further design an input-aware steering mechanism that adaptively determines when and how the erasure direction should be applied, preserving utility on retained knowledge while enforcing forgetting on designated content. Experiments on LLaVA-1.5 and Qwen-2.5-VL demonstrate that MLLMEraser consistently outperforms state-of-the-art MLLM unlearning baselines, achieving stronger forgetting performance with lower computational cost and minimal utility degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
Ding, Chenlu
Wu, Jiancan
Sheng, Leheng
Zhang, Fan
Yuan, Yancheng
Wang, Xiang
He, Xiangnan
Machine Learning
Artificial Intelligence
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities across vision-language tasks, yet their large-scale deployment raises pressing concerns about memorized private data, outdated knowledge, and harmful content. Existing unlearning approaches for MLLMs typically adapt training-based strategies such as gradient ascent or preference optimization, but these methods are computationally expensive, irreversible, and often distort retained knowledge. In this work, we propose MLLMEraser, an input-aware, training-free framework for test-time unlearning. Our approach leverages activation steering to enable dynamic knowledge erasure without parameter updates. Specifically, we construct a multimodal erasure direction by contrasting adversarially perturbed, knowledge-recall image-text pairs with knowledge-erasure counterparts, capturing both textual and visual discrepancies. To prevent unnecessary interference, we further design an input-aware steering mechanism that adaptively determines when and how the erasure direction should be applied, preserving utility on retained knowledge while enforcing forgetting on designated content. Experiments on LLaVA-1.5 and Qwen-2.5-VL demonstrate that MLLMEraser consistently outperforms state-of-the-art MLLM unlearning baselines, achieving stronger forgetting performance with lower computational cost and minimal utility degradation.
title MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.04217