Resolving Lexical Bias in Model Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rizwan, Hammad, Rosati, Domenic, Wu, Ga, Sajjad, Hassan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918156783583232
author Rizwan, Hammad
Rosati, Domenic
Wu, Ga
Sajjad, Hassan
author_facet Rizwan, Hammad
Rosati, Domenic
Wu, Ga
Sajjad, Hassan
contents Model editing aims to modify the outputs of large language models after they are trained. Previous approaches have often involved direct alterations to model weights, which can result in model degradation. Recent techniques avoid making modifications to the model's weights by using an adapter that applies edits to the model when triggered by semantic similarity in the representation space. We demonstrate that current adapter methods are critically vulnerable to strong lexical biases, leading to issues such as applying edits to irrelevant prompts with overlapping words. This paper presents a principled approach to learning a disentangled representation space that facilitates precise localization of edits by maintaining distance between irrelevant prompts while preserving proximity among paraphrases. In our empirical study, we show that our method (Projector Editor Networks for Model Editing - PENME) achieves state-of-the-art model editing results while being more computationally efficient during inference than previous methods and adaptable across different architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10411
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Resolving Lexical Bias in Model Editing
Rizwan, Hammad
Rosati, Domenic
Wu, Ga
Sajjad, Hassan
Computation and Language
Model editing aims to modify the outputs of large language models after they are trained. Previous approaches have often involved direct alterations to model weights, which can result in model degradation. Recent techniques avoid making modifications to the model's weights by using an adapter that applies edits to the model when triggered by semantic similarity in the representation space. We demonstrate that current adapter methods are critically vulnerable to strong lexical biases, leading to issues such as applying edits to irrelevant prompts with overlapping words. This paper presents a principled approach to learning a disentangled representation space that facilitates precise localization of edits by maintaining distance between irrelevant prompts while preserving proximity among paraphrases. In our empirical study, we show that our method (Projector Editor Networks for Model Editing - PENME) achieves state-of-the-art model editing results while being more computationally efficient during inference than previous methods and adaptable across different architectures.
title Resolving Lexical Bias in Model Editing
topic Computation and Language
url https://arxiv.org/abs/2408.10411