Why Does New Knowledge Create Messy Ripple Effects in LLMs?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Jiaxin, Zhang, Zixuan, Li, Manling, Yu, Pengfei, Ji, Heng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911065013485568
author Qin, Jiaxin
Zhang, Zixuan
Li, Manling
Yu, Pengfei
Ji, Heng
author_facet Qin, Jiaxin
Zhang, Zixuan
Li, Manling
Yu, Pengfei
Ji, Heng
contents Extensive previous research has focused on post-training knowledge editing (KE) for language models (LMs) to ensure that knowledge remains accurate and up-to-date. One desired property and open question in KE is to let edited LMs correctly handle ripple effects, where LM is expected to answer its logically related knowledge accurately. In this paper, we answer the question of why most KE methods still create messy ripple effects. We conduct extensive analysis and identify a salient indicator, GradSim, that effectively reveals when and why updated knowledge ripples in LMs. GradSim is computed by the cosine similarity between gradients of the original fact and its related knowledge. We observe a strong positive correlation between ripple effect performance and GradSim across different LMs, KE methods, and evaluation metrics. Further investigations into three counter-intuitive failure cases (Negation, Over-Ripple, Multi-Lingual) of ripple effects demonstrate that these failures are often associated with very low GradSim. This finding validates that GradSim is an effective indicator of when knowledge ripples in LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12828
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Why Does New Knowledge Create Messy Ripple Effects in LLMs?
Qin, Jiaxin
Zhang, Zixuan
Li, Manling
Yu, Pengfei
Ji, Heng
Computation and Language
Artificial Intelligence
Extensive previous research has focused on post-training knowledge editing (KE) for language models (LMs) to ensure that knowledge remains accurate and up-to-date. One desired property and open question in KE is to let edited LMs correctly handle ripple effects, where LM is expected to answer its logically related knowledge accurately. In this paper, we answer the question of why most KE methods still create messy ripple effects. We conduct extensive analysis and identify a salient indicator, GradSim, that effectively reveals when and why updated knowledge ripples in LMs. GradSim is computed by the cosine similarity between gradients of the original fact and its related knowledge. We observe a strong positive correlation between ripple effect performance and GradSim across different LMs, KE methods, and evaluation metrics. Further investigations into three counter-intuitive failure cases (Negation, Over-Ripple, Multi-Lingual) of ripple effects demonstrate that these failures are often associated with very low GradSim. This finding validates that GradSim is an effective indicator of when knowledge ripples in LMs.
title Why Does New Knowledge Create Messy Ripple Effects in LLMs?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.12828