Analysing the Residual Stream of Language Models Under Knowledge Conflicts

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Yu, Du, Xiaotang, Hong, Giwon, Gema, Aryo Pradipta, Devoto, Alessio, Wang, Hongru, He, Xuanli, Wong, Kam-Fai, Minervini, Pasquale
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916605437411328
author Zhao, Yu
Du, Xiaotang
Hong, Giwon
Gema, Aryo Pradipta
Devoto, Alessio
Wang, Hongru
He, Xuanli
Wong, Kam-Fai
Minervini, Pasquale
author_facet Zhao, Yu
Du, Xiaotang
Hong, Giwon
Gema, Aryo Pradipta
Devoto, Alessio
Wang, Hongru
He, Xuanli
Wong, Kam-Fai
Minervini, Pasquale
contents Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided in the context. Such conflicts can lead to undesirable model behaviour, such as reliance on outdated or incorrect information. In this work, we investigate whether LLMs can identify knowledge conflicts and whether it is possible to know which source of knowledge the model will rely on by analysing the residual stream of the LLM. Through probing tasks, we find that LLMs can internally register the signal of knowledge conflict in the residual stream, which can be accurately detected by probing the intermediate model activations. This allows us to detect conflicts within the residual stream before generating the answers without modifying the input or model parameters. Moreover, we find that the residual stream shows significantly different patterns when the model relies on contextual knowledge versus parametric knowledge to resolve conflicts. This pattern can be employed to estimate the behaviour of LLMs when conflict happens and prevent unexpected answers before producing the answers. Our analysis offers insights into how LLMs internally manage knowledge conflicts and provides a foundation for developing methods to control the knowledge selection processes.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16090
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Analysing the Residual Stream of Language Models Under Knowledge Conflicts
Zhao, Yu
Du, Xiaotang
Hong, Giwon
Gema, Aryo Pradipta
Devoto, Alessio
Wang, Hongru
He, Xuanli
Wong, Kam-Fai
Minervini, Pasquale
Computation and Language
Large language models (LLMs) can store a significant amount of factual knowledge in their parameters. However, their parametric knowledge may conflict with the information provided in the context. Such conflicts can lead to undesirable model behaviour, such as reliance on outdated or incorrect information. In this work, we investigate whether LLMs can identify knowledge conflicts and whether it is possible to know which source of knowledge the model will rely on by analysing the residual stream of the LLM. Through probing tasks, we find that LLMs can internally register the signal of knowledge conflict in the residual stream, which can be accurately detected by probing the intermediate model activations. This allows us to detect conflicts within the residual stream before generating the answers without modifying the input or model parameters. Moreover, we find that the residual stream shows significantly different patterns when the model relies on contextual knowledge versus parametric knowledge to resolve conflicts. This pattern can be employed to estimate the behaviour of LLMs when conflict happens and prevent unexpected answers before producing the answers. Our analysis offers insights into how LLMs internally manage knowledge conflicts and provides a foundation for developing methods to control the knowledge selection processes.
title Analysing the Residual Stream of Language Models Under Knowledge Conflicts
topic Computation and Language
url https://arxiv.org/abs/2410.16090