Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Seongmin, Cho, Aeree, Kim, Grace C., Peng, ShengYun, Phute, Mansi, Chau, Duen Horng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913878895493120
author Lee, Seongmin
Cho, Aeree
Kim, Grace C.
Peng, ShengYun
Phute, Mansi
Chau, Duen Horng
author_facet Lee, Seongmin
Cho, Aeree
Kim, Grace C.
Peng, ShengYun
Phute, Mansi
Chau, Duen Horng
contents As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe outputs and guide safety, but such connections with safety are often overlooked in prior surveys. We present the first survey that bridges this gap, introducing a unified framework that connects safety-focused interpretation methods, the safety enhancements they inform, and the tools that operationalize them. Our novel taxonomy, organized by LLM workflow stages, summarizes nearly 70 works at their intersections. We conclude with open challenges and future directions. This timely survey helps researchers and practitioners navigate key advancements for safer, more interpretable LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05451
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
Lee, Seongmin
Cho, Aeree
Kim, Grace C.
Peng, ShengYun
Phute, Mansi
Chau, Duen Horng
Software Engineering
Artificial Intelligence
Computation and Language
As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe outputs and guide safety, but such connections with safety are often overlooked in prior surveys. We present the first survey that bridges this gap, introducing a unified framework that connects safety-focused interpretation methods, the safety enhancements they inform, and the tools that operationalize them. Our novel taxonomy, organized by LLM workflow stages, summarizes nearly 70 works at their intersections. We conclude with open challenges and future directions. This timely survey helps researchers and practitioners navigate key advancements for safer, more interpretable LLMs.
title Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.05451