Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Jinqi, Ding, Tianjiao, Chan, Kwan Ho Ryan, Thaker, Darshan, Chattopadhyay, Aditya, Callison-Burch, Chris, Vidal, René
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2406.04331
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912105143205888
author Luo, Jinqi
Ding, Tianjiao
Chan, Kwan Ho Ryan
Thaker, Darshan
Chattopadhyay, Aditya
Callison-Burch, Chris
Vidal, René
author_facet Luo, Jinqi
Ding, Tianjiao
Chan, Kwan Ho Ryan
Thaker, Darshan
Chattopadhyay, Aditya
Callison-Burch, Chris
Vidal, René
contents Large Language Models (LLMs) are being used for a wide variety of tasks. While they are capable of generating human-like responses, they can also produce undesirable output including potentially harmful information, racist or sexist language, and hallucinations. Alignment methods are designed to reduce such undesirable outputs via techniques such as fine-tuning, prompt engineering, and representation engineering. However, existing methods face several challenges: some require costly fine-tuning for every alignment task; some do not adequately remove undesirable concepts, failing alignment; some remove benign concepts, lowering the linguistic capabilities of LLMs. To address these issues, we propose Parsimonious Concept Engineering (PaCE), a novel activation engineering framework for alignment. First, to sufficiently model the concepts, we construct a large-scale concept dictionary in the activation space, in which each atom corresponds to a semantic concept. Given any alignment task, we instruct a concept partitioner to efficiently annotate the concepts as benign or undesirable. Then, at inference time, we decompose the LLM activations along the concept dictionary via sparse coding, to accurately represent the activations as linear combinations of benign and undesirable components. By removing the latter ones from the activations, we reorient the behavior of the LLM towards the alignment goal. We conduct experiments on tasks such as response detoxification, faithfulness enhancement, and sentiment revising, and show that PaCE achieves state-of-the-art alignment performance while maintaining linguistic capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PaCE: Parsimonious Concept Engineering for Large Language Models
Luo, Jinqi
Ding, Tianjiao
Chan, Kwan Ho Ryan
Thaker, Darshan
Chattopadhyay, Aditya
Callison-Burch, Chris
Vidal, René
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Large Language Models (LLMs) are being used for a wide variety of tasks. While they are capable of generating human-like responses, they can also produce undesirable output including potentially harmful information, racist or sexist language, and hallucinations. Alignment methods are designed to reduce such undesirable outputs via techniques such as fine-tuning, prompt engineering, and representation engineering. However, existing methods face several challenges: some require costly fine-tuning for every alignment task; some do not adequately remove undesirable concepts, failing alignment; some remove benign concepts, lowering the linguistic capabilities of LLMs. To address these issues, we propose Parsimonious Concept Engineering (PaCE), a novel activation engineering framework for alignment. First, to sufficiently model the concepts, we construct a large-scale concept dictionary in the activation space, in which each atom corresponds to a semantic concept. Given any alignment task, we instruct a concept partitioner to efficiently annotate the concepts as benign or undesirable. Then, at inference time, we decompose the LLM activations along the concept dictionary via sparse coding, to accurately represent the activations as linear combinations of benign and undesirable components. By removing the latter ones from the activations, we reorient the behavior of the LLM towards the alignment goal. We conduct experiments on tasks such as response detoxification, faithfulness enhancement, and sentiment revising, and show that PaCE achieves state-of-the-art alignment performance while maintaining linguistic capabilities.
title PaCE: Parsimonious Concept Engineering for Large Language Models
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2406.04331