Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhipeng, Zhou, Kun, Zhao, Wayne Xin, Wang, Jingyuan, Wen, Ji-Rong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914963218497536
author Chen, Zhipeng
Zhou, Kun
Zhao, Wayne Xin
Wang, Jingyuan
Wen, Ji-Rong
author_facet Chen, Zhipeng
Zhou, Kun
Zhao, Wayne Xin
Wang, Jingyuan
Wen, Ji-Rong
contents Large language models (LLMs) are still struggling in aligning with human preference in complex tasks and scenarios. They are prone to overfit into the unexpected patterns or superficial styles in the training data. We conduct an empirical study that only selects the top-10\% most updated parameters in LLMs for alignment training, and see improvements in the convergence process and final performance. It indicates the existence of redundant neurons in LLMs for alignment training. To reduce its influence, we propose a low-redundant alignment method named \textbf{ALLO}, focusing on optimizing the most related neurons with the most useful supervised signals. Concretely, we first identify the neurons that are related to the human preference data by a gradient-based strategy, then identify the alignment-related key tokens by reward models for computing loss. Besides, we also decompose the alignment process into the forgetting and learning stages, where we first forget the tokens with unaligned knowledge and then learn aligned knowledge, by updating different ratios of neurons, respectively. Experimental results on 10 datasets have shown the effectiveness of ALLO. Our code and data are available at \url{https://github.com/RUCAIBox/ALLO}.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12606
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment
Chen, Zhipeng
Zhou, Kun
Zhao, Wayne Xin
Wang, Jingyuan
Wen, Ji-Rong
Computation and Language
Large language models (LLMs) are still struggling in aligning with human preference in complex tasks and scenarios. They are prone to overfit into the unexpected patterns or superficial styles in the training data. We conduct an empirical study that only selects the top-10\% most updated parameters in LLMs for alignment training, and see improvements in the convergence process and final performance. It indicates the existence of redundant neurons in LLMs for alignment training. To reduce its influence, we propose a low-redundant alignment method named \textbf{ALLO}, focusing on optimizing the most related neurons with the most useful supervised signals. Concretely, we first identify the neurons that are related to the human preference data by a gradient-based strategy, then identify the alignment-related key tokens by reward models for computing loss. Besides, we also decompose the alignment process into the forgetting and learning stages, where we first forget the tokens with unaligned knowledge and then learn aligned knowledge, by updating different ratios of neurons, respectively. Experimental results on 10 datasets have shown the effectiveness of ALLO. Our code and data are available at \url{https://github.com/RUCAIBox/ALLO}.
title Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment
topic Computation and Language
url https://arxiv.org/abs/2406.12606