LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Junsong, Zhou, Jie, Zhan, Bihao, Yang, Yutao, Pan, Qianjun, Chen, Shilian, Huai, Tianyu, Li, Xin, Chen, Qin, He, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918432722649088
author Li, Junsong
Zhou, Jie
Zhan, Bihao
Yang, Yutao
Pan, Qianjun
Chen, Shilian
Huai, Tianyu
Li, Xin
Chen, Qin
He, Liang
author_facet Li, Junsong
Zhou, Jie
Zhan, Bihao
Yang, Yutao
Pan, Qianjun
Chen, Shilian
Huai, Tianyu
Li, Xin
Chen, Qin
He, Liang
contents Alignment plays a crucial role in Large Language Models (LLMs) in aligning with human preferences on a specific task/domain. Traditional alignment methods suffer from catastrophic forgetting, where models lose previously acquired knowledge when adapting to new preferences or domains. We introduce LifeAlign, a novel framework for lifelong alignment that enables LLMs to maintain consistent human preference alignment across sequential learning tasks without forgetting previously learned knowledge. Our approach consists of two key innovations. First, we propose a focalized preference optimization strategy that aligns LLMs with new preferences while preventing the erosion of knowledge acquired from previous tasks. Second, we develop a short-to-long memory consolidation mechanism that merges denoised short-term preference representations into stable long-term memory using intrinsic dimensionality reduction, enabling efficient storage and retrieval of alignment patterns across diverse domains. We evaluate LifeAlign across multiple sequential alignment tasks spanning different domains and preference types. Experimental results demonstrate that our method achieves superior performance in maintaining both preference alignment quality and knowledge retention compared to existing lifelong learning approaches. The codes and datasets have been released on https://github.com/real-ljs/LifeAlign.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
Li, Junsong
Zhou, Jie
Zhan, Bihao
Yang, Yutao
Pan, Qianjun
Chen, Shilian
Huai, Tianyu
Li, Xin
Chen, Qin
He, Liang
Computation and Language
Artificial Intelligence
Machine Learning
Alignment plays a crucial role in Large Language Models (LLMs) in aligning with human preferences on a specific task/domain. Traditional alignment methods suffer from catastrophic forgetting, where models lose previously acquired knowledge when adapting to new preferences or domains. We introduce LifeAlign, a novel framework for lifelong alignment that enables LLMs to maintain consistent human preference alignment across sequential learning tasks without forgetting previously learned knowledge. Our approach consists of two key innovations. First, we propose a focalized preference optimization strategy that aligns LLMs with new preferences while preventing the erosion of knowledge acquired from previous tasks. Second, we develop a short-to-long memory consolidation mechanism that merges denoised short-term preference representations into stable long-term memory using intrinsic dimensionality reduction, enabling efficient storage and retrieval of alignment patterns across diverse domains. We evaluate LifeAlign across multiple sequential alignment tasks spanning different domains and preference types. Experimental results demonstrate that our method achieves superior performance in maintaining both preference alignment quality and knowledge retention compared to existing lifelong learning approaches. The codes and datasets have been released on https://github.com/real-ljs/LifeAlign.
title LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.17183