LongSafety: Enhance Safety for Long-Context LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Mianqiu, Liu, Xiaoran, Zhou, Shaojun, Zhang, Mozhi, Guo, Qipeng, Li, Linyang, Tan, Chenkun, Gao, Yang, Wang, Pengyu, Li, Linlin, Liu, Qun, Zhou, Yaqian, Qiu, Xipeng, Huang, Xuanjing
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929733393973248
author Huang, Mianqiu
Liu, Xiaoran
Zhou, Shaojun
Zhang, Mozhi
Guo, Qipeng
Li, Linyang
Tan, Chenkun
Gao, Yang
Wang, Pengyu
Li, Linlin
Liu, Qun
Zhou, Yaqian
Qiu, Xipeng
Huang, Xuanjing
author_facet Huang, Mianqiu
Liu, Xiaoran
Zhou, Shaojun
Zhang, Mozhi
Guo, Qipeng
Li, Linyang
Tan, Chenkun
Gao, Yang
Wang, Pengyu
Li, Linlin
Liu, Qun
Zhou, Yaqian
Qiu, Xipeng
Huang, Xuanjing
contents Recent advancements in model architectures and length extrapolation techniques have significantly extended the context length of large language models (LLMs), paving the way for their application in increasingly complex tasks. However, despite the growing capabilities of long-context LLMs, the safety issues in long-context scenarios remain underexplored. While safety alignment in short context has been widely studied, the safety concerns of long-context LLMs have not been adequately addressed. In this work, we introduce \textbf{LongSafety}, a comprehensive safety alignment dataset for long-context LLMs, containing 10 tasks and 17k samples, with an average length of 40.9k tokens. Our experiments demonstrate that training with LongSafety can enhance long-context safety performance while enhancing short-context safety and preserving general capabilities. Furthermore, we demonstrate that long-context safety does not equal long-context alignment with short-context safety data and LongSafety has generalizing capabilities in context length and long-context safety scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06899
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LongSafety: Enhance Safety for Long-Context LLMs
Huang, Mianqiu
Liu, Xiaoran
Zhou, Shaojun
Zhang, Mozhi
Guo, Qipeng
Li, Linyang
Tan, Chenkun
Gao, Yang
Wang, Pengyu
Li, Linlin
Liu, Qun
Zhou, Yaqian
Qiu, Xipeng
Huang, Xuanjing
Computation and Language
Artificial Intelligence
Machine Learning
Recent advancements in model architectures and length extrapolation techniques have significantly extended the context length of large language models (LLMs), paving the way for their application in increasingly complex tasks. However, despite the growing capabilities of long-context LLMs, the safety issues in long-context scenarios remain underexplored. While safety alignment in short context has been widely studied, the safety concerns of long-context LLMs have not been adequately addressed. In this work, we introduce \textbf{LongSafety}, a comprehensive safety alignment dataset for long-context LLMs, containing 10 tasks and 17k samples, with an average length of 40.9k tokens. Our experiments demonstrate that training with LongSafety can enhance long-context safety performance while enhancing short-context safety and preserving general capabilities. Furthermore, we demonstrate that long-context safety does not equal long-context alignment with short-context safety data and LongSafety has generalizing capabilities in context length and long-context safety scenarios.
title LongSafety: Enhance Safety for Long-Context LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.06899