Advancing LLM Safe Alignment with Safety Representation Ranking
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Tianqi, Wei, Zeming, Chen, Quan, Zhang, Chenheng, Wang, Yisen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Long-Short Alignment for Effective Long-Context Modeling in LLMs
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
A Theoretical Understanding of Self-Correction through In-context Alignment
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
Language Ranker: A Lightweight Ranking framework for LLM Decoding
by: Zhang, Chenheng, et al.
Published: (2025)
by: Zhang, Chenheng, et al.
Published: (2025)
Autoregressive Models Rival Diffusion Models at ANY-ORDER Generation
by: Du, Tianqi, et al.
Published: (2026)
by: Du, Tianqi, et al.
Published: (2026)
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
What is Wrong with Perplexity for Long-context Language Modeling?
by: Fang, Lizhe, et al.
Published: (2024)
by: Fang, Lizhe, et al.
Published: (2024)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
On the Role of Discrete Tokenization in Visual Representation Learning
by: Du, Tianqi, et al.
Published: (2024)
by: Du, Tianqi, et al.
Published: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
by: Wei, Zeming, et al.
Published: (2023)
by: Wei, Zeming, et al.
Published: (2023)
CrossTrafficLLM: A Human-Centric Framework for Interpretable Traffic Intelligence via Large Language Model
by: Du, Zeming, et al.
Published: (2025)
by: Du, Zeming, et al.
Published: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
by: Wu, Chengcan, et al.
Published: (2025)
by: Wu, Chengcan, et al.
Published: (2025)
Unintended Impacts of LLM Alignment on Global Representation
by: Ryan, Michael J., et al.
Published: (2024)
by: Ryan, Michael J., et al.
Published: (2024)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
by: Wang, Dianyun, et al.
Published: (2025)
by: Wang, Dianyun, et al.
Published: (2025)
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
by: Perin, Gabriel J., et al.
Published: (2025)
by: Perin, Gabriel J., et al.
Published: (2025)
RAPO: Risk-Aware Preference Optimization for Generalizable Safe Reasoning
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
by: Sun, Guanglong, et al.
Published: (2026)
by: Sun, Guanglong, et al.
Published: (2026)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
by: Yang, Junxiao, et al.
Published: (2026)
by: Yang, Junxiao, et al.
Published: (2026)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
CSULoRA: Closest Safe Update Low-Rank Adaptation
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
by: Ji, Xiaotong, et al.
Published: (2025)
by: Ji, Xiaotong, et al.
Published: (2025)
Theoretical Understanding of In-Context Learning in Shallow Transformers with Unstructured Data
by: Xing, Yue, et al.
Published: (2024)
by: Xing, Yue, et al.
Published: (2024)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
by: Wang, Guoli, et al.
Published: (2026)
by: Wang, Guoli, et al.
Published: (2026)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
by: Dai, Wei, et al.
Published: (2024)
by: Dai, Wei, et al.
Published: (2024)
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
by: Li, Mingjie, et al.
Published: (2025)
by: Li, Mingjie, et al.
Published: (2025)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
by: Ai, Mengting, et al.
Published: (2023)
by: Ai, Mengting, et al.
Published: (2023)
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
by: Chen, Menglan, et al.
Published: (2025)
by: Chen, Menglan, et al.
Published: (2025)
Decoding Large Language Diffusion Models with Foreseeing Movement
by: Mo, Yichuan, et al.
Published: (2025)
by: Mo, Yichuan, et al.
Published: (2025)
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
DLM-One: Diffusion Language Models for One-Step Sequence Generation
by: Chen, Tianqi, et al.
Published: (2025)
by: Chen, Tianqi, et al.
Published: (2025)
Automata Extraction from Transformers
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
by: Wang, Duo, et al.
Published: (2024)
by: Wang, Duo, et al.
Published: (2024)
Similar Items
-
Long-Short Alignment for Effective Long-Context Modeling in LLMs
by: Du, Tianqi, et al.
Published: (2025) -
A Theoretical Understanding of Self-Correction through In-context Alignment
by: Wang, Yifei, et al.
Published: (2024) -
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025) -
Language Ranker: A Lightweight Ranking framework for LLM Decoding
by: Zhang, Chenheng, et al.
Published: (2025) -
Autoregressive Models Rival Diffusion Models at ANY-ORDER Generation
by: Du, Tianqi, et al.
Published: (2026)