The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Wenbo, Liu, Zhichao, Chen, Qiguang, Zhou, Xiangyang, Yu, Haining, Jia, Xiaohua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
di: Pan, Wenbo, et al.
Pubblicazione: (2025)
WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation
di: Liu, Zhichao, et al.
Pubblicazione: (2026)
di: Liu, Zhichao, et al.
Pubblicazione: (2026)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
di: Bennion, Jonathan, et al.
Pubblicazione: (2025)
di: Bennion, Jonathan, et al.
Pubblicazione: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
di: Brito, Iago Alves, et al.
Pubblicazione: (2026)
di: Brito, Iago Alves, et al.
Pubblicazione: (2026)
Orthogonal Finetuning for Direct Preference Optimization
di: Yang, Chenxu, et al.
Pubblicazione: (2024)
di: Yang, Chenxu, et al.
Pubblicazione: (2024)
Safety Alignment via Constrained Knowledge Unlearning
di: Shi, Zesheng, et al.
Pubblicazione: (2025)
di: Shi, Zesheng, et al.
Pubblicazione: (2025)
DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspective
di: Peng, Dengyun, et al.
Pubblicazione: (2025)
di: Peng, Dengyun, et al.
Pubblicazione: (2025)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
di: Li, Shilong, et al.
Pubblicazione: (2024)
di: Li, Shilong, et al.
Pubblicazione: (2024)
COSMIC: Generalized Refusal Direction Identification in LLM Activations
di: Siu, Vincent, et al.
Pubblicazione: (2025)
di: Siu, Vincent, et al.
Pubblicazione: (2025)
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
Uncertainty Quantification of Large Language Models through Multi-Dimensional Responses
di: Chen, Tiejin, et al.
Pubblicazione: (2025)
di: Chen, Tiejin, et al.
Pubblicazione: (2025)
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
di: Zhou, Zhanhui, et al.
Pubblicazione: (2024)
di: Zhou, Zhanhui, et al.
Pubblicazione: (2024)
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
di: Sun, Chenkai, et al.
Pubblicazione: (2025)
di: Sun, Chenkai, et al.
Pubblicazione: (2025)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
Adaptive Stopping for Multi-Turn LLM Reasoning
di: Zhou, Xiaofan, et al.
Pubblicazione: (2026)
di: Zhou, Xiaofan, et al.
Pubblicazione: (2026)
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
di: Gao, Chengqian, et al.
Pubblicazione: (2025)
di: Gao, Chengqian, et al.
Pubblicazione: (2025)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
di: Li, Mingjie, et al.
Pubblicazione: (2026)
di: Li, Mingjie, et al.
Pubblicazione: (2026)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
di: Zhu, Yubo, et al.
Pubblicazione: (2025)
di: Zhu, Yubo, et al.
Pubblicazione: (2025)
What Matters For Safety Alignment?
di: Li, Xing, et al.
Pubblicazione: (2026)
di: Li, Xing, et al.
Pubblicazione: (2026)
CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
di: Kwon, Deuksin, et al.
Pubblicazione: (2025)
di: Kwon, Deuksin, et al.
Pubblicazione: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
di: Wang, Yanbo, et al.
Pubblicazione: (2026)
KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework
di: Li, Cheng, et al.
Pubblicazione: (2025)
di: Li, Cheng, et al.
Pubblicazione: (2025)
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization
di: Wadhwa, Sahil, et al.
Pubblicazione: (2024)
di: Wadhwa, Sahil, et al.
Pubblicazione: (2024)
Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
di: Wang, Xi, et al.
Pubblicazione: (2025)
di: Wang, Xi, et al.
Pubblicazione: (2025)
Flames: Benchmarking Value Alignment of LLMs in Chinese
di: Huang, Kexin, et al.
Pubblicazione: (2023)
di: Huang, Kexin, et al.
Pubblicazione: (2023)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
di: Zhang, Jingyu, et al.
Pubblicazione: (2024)
di: Zhang, Jingyu, et al.
Pubblicazione: (2024)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
di: Cheng, Ruoxi, et al.
Pubblicazione: (2025)
Multiple LLM Agents Debate for Equitable Cultural Alignment
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
di: Ki, Dayeon, et al.
Pubblicazione: (2025)
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
di: Li, Yisen, et al.
Pubblicazione: (2025)
di: Li, Yisen, et al.
Pubblicazione: (2025)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
di: Hao, Haochang, et al.
Pubblicazione: (2026)
di: Hao, Haochang, et al.
Pubblicazione: (2026)
Course-Correction: Safety Alignment Using Synthetic Preferences
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Cat-DPO: Category-Adaptive Safety Alignment
di: Yang, Tiankai, et al.
Pubblicazione: (2026)
di: Yang, Tiankai, et al.
Pubblicazione: (2026)
SafeWorld: Geo-Diverse Safety Alignment
di: Yin, Da, et al.
Pubblicazione: (2024)
di: Yin, Da, et al.
Pubblicazione: (2024)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
di: Alssum, Lama, et al.
Pubblicazione: (2025)
di: Alssum, Lama, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
di: Pan, Wenbo, et al.
Pubblicazione: (2025) -
WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation
di: Liu, Zhichao, et al.
Pubblicazione: (2026) -
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
di: Bennion, Jonathan, et al.
Pubblicazione: (2025) -
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024) -
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
di: Brito, Iago Alves, et al.
Pubblicazione: (2026)