The Blessing and Curse of Dimensionality in Safety Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Teo, Rachel S. Y., Abdullaev, Laziz U., Nguyen, Tan M. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Elliptical Attention
par: Nielsen, Stefan K., et autres
Publié: (2024)
par: Nielsen, Stefan K., et autres
Publié: (2024)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
Tight Clusters Make Specialized Experts
par: Nielsen, Stefan K., et autres
Publié: (2025)
par: Nielsen, Stefan K., et autres
Publié: (2025)
Concept Heterogeneity-aware Representation Steering
par: Abdullaev, Laziz U., et autres
Publié: (2026)
par: Abdullaev, Laziz U., et autres
Publié: (2026)
Transformer Meets Twicing: Harnessing Unattended Residual Information
par: Abdullaev, Laziz, et autres
Publié: (2025)
par: Abdullaev, Laziz, et autres
Publié: (2025)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
par: Cao, Hengjie, et autres
Publié: (2026)
par: Cao, Hengjie, et autres
Publié: (2026)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
par: Teo, Rachel S. Y., et autres
Publié: (2025)
par: Teo, Rachel S. Y., et autres
Publié: (2025)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
par: Teo, Rachel S. Y., et autres
Publié: (2024)
par: Teo, Rachel S. Y., et autres
Publié: (2024)
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
par: Teo, Rachel S. Y., et autres
Publié: (2024)
par: Teo, Rachel S. Y., et autres
Publié: (2024)
Revisiting Transformers with Insights from Image Filtering and Boosting
par: Abdullaev, Laziz U., et autres
Publié: (2025)
par: Abdullaev, Laziz U., et autres
Publié: (2025)
The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
par: Liang, Qiyao, et autres
Publié: (2026)
par: Liang, Qiyao, et autres
Publié: (2026)
How DNNs break the Curse of Dimensionality: Compositionality and Symmetry Learning
par: Jacot, Arthur, et autres
Publié: (2024)
par: Jacot, Arthur, et autres
Publié: (2024)
BSO: Safety Alignment Is Density Ratio Matching
par: Nguyen, Tien-Phat, et autres
Publié: (2026)
par: Nguyen, Tien-Phat, et autres
Publié: (2026)
Mixture of Experts Softens the Curse of Dimensionality in Operator Learning
par: Kratsios, Anastasis, et autres
Publié: (2024)
par: Kratsios, Anastasis, et autres
Publié: (2024)
CNNs Avoid Curse of Dimensionality by Learning on Patches
par: Madala, Vamshi C., et autres
Publié: (2022)
par: Madala, Vamshi C., et autres
Publié: (2022)
The Curse of Depth in Large Language Models
par: Sun, Wenfang, et autres
Publié: (2025)
par: Sun, Wenfang, et autres
Publié: (2025)
Dispelling the Curse of Singularities in Neural Network Optimizations
par: Cao, Hengjie, et autres
Publié: (2026)
par: Cao, Hengjie, et autres
Publié: (2026)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
par: Kitouni, Ouail, et autres
Publié: (2024)
par: Kitouni, Ouail, et autres
Publié: (2024)
Curriculum Learning for Safety Alignment
par: Kumar, Sandeep, et autres
Publié: (2026)
par: Kumar, Sandeep, et autres
Publié: (2026)
An Analysis and Mitigation of the Reversal Curse
par: Lv, Ang, et autres
Publié: (2023)
par: Lv, Ang, et autres
Publié: (2023)
Completion of the DrugMatrix Toxicogenomics Database using 3-Dimensional Tensors
par: Nguyen, Tan, et autres
Publié: (2025)
par: Nguyen, Tan, et autres
Publié: (2025)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
par: Zhang, Jiawei, et autres
Publié: (2025)
par: Zhang, Jiawei, et autres
Publié: (2025)
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
par: Guo, Qingyan, et autres
Publié: (2024)
par: Guo, Qingyan, et autres
Publié: (2024)
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
par: Lu, Haoran, et autres
Publié: (2025)
par: Lu, Haoran, et autres
Publié: (2025)
On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation
par: Zhang, Yuheng, et autres
Publié: (2024)
par: Zhang, Yuheng, et autres
Publié: (2024)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
par: Tan, Zhewen, et autres
Publié: (2026)
par: Tan, Zhewen, et autres
Publié: (2026)
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
par: Springer, Max, et autres
Publié: (2026)
par: Springer, Max, et autres
Publié: (2026)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
par: Feng, Jingyuan, et autres
Publié: (2026)
par: Feng, Jingyuan, et autres
Publié: (2026)
A Principled Loss Function for Direct Language Model Alignment
par: Tan, Yuandong
Publié: (2025)
par: Tan, Yuandong
Publié: (2025)
Test-Time Safety Alignment
par: Saglam, Baturay, et autres
Publié: (2026)
par: Saglam, Baturay, et autres
Publié: (2026)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
par: Xu, Yang, et autres
Publié: (2026)
par: Xu, Yang, et autres
Publié: (2026)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
par: Krishna, Kundan, et autres
Publié: (2025)
par: Krishna, Kundan, et autres
Publié: (2025)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
par: Jiang, Jie, et autres
Publié: (2026)
par: Jiang, Jie, et autres
Publié: (2026)
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
par: Cao, Chentao, et autres
Publié: (2025)
par: Cao, Chentao, et autres
Publié: (2025)
Safety-Utility Conflicts Are Not Global: Surgical Alignment via Head-Level Diagnosis
par: Cai, Wang, et autres
Publié: (2026)
par: Cai, Wang, et autres
Publié: (2026)
Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
par: Vega, Jason, et autres
Publié: (2024)
par: Vega, Jason, et autres
Publié: (2024)
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
par: Verma, Richa, et autres
Publié: (2026)
par: Verma, Richa, et autres
Publié: (2026)
Angular Steering: Behavior Control via Rotation in Activation Space
par: Vu, Hieu M., et autres
Publié: (2025)
par: Vu, Hieu M., et autres
Publié: (2025)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
par: Ke, Yekun, et autres
Publié: (2024)
par: Ke, Yekun, et autres
Publié: (2024)
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
par: Hossain, Saad, et autres
Publié: (2025)
par: Hossain, Saad, et autres
Publié: (2025)
Documents similaires
-
Elliptical Attention
par: Nielsen, Stefan K., et autres
Publié: (2024) -
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025) -
Tight Clusters Make Specialized Experts
par: Nielsen, Stefan K., et autres
Publié: (2025) -
Concept Heterogeneity-aware Representation Steering
par: Abdullaev, Laziz U., et autres
Publié: (2026) -
Transformer Meets Twicing: Harnessing Unattended Residual Information
par: Abdullaev, Laziz, et autres
Publié: (2025)