Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jingyu, Elgohary, Ahmed, Magooda, Ahmed, Khashabi, Daniel, Van Durme, Benjamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jailbreak Distillation: Renewable Safety Benchmarking
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
Many-Tier Instruction Hierarchy in LLM Agents
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
von: Weller, Orion, et al.
Veröffentlicht: (2023)
von: Weller, Orion, et al.
Veröffentlicht: (2023)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
von: Shen, Lingfeng, et al.
Veröffentlicht: (2024)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2024)
RE-Adapt: Reverse Engineered Adaptation of Large Language Models
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
SafeWorld: Geo-Diverse Safety Alignment
von: Yin, Da, et al.
Veröffentlicht: (2024)
von: Yin, Da, et al.
Veröffentlicht: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
Test-Time Safety Alignment
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
von: Krishna, Kundan, et al.
Veröffentlicht: (2025)
von: Krishna, Kundan, et al.
Veröffentlicht: (2025)
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
von: Brito, Iago Alves, et al.
Veröffentlicht: (2026)
von: Brito, Iago Alves, et al.
Veröffentlicht: (2026)
SEQR: Secure and Efficient QR-based LoRA Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
LoRA-Augmented Generation (LAG) for Knowledge-Intensive Language Tasks
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
SpectR: Dynamically Composing LM Experts with Spectral Routing
von: Fleshman, William, et al.
Veröffentlicht: (2025)
von: Fleshman, William, et al.
Veröffentlicht: (2025)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
von: Xu, Kevin, et al.
Veröffentlicht: (2024)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
Certified Mitigation of Worst-Case LLM Copyright Infringement
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
LM Agents for Coordinating Multi-User Information Gathering
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
von: Jhamtani, Harsh, et al.
Veröffentlicht: (2025)
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2025)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
von: Alssum, Lama, et al.
Veröffentlicht: (2025)
STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules
von: Wu, Di, et al.
Veröffentlicht: (2026)
von: Wu, Di, et al.
Veröffentlicht: (2026)
Language Models and Logic Programs for Trustworthy Tax Reasoning
von: Jurayj, William, et al.
Veröffentlicht: (2025)
von: Jurayj, William, et al.
Veröffentlicht: (2025)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
von: Fleshman, William, et al.
Veröffentlicht: (2024)
von: Fleshman, William, et al.
Veröffentlicht: (2024)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
von: Wang, Dianyun, et al.
Veröffentlicht: (2025)
von: Wang, Dianyun, et al.
Veröffentlicht: (2025)
Cat-DPO: Category-Adaptive Safety Alignment
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
Safety Alignment via Constrained Knowledge Unlearning
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
What Matters For Safety Alignment?
von: Li, Xing, et al.
Veröffentlicht: (2026)
von: Li, Xing, et al.
Veröffentlicht: (2026)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
Reframing Tax Law Entailment as Analogical Reasoning
von: Zou, Xinrui, et al.
Veröffentlicht: (2024)
von: Zou, Xinrui, et al.
Veröffentlicht: (2024)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
Towards Context-Invariant Safety Alignment for Large Language Models
von: Wang, Yixu, et al.
Veröffentlicht: (2026)
von: Wang, Yixu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Jailbreak Distillation: Renewable Safety Benchmarking
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025) -
Many-Tier Instruction Hierarchy in LLM Agents
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026) -
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026) -
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024) -
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)