Test-Time Safety Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Saglam, Baturay, Kalogerias, Dionysis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Time Detoxification without Training or Learning Anything
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Compatible Gradient Approximations for Actor-Critic Algorithms
by: Saglam, Baturay, et al.
Published: (2024)
by: Saglam, Baturay, et al.
Published: (2024)
Learning Task Representations from In-Context Learning
by: Saglam, Baturay, et al.
Published: (2025)
by: Saglam, Baturay, et al.
Published: (2025)
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
by: Lee, Jane H., et al.
Published: (2025)
by: Lee, Jane H., et al.
Published: (2025)
Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
by: Vijay, Supriti, et al.
Published: (2025)
by: Vijay, Supriti, et al.
Published: (2025)
FEDSTR: Money-In AI-Out | A Decentralized Marketplace for Federated Learning and LLM Training on the NOSTR Protocol
by: Nikolakakis, Konstantinos E., et al.
Published: (2024)
by: Nikolakakis, Konstantinos E., et al.
Published: (2024)
Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
by: Saglam, Baturay, et al.
Published: (2025)
by: Saglam, Baturay, et al.
Published: (2025)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
by: Krishna, Kundan, et al.
Published: (2025)
by: Krishna, Kundan, et al.
Published: (2025)
Multilingual Safety Alignment via Self-Distillation
by: Qin, Ruiyang, et al.
Published: (2026)
by: Qin, Ruiyang, et al.
Published: (2026)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
by: Chen, Jianhui, et al.
Published: (2024)
by: Chen, Jianhui, et al.
Published: (2024)
Course-Correction: Safety Alignment Using Synthetic Preferences
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
FedAVOT: Exact Distribution Alignment in Federated Learning via Masked Optimal Transport
by: Herlock, et al.
Published: (2025)
by: Herlock, et al.
Published: (2025)
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
by: Liu, Qin, et al.
Published: (2024)
by: Liu, Qin, et al.
Published: (2024)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
by: Yang, Junxiao, et al.
Published: (2026)
by: Yang, Junxiao, et al.
Published: (2026)
MESA: Improving MoE Safety Alignment via Decentralized Expertise
by: Sun, Yitong, et al.
Published: (2026)
by: Sun, Yitong, et al.
Published: (2026)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
by: Zhou, Zhanhui, et al.
Published: (2024)
by: Zhou, Zhanhui, et al.
Published: (2024)
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models
by: Kazemi, Hamid, et al.
Published: (2026)
by: Kazemi, Hamid, et al.
Published: (2026)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
by: Wang, Dianyun, et al.
Published: (2025)
by: Wang, Dianyun, et al.
Published: (2025)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
by: Khattar, Vanshaj, et al.
Published: (2026)
by: Khattar, Vanshaj, et al.
Published: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
by: Giordani, Jeremiah
Published: (2025)
by: Giordani, Jeremiah
Published: (2025)
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
by: Li, Guanlin, et al.
Published: (2025)
by: Li, Guanlin, et al.
Published: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
by: Pan, Birong, et al.
Published: (2025)
by: Pan, Birong, et al.
Published: (2025)
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
by: Cai, Yang, et al.
Published: (2026)
by: Cai, Yang, et al.
Published: (2026)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
In-Place Test-Time Training
by: Feng, Guhao, et al.
Published: (2026)
by: Feng, Guhao, et al.
Published: (2026)
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
by: Bobbili, Sarat Chandra, et al.
Published: (2025)
by: Bobbili, Sarat Chandra, et al.
Published: (2025)
It's Not That Simple. An Analysis of Simple Test-Time Scaling
by: Wu, Guojun
Published: (2025)
by: Wu, Guojun
Published: (2025)
Self-Improving LLM Agents at Test-Time
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
Similar Items
-
Test-Time Detoxification without Training or Learning Anything
by: Saglam, Baturay, et al.
Published: (2026) -
Compatible Gradient Approximations for Actor-Critic Algorithms
by: Saglam, Baturay, et al.
Published: (2024) -
Learning Task Representations from In-Context Learning
by: Saglam, Baturay, et al.
Published: (2025) -
Risk-Averse Constrained Reinforcement Learning with Optimized Certainty Equivalents
by: Lee, Jane H., et al.
Published: (2025) -
Think Before You Retrieve: Learning Test-Time Adaptive Search with Small Language Models
by: Vijay, Supriti, et al.
Published: (2025)