Safety Alignment of LMs via Non-cooperative Games
Fuente:
arXiv
Saved in:
| Main Authors: | Paulus, Anselm, Kulikov, Ilia, Amos, Brandon, Munos, Rémi, Evtimov, Ivan, Chaudhuri, Kamalika, Zharmagambetov, Arman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025)
by: Evtimov, Ivan, et al.
Published: (2025)
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
by: Zharmagambetov, Arman, et al.
Published: (2025)
by: Zharmagambetov, Arman, et al.
Published: (2025)
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
by: Paulus, Anselm, et al.
Published: (2024)
by: Paulus, Anselm, et al.
Published: (2024)
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
by: Tomani, Christian, et al.
Published: (2024)
by: Tomani, Christian, et al.
Published: (2024)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
by: Zhu, Sicheng, et al.
Published: (2024)
by: Zhu, Sicheng, et al.
Published: (2024)
Super-Exponential Regret for UCT, AlphaGo and Variants
by: Orseau, Laurent, et al.
Published: (2024)
by: Orseau, Laurent, et al.
Published: (2024)
Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment
by: Peng, Fred Zhangzhi, et al.
Published: (2026)
by: Peng, Fred Zhangzhi, et al.
Published: (2026)
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
by: Chen, Sizhe, et al.
Published: (2025)
by: Chen, Sizhe, et al.
Published: (2025)
Influence-based Attributions can be Manipulated
by: Yadav, Chhavi, et al.
Published: (2024)
by: Yadav, Chhavi, et al.
Published: (2024)
On amortizing convex conjugates for optimal transport
by: Amos, Brandon
Published: (2022)
by: Amos, Brandon
Published: (2022)
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
by: Kirichenko, Polina, et al.
Published: (2025)
by: Kirichenko, Polina, et al.
Published: (2025)
Positional Encoding via Token-Aware Phase Attention
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Tutorial on amortized optimization
by: Amos, Brandon
Published: (2022)
by: Amos, Brandon
Published: (2022)
Distilling LLM Feedback for Lean Theorem Proving
by: Narozniak, Gaetan, et al.
Published: (2026)
by: Narozniak, Gaetan, et al.
Published: (2026)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
by: Koga, Tatsuki, et al.
Published: (2024)
by: Koga, Tatsuki, et al.
Published: (2024)
Distilling System 2 into System 1
by: Yu, Ping, et al.
Published: (2024)
by: Yu, Ping, et al.
Published: (2024)
SecAlign: Defending Against Prompt Injection with Preference Optimization
by: Chen, Sizhe, et al.
Published: (2024)
by: Chen, Sizhe, et al.
Published: (2024)
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
by: Yadav, Chhavi, et al.
Published: (2025)
by: Yadav, Chhavi, et al.
Published: (2025)
FairProof : Confidential and Certifiable Fairness for Neural Networks
by: Yadav, Chhavi, et al.
Published: (2024)
by: Yadav, Chhavi, et al.
Published: (2024)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Gradient-based Jailbreak Images for Multimodal Fusion Models
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
by: Dziemian, Mateusz, et al.
Published: (2026)
by: Dziemian, Mateusz, et al.
Published: (2026)
Generalized Preference Optimization: A Unified Approach to Offline Alignment
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Cross-Domain Imitation Learning via Optimal Transport
by: Fickinger, Arnaud, et al.
Published: (2021)
by: Fickinger, Arnaud, et al.
Published: (2021)
Automatic Textbook Formalization
by: Gloeckle, Fabian, et al.
Published: (2026)
by: Gloeckle, Fabian, et al.
Published: (2026)
Temporal Difference Flows
by: Farebrother, Jesse, et al.
Published: (2025)
by: Farebrother, Jesse, et al.
Published: (2025)
BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
by: Lee, Sangyun, et al.
Published: (2025)
by: Lee, Sangyun, et al.
Published: (2025)
Bridging Bots: from Perception to Action via Multimodal-LMs and Knowledge Graphs
by: Martorana, Margherita, et al.
Published: (2025)
by: Martorana, Margherita, et al.
Published: (2025)
Safety-Preserving PTQ via Contrastive Alignment Loss
by: Wee, Sunghyun, et al.
Published: (2025)
by: Wee, Sunghyun, et al.
Published: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
by: Mireshghallah, Niloofar, et al.
Published: (2025)
by: Mireshghallah, Niloofar, et al.
Published: (2025)
Formalizing Mathematics at Scale
by: Rammal, Ahmad, et al.
Published: (2026)
by: Rammal, Ahmad, et al.
Published: (2026)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
by: Whitehouse, Chenxi, et al.
Published: (2025)
by: Whitehouse, Chenxi, et al.
Published: (2025)
Improving Safety Alignment via Balanced Direct Preference Optimization
by: Zhao, Shiji, et al.
Published: (2026)
by: Zhao, Shiji, et al.
Published: (2026)
Customizing Language Model Responses with Contrastive In-Context Learning
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
References Improve LLM Alignment in Non-Verifiable Domains
by: Shi, Kejian, et al.
Published: (2026)
by: Shi, Kejian, et al.
Published: (2026)
Agent Safety Alignment via Reinforcement Learning
by: Sha, Zeyang, et al.
Published: (2025)
by: Sha, Zeyang, et al.
Published: (2025)
Safety Alignment via Constrained Knowledge Unlearning
by: Shi, Zesheng, et al.
Published: (2025)
by: Shi, Zesheng, et al.
Published: (2025)
One-shot Optimized Steering Vectors Mediate Safety-relevant Behaviors in LLMs
by: Dunefsky, Jacob, et al.
Published: (2025)
by: Dunefsky, Jacob, et al.
Published: (2025)
Similar Items
-
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025) -
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
by: Zharmagambetov, Arman, et al.
Published: (2025) -
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
by: Paulus, Anselm, et al.
Published: (2024) -
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
by: Tomani, Christian, et al.
Published: (2024) -
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
by: Wen, Yuxin, et al.
Published: (2025)