AI Risk Management Should Incorporate Both Safety and Security
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Xiangyu, Huang, Yangsibo, Zeng, Yi, Debenedetti, Edoardo, Geiping, Jonas, He, Luxi, Huang, Kaixuan, Madhushani, Udari, Sehwag, Vikash, Shi, Weijia, Wei, Boyi, Xie, Tinghao, Chen, Danqi, Chen, Pin-Yu, Ding, Jeffrey, Jia, Ruoxi, Ma, Jiaqi, Narayanan, Arvind, Su, Weijie J, Wang, Mengdi, Xiao, Chaowei, Li, Bo, Song, Dawn, Henderson, Peter, Mittal, Prateek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
by: Xie, Tinghao, et al.
Published: (2024)
by: Xie, Tinghao, et al.
Published: (2024)
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
In-Context Learning with Topological Information for Knowledge Graph Completion
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
Fantastic Copyrighted Beasts and How (Not) to Generate Them
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
by: He, Luxi, et al.
Published: (2025)
by: He, Luxi, et al.
Published: (2025)
Adapting to Evolving Adversaries with Regularized Continual Robust Training
by: Dai, Sihui, et al.
Published: (2025)
by: Dai, Sihui, et al.
Published: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
Can LLMs be Scammed? A Baseline Measurement Study
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
Data Shapley in One Training Run
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
by: Cook, Thomas, et al.
Published: (2025)
by: Cook, Thomas, et al.
Published: (2025)
Does More Inference-Time Compute Really Help Robustness?
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Capturing the Temporal Dependence of Training Data Influence
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
Efficient Data Shapley for Weighted Nearest Neighbor Algorithms
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
by: Karthikeyan, Harish, et al.
Published: (2025)
by: Karthikeyan, Harish, et al.
Published: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
Parrondo's paradox in quantum walks with inhomogeneous coins
by: Mittal, Vikash, et al.
Published: (2024)
by: Mittal, Vikash, et al.
Published: (2024)
Quantum Magic in Discrete-Time Quantum Walk
by: Mittal, Vikash, et al.
Published: (2025)
by: Mittal, Vikash, et al.
Published: (2025)
Detecting Pretraining Data from Large Language Models
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
Averaging quadratically twisted $L$-values and their derivatives
by: Huang, Tinghao
Published: (2025)
by: Huang, Tinghao
Published: (2025)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
by: Wang, Jiachen T., et al.
Published: (2025)
by: Wang, Jiachen T., et al.
Published: (2025)
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026)
by: Pu, George, et al.
Published: (2026)
Evaluating Copyright Takedown Methods for Language Models
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
by: Huang, Kaixuan, et al.
Published: (2024)
by: Huang, Kaixuan, et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
Discrete time quantum walk of locally interacting walkers
by: Mittal, Vikash, et al.
Published: (2025)
by: Mittal, Vikash, et al.
Published: (2025)
Certifiably Robust RAG against Retrieval Corruption
by: Xiang, Chong, et al.
Published: (2024)
by: Xiang, Chong, et al.
Published: (2024)
A Heterogeneous Agent Model of Mortgage Servicing: An Income-based Relief Analysis
by: Garg, Deepeka, et al.
Published: (2024)
by: Garg, Deepeka, et al.
Published: (2024)
On Ramanujan Primes for Hecke-Maass Cusp Forms
by: Huang, Tinghao, et al.
Published: (2026)
by: Huang, Tinghao, et al.
Published: (2026)
A Theoretical Perspective for Speculative Decoding Algorithm
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI
by: Cheng, Siyuan, et al.
Published: (2025)
by: Cheng, Siyuan, et al.
Published: (2025)
Evaluating and Mitigating IP Infringement in Visual Generative AI
by: Wang, Zhenting, et al.
Published: (2024)
by: Wang, Zhenting, et al.
Published: (2024)
Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget
by: Sehwag, Vikash, et al.
Published: (2024)
by: Sehwag, Vikash, et al.
Published: (2024)
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
by: Xiao, Yuchen, et al.
Published: (2023)
by: Xiao, Yuchen, et al.
Published: (2023)
Similar Items
-
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
by: Xie, Tinghao, et al.
Published: (2024) -
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024) -
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025) -
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024) -
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)