Saffron-1: Safety Inference Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Ruizhong, Li, Gaotang, Wei, Tianxin, He, Jingrui, Tong, Hanghang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
Graph homophily booster: Reimagining the role of discrete features in heterophilic graph learning
by: Qiu, Ruizhong, et al.
Published: (2026)
by: Qiu, Ruizhong, et al.
Published: (2026)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024)
by: Liu, Yule, et al.
Published: (2024)
Membership Inference Attack with Partial Features
by: Wang, Xurun, et al.
Published: (2025)
by: Wang, Xurun, et al.
Published: (2025)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
Privacy Auditing of Multi-domain Graph Pre-trained Model under Membership Inference Attacks
by: Luo, Jiayi, et al.
Published: (2025)
by: Luo, Jiayi, et al.
Published: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Dissecting Distribution Inference
by: Suri, Anshuman, et al.
Published: (2022)
by: Suri, Anshuman, et al.
Published: (2022)
Similarity-based Label Inference Attack against Training and Inference of Split Learning
by: Liu, Junlin, et al.
Published: (2022)
by: Liu, Junlin, et al.
Published: (2022)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
by: Liu, Guozhi, et al.
Published: (2025)
by: Liu, Guozhi, et al.
Published: (2025)
Analyzing Inference Privacy Risks Through Gradients in Machine Learning
by: Li, Zhuohang, et al.
Published: (2024)
by: Li, Zhuohang, et al.
Published: (2024)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision
by: Mukherjee, Manisha, et al.
Published: (2026)
by: Mukherjee, Manisha, et al.
Published: (2026)
What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
by: Chang, Jiamin, et al.
Published: (2025)
by: Chang, Jiamin, et al.
Published: (2025)
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Breaking Silos: Adaptive Model Fusion Unlocks Better Time Series Forecasting
by: Liu, Zhining, et al.
Published: (2025)
by: Liu, Zhining, et al.
Published: (2025)
On Membership Inference Attacks in Knowledge Distillation
by: Cui, Ziyao, et al.
Published: (2025)
by: Cui, Ziyao, et al.
Published: (2025)
Feature Inference Attack on Shapley Values
by: Luo, Xinjian, et al.
Published: (2024)
by: Luo, Xinjian, et al.
Published: (2024)
DCMI: A Differential Calibration Membership Inference Attack Against Retrieval-Augmented Generation
by: Gao, Xinyu, et al.
Published: (2025)
by: Gao, Xinyu, et al.
Published: (2025)
Why Safety Probes Catch Liars But Miss Fanatics
by: Haralambiev, Kristiyan
Published: (2026)
by: Haralambiev, Kristiyan
Published: (2026)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Fair Finetuning Mitigates Distribution Inference Attacks
by: Naidu, Rakshit
Published: (2026)
by: Naidu, Rakshit
Published: (2026)
Attribute Inference Attacks for Federated Regression Tasks
by: Diana, Francesco, et al.
Published: (2024)
by: Diana, Francesco, et al.
Published: (2024)
A Safety and Security Framework for Real-World Agentic Systems
by: Ghosh, Shaona, et al.
Published: (2025)
by: Ghosh, Shaona, et al.
Published: (2025)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
by: Lai, Zhenglin, et al.
Published: (2025)
by: Lai, Zhenglin, et al.
Published: (2025)
RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models
by: Liang, Jiacheng, et al.
Published: (2026)
by: Liang, Jiacheng, et al.
Published: (2026)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
by: Rozenfeld, Shir, et al.
Published: (2026)
by: Rozenfeld, Shir, et al.
Published: (2026)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
Linearizing Models for Efficient yet Robust Private Inference
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
by: Lin, Shuyi, et al.
Published: (2025)
by: Lin, Shuyi, et al.
Published: (2025)
Similar Items
-
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
by: Lin, Xiao, et al.
Published: (2024) -
Graph homophily booster: Reimagining the role of discrete features in heterophilic graph learning
by: Qiu, Ruizhong, et al.
Published: (2026) -
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
by: Ghosal, Soumya Suvra, et al.
Published: (2024) -
Quantized Delta Weight Is Safety Keeper
by: Liu, Yule, et al.
Published: (2024) -
Membership Inference Attack with Partial Features
by: Wang, Xurun, et al.
Published: (2025)