Robust LLM safeguarding via refusal feature adversarial training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Lei, Do, Virginie, Hambardzumyan, Karen, Cancedda, Nicola |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM Dataset Inference: Did you train on my dataset?
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
von: Maini, Pratyush, et al.
Veröffentlicht: (2024)
Are aligned neural networks adversarially aligned?
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024)
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
Proving membership in LLM pretraining data via data watermarks
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024)
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
von: Hsiung, Lei, et al.
Veröffentlicht: (2025)
von: Hsiung, Lei, et al.
Veröffentlicht: (2025)
Enhancing Robustness of AI Offensive Code Generators via Data Augmentation
von: Improta, Cristina, et al.
Veröffentlicht: (2023)
von: Improta, Cristina, et al.
Veröffentlicht: (2023)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
A Framework for Cost-Effective and Self-Adaptive LLM Shaking and Recovery Mechanism
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyu, et al.
Veröffentlicht: (2024)
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
von: Zhang, Ruisi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruisi, et al.
Veröffentlicht: (2025)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
von: Park, Seong-Gyu, et al.
Veröffentlicht: (2026)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
von: Li, Nathaniel, et al.
Veröffentlicht: (2024)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!
von: Yoon, Do-hyeon, et al.
Veröffentlicht: (2025)
von: Yoon, Do-hyeon, et al.
Veröffentlicht: (2025)
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
von: Rastogi, Saksham, et al.
Veröffentlicht: (2024)
von: Rastogi, Saksham, et al.
Veröffentlicht: (2024)
LLM Unlearning Should Be Form-Independent
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
von: Ye, Xiaotian, et al.
Veröffentlicht: (2025)
GCG Attack On A Diffusion LLM
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
von: Neyroud, Ruben, et al.
Veröffentlicht: (2025)
Certifiably Robust RAG against Retrieval Corruption
von: Xiang, Chong, et al.
Veröffentlicht: (2024)
von: Xiang, Chong, et al.
Veröffentlicht: (2024)
Robust Distortion-free Watermarks for Language Models
von: Kuditipudi, Rohith, et al.
Veröffentlicht: (2023)
von: Kuditipudi, Rohith, et al.
Veröffentlicht: (2023)
LLMGuard: Guarding Against Unsafe LLM Behavior
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
von: Goyal, Shubh, et al.
Veröffentlicht: (2024)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
von: Assogba, Yannick, et al.
Veröffentlicht: (2026)
Localizing Malicious Outputs from CodeLLM
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
von: Borana, Mayukh, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
von: Krishna, Arjun, et al.
Veröffentlicht: (2025)
PVMark: Enabling Public Verifiability for LLM Watermarking Schemes
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
von: Duan, Haohua, et al.
Veröffentlicht: (2025)
Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
von: Zhang, Boyang, et al.
Veröffentlicht: (2026)
CERT-ED: Certifiably Robust Text Classification for Edit Distance
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2024)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
von: Cui, Xinyue, et al.
Veröffentlicht: (2025)
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
von: Shafee, Samaneh, et al.
Veröffentlicht: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
von: Hu, Hanjiang, et al.
Veröffentlicht: (2025)
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2025)
von: Huang, Zhuoqun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM Dataset Inference: Did you train on my dataset?
von: Maini, Pratyush, et al.
Veröffentlicht: (2024) -
Are aligned neural networks adversarially aligned?
von: Carlini, Nicholas, et al.
Veröffentlicht: (2023) -
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
von: Liang, Jiacheng, et al.
Veröffentlicht: (2024) -
Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens
von: Zhang, Anqi, et al.
Veröffentlicht: (2024) -
Proving membership in LLM pretraining data via data watermarks
von: Wei, Johnny Tian-Zheng, et al.
Veröffentlicht: (2024)