Saved in:
| Main Authors: | Li, Tung-Ling, Liu, Hongliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.24056 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
by: Hu, Hanjiang, et al.
Published: (2025)
by: Hu, Hanjiang, et al.
Published: (2025)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
by: Wang, Linlin, et al.
Published: (2025)
by: Wang, Linlin, et al.
Published: (2025)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
by: Chen, Yen-Shan, et al.
Published: (2025)
by: Chen, Yen-Shan, et al.
Published: (2025)
AirGapAgent: Protecting Privacy-Conscious Conversational Agents
by: Bagdasarian, Eugene, et al.
Published: (2024)
by: Bagdasarian, Eugene, et al.
Published: (2024)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
by: Cao, Bochuan, et al.
Published: (2023)
by: Cao, Bochuan, et al.
Published: (2023)
Improving LLM Safety Alignment with Dual-Objective Optimization
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
by: Zhu, Kaijie, et al.
Published: (2023)
by: Zhu, Kaijie, et al.
Published: (2023)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
by: Zhang, Fangyuan, et al.
Published: (2024)
by: Zhang, Fangyuan, et al.
Published: (2024)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
by: Liang, Jiacheng, et al.
Published: (2024)
by: Liang, Jiacheng, et al.
Published: (2024)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Revisiting the Robustness of Watermarking to Paraphrasing Attacks
by: Rastogi, Saksham, et al.
Published: (2024)
by: Rastogi, Saksham, et al.
Published: (2024)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
Robust Distortion-free Watermarks for Language Models
by: Kuditipudi, Rohith, et al.
Published: (2023)
by: Kuditipudi, Rohith, et al.
Published: (2023)
Certifiably Robust RAG against Retrieval Corruption
by: Xiang, Chong, et al.
Published: (2024)
by: Xiang, Chong, et al.
Published: (2024)
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Promoting Data and Model Privacy in Federated Learning through Quantized LoRA
by: Zhu, JianHao, et al.
Published: (2024)
by: Zhu, JianHao, et al.
Published: (2024)
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models
by: Dang, Cuong, et al.
Published: (2024)
by: Dang, Cuong, et al.
Published: (2024)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
by: Cui, Xinyue, et al.
Published: (2025)
by: Cui, Xinyue, et al.
Published: (2025)
CERT-ED: Certifiably Robust Text Classification for Edit Distance
by: Huang, Zhuoqun, et al.
Published: (2024)
by: Huang, Zhuoqun, et al.
Published: (2024)
Robust LLM safeguarding via refusal feature adversarial training
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs
by: Liu, Hongliang, et al.
Published: (2026)
by: Liu, Hongliang, et al.
Published: (2026)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
by: Li, Jianwei, et al.
Published: (2025)
by: Li, Jianwei, et al.
Published: (2025)
AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
by: Huang, Zhuoqun, et al.
Published: (2025)
by: Huang, Zhuoqun, et al.
Published: (2025)
Enhancing Robustness of AI Offensive Code Generators via Data Augmentation
by: Improta, Cristina, et al.
Published: (2023)
by: Improta, Cristina, et al.
Published: (2023)
Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign
by: Zhang, Ruisi, et al.
Published: (2025)
by: Zhang, Ruisi, et al.
Published: (2025)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
by: Lou, Qian, et al.
Published: (2024)
by: Lou, Qian, et al.
Published: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
by: Liu, Ruixuan, et al.
Published: (2026)
by: Liu, Ruixuan, et al.
Published: (2026)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
Lifelong Safety Alignment for Language Models
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic Writing
by: Liu, Zeyan, et al.
Published: (2023)
by: Liu, Zeyan, et al.
Published: (2023)
Digger: Detecting Copyright Content Mis-usage in Large Language Model Training
by: Li, Haodong, et al.
Published: (2024)
by: Li, Haodong, et al.
Published: (2024)
Exploring Vulnerabilities and Protections in Large Language Models: A Survey
by: Liu, Frank Weizhen, et al.
Published: (2024)
by: Liu, Frank Weizhen, et al.
Published: (2024)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
Directional Embedding Smoothing for Robust Vision Language Models
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Similar Items
-
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
by: Li, Tung-Ling, et al.
Published: (2025) -
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
by: Hu, Hanjiang, et al.
Published: (2025) -
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
by: Wang, Linlin, et al.
Published: (2025) -
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
by: Li, Yuxi, et al.
Published: (2024) -
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
by: Chen, Yen-Shan, et al.
Published: (2025)