Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Zi, Hu, Haibo, Ye, Qingqing, Xiao, Yaxin, Li, Ronghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection
by: Liang, Zi, et al.
Published: (2026)
by: Liang, Zi, et al.
Published: (2026)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
by: Liang, Zi, et al.
Published: (2025)
by: Liang, Zi, et al.
Published: (2025)
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
by: Liang, Zi, et al.
Published: (2024)
by: Liang, Zi, et al.
Published: (2024)
FIT to Forget: Robust Continual Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation
by: Liang, Zi, et al.
Published: (2024)
by: Liang, Zi, et al.
Published: (2024)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
by: Jiang, Tanqiu, et al.
Published: (2024)
by: Jiang, Tanqiu, et al.
Published: (2024)
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
by: Liang, Zi, et al.
Published: (2025)
by: Liang, Zi, et al.
Published: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
by: Luo, Weidi, et al.
Published: (2024)
by: Luo, Weidi, et al.
Published: (2024)
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
by: Yu, Ye, et al.
Published: (2026)
by: Yu, Ye, et al.
Published: (2026)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Decoupled Alignment for Robust Plug-and-Play Adaptation
by: Luo, Haozheng, et al.
Published: (2024)
by: Luo, Haozheng, et al.
Published: (2024)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
by: Xu, Xiaoyu, et al.
Published: (2025)
by: Xu, Xiaoyu, et al.
Published: (2025)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
by: Wei, Jiali, et al.
Published: (2026)
by: Wei, Jiali, et al.
Published: (2026)
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)
by: Xu, Feiyue, et al.
Published: (2026)
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
by: Xu, Naen, et al.
Published: (2026)
by: Xu, Naen, et al.
Published: (2026)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
by: Zhang, Jinchuan, et al.
Published: (2024)
by: Zhang, Jinchuan, et al.
Published: (2024)
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
SOLIDO: A Robust Watermarking Method for Speech Synthesis via Low-Rank Adaptation
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
SEAL: Entangled White-box Watermarks on Low-Rank Adaptation
by: Oh, Giyeong, et al.
Published: (2025)
by: Oh, Giyeong, et al.
Published: (2025)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
by: Li, Xiao, et al.
Published: (2024)
by: Li, Xiao, et al.
Published: (2024)
Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents
by: Liang, Zhibo, et al.
Published: (2025)
by: Liang, Zhibo, et al.
Published: (2025)
Differentially Private Federated Low Rank Adaptation Beyond Fixed-Matrix
by: Wen, Ming, et al.
Published: (2025)
by: Wen, Ming, et al.
Published: (2025)
Distract Large Language Models for Automatic Jailbreak Attack
by: Xiao, Zeguan, et al.
Published: (2024)
by: Xiao, Zeguan, et al.
Published: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
FedCC: Robust Federated Learning against Model Poisoning Attacks
by: Jeong, Hyejun, et al.
Published: (2022)
by: Jeong, Hyejun, et al.
Published: (2022)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
by: Yan, Dong, et al.
Published: (2026)
by: Yan, Dong, et al.
Published: (2026)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
by: Hu, Xiaomeng, et al.
Published: (2025)
by: Hu, Xiaomeng, et al.
Published: (2025)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
by: Pu, Rui, et al.
Published: (2024)
by: Pu, Rui, et al.
Published: (2024)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
by: Zhu, Xiaoyuan, et al.
Published: (2025)
by: Zhu, Xiaoyuan, et al.
Published: (2025)
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
by: Zhao, Jin, et al.
Published: (2026)
by: Zhao, Jin, et al.
Published: (2026)
MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
Does Few-shot Learning Suffer from Backdoor Attacks?
by: Liu, Xinwei, et al.
Published: (2023)
by: Liu, Xinwei, et al.
Published: (2023)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
HauntAttack: When Attack Follows Reasoning as a Shadow
by: Ma, Jingyuan, et al.
Published: (2025)
by: Ma, Jingyuan, et al.
Published: (2025)
Interactive Trimming against Evasive Online Data Manipulation Attacks: A Game-Theoretic Approach
by: Fu, Yue, et al.
Published: (2024)
by: Fu, Yue, et al.
Published: (2024)
DiscourseFlip: An Oblique Discourse-Level Opinion Manipulation Attack against Black-box Retrieval-Augmented Generation
by: Gong, Yuyang, et al.
Published: (2026)
by: Gong, Yuyang, et al.
Published: (2026)
Similar Items
-
Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection
by: Liang, Zi, et al.
Published: (2026) -
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
by: Liang, Zi, et al.
Published: (2025) -
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
by: Liang, Zi, et al.
Published: (2024) -
FIT to Forget: Robust Continual Unlearning for Large Language Models
by: Xu, Xiaoyu, et al.
Published: (2026) -
"Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced Distillation
by: Liang, Zi, et al.
Published: (2024)