Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yugeng, Cong, Tianshuo, Zhao, Zhengyu, Backes, Michael, Shen, Yun, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Watermarking LLM-Generated Datasets in Downstream Tasks
by: Liu, Yugeng, et al.
Published: (2025)
by: Liu, Yugeng, et al.
Published: (2025)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Amplifying Machine Learning Attacks Through Strategic Compositions
by: Liu, Yugeng, et al.
Published: (2025)
by: Liu, Yugeng, et al.
Published: (2025)
Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples
by: Han, Sicong, et al.
Published: (2025)
by: Han, Sicong, et al.
Published: (2025)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
The Challenge of Identifying the Origin of Black-Box Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Peering Behind the Shield: Guardrail Identification in Large Language Models
by: Yang, Ziqing, et al.
Published: (2025)
by: Yang, Ziqing, et al.
Published: (2025)
IrisFP: Adversarial-Example-based Model Fingerprinting with Enhanced Uniqueness and Robustness
by: Geng, Ziye, et al.
Published: (2026)
by: Geng, Ziye, et al.
Published: (2026)
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
by: Jiang, Yukun, et al.
Published: (2025)
by: Jiang, Yukun, et al.
Published: (2025)
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
by: Wu, Yixin, et al.
Published: (2023)
by: Wu, Yixin, et al.
Published: (2023)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
Effectiveness of Adversarial Benign and Malware Examples in Evasion and Poisoning Attacks
by: Kozák, Matouš, et al.
Published: (2025)
by: Kozák, Matouš, et al.
Published: (2025)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
Adversarial Example Based Fingerprinting for Robust Copyright Protection in Split Learning
by: Lin, Zhangting, et al.
Published: (2025)
by: Lin, Zhangting, et al.
Published: (2025)
Quantization Aware Attack: Enhancing Transferable Adversarial Attacks by Model Quantization
by: Yang, Yulong, et al.
Published: (2023)
by: Yang, Yulong, et al.
Published: (2023)
Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path
by: Ren, Yuchen, et al.
Published: (2024)
by: Ren, Yuchen, et al.
Published: (2024)
Adversarially Robust Assembly Language Model for Packed Executables Detection
by: Li, Shijia, et al.
Published: (2025)
by: Li, Shijia, et al.
Published: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
Have You Merged My Model? On The Robustness of Large Language Model IP Protection Methods Against Model Merging
by: Cong, Tianshuo, et al.
Published: (2024)
by: Cong, Tianshuo, et al.
Published: (2024)
A Survey on Adversarial Machine Learning for Code Data: Realistic Threats, Countermeasures, and Interpretations
by: Yang, Yulong, et al.
Published: (2024)
by: Yang, Yulong, et al.
Published: (2024)
Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data
by: Song, Tianle, et al.
Published: (2025)
by: Song, Tianle, et al.
Published: (2025)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Efficient Data-Free Model Stealing with Label Diversity
by: Liu, Yiyong, et al.
Published: (2024)
by: Liu, Yiyong, et al.
Published: (2024)
SoK: Understanding Vulnerabilities in the Large Language Model Supply Chain
by: Wang, Shenao, et al.
Published: (2025)
by: Wang, Shenao, et al.
Published: (2025)
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
by: Qu, Yiting, et al.
Published: (2025)
by: Qu, Yiting, et al.
Published: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
by: Guan, Shaowei, et al.
Published: (2025)
by: Guan, Shaowei, et al.
Published: (2025)
NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples
by: Pu, Shi, et al.
Published: (2025)
by: Pu, Shi, et al.
Published: (2025)
Improving Sustainability of Adversarial Examples in Class-Incremental Learning
by: Liu, Taifeng, et al.
Published: (2025)
by: Liu, Taifeng, et al.
Published: (2025)
Masked Language Model Based Textual Adversarial Example Detection
by: Zhang, Xiaomei, et al.
Published: (2023)
by: Zhang, Xiaomei, et al.
Published: (2023)
Matrix Kloosterman Sums, Random Matrix Statistics, and Cryptography
by: Yang, Tianshuo
Published: (2026)
by: Yang, Tianshuo
Published: (2026)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
by: Yang, Ziqing, et al.
Published: (2024)
by: Yang, Ziqing, et al.
Published: (2024)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
by: Wen, Rui, et al.
Published: (2024)
by: Wen, Rui, et al.
Published: (2024)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
by: Liu, Zesen, et al.
Published: (2024)
by: Liu, Zesen, et al.
Published: (2024)
Improving Adversarial Robustness in Android Malware Detection by Reducing the Impact of Spurious Correlations
by: Bostani, Hamid, et al.
Published: (2024)
by: Bostani, Hamid, et al.
Published: (2024)
Prompt Stealing Attacks Against Text-to-Image Generation Models
by: Shen, Xinyue, et al.
Published: (2023)
by: Shen, Xinyue, et al.
Published: (2023)
On the Adversarial Robustness of Instruction-Tuned Large Language Models for Code
by: Hossen, Md Imran, et al.
Published: (2024)
by: Hossen, Md Imran, et al.
Published: (2024)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
by: Ran, Delong, et al.
Published: (2024)
by: Ran, Delong, et al.
Published: (2024)
Similar Items
-
Watermarking LLM-Generated Datasets in Downstream Tasks
by: Liu, Yugeng, et al.
Published: (2025) -
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
by: Jiang, Yukun, et al.
Published: (2024) -
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023) -
Amplifying Machine Learning Attacks Through Strategic Compositions
by: Liu, Yugeng, et al.
Published: (2025) -
Prediction Inconsistency Helps Achieve Generalizable Detection of Adversarial Examples
by: Han, Sicong, et al.
Published: (2025)