Saved in:
| Main Authors: | Grimes, Keltin, Christiani, Marco, Shriver, David, Connor, Marissa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.13341 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
by: Tran, Khang, et al.
Published: (2026)
by: Tran, Khang, et al.
Published: (2026)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
by: Sinha, Anusha, et al.
Published: (2025)
by: Sinha, Anusha, et al.
Published: (2025)
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
A Systematic Review of Poisoning Attacks Against Large Language Models
by: Fendley, Neil, et al.
Published: (2025)
by: Fendley, Neil, et al.
Published: (2025)
Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
by: Zhang, Boyang, et al.
Published: (2025)
by: Zhang, Boyang, et al.
Published: (2025)
Learning to Poison Large Language Models for Downstream Manipulation
by: Zhou, Xiangyu, et al.
Published: (2024)
by: Zhou, Xiangyu, et al.
Published: (2024)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
by: Zou, Wei, et al.
Published: (2024)
by: Zou, Wei, et al.
Published: (2024)
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
by: Li, Yuxi, et al.
Published: (2024)
by: Li, Yuxi, et al.
Published: (2024)
Concept Drift Detection using Ensemble of Integrally Private Models
by: Varshney, Ayush K., et al.
Published: (2024)
by: Varshney, Ayush K., et al.
Published: (2024)
Adaptive and Robust Data Poisoning Detection and Sanitization in Wearable IoT Systems using Large Language Models
by: Mithsara, W. K. M, et al.
Published: (2025)
by: Mithsara, W. K. M, et al.
Published: (2025)
Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts
by: Zarlenga, Mateo Espinosa, et al.
Published: (2025)
by: Zarlenga, Mateo Espinosa, et al.
Published: (2025)
TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
by: Aghakhani, Hojjat, et al.
Published: (2023)
by: Aghakhani, Hojjat, et al.
Published: (2023)
The 'Sure' Trap: Multi-Scale Poisoning Analysis of Stealthy Compliance-Only Backdoors in Fine-Tuned Large Language Models
by: Tan, Yuting, et al.
Published: (2025)
by: Tan, Yuting, et al.
Published: (2025)
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2025)
by: Bouaziz, Wassim, et al.
Published: (2025)
Understanding Concept Drift with Deprecated Permissions in Android Malware Detection
by: Sabbah, Ahmed, et al.
Published: (2025)
by: Sabbah, Ahmed, et al.
Published: (2025)
Awakening the Hydra: Stabilizing Multi-Concept Backdoor Injection in Text-to-Image Diffusion Models
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
Poisoned Source Code Detection in Code Models
by: Ghannoum, Ehab, et al.
Published: (2025)
by: Ghannoum, Ehab, et al.
Published: (2025)
Thwarting Cybersecurity Attacks with Explainable Concept Drift
by: Shaer, Ibrahim, et al.
Published: (2024)
by: Shaer, Ibrahim, et al.
Published: (2024)
Cluster Analysis and Concept Drift Detection in Malware
by: Mishra, Aniket, et al.
Published: (2025)
by: Mishra, Aniket, et al.
Published: (2025)
Adaptive Domain Inference Attack with Concept Hierarchy
by: Gu, Yuechun, et al.
Published: (2023)
by: Gu, Yuechun, et al.
Published: (2023)
How to Defend Against Large-scale Model Poisoning Attacks in Federated Learning: A Vertical Solution
by: Wang, Jinbo, et al.
Published: (2024)
by: Wang, Jinbo, et al.
Published: (2024)
Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models
by: Lai, Songning, et al.
Published: (2024)
by: Lai, Songning, et al.
Published: (2024)
Prompt Obfuscation for Large Language Models
by: Pape, David, et al.
Published: (2024)
by: Pape, David, et al.
Published: (2024)
FedRecAttack: Model Poisoning Attack to Federated Recommendation
by: Rong, Dazhong, et al.
Published: (2022)
by: Rong, Dazhong, et al.
Published: (2022)
Adversarial Update-Based Federated Unlearning for Poisoned Model Recovery
by: Zhao, Wenwei, et al.
Published: (2026)
by: Zhao, Wenwei, et al.
Published: (2026)
Optimized Deep Learning Models for Malware Detection under Concept Drift
by: Maillet, William, et al.
Published: (2023)
by: Maillet, William, et al.
Published: (2023)
State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
by: Guo, Ji, et al.
Published: (2026)
by: Guo, Ji, et al.
Published: (2026)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
JULI: Jailbreak Large Language Models by Self-Introspection
by: Wang, Jesson, et al.
Published: (2025)
by: Wang, Jesson, et al.
Published: (2025)
DRMD: Deep Reinforcement Learning for Malware Detection under Concept Drift
by: McFadden, Shae, et al.
Published: (2025)
by: McFadden, Shae, et al.
Published: (2025)
LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift Analysis
by: Haque, Md Ahsanul, et al.
Published: (2025)
by: Haque, Md Ahsanul, et al.
Published: (2025)
Binary Anomaly Detection in Streaming IoT Traffic under Concept Drift
by: Carnier, Rodrigo Matos, et al.
Published: (2025)
by: Carnier, Rodrigo Matos, et al.
Published: (2025)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Identifying Predictions That Influence the Future: Detecting Performative Concept Drift in Data Streams
by: Gower-Winter, Brandon, et al.
Published: (2024)
by: Gower-Winter, Brandon, et al.
Published: (2024)
MADCAT: Combating Malware Detection Under Concept Drift with Test-Time Adaptation
by: Roh, Eunjin, et al.
Published: (2025)
by: Roh, Eunjin, et al.
Published: (2025)
ADAPT: A Pseudo-labeling Approach to Combat Concept Drift in Malware Detection
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
FARM: Few-shot Adaptive Malware Family Classification under Concept Drift
by: Guldemir, Numan Halit, et al.
Published: (2026)
by: Guldemir, Numan Halit, et al.
Published: (2026)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Similar Items
-
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
by: Tran, Khang, et al.
Published: (2026) -
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
by: Sinha, Anusha, et al.
Published: (2025) -
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025) -
A Systematic Review of Poisoning Attacks Against Large Language Models
by: Fendley, Neil, et al.
Published: (2025) -
Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
by: Zhang, Boyang, et al.
Published: (2025)