Cascading Adversarial Bias from Injection to Distillation in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chaudhari, Harsh, Hayes, Jamie, Jagielski, Matthew, Shumailov, Ilia, Nasr, Milad, Oprea, Alina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024)
by: Foerster, Hanna, et al.
Published: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025)
by: Suri, Anshuman, et al.
Published: (2025)
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Soft Instruction De-escalation Defense
by: Walter, Nils Philipp, et al.
Published: (2025)
by: Walter, Nils Philipp, et al.
Published: (2025)
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
UTrace: Poisoning Forensics for Private Collaborative Learning
by: Rose, Evan, et al.
Published: (2024)
by: Rose, Evan, et al.
Published: (2024)
Black-Box Privacy Attacks on Shared Representations in Multitask Learning
by: Abascal, John, et al.
Published: (2025)
by: Abascal, John, et al.
Published: (2025)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
by: Foerster, Hanna, et al.
Published: (2026)
by: Foerster, Hanna, et al.
Published: (2026)
Remote Timing Attacks on Efficient Language Model Inference
by: Carlini, Nicholas, et al.
Published: (2024)
by: Carlini, Nicholas, et al.
Published: (2024)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
by: Steinke, Thomas, et al.
Published: (2024)
by: Steinke, Thomas, et al.
Published: (2024)
Adversarial Inception Backdoor Attacks against Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
Identifying Models Behind Text-to-Image Leaderboards
by: Naseh, Ali, et al.
Published: (2026)
by: Naseh, Ali, et al.
Published: (2026)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
by: Gao, Yue, et al.
Published: (2023)
by: Gao, Yue, et al.
Published: (2023)
Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries
by: Laws, Matthew D., et al.
Published: (2026)
by: Laws, Matthew D., et al.
Published: (2026)
Locking Machine Learning Models into Hardware
by: Clifford, Eleanor, et al.
Published: (2024)
by: Clifford, Eleanor, et al.
Published: (2024)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Reconstruction of Personally Identifiable Information from Supervised Finetuned Models
by: Furukawa, Sae, et al.
Published: (2026)
by: Furukawa, Sae, et al.
Published: (2026)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
by: Huang, Yangsibo, et al.
Published: (2025)
by: Huang, Yangsibo, et al.
Published: (2025)
Exploring the limits of strong membership inference attacks on large language models
by: Hayes, Jamie, et al.
Published: (2025)
by: Hayes, Jamie, et al.
Published: (2025)
ceLLMate: Sandboxing Browser AI Agents
by: Meng, Luoxi, et al.
Published: (2025)
by: Meng, Luoxi, et al.
Published: (2025)
TMI! Finetuned Models Leak Private Information from their Pretraining Data
by: Abascal, John, et al.
Published: (2023)
by: Abascal, John, et al.
Published: (2023)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
by: Dahiya, Pranav, et al.
Published: (2023)
by: Dahiya, Pranav, et al.
Published: (2023)
Beyond Labeling Oracles: What does it mean to steal ML models?
by: Shafran, Avital, et al.
Published: (2023)
by: Shafran, Avital, et al.
Published: (2023)
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
by: Küchler, Nicolas, et al.
Published: (2025)
by: Küchler, Nicolas, et al.
Published: (2025)
Machine Learning Models Have a Supply Chain Problem
by: Meiklejohn, Sarah, et al.
Published: (2025)
by: Meiklejohn, Sarah, et al.
Published: (2025)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
by: Rinberg, Roy, et al.
Published: (2025)
by: Rinberg, Roy, et al.
Published: (2025)
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
by: Clifford, Eleanor, et al.
Published: (2022)
by: Clifford, Eleanor, et al.
Published: (2022)
When Vision Fails: Text Attacks Against ViT and OCR
by: Boucher, Nicholas, et al.
Published: (2023)
by: Boucher, Nicholas, et al.
Published: (2023)
Similar Items
-
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026) -
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024) -
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024) -
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025) -
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)