Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
Fuente:
arXiv
Saved in:
| Main Authors: | Foerster, Hanna, Shumailov, Ilia, Zhao, Yiren, Chaudhari, Harsh, Hayes, Jamie, Mullins, Robert, Gal, Yarin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024)
by: Foerster, Hanna, et al.
Published: (2024)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
by: Foerster, Hanna, et al.
Published: (2026)
by: Foerster, Hanna, et al.
Published: (2026)
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
by: Clifford, Eleanor, et al.
Published: (2022)
by: Clifford, Eleanor, et al.
Published: (2022)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Locking Machine Learning Models into Hardware
by: Clifford, Eleanor, et al.
Published: (2024)
by: Clifford, Eleanor, et al.
Published: (2024)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Architectural Neural Backdoors from First Principles
by: Langford, Harry, et al.
Published: (2024)
by: Langford, Harry, et al.
Published: (2024)
Watermarking Needs Input Repetition Masking
by: Khachaturov, David, et al.
Published: (2025)
by: Khachaturov, David, et al.
Published: (2025)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
Soft Instruction De-escalation Defense
by: Walter, Nils Philipp, et al.
Published: (2025)
by: Walter, Nils Philipp, et al.
Published: (2025)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
by: Gao, Yue, et al.
Published: (2023)
by: Gao, Yue, et al.
Published: (2023)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
Kraken: Higher-order EM Side-Channel Attacks on DNNs in Near and Far Field
by: Horvath, Peter, et al.
Published: (2026)
by: Horvath, Peter, et al.
Published: (2026)
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
When Vision Fails: Text Attacks Against ViT and OCR
by: Boucher, Nicholas, et al.
Published: (2023)
by: Boucher, Nicholas, et al.
Published: (2023)
UTrace: Poisoning Forensics for Private Collaborative Learning
by: Rose, Evan, et al.
Published: (2024)
by: Rose, Evan, et al.
Published: (2024)
Inverting Gradient Attacks Makes Powerful Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2024)
by: Bouaziz, Wassim, et al.
Published: (2024)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
by: Dahiya, Pranav, et al.
Published: (2023)
by: Dahiya, Pranav, et al.
Published: (2023)
ceLLMate: Sandboxing Browser AI Agents
by: Meng, Luoxi, et al.
Published: (2025)
by: Meng, Luoxi, et al.
Published: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Defeating Prompt Injections by Design
by: Debenedetti, Edoardo, et al.
Published: (2025)
by: Debenedetti, Edoardo, et al.
Published: (2025)
PoisonCatcher: Revealing and Identifying LDP Poisoning Attacks in IIoT
by: Shuai, Lisha, et al.
Published: (2024)
by: Shuai, Lisha, et al.
Published: (2024)
Beyond Labeling Oracles: What does it mean to steal ML models?
by: Shafran, Avital, et al.
Published: (2023)
by: Shafran, Avital, et al.
Published: (2023)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
by: Vero, Mark, et al.
Published: (2026)
by: Vero, Mark, et al.
Published: (2026)
Sharpness-Aware Data Poisoning Attack
by: He, Pengfei, et al.
Published: (2023)
by: He, Pengfei, et al.
Published: (2023)
Introducing a New Alert Data Set for Multi-Step Attack Analysis
by: Landauer, Max, et al.
Published: (2023)
by: Landauer, Max, et al.
Published: (2023)
Fixing 7,400 Bugs for 1$: Cheap Crash-Site Program Repair
by: Zheng, Han, et al.
Published: (2025)
by: Zheng, Han, et al.
Published: (2025)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
by: Rinberg, Roy, et al.
Published: (2025)
by: Rinberg, Roy, et al.
Published: (2025)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
by: Küchler, Nicolas, et al.
Published: (2025)
by: Küchler, Nicolas, et al.
Published: (2025)
Are We There Yet? Timing and Floating-Point Attacks on Differential Privacy Systems
by: Jin, Jiankai, et al.
Published: (2021)
by: Jin, Jiankai, et al.
Published: (2021)
Mitigating Data Poisoning Attacks to Local Differential Privacy
by: Li, Xiaolin, et al.
Published: (2025)
by: Li, Xiaolin, et al.
Published: (2025)
Poisoning Attacks to Local Differential Privacy for Ranking Estimation
by: Zhan, Pei, et al.
Published: (2025)
by: Zhan, Pei, et al.
Published: (2025)
Poisoning the Pixels: Revisiting Backdoor Attacks on Semantic Segmentation
by: Zhang, Guangsheng, et al.
Published: (2026)
by: Zhang, Guangsheng, et al.
Published: (2026)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
by: Sneh, Jonathan, et al.
Published: (2025)
by: Sneh, Jonathan, et al.
Published: (2025)
Transferable Availability Poisoning Attacks
by: Liu, Yiyong, et al.
Published: (2023)
by: Liu, Yiyong, et al.
Published: (2023)
Effectiveness of Adversarial Benign and Malware Examples in Evasion and Poisoning Attacks
by: Kozák, Matouš, et al.
Published: (2025)
by: Kozák, Matouš, et al.
Published: (2025)
Fake Resume Attacks: Data Poisoning on Online Job Platforms
by: Yamashita, Michiharu, et al.
Published: (2024)
by: Yamashita, Michiharu, et al.
Published: (2024)
Similar Items
-
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026) -
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024) -
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
by: Foerster, Hanna, et al.
Published: (2026) -
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025) -
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
by: Clifford, Eleanor, et al.
Published: (2022)