Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chaudhari, Harsh, Rathbun, Ethan, Foerster, Hanna, Hayes, Jamie, Jagielski, Matthew, Nasr, Milad, Shumailov, Ilia, Oprea, Alina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024)
by: Foerster, Hanna, et al.
Published: (2024)
Adversarial Inception Backdoor Attacks against Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2024)
by: Rathbun, Ethan, et al.
Published: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
by: Foerster, Hanna, et al.
Published: (2026)
by: Foerster, Hanna, et al.
Published: (2026)
Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2026)
by: Rathbun, Ethan, et al.
Published: (2026)
UTrace: Poisoning Forensics for Private Collaborative Learning
by: Rose, Evan, et al.
Published: (2024)
by: Rose, Evan, et al.
Published: (2024)
Buffer Overflow in Mixture of Experts
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
by: Nasr, Milad, et al.
Published: (2025)
by: Nasr, Milad, et al.
Published: (2025)
Black-Box Privacy Attacks on Shared Representations in Multitask Learning
by: Abascal, John, et al.
Published: (2025)
by: Abascal, John, et al.
Published: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
by: Shi, Chongyang, et al.
Published: (2025)
by: Shi, Chongyang, et al.
Published: (2025)
Stealing User Prompts from Mixture of Experts
by: Yona, Itay, et al.
Published: (2024)
by: Yona, Itay, et al.
Published: (2024)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024)
by: Chadha, Karan, et al.
Published: (2024)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
by: Shumailov, Ilia, et al.
Published: (2024)
by: Shumailov, Ilia, et al.
Published: (2024)
Soft Instruction De-escalation Defense
by: Walter, Nils Philipp, et al.
Published: (2025)
by: Walter, Nils Philipp, et al.
Published: (2025)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
by: Gao, Yue, et al.
Published: (2023)
by: Gao, Yue, et al.
Published: (2023)
Interpreting the Repeated Token Phenomenon in Large Language Models
by: Yona, Itay, et al.
Published: (2025)
by: Yona, Itay, et al.
Published: (2025)
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025)
by: Suri, Anshuman, et al.
Published: (2025)
Remote Timing Attacks on Efficient Language Model Inference
by: Carlini, Nicholas, et al.
Published: (2024)
by: Carlini, Nicholas, et al.
Published: (2024)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Thought Purity: A Defense Framework For Chain-of-Thought Attack
by: Xue, Zihao, et al.
Published: (2025)
by: Xue, Zihao, et al.
Published: (2025)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
by: Liu, Shuaitong, et al.
Published: (2025)
by: Liu, Shuaitong, et al.
Published: (2025)
Identifying Models Behind Text-to-Image Leaderboards
by: Naseh, Ali, et al.
Published: (2026)
by: Naseh, Ali, et al.
Published: (2026)
Hierarchical Multi-agent Reinforcement Learning for Cyber Network Defense
by: Singh, Aditya Vikram, et al.
Published: (2024)
by: Singh, Aditya Vikram, et al.
Published: (2024)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
by: Steinke, Thomas, et al.
Published: (2024)
by: Steinke, Thomas, et al.
Published: (2024)
Chameleon: Increasing Label-Only Membership Leakage with Adaptive Poisoning
by: Chaudhari, Harsh, et al.
Published: (2023)
by: Chaudhari, Harsh, et al.
Published: (2023)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
by: Syros, Georgios, et al.
Published: (2026)
by: Syros, Georgios, et al.
Published: (2026)
Synthetic Query Generation for Privacy-Preserving Deep Retrieval Systems using Differentially Private Language Models
by: Carranza, Aldo Gael, et al.
Published: (2023)
by: Carranza, Aldo Gael, et al.
Published: (2023)
Kraken: Higher-order EM Side-Channel Attacks on DNNs in Near and Far Field
by: Horvath, Peter, et al.
Published: (2026)
by: Horvath, Peter, et al.
Published: (2026)
Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries
by: Laws, Matthew D., et al.
Published: (2026)
by: Laws, Matthew D., et al.
Published: (2026)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
When Vision Fails: Text Attacks Against ViT and OCR
by: Boucher, Nicholas, et al.
Published: (2023)
by: Boucher, Nicholas, et al.
Published: (2023)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
by: Zolkowski, Artur, et al.
Published: (2025)
by: Zolkowski, Artur, et al.
Published: (2025)
Exploring the limits of strong membership inference attacks on large language models
by: Hayes, Jamie, et al.
Published: (2025)
by: Hayes, Jamie, et al.
Published: (2025)
Transferable Availability Poisoning Attacks
by: Liu, Yiyong, et al.
Published: (2023)
by: Liu, Yiyong, et al.
Published: (2023)
Similar Items
-
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025) -
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025) -
SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning Agents
by: Rathbun, Ethan, et al.
Published: (2024) -
Beyond Slow Signs in High-fidelity Model Extraction
by: Foerster, Hanna, et al.
Published: (2024) -
Adversarial Inception Backdoor Attacks against Reinforcement Learning
by: Rathbun, Ethan, et al.
Published: (2024)