Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaudhari, Harsh, Rathbun, Ethan, Foerster, Hanna, Hayes, Jamie, Jagielski, Matthew, Nasr, Milad, Shumailov, Ilia, Oprea, Alina
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911404502548480
author Chaudhari, Harsh
Rathbun, Ethan
Foerster, Hanna
Hayes, Jamie
Jagielski, Matthew
Nasr, Milad
Shumailov, Ilia
Oprea, Alina
author_facet Chaudhari, Harsh
Rathbun, Ethan
Foerster, Hanna
Hayes, Jamie
Jagielski, Matthew
Nasr, Milad
Shumailov, Ilia
Oprea, Alina
contents Chain-of-Thought (CoT) reasoning has emerged as a powerful technique for enhancing large language models' capabilities by generating intermediate reasoning steps for complex tasks. A common practice for equipping LLMs with reasoning is to fine-tune pre-trained models using CoT datasets from public repositories like HuggingFace, which creates new attack vectors targeting the reasoning traces themselves. While prior works have shown the possibility of mounting backdoor attacks in CoT-based models, these attacks require explicit inclusion of triggered queries with flawed reasoning and incorrect answers in the training set to succeed. Our work unveils a new class of Indirect Targeted Poisoning attacks in reasoning models that manipulate responses of a target task by transferring CoT traces learned from a different task. Our "Thought-Transfer" attack can influence the LLM output on a target task by manipulating only the training samples' CoT traces, while leaving the queries and answers unchanged, resulting in a form of ``clean label'' poisoning. Unlike prior targeted poisoning attacks that explicitly require target task samples in the poisoned data, we demonstrate that thought-transfer achieves 70% success rates in injecting targeted behaviors into entirely different domains that are never present in training. Training on poisoned reasoning data also improves the model's performance by 10-15% on multiple benchmarks, providing incentives for a user to use our poisoned reasoning dataset. Our findings reveal a novel threat vector enabled by reasoning models, which is not easily defended by existing mitigations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19061
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
Chaudhari, Harsh
Rathbun, Ethan
Foerster, Hanna
Hayes, Jamie
Jagielski, Matthew
Nasr, Milad
Shumailov, Ilia
Oprea, Alina
Cryptography and Security
Machine Learning
Chain-of-Thought (CoT) reasoning has emerged as a powerful technique for enhancing large language models' capabilities by generating intermediate reasoning steps for complex tasks. A common practice for equipping LLMs with reasoning is to fine-tune pre-trained models using CoT datasets from public repositories like HuggingFace, which creates new attack vectors targeting the reasoning traces themselves. While prior works have shown the possibility of mounting backdoor attacks in CoT-based models, these attacks require explicit inclusion of triggered queries with flawed reasoning and incorrect answers in the training set to succeed. Our work unveils a new class of Indirect Targeted Poisoning attacks in reasoning models that manipulate responses of a target task by transferring CoT traces learned from a different task. Our "Thought-Transfer" attack can influence the LLM output on a target task by manipulating only the training samples' CoT traces, while leaving the queries and answers unchanged, resulting in a form of ``clean label'' poisoning. Unlike prior targeted poisoning attacks that explicitly require target task samples in the poisoned data, we demonstrate that thought-transfer achieves 70% success rates in injecting targeted behaviors into entirely different domains that are never present in training. Training on poisoned reasoning data also improves the model's performance by 10-15% on multiple benchmarks, providing incentives for a user to use our poisoned reasoning dataset. Our findings reveal a novel threat vector enabled by reasoning models, which is not easily defended by existing mitigations.
title Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2601.19061