Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Jiacheng, Li, Yiming, Song, Tao, Wang, Weijian, Qu, Wenjie, Guan, Haibing, Zhang, Jiaheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
di: Yang, Wenyuan, et al.
Pubblicazione: (2025)
di: Yang, Wenyuan, et al.
Pubblicazione: (2025)
Self-Sovereign Agent
di: Qu, Wenjie, et al.
Pubblicazione: (2026)
di: Qu, Wenjie, et al.
Pubblicazione: (2026)
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
Stealthy Yet Effective: Distribution-Preserving Backdoor Attacks on Graph Classification
di: Wang, Xiaobao, et al.
Pubblicazione: (2025)
di: Wang, Xiaobao, et al.
Pubblicazione: (2025)
SettleFL: Trustless and Scalable Reward Settlement Protocol for Federated Learning on Permissionless Blockchains (Extended version)
di: Liang, Shuang, et al.
Pubblicazione: (2026)
di: Liang, Shuang, et al.
Pubblicazione: (2026)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
di: Li, Yuexin, et al.
Pubblicazione: (2026)
di: Li, Yuexin, et al.
Pubblicazione: (2026)
LoRAGuard: An Effective Black-box Watermarking Approach for LoRAs
di: Lv, Peizhuo, et al.
Pubblicazione: (2025)
di: Lv, Peizhuo, et al.
Pubblicazione: (2025)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
di: Guo, Zhen, et al.
Pubblicazione: (2025)
di: Guo, Zhen, et al.
Pubblicazione: (2025)
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
di: Liang, Jiacheng, et al.
Pubblicazione: (2024)
di: Liang, Jiacheng, et al.
Pubblicazione: (2024)
Hashed Watermark as a Filter: Defeating Forging and Overwriting Attacks in Weight-based Neural Network Watermarking
di: Yao, Yuan, et al.
Pubblicazione: (2025)
di: Yao, Yuan, et al.
Pubblicazione: (2025)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
di: Shao, Shuo, et al.
Pubblicazione: (2024)
di: Shao, Shuo, et al.
Pubblicazione: (2024)
Ideal Attribution and Faithful Watermarks for Language Models
di: Song, Min Jae, et al.
Pubblicazione: (2025)
di: Song, Min Jae, et al.
Pubblicazione: (2025)
SWaRL: Safeguard Code Watermarking via Reinforcement Learning
di: Javidnia, Neusha, et al.
Pubblicazione: (2026)
di: Javidnia, Neusha, et al.
Pubblicazione: (2026)
Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models
di: Zhang, Yiyang, et al.
Pubblicazione: (2026)
di: Zhang, Yiyang, et al.
Pubblicazione: (2026)
Stealthy Adversarial Attacks on Stochastic Multi-Armed Bandits
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
LLM Fingerprinting via Semantically Conditioned Watermarks
di: Gloaguen, Thibaud, et al.
Pubblicazione: (2025)
di: Gloaguen, Thibaud, et al.
Pubblicazione: (2025)
Robust Spectral Watermark for Synthetic Tabular Data
di: Zhao, Yizhou, et al.
Pubblicazione: (2025)
di: Zhao, Yizhou, et al.
Pubblicazione: (2025)
R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models
di: Zhang, Ziming, et al.
Pubblicazione: (2026)
di: Zhang, Ziming, et al.
Pubblicazione: (2026)
Distortion-free Watermarks are not Truly Distortion-free under Watermark Key Collisions
di: Wu, Yihan, et al.
Pubblicazione: (2024)
di: Wu, Yihan, et al.
Pubblicazione: (2024)
State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
di: Guo, Ji, et al.
Pubblicazione: (2026)
di: Guo, Ji, et al.
Pubblicazione: (2026)
Lurking in the shadows: Unveiling Stealthy Backdoor Attacks against Personalized Federated Learning
di: Lyu, Xiaoting, et al.
Pubblicazione: (2024)
di: Lyu, Xiaoting, et al.
Pubblicazione: (2024)
Stealthy Imitation: Reward-guided Environment-free Policy Stealing
di: Zhuang, Zhixiong, et al.
Pubblicazione: (2024)
di: Zhuang, Zhixiong, et al.
Pubblicazione: (2024)
GESR: Graph-Based Edge Semantic Reconstruction for Stealthy Communication Detection with Benign-Only Training
di: Xu, Henghui, et al.
Pubblicazione: (2026)
di: Xu, Henghui, et al.
Pubblicazione: (2026)
Graph-Aware Stealthy Poison-Text Backdoors for Text-Attributed Graphs
di: Luo, Qi, et al.
Pubblicazione: (2026)
di: Luo, Qi, et al.
Pubblicazione: (2026)
SWA-LDM: Toward Stealthy Watermarks for Latent Diffusion Models
di: Yang, Zhonghao, et al.
Pubblicazione: (2025)
di: Yang, Zhonghao, et al.
Pubblicazione: (2025)
Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
di: Libon, Lena, et al.
Pubblicazione: (2025)
di: Libon, Lena, et al.
Pubblicazione: (2025)
Robust GNN Watermarking via Implicit Perception of Topological Invariants
di: Li, Jipeng, et al.
Pubblicazione: (2025)
di: Li, Jipeng, et al.
Pubblicazione: (2025)
Unforgeable Watermarks for Language Models via Robust Signatures
di: Lin, Huijia, et al.
Pubblicazione: (2026)
di: Lin, Huijia, et al.
Pubblicazione: (2026)
AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
di: Liang, Jiacheng, et al.
Pubblicazione: (2025)
di: Liang, Jiacheng, et al.
Pubblicazione: (2025)
SDBA: A Stealthy and Long-Lasting Durable Backdoor Attack in Federated Learning
di: Choe, Minyeong, et al.
Pubblicazione: (2024)
di: Choe, Minyeong, et al.
Pubblicazione: (2024)
SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
di: Xu, Haotian, et al.
Pubblicazione: (2025)
di: Xu, Haotian, et al.
Pubblicazione: (2025)
Traceable Black-box Watermarks for Federated Learning
di: Xu, Jiahao, et al.
Pubblicazione: (2025)
di: Xu, Jiahao, et al.
Pubblicazione: (2025)
Provably Robust Multi-bit Watermarking for AI-generated Text
di: Qu, Wenjie, et al.
Pubblicazione: (2024)
di: Qu, Wenjie, et al.
Pubblicazione: (2024)
DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks
di: Yang, Yunfei, et al.
Pubblicazione: (2025)
di: Yang, Yunfei, et al.
Pubblicazione: (2025)
Poisoning with A Pill: Circumventing Detection in Federated Learning
di: Guo, Hanxi, et al.
Pubblicazione: (2024)
di: Guo, Hanxi, et al.
Pubblicazione: (2024)
Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
di: Hsu, Chia-Yi, et al.
Pubblicazione: (2026)
di: Hsu, Chia-Yi, et al.
Pubblicazione: (2026)
Output Supervision Can Obfuscate the Chain of Thought
di: Drori, Jacob, et al.
Pubblicazione: (2025)
di: Drori, Jacob, et al.
Pubblicazione: (2025)
Watermarking Generative Categorical Data
di: Gu, Bochao, et al.
Pubblicazione: (2024)
di: Gu, Bochao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
di: Yang, Wenyuan, et al.
Pubblicazione: (2025) -
Self-Sovereign Agent
di: Qu, Wenjie, et al.
Pubblicazione: (2026) -
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
di: Wu, Linyu, et al.
Pubblicazione: (2025) -
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026) -
Stealthy Yet Effective: Distribution-Preserving Backdoor Attacks on Graph Classification
di: Wang, Xiaobao, et al.
Pubblicazione: (2025)