Silent Sabotage During Fine-Tuning: Few-Shot Rationale Poisoning of Compact Medical LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Jingyuan, Wang, Wenjie, Wu, Ji, Gao, Jiandong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs
by: Arachchige, Shashie Dilhara Batan, et al.
Published: (2025)
by: Arachchige, Shashie Dilhara Batan, et al.
Published: (2025)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Scaling Trends for Data Poisoning in LLMs
by: Bowen, Dillon, et al.
Published: (2024)
by: Bowen, Dillon, et al.
Published: (2024)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
by: Lu, Ning, et al.
Published: (2025)
by: Lu, Ning, et al.
Published: (2025)
Silent Commitment Failure in Instruction-Tuned Language Models: Evidence of Governability Divergence Across Architectures
by: Ruddell, Gregory M.
Published: (2026)
by: Ruddell, Gregory M.
Published: (2026)
In-Context Unlearning: Language Models as Few Shot Unlearners
by: Pawelczyk, Martin, et al.
Published: (2023)
by: Pawelczyk, Martin, et al.
Published: (2023)
From Flows to Words: Can Zero-/Few-Shot LLMs Detect Network Intrusions? A Grammar-Constrained, Calibrated Evaluation on UNSW-NB15
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Have You Poisoned My Data? Defending Neural Networks against Data Poisoning
by: De Gaspari, Fabio, et al.
Published: (2024)
by: De Gaspari, Fabio, et al.
Published: (2024)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
by: Poppi, Samuele, et al.
Published: (2024)
by: Poppi, Samuele, et al.
Published: (2024)
PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud Collaboration
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
by: Hasan, Adib, et al.
Published: (2024)
by: Hasan, Adib, et al.
Published: (2024)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
PACZero: PAC-Private Fine-Tuning of Language Models via Sign Quantization
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
by: Frikha, Ahmed, et al.
Published: (2024)
by: Frikha, Ahmed, et al.
Published: (2024)
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
by: Konrad, Phongsakon Mark, et al.
Published: (2026)
Semantic-Aware Contrastive Fine-Tuning: Boosting Multimodal Malware Classification with Discriminative Embeddings
by: Sanchez, Ivan Montoya, et al.
Published: (2025)
by: Sanchez, Ivan Montoya, et al.
Published: (2025)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions
by: Li, Wenjuan, et al.
Published: (2026)
by: Li, Wenjuan, et al.
Published: (2026)
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
by: Du, Hao, et al.
Published: (2025)
by: Du, Hao, et al.
Published: (2025)
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation
by: Chen, Qizhi, et al.
Published: (2026)
by: Chen, Qizhi, et al.
Published: (2026)
Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
by: Foerster, Hanna, et al.
Published: (2025)
by: Foerster, Hanna, et al.
Published: (2025)
FedNIA: Noise-Induced Activation Analysis for Mitigating Data Poisoning in FL
by: Hallaji, Ehsan, et al.
Published: (2025)
by: Hallaji, Ehsan, et al.
Published: (2025)
Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning
by: Xu, Yinglun, et al.
Published: (2024)
by: Xu, Yinglun, et al.
Published: (2024)
Concealing Backdoor Model Updates in Federated Learning by Trigger-Optimized Data Poisoning
by: Zhang, Yujie, et al.
Published: (2024)
by: Zhang, Yujie, et al.
Published: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
FedReview: A Review Mechanism for Rejecting Poisoned Updates in Federated Learning
by: Zheng, Tianhang, et al.
Published: (2024)
by: Zheng, Tianhang, et al.
Published: (2024)
DP-Adam-AC: Privacy-preserving Fine-Tuning of Localizable Language Models Using Adam Optimization with Adaptive Clipping
by: Yang, Ruoxing
Published: (2025)
by: Yang, Ruoxing
Published: (2025)
BadSampler: Harnessing the Power of Catastrophic Forgetting to Poison Byzantine-robust Federated Learning
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
SAFELOC: Overcoming Data Poisoning Attacks in Heterogeneous Federated Machine Learning for Indoor Localization
by: Singampalli, Akhil, et al.
Published: (2024)
by: Singampalli, Akhil, et al.
Published: (2024)
EAB-FL: Exacerbating Algorithmic Bias through Model Poisoning Attacks in Federated Learning
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
by: Meerza, Syed Irfan Ali, et al.
Published: (2024)
Similar Items
-
CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs
by: Arachchige, Shashie Dilhara Batan, et al.
Published: (2025) -
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025) -
Scaling Trends for Data Poisoning in LLMs
by: Bowen, Dillon, et al.
Published: (2024) -
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
by: Lu, Ning, et al.
Published: (2025) -
Silent Commitment Failure in Instruction-Tuned Language Models: Evidence of Governability Divergence Across Architectures
by: Ruddell, Gregory M.
Published: (2026)