Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
Fuente:
arXiv
Saved in:
| Main Authors: | Churina, Svetlana, Chebrolu, Niranjan, Jaidka, Kokil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2025)
by: Bouaziz, Wassim, et al.
Published: (2025)
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
by: Churina, Svetlana, et al.
Published: (2024)
by: Churina, Svetlana, et al.
Published: (2024)
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
by: Yuan, Shuai, et al.
Published: (2025)
by: Yuan, Shuai, et al.
Published: (2025)
Indiscriminate Data Poisoning Attacks on Pre-trained Feature Extractors
by: Lu, Yiwei, et al.
Published: (2024)
by: Lu, Yiwei, et al.
Published: (2024)
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Gray-Box Poisoning of Continuous Malware Ingestion Pipelines
by: Dolejš, Jan, et al.
Published: (2026)
by: Dolejš, Jan, et al.
Published: (2026)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
On Evaluating the Poisoning Robustness of Federated Learning under Local Differential Privacy
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Online Poisoning Attack Against Reinforcement Learning under Black-box Environments
by: Li, Jianhui, et al.
Published: (2024)
by: Li, Jianhui, et al.
Published: (2024)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
by: Tran, Khang, et al.
Published: (2026)
by: Tran, Khang, et al.
Published: (2026)
FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
by: Long, Yuzhen, et al.
Published: (2025)
by: Long, Yuzhen, et al.
Published: (2025)
Timber! Poisoning Decision Trees
by: Calzavara, Stefano, et al.
Published: (2024)
by: Calzavara, Stefano, et al.
Published: (2024)
Transferable Availability Poisoning Attacks
by: Liu, Yiyong, et al.
Published: (2023)
by: Liu, Yiyong, et al.
Published: (2023)
Provable Watermarking for Data Poisoning Attacks
by: Zhu, Yifan, et al.
Published: (2025)
by: Zhu, Yifan, et al.
Published: (2025)
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2025)
Lower Bounds for Public-Private Learning under Distribution Shift
by: Setlur, Amrith, et al.
Published: (2025)
by: Setlur, Amrith, et al.
Published: (2025)
GShield: Mitigating Poisoning Attacks in Federated Learning
by: M., Sameera K., et al.
Published: (2025)
by: M., Sameera K., et al.
Published: (2025)
Poisoned Source Code Detection in Code Models
by: Ghannoum, Ehab, et al.
Published: (2025)
by: Ghannoum, Ehab, et al.
Published: (2025)
Indiscriminate Data Poisoning Attacks on Neural Networks
by: Lu, Yiwei, et al.
Published: (2022)
by: Lu, Yiwei, et al.
Published: (2022)
UTrace: Poisoning Forensics for Private Collaborative Learning
by: Rose, Evan, et al.
Published: (2024)
by: Rose, Evan, et al.
Published: (2024)
Privacy-Preserving 3-Layer Neural Network Training
by: Chiang, John
Published: (2023)
by: Chiang, John
Published: (2023)
On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift
by: Thaker, Pratiksha, et al.
Published: (2023)
by: Thaker, Pratiksha, et al.
Published: (2023)
Local Environment Poisoning Attacks on Federated Reinforcement Learning
by: Ma, Evelyn, et al.
Published: (2023)
by: Ma, Evelyn, et al.
Published: (2023)
TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
by: Aghakhani, Hojjat, et al.
Published: (2023)
by: Aghakhani, Hojjat, et al.
Published: (2023)
Inverting Gradient Attacks Makes Powerful Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2024)
by: Bouaziz, Wassim, et al.
Published: (2024)
Data Poisoning: An Overlooked Threat to Power Grid Resilience
by: Agah, Nora, et al.
Published: (2024)
by: Agah, Nora, et al.
Published: (2024)
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Sybil-based Virtual Data Poisoning Attacks in Federated Learning
by: Zhu, Changxun, et al.
Published: (2025)
by: Zhu, Changxun, et al.
Published: (2025)
Poisoning Attacks to Local Differential Privacy Protocols for Trajectory Data
by: Hsu, I-Jung, et al.
Published: (2025)
by: Hsu, I-Jung, et al.
Published: (2025)
Salsa Fresca: Angular Embeddings and Pre-Training for ML Attacks on Learning With Errors
by: Stevens, Samuel, et al.
Published: (2024)
by: Stevens, Samuel, et al.
Published: (2024)
FedRecAttack: Model Poisoning Attack to Federated Recommendation
by: Rong, Dazhong, et al.
Published: (2022)
by: Rong, Dazhong, et al.
Published: (2022)
Comments on "Privacy-Enhanced Federated Learning Against Poisoning Adversaries"
by: Schneider, Thomas, et al.
Published: (2024)
by: Schneider, Thomas, et al.
Published: (2024)
Exact Certification of (Graph) Neural Networks Against Label Poisoning
by: Sabanayagam, Mahalakshmi, et al.
Published: (2024)
by: Sabanayagam, Mahalakshmi, et al.
Published: (2024)
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
by: Hu, Haoyang, et al.
Published: (2026)
by: Hu, Haoyang, et al.
Published: (2026)
Adversarial Update-Based Federated Unlearning for Poisoned Model Recovery
by: Zhao, Wenwei, et al.
Published: (2026)
by: Zhao, Wenwei, et al.
Published: (2026)
Data Poisoning Attacks in Intelligent Transportation Systems: A Survey
by: Wang, Feilong, et al.
Published: (2024)
by: Wang, Feilong, et al.
Published: (2024)
Debiased Graph Poisoning Attack via Contrastive Surrogate Objective
by: Yoon, Kanghoon, et al.
Published: (2024)
by: Yoon, Kanghoon, et al.
Published: (2024)
Enhancing the Antidote: Improved Pointwise Certifications against Poisoning Attacks
by: Liu, Shijie, et al.
Published: (2023)
by: Liu, Shijie, et al.
Published: (2023)
Continuous Multi-Task Pre-training for Malicious URL Detection and Webpage Classification
by: Li, Yujie, et al.
Published: (2024)
by: Li, Yujie, et al.
Published: (2024)
Similar Items
-
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
by: Bouaziz, Wassim, et al.
Published: (2025) -
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
by: Churina, Svetlana, et al.
Published: (2024) -
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
by: Yuan, Shuai, et al.
Published: (2025) -
Indiscriminate Data Poisoning Attacks on Pre-trained Feature Extractors
by: Lu, Yiwei, et al.
Published: (2024) -
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)