Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Kusaka, Shigeki, Saito, Keita, Kudo, Mikoto, Tanabe, Takumi, Wachi, Akifumi, Akimoto, Youhei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026)
by: Kudo, Mikoto, et al.
Published: (2026)
Stepwise Alignment for Constrained Language Model Policy Optimization
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025)
by: Wachi, Akifumi, et al.
Published: (2025)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
by: Tran, Thien Q., et al.
Published: (2025)
by: Tran, Thien Q., et al.
Published: (2025)
Policy Iteration for Two-Player General-Sum Stochastic Stackelberg Games
by: Kudo, Mikoto, et al.
Published: (2024)
by: Kudo, Mikoto, et al.
Published: (2024)
Target Return Optimizer for Multi-Game Decision Transformer
by: Tatematsu, Kensuke, et al.
Published: (2025)
by: Tatematsu, Kensuke, et al.
Published: (2025)
A Survey of Constraint Formulations in Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
Long-term Safe Reinforcement Learning with Binary Feedback
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
by: Wachi, Akifumi, et al.
Published: (2026)
by: Wachi, Akifumi, et al.
Published: (2026)
Statistically Significant Concept-based Explanation of Image Classifiers via Model Knockoffs
by: Xu, Kaiwen, et al.
Published: (2023)
by: Xu, Kaiwen, et al.
Published: (2023)
Poisoning the Inner Prediction Logic of Graph Neural Networks for Clean-Label Backdoor Attacks
by: Zhang, Yuxiang, et al.
Published: (2026)
by: Zhang, Yuxiang, et al.
Published: (2026)
Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
by: Fukasaku, Ryoya, et al.
Published: (2025)
by: Fukasaku, Ryoya, et al.
Published: (2025)
Topology-Independent Robustness of the Weighted Mean under Label Poisoning Attacks in Heterogeneous Decentralized Learning
by: Peng, Jie, et al.
Published: (2026)
by: Peng, Jie, et al.
Published: (2026)
Relabeling Minimal Training Subset to Flip a Prediction
by: Yang, Jinghan, et al.
Published: (2023)
by: Yang, Jinghan, et al.
Published: (2023)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
A Backdoor Approach with Inverted Labels Using Dirty Label-Flipping Attacks
by: Mengara, Orson
Published: (2024)
by: Mengara, Orson
Published: (2024)
Peak-Controlled Logits Poisoning Attack in Federated Distillation
by: Tang, Yuhan, et al.
Published: (2024)
by: Tang, Yuhan, et al.
Published: (2024)
Perfect Alignment May be Poisonous to Graph Contrastive Learning
by: Liu, Jingyu, et al.
Published: (2023)
by: Liu, Jingyu, et al.
Published: (2023)
Flipping-based Policy for Chance-Constrained Markov Decision Processes
by: Shen, Xun, et al.
Published: (2024)
by: Shen, Xun, et al.
Published: (2024)
Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing
by: Broadwater, Keita
Published: (2026)
by: Broadwater, Keita
Published: (2026)
Learning to Attack: A Bandit Approach to Adversarial Context Poisoning
by: Telikani, Ray, et al.
Published: (2026)
by: Telikani, Ray, et al.
Published: (2026)
DMPA: Model Poisoning Attacks on Decentralized Federated Learning for Model Differences
by: Feng, Chao, et al.
Published: (2025)
by: Feng, Chao, et al.
Published: (2025)
Exploiting Meta-Learning-based Poisoning Attacks for Graph Link Prediction
by: Li, Mingchen, et al.
Published: (2025)
by: Li, Mingchen, et al.
Published: (2025)
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
by: Halloran, John T., et al.
Published: (2026)
by: Halloran, John T., et al.
Published: (2026)
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
Multi-level Certified Defense Against Poisoning Attacks in Offline Reinforcement Learning
by: Liu, Shijie, et al.
Published: (2025)
by: Liu, Shijie, et al.
Published: (2025)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
by: Di, Jimmy Z., et al.
Published: (2022)
by: Di, Jimmy Z., et al.
Published: (2022)
DeepMedcast: A Deep Learning Method for Generating Intermediate Weather Forecasts among Multiple NWP Models
by: Kudo, Atsushi
Published: (2024)
by: Kudo, Atsushi
Published: (2024)
Verification of Bit-Flip Attacks against Quantized Neural Networks
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
FLAegis: A Two-Layer Defense Framework for Federated Learning Against Poisoning Attacks
by: Campos, Enrique Mármol, et al.
Published: (2025)
by: Campos, Enrique Mármol, et al.
Published: (2025)
Mitigating Label Flipping Attacks in Malicious URL Detectors Using Ensemble Trees
by: Nowroozi, Ehsan, et al.
Published: (2024)
by: Nowroozi, Ehsan, et al.
Published: (2024)
PoiCGAN: A Targeted Poisoning Based on Feature-Label Joint Perturbation in Federated Learning
by: Liu, Tao, et al.
Published: (2026)
by: Liu, Tao, et al.
Published: (2026)
Stealthy Poisoning Attacks Bypass Defenses in Regression Settings
by: Carnerero-Cano, Javier, et al.
Published: (2026)
by: Carnerero-Cano, Javier, et al.
Published: (2026)
Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
by: Fujii, Kazuki, et al.
Published: (2025)
by: Fujii, Kazuki, et al.
Published: (2025)
Collapsing Sequence-Level Data-Policy Coverage via Poisoning Attack in Offline Reinforcement Learning
by: Zhou, Xue, et al.
Published: (2025)
by: Zhou, Xue, et al.
Published: (2025)
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
by: Das, Sanjay, et al.
Published: (2024)
by: Das, Sanjay, et al.
Published: (2024)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
Label Smoothing is a Pragmatic Information Bottleneck
by: Kudo, Sota
Published: (2025)
by: Kudo, Sota
Published: (2025)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
by: Zhang, Heyi, et al.
Published: (2025)
by: Zhang, Heyi, et al.
Published: (2025)
Similar Items
-
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026) -
Stepwise Alignment for Constrained Language Model Policy Optimization
by: Wachi, Akifumi, et al.
Published: (2024) -
A Provable Approach for End-to-End Safe Reinforcement Learning
by: Wachi, Akifumi, et al.
Published: (2025) -
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
by: Tran, Thien Q., et al.
Published: (2025) -
Policy Iteration for Two-Player General-Sum Stochastic Stackelberg Games
by: Kudo, Mikoto, et al.
Published: (2024)