RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jiongxiao, Wu, Junlin, Chen, Muhao, Vorobeychik, Yevgeniy, Xiao, Chaowei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024)
by: Wu, Junlin, et al.
Published: (2024)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
Human-in-the-Loop Generation of Adversarial Texts: A Case Study on Tibetan Script
by: Cao, Xi, et al.
Published: (2024)
by: Cao, Xi, et al.
Published: (2024)
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
by: Esposito, Matteo, et al.
Published: (2024)
by: Esposito, Matteo, et al.
Published: (2024)
Language Model Agents Under Attack: A Cross Model-Benchmark of Profit-Seeking Behaviors in Customer Service
by: Zhang, Jingyu
Published: (2025)
by: Zhang, Jingyu
Published: (2025)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
Risk Psychology & Cyber-Attack Tactics
by: Kim, Rubens, et al.
Published: (2025)
by: Kim, Rubens, et al.
Published: (2025)
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
by: Inie, Nanna, et al.
Published: (2023)
by: Inie, Nanna, et al.
Published: (2023)
Jaco: An Offline Running Privacy-aware Voice Assistant
by: Bermuth, Daniel, et al.
Published: (2022)
by: Bermuth, Daniel, et al.
Published: (2022)
Test Security in Remote Testing Age: Perspectives from Process Data Analytics and AI
by: Hao, Jiangang, et al.
Published: (2024)
by: Hao, Jiangang, et al.
Published: (2024)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
by: Wang, Jiongxiao, et al.
Published: (2024)
by: Wang, Jiongxiao, et al.
Published: (2024)
SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
Human-Centered Privacy Research in the Age of Large Language Models
by: Li, Tianshi, et al.
Published: (2024)
by: Li, Tianshi, et al.
Published: (2024)
Can Large Language Models Automate Phishing Warning Explanations? A Controlled Experiment on Effectiveness and User Perception
by: Cau, Federico Maria, et al.
Published: (2025)
by: Cau, Federico Maria, et al.
Published: (2025)
MORPHEUS: A Multidimensional Framework for Modeling, Measuring, and Mitigating Human Factors in Cybersecurity
by: Desolda, Giuseppe, et al.
Published: (2025)
by: Desolda, Giuseppe, et al.
Published: (2025)
Professor X: Manipulating EEG BCI with Invisible and Robust Backdoor Attack
by: Liu, Xuan-Hao, et al.
Published: (2024)
by: Liu, Xuan-Hao, et al.
Published: (2024)
A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models
by: Derner, Erik, et al.
Published: (2023)
by: Derner, Erik, et al.
Published: (2023)
Comparative Simulation of Phishing Attacks on a Critical Information Infrastructure Organization: An Empirical Study
by: Sirawongphatsara, Patsita, et al.
Published: (2024)
by: Sirawongphatsara, Patsita, et al.
Published: (2024)
Human Factors in the LastPass Breach
by: Sugunaraj, Niroop
Published: (2024)
by: Sugunaraj, Niroop
Published: (2024)
Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning
by: Qiu, Wenjun, et al.
Published: (2024)
by: Qiu, Wenjun, et al.
Published: (2024)
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
by: Prakash, Vijay, et al.
Published: (2024)
by: Prakash, Vijay, et al.
Published: (2024)
Watch Your Language: Investigating Content Moderation with Large Language Models
by: Kumar, Deepak, et al.
Published: (2023)
by: Kumar, Deepak, et al.
Published: (2023)
Human-AI Collaboration in Cloud Security: Cognitive Hierarchy-Driven Deep Reinforcement Learning
by: Aref, Zahra, et al.
Published: (2025)
by: Aref, Zahra, et al.
Published: (2025)
SoK: Come Together -- Unifying Security, Information Theory, and Cognition for a Mixed Reality Deception Attack Ontology & Analysis Framework
by: Teymourian, Ali, et al.
Published: (2025)
by: Teymourian, Ali, et al.
Published: (2025)
False Reality: Uncovering Sensor-induced Human-VR Interaction Vulnerability
by: Jiang, Yancheng, et al.
Published: (2025)
by: Jiang, Yancheng, et al.
Published: (2025)
Security Risks of AI Agents Hiring Humans: An Empirical Marketplace Study
by: Mehta, Pulak
Published: (2026)
by: Mehta, Pulak
Published: (2026)
REVERSIM: An Open-Source Environment for the Controlled Study of Human Aspects in Hardware Reverse Engineering
by: Becker, Steffen, et al.
Published: (2023)
by: Becker, Steffen, et al.
Published: (2023)
SECURE: Benchmarking Large Language Models for Cybersecurity
by: Bhusal, Dipkamal, et al.
Published: (2024)
by: Bhusal, Dipkamal, et al.
Published: (2024)
Anti-Sensing: Defense against Unauthorized Radar-based Human Vital Sign Sensing with Physically Realizable Wearable Oscillators
by: Oshim, Md Farhan Tasnim, et al.
Published: (2025)
by: Oshim, Md Farhan Tasnim, et al.
Published: (2025)
Personalised Feedback Framework for Online Education Programmes Using Generative AI
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
by: Kuzminykh, Ievgeniia, et al.
Published: (2024)
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
by: La Torre, Antonio, et al.
Published: (2025)
by: La Torre, Antonio, et al.
Published: (2025)
Towards Scalable Defenses against Intimate Partner Infiltrations
by: Yang, Weisi, et al.
Published: (2025)
by: Yang, Weisi, et al.
Published: (2025)
Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale
by: Rozema, Andrew T., et al.
Published: (2025)
by: Rozema, Andrew T., et al.
Published: (2025)
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
by: Howard, Austin
Published: (2025)
by: Howard, Austin
Published: (2025)
Hacc-Man: An Arcade Game for Jailbreaking LLMs
by: Valentim, Matheus, et al.
Published: (2024)
by: Valentim, Matheus, et al.
Published: (2024)
Emergent misalignment as prompt sensitivity: A research note
by: Wyse, Tim, et al.
Published: (2025)
by: Wyse, Tim, et al.
Published: (2025)
Usability Study of Security Features in Programmable Logic Controllers
by: Li, Karen, et al.
Published: (2022)
by: Li, Karen, et al.
Published: (2022)
Evaluating the Usability of LLMs in Threat Intelligence Enrichment
by: Srikanth, Sanchana, et al.
Published: (2024)
by: Srikanth, Sanchana, et al.
Published: (2024)
Similar Items
-
Preference Poisoning Attacks on Reward Model Learning
by: Wu, Junlin, et al.
Published: (2024) -
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024) -
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
by: Chen, Chaoran, et al.
Published: (2025) -
FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks
by: Wang, Jiongxiao, et al.
Published: (2024) -
Human-in-the-Loop Generation of Adversarial Texts: A Case Study on Tibetan Script
by: Cao, Xi, et al.
Published: (2024)