Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
Fuente:
arXiv
Saved in:
| Main Authors: | Schulhoff, Sander, Pinto, Jeremy, Khan, Anaum, Bouchard, Louis-François, Si, Chenglei, Anati, Svetlina, Tagliabue, Valen, Kost, Anson Liu, Carnahan, Christopher, Boyd-Graber, Jordan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
by: Tagliabue, Valen, et al.
Published: (2025)
by: Tagliabue, Valen, et al.
Published: (2025)
Prompt-Hacking: The New p-Hacking?
by: Kosch, Thomas, et al.
Published: (2025)
by: Kosch, Thomas, et al.
Published: (2025)
Hacking Predictors Means Hacking Cars: Using Sensitivity Analysis to Identify Trajectory Prediction Vulnerabilities for Autonomous Driving Security
by: Gibson, Marsalis, et al.
Published: (2024)
by: Gibson, Marsalis, et al.
Published: (2024)
Ethical Hacking
by: Maurushat, Alana
Published: (2024)
by: Maurushat, Alana
Published: (2024)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
ShadowHack: Hacking Shadows via Luminance-Color Divide and Conquer
by: Hu, Jin, et al.
Published: (2024)
by: Hu, Jin, et al.
Published: (2024)
Landau Damping of Collective Neutrino Oscillation Waves
by: Kost, Anson, et al.
Published: (2026)
by: Kost, Anson, et al.
Published: (2026)
Local-equilibrium theory of neutrino oscillations
by: Johns, Lucas, et al.
Published: (2025)
by: Johns, Lucas, et al.
Published: (2025)
Ethical Hacking and Cybersecurity
by: Mr. Prasadu Gurram Mrs Veena Dhavalgi Mrs. K. Sundareswari Dr. Suraya Mubeen
Published: (2025)
by: Mr. Prasadu Gurram Mrs Veena Dhavalgi Mrs. K. Sundareswari Dr. Suraya Mubeen
Published: (2025)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
by: Taylor, Mia, et al.
Published: (2025)
by: Taylor, Mia, et al.
Published: (2025)
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
by: Ren, Xiaoxue, et al.
Published: (2025)
by: Ren, Xiaoxue, et al.
Published: (2025)
Inferring Discussion Topics about Exploitation of Vulnerabilities from Underground Hacking Forums
by: Moreno-Vera, Felipe
Published: (2024)
by: Moreno-Vera, Felipe
Published: (2024)
Hacking CTFs with Plain Agents
by: Turtayev, Rustem, et al.
Published: (2024)
by: Turtayev, Rustem, et al.
Published: (2024)
Defining and Characterizing Reward Hacking
by: Skalse, Joar, et al.
Published: (2022)
by: Skalse, Joar, et al.
Published: (2022)
Hacking Gender and Technology in Journalism
by: De Vuyst, Sara
Published: (2020)
by: De Vuyst, Sara
Published: (2020)
Hacking the system: Open access
by: Abraham San Pedro Salazar
Published: (2018)
by: Abraham San Pedro Salazar
Published: (2018)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)
by: Si, Chenglei, et al.
Published: (2023)
Ultra-Fast Wireless Power Hacking
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
The Power of Tests for Detecting $p$-Hacking
by: Elliott, Graham, et al.
Published: (2022)
by: Elliott, Graham, et al.
Published: (2022)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Explanation Hacking: The perils of algorithmic recourse
by: Sullivan, Emily, et al.
Published: (2024)
by: Sullivan, Emily, et al.
Published: (2024)
Hacking Task Confounder in Meta-Learning
by: Wang, Jingyao, et al.
Published: (2023)
by: Wang, Jingyao, et al.
Published: (2023)
Ethical Hacking and its role in Cybersecurity
by: Asif, Fatima, et al.
Published: (2024)
by: Asif, Fatima, et al.
Published: (2024)
Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents
by: Jeurissen, Dominik, et al.
Published: (2024)
by: Jeurissen, Dominik, et al.
Published: (2024)
Once-in-a-lifetime encounter models for neutrino media: From coherent oscillations to flavor equilibration
by: Kost, Anson, et al.
Published: (2024)
by: Kost, Anson, et al.
Published: (2024)
Once-in-a-lifetime encounter models for neutrino media II: Quasi-steady states and miscidynamic flavor evolution
by: Kost, Anson, et al.
Published: (2025)
by: Kost, Anson, et al.
Published: (2025)
Do Synthetic Trajectories Reflect Real Reward Hacking? A Systematic Study on Monitoring In-the-Wild Hacking in Code Generation
by: Li, Lichen, et al.
Published: (2026)
by: Li, Lichen, et al.
Published: (2026)
EvilGenie: A Reward Hacking Benchmark
by: Gabor, Jonathan, et al.
Published: (2025)
by: Gabor, Jonathan, et al.
Published: (2025)
LLM Agents can Autonomously Hack Websites
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
Spontaneous Reward Hacking in Iterative Self-Refinement
by: Pan, Jane, et al.
Published: (2024)
by: Pan, Jane, et al.
Published: (2024)
Hacking, The Lazy Way: LLM Augmented Pentesting
by: Goyal, Dhruva, et al.
Published: (2024)
by: Goyal, Dhruva, et al.
Published: (2024)
Hacking a surrogate model approach to XAI
by: Wilhelm, Alexander, et al.
Published: (2024)
by: Wilhelm, Alexander, et al.
Published: (2024)
Hacking quantum computers with row hammer attack
by: Almaguer-Angeles, Fernando, et al.
Published: (2025)
by: Almaguer-Angeles, Fernando, et al.
Published: (2025)
You Have Been Hacked—Now What
by: Rick Williams
Published: (2026)
by: Rick Williams
Published: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Reward Hacking as Equilibrium under Finite Evaluation
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
Hacked in Translation -- from Subtitles to Complete Takeover
by: Herscovici, Omri, et al.
Published: (2024)
by: Herscovici, Omri, et al.
Published: (2024)
Hacking Cryptographic Protocols with Tensor Network Attacks
by: Aizpurua, Borja, et al.
Published: (2024)
by: Aizpurua, Borja, et al.
Published: (2024)
X Hacking: The Threat of Misguided AutoML
by: Sharma, Rahul, et al.
Published: (2024)
by: Sharma, Rahul, et al.
Published: (2024)
Similar Items
-
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
by: Tagliabue, Valen, et al.
Published: (2025) -
Prompt-Hacking: The New p-Hacking?
by: Kosch, Thomas, et al.
Published: (2025) -
Hacking Predictors Means Hacking Cars: Using Sensitivity Analysis to Identify Trajectory Prediction Vulnerabilities for Autonomous Driving Security
by: Gibson, Marsalis, et al.
Published: (2024) -
Ethical Hacking
by: Maurushat, Alana
Published: (2024) -
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)