Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
Fuente:
arXiv
Saved in:
| Main Authors: | Yamamura, Atsushi, Ganguli, Surya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Dagger Behind Smile: Fool LLMs with a Happy Ending Story
by: Song, Xurui, et al.
Published: (2025)
by: Song, Xurui, et al.
Published: (2025)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
by: Yao, Yang, et al.
Published: (2025)
by: Yao, Yang, et al.
Published: (2025)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
by: Timmer, Roelien C., et al.
Published: (2024)
by: Timmer, Roelien C., et al.
Published: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
by: Saiem, Bijoy Ahmed, et al.
Published: (2024)
Emergent misalignment as prompt sensitivity: A research note
by: Wyse, Tim, et al.
Published: (2025)
by: Wyse, Tim, et al.
Published: (2025)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
Goal-guided Generative Prompt Injection Attack on Large Language Models
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
Adversarial Attacks Against Automated Fact-Checking: A Survey
by: Liu, Fanzhen, et al.
Published: (2025)
by: Liu, Fanzhen, et al.
Published: (2025)
LLM in the Shell: Generative Honeypots
by: Sladić, Muris, et al.
Published: (2023)
by: Sladić, Muris, et al.
Published: (2023)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Contextualized Privacy Defense for LLM Agents
by: Wen, Yule, et al.
Published: (2026)
by: Wen, Yule, et al.
Published: (2026)
LLM Jailbreak Detection for (Almost) Free!
by: Chen, Guorui, et al.
Published: (2025)
by: Chen, Guorui, et al.
Published: (2025)
OneShield -- the Next Generation of LLM Guardrails
by: DeLuca, Chad, et al.
Published: (2025)
by: DeLuca, Chad, et al.
Published: (2025)
RedSage: A Cybersecurity Generalist LLM
by: Suryanto, Naufal, et al.
Published: (2026)
by: Suryanto, Naufal, et al.
Published: (2026)
FLAME: Flexible LLM-Assisted Moderation Engine
by: Bakulin, Ivan, et al.
Published: (2025)
by: Bakulin, Ivan, et al.
Published: (2025)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
by: Wu, Tianyi, et al.
Published: (2025)
by: Wu, Tianyi, et al.
Published: (2025)
LLM for SoC Security: A Paradigm Shift
by: Saha, Dipayan, et al.
Published: (2023)
by: Saha, Dipayan, et al.
Published: (2023)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
by: Gu, Chenxi, et al.
Published: (2026)
by: Gu, Chenxi, et al.
Published: (2026)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
by: Das, Saswat, et al.
Published: (2025)
by: Das, Saswat, et al.
Published: (2025)
NeuroFilter: Privacy Guardrails for Conversational LLM Agents
by: Das, Saswat, et al.
Published: (2026)
by: Das, Saswat, et al.
Published: (2026)
Searching for Privacy Risks in LLM Agents via Simulation
by: Zhang, Yanzhe, et al.
Published: (2025)
by: Zhang, Yanzhe, et al.
Published: (2025)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Large Language Model Sentinel: LLM Agent for Adversarial Purification
by: Lin, Guang, et al.
Published: (2024)
by: Lin, Guang, et al.
Published: (2024)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024)
by: Liang, Jiashuo, et al.
Published: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
by: Yu, Miao, et al.
Published: (2024)
by: Yu, Miao, et al.
Published: (2024)
Dynamic Fog Computing for Enhanced LLM Execution in Medical Applications
by: Zagar, Philipp, et al.
Published: (2024)
by: Zagar, Philipp, et al.
Published: (2024)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
by: Ding, Ruomeng, et al.
Published: (2026)
by: Ding, Ruomeng, et al.
Published: (2026)
XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts
by: Xu, Jiahao, et al.
Published: (2026)
by: Xu, Jiahao, et al.
Published: (2026)
Token-level Data Selection for Safe LLM Fine-tuning
by: Li, Yanping, et al.
Published: (2026)
by: Li, Yanping, et al.
Published: (2026)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
by: Jiang, Bo
Published: (2026)
by: Jiang, Bo
Published: (2026)
Exfiltration of personal information from ChatGPT via prompt injection
by: Schwartzman, Gregory
Published: (2024)
by: Schwartzman, Gregory
Published: (2024)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
by: Li, Xirui, et al.
Published: (2024)
by: Li, Xirui, et al.
Published: (2024)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
by: Agarwal, Divyansh, et al.
Published: (2024)
by: Agarwal, Divyansh, et al.
Published: (2024)
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
by: Wang, Liwen, et al.
Published: (2025)
by: Wang, Liwen, et al.
Published: (2025)
Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling
by: Llewellyn, Mary, et al.
Published: (2025)
by: Llewellyn, Mary, et al.
Published: (2025)
Similar Items
-
Are aligned neural networks adversarially aligned?
by: Carlini, Nicholas, et al.
Published: (2023) -
Dagger Behind Smile: Fool LLMs with a Happy Ending Story
by: Song, Xurui, et al.
Published: (2025) -
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
by: Yao, Yang, et al.
Published: (2025) -
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
by: Nakka, Krishna Kanth, et al.
Published: (2024) -
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
by: Timmer, Roelien C., et al.
Published: (2024)