Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Bang, Zhu, Sicheng, Zhang, Ruiyi, Panaitescu-Liess, Michael-Andrei, Xu, Yuancheng, Huang, Furong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
von: Panaitescu-Liess, Michael-Andrei, et al.
Veröffentlicht: (2025)
von: Panaitescu-Liess, Michael-Andrei, et al.
Veröffentlicht: (2025)
Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
von: Panaitescu-Liess, Michael-Andrei, et al.
Veröffentlicht: (2025)
von: Panaitescu-Liess, Michael-Andrei, et al.
Veröffentlicht: (2025)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
von: Akiri, Charankumar, et al.
Veröffentlicht: (2025)
von: Akiri, Charankumar, et al.
Veröffentlicht: (2025)
Prompt Injection Vulnerability of Consensus Generating Applications in Digital Democracy
von: Gudiño-Rosero, Jairo, et al.
Veröffentlicht: (2025)
von: Gudiño-Rosero, Jairo, et al.
Veröffentlicht: (2025)
What's Privacy Good for? Measuring Privacy as a Shield from Harms due to Personal Data Use
von: Gajavalli, Sri Harsha, et al.
Veröffentlicht: (2025)
von: Gajavalli, Sri Harsha, et al.
Veröffentlicht: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
von: Zhu, Changjia, et al.
Veröffentlicht: (2025)
von: Zhu, Changjia, et al.
Veröffentlicht: (2025)
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
von: An, Bang, et al.
Veröffentlicht: (2023)
von: An, Bang, et al.
Veröffentlicht: (2023)
RealHarm: A Collection of Real-World Language Model Application Failures
von: Jeune, Pierre Le, et al.
Veröffentlicht: (2025)
von: Jeune, Pierre Le, et al.
Veröffentlicht: (2025)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
von: Jiang, Yukun, et al.
Veröffentlicht: (2024)
von: Jiang, Yukun, et al.
Veröffentlicht: (2024)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
von: Liu, Houjun, et al.
Veröffentlicht: (2026)
von: Liu, Houjun, et al.
Veröffentlicht: (2026)
Large Language Models as a (Bad) Security Norm in the Context of Regulation and Compliance
von: Ludvigsen, Kaspar Rosager
Veröffentlicht: (2025)
von: Ludvigsen, Kaspar Rosager
Veröffentlicht: (2025)
WAVES: Benchmarking the Robustness of Image Watermarks
von: An, Bang, et al.
Veröffentlicht: (2024)
von: An, Bang, et al.
Veröffentlicht: (2024)
Prompt injections as a tool for preserving identity in GAI image descriptions
von: Glazko, Kate, et al.
Veröffentlicht: (2025)
von: Glazko, Kate, et al.
Veröffentlicht: (2025)
MM-AttacKG: A Multimodal Approach to Attack Graph Construction with Large Language Models
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2025)
Data Defenses Against Large Language Models
von: Agnew, William, et al.
Veröffentlicht: (2024)
von: Agnew, William, et al.
Veröffentlicht: (2024)
The TCF doesn't really A(A)ID -- Automatic Privacy Analysis and Legal Compliance of TCF-based Android Applications
von: Morel, Victor, et al.
Veröffentlicht: (2026)
von: Morel, Victor, et al.
Veröffentlicht: (2026)
How Susceptible are Large Language Models to Ideological Manipulation?
von: Chen, Kai, et al.
Veröffentlicht: (2024)
von: Chen, Kai, et al.
Veröffentlicht: (2024)
An Investigation into Misuse of Java Security APIs by Large Language Models
von: Mousavi, Zahra, et al.
Veröffentlicht: (2024)
von: Mousavi, Zahra, et al.
Veröffentlicht: (2024)
AI Agents May Always Fall for Prompt Injections
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
von: Burnat, Florian A. D., et al.
Veröffentlicht: (2026)
von: Burnat, Florian A. D., et al.
Veröffentlicht: (2026)
A Proposal for Evaluating the Operational Risk for ChatBots based on Large Language Models
von: Pinacho-Davidson, Pedro, et al.
Veröffentlicht: (2025)
von: Pinacho-Davidson, Pedro, et al.
Veröffentlicht: (2025)
An Evaluation of Chat Safety Moderations in Roblox
von: Kaushik, Priya, et al.
Veröffentlicht: (2026)
von: Kaushik, Priya, et al.
Veröffentlicht: (2026)
Evaluating the Impacts of Swapping on the US Decennial Census
von: Ballesteros, Maria, et al.
Veröffentlicht: (2025)
von: Ballesteros, Maria, et al.
Veröffentlicht: (2025)
Defending Against Intelligent Attackers at Large Scales
von: Lohn, Andrew J.
Veröffentlicht: (2025)
von: Lohn, Andrew J.
Veröffentlicht: (2025)
A Large-Scale Study of Telegram Bots
von: Tsuchiya, Taro, et al.
Veröffentlicht: (2026)
von: Tsuchiya, Taro, et al.
Veröffentlicht: (2026)
Safeguarding Efficacy in Large Language Models: Evaluating Resistance to Human-Written and Algorithmic Adversarial Prompts
von: Downey-Webb, Tiarnaigh, et al.
Veröffentlicht: (2025)
von: Downey-Webb, Tiarnaigh, et al.
Veröffentlicht: (2025)
Automatic Generation of Web Censorship Probe Lists
von: Tang, Jenny, et al.
Veröffentlicht: (2024)
von: Tang, Jenny, et al.
Veröffentlicht: (2024)
Evaluating Privacy Measures in Healthcare Apps Predominantly Used by Older Adults
von: Saka, Suleiman, et al.
Veröffentlicht: (2024)
von: Saka, Suleiman, et al.
Veröffentlicht: (2024)
Analysing Multidisciplinary Approaches to Fight Large-Scale Digital Influence Operations
von: Arroyo, David, et al.
Veröffentlicht: (2025)
von: Arroyo, David, et al.
Veröffentlicht: (2025)
Evaluating Organization Security: User Stories of European Union NIS2 Directive
von: Seeba, Mari, et al.
Veröffentlicht: (2025)
von: Seeba, Mari, et al.
Veröffentlicht: (2025)
Attacks on Third-Party APIs of Large Language Models
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
von: Krishna, Satyapriya, et al.
Veröffentlicht: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
INR-Based Generative Steganography by Point Cloud Representation
von: Yangjie, Zhong, et al.
Veröffentlicht: (2024)
von: Yangjie, Zhong, et al.
Veröffentlicht: (2024)
Abusing the Internet of Medical Things: Evaluating Threat Models and Forensic Readiness for Multi-Vector Attacks on Connected Healthcare Devices
von: Straw, Isabel, et al.
Veröffentlicht: (2026)
von: Straw, Isabel, et al.
Veröffentlicht: (2026)
Proof of Authenticity of General IoT Information with Tamper-Evident Sensors and Blockchain
von: Saito, Kenji
Veröffentlicht: (2025)
von: Saito, Kenji
Veröffentlicht: (2025)
A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
von: Hong, Rachel, et al.
Veröffentlicht: (2025)
von: Hong, Rachel, et al.
Veröffentlicht: (2025)
Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?
von: Xu, Naen, et al.
Veröffentlicht: (2025)
von: Xu, Naen, et al.
Veröffentlicht: (2025)
Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-assisted Instruction
von: Patel, Karan, et al.
Veröffentlicht: (2025)
von: Patel, Karan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
von: Panaitescu-Liess, Michael-Andrei, et al.
Veröffentlicht: (2025) -
Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
von: Panaitescu-Liess, Michael-Andrei, et al.
Veröffentlicht: (2025) -
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
von: Akiri, Charankumar, et al.
Veröffentlicht: (2025) -
Prompt Injection Vulnerability of Consensus Generating Applications in Digital Democracy
von: Gudiño-Rosero, Jairo, et al.
Veröffentlicht: (2025) -
What's Privacy Good for? Measuring Privacy as a Shield from Harms due to Personal Data Use
von: Gajavalli, Sri Harsha, et al.
Veröffentlicht: (2025)