Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bisconti, Piercosma, Prandi, Matteo, Pierucci, Federico, Giarrusso, Francesco, Syrnikov, Marcantonio Bracale, Galisai, Marcello, Suriani, Vincenzo, Sorokoletova, Olga, Sartore, Federico, Nardi, Daniele |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Institutional AI: A Governance Framework for Distributional AGI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
by: Galisai, Marcello, et al.
Published: (2026)
by: Galisai, Marcello, et al.
Published: (2026)
Metaphor Is Not All Attention Needs
by: Sorokoletova, Olga, et al.
Published: (2026)
by: Sorokoletova, Olga, et al.
Published: (2026)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
by: Prandi, Matteo, et al.
Published: (2025)
by: Prandi, Matteo, et al.
Published: (2025)
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
by: Bisconti, Piercosma, et al.
Published: (2026)
by: Bisconti, Piercosma, et al.
Published: (2026)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
by: Giarrusso, Francesco, et al.
Published: (2025)
by: Giarrusso, Francesco, et al.
Published: (2025)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
by: Sorokoletova, Olga E., et al.
Published: (2026)
by: Sorokoletova, Olga E., et al.
Published: (2026)
Standards for trustworthy AI in the European Union: technical rationale, structural challenges, and an implementation path
by: Bisconti, Piercosma, et al.
Published: (2026)
by: Bisconti, Piercosma, et al.
Published: (2026)
Play Everywhere: A Temporal Logic based Game Environment Independent Approach for Playing Soccer with Robots
by: Suriani, Vincenzo, et al.
Published: (2024)
by: Suriani, Vincenzo, et al.
Published: (2024)
Shape and Style GAN-based Multispectral Data Augmentation for Crop/Weed Segmentation in Precision Farming
by: Fawakherji, Mulham, et al.
Published: (2024)
by: Fawakherji, Mulham, et al.
Published: (2024)
Multi Robot Coordination in Highly Dynamic Environments: Tackling Asymmetric Obstacles and Limited Communication
by: Suriani, Vincenzo, et al.
Published: (2025)
by: Suriani, Vincenzo, et al.
Published: (2025)
EMPOWER: Embodied Multi-role Open-vocabulary Planning with Online Grounding and Execution
by: Argenziano, Francesco, et al.
Published: (2024)
by: Argenziano, Francesco, et al.
Published: (2024)
Multi-Agent Planning Using Visual Language Models
by: Brienza, Michele, et al.
Published: (2024)
by: Brienza, Michele, et al.
Published: (2024)
Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
by: Drid, Abdel Hakim, et al.
Published: (2025)
by: Drid, Abdel Hakim, et al.
Published: (2025)
LLM Based Multi-Agent Generation of Semi-structured Documents from Semantic Templates in the Public Administration Domain
by: Musumeci, Emanuele, et al.
Published: (2024)
by: Musumeci, Emanuele, et al.
Published: (2024)
A Participatory Strategy for AI Ethics in Education and Rehabilitation grounded in the Capability Approach
by: Cesaroni, Valeria, et al.
Published: (2025)
by: Cesaroni, Valeria, et al.
Published: (2025)
Multi-Agent Coordination for a Partially Observable and Dynamic Robot Soccer Environment with Limited Communication
by: Affinita, Daniele, et al.
Published: (2024)
by: Affinita, Daniele, et al.
Published: (2024)
LLCoach: Generating Robot Soccer Plans using Multi-Role Large Language Models
by: Brienza, Michele, et al.
Published: (2024)
by: Brienza, Michele, et al.
Published: (2024)
Boosting Deep Reinforcement Learning with Semantic Knowledge for Robotic Manipulators
by: Güitta-López, Lucía, et al.
Published: (2026)
by: Güitta-López, Lucía, et al.
Published: (2026)
Self-supervised Feature Extraction for Enhanced Ball Detection on Soccer Robots
by: Lin, Can, et al.
Published: (2025)
by: Lin, Can, et al.
Published: (2025)
Context Matters! Relaxing Goals with LLMs for Feasible 3D Scene Planning
by: Musumeci, Emanuele, et al.
Published: (2025)
by: Musumeci, Emanuele, et al.
Published: (2025)
LOST-3DSG: Lightweight Open-Vocabulary 3D Scene Graphs with Semantic Tracking in Dynamic Environments
by: Ferraina, Sara Micol, et al.
Published: (2026)
by: Ferraina, Sara Micol, et al.
Published: (2026)
Enhancing robot reliability for health-care facilities by means of Human-Aware Navigation Planning
by: Sorokoletova, Olga E., et al.
Published: (2024)
by: Sorokoletova, Olga E., et al.
Published: (2024)
Equilibrium and Pricing in Consumer Networks with Nonlinear Utilities: An Online Shape-Constrained Learning Approach
by: Bracale, Daniele, et al.
Published: (2026)
by: Bracale, Daniele, et al.
Published: (2026)
Towards a scalable AI-driven framework for data-independent Cyber Threat Intelligence Information Extraction
by: Sorokoletova, Olga, et al.
Published: (2025)
by: Sorokoletova, Olga, et al.
Published: (2025)
R2F: Repurposing Ray Frontiers for LLM-free Object Navigation
by: Argenziano, Francesco, et al.
Published: (2026)
by: Argenziano, Francesco, et al.
Published: (2026)
Real-Time Multimodal Signal Processing for HRI in RoboCup: Understanding a Human Referee
by: Ansalone, Filippo, et al.
Published: (2024)
by: Ansalone, Filippo, et al.
Published: (2024)
Descripción morfológica de amebas testadas del género Arcella sp., a través del uso de técnicas multiples
by: Karen Bracale
Published: (2019)
by: Karen Bracale
Published: (2019)
Many-Turn Jailbreaking
by: Yang, Xianjun, et al.
Published: (2025)
by: Yang, Xianjun, et al.
Published: (2025)
Do Finite Density Effects Jeopardize Axion Nucleophobia in Supernovae?
by: Di Luzio, Luca, et al.
Published: (2024)
by: Di Luzio, Luca, et al.
Published: (2024)
íCuantos inconvenientes!
by: Sartore Joel
Published: (2000)
by: Sartore Joel
Published: (2000)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
by: Reddy, Aashray, et al.
Published: (2025)
by: Reddy, Aashray, et al.
Published: (2025)
Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
by: Trippodo, Matteo, et al.
Published: (2025)
by: Trippodo, Matteo, et al.
Published: (2025)
Fatty acid amide hydrolase (FAAH) and the endocannabinoid system in obesity: Mechanistic insights and pharmacological opportunities beyond incretin‐based therapies
by: Ilaria Serra, et al.
Published: (2026)
by: Ilaria Serra, et al.
Published: (2026)
Vanishing viscosity limit for the compressible Navier-Stokes equations with non-linear density dependent viscosities
by: Bisconti, Luca, et al.
Published: (2024)
by: Bisconti, Luca, et al.
Published: (2024)
Similar Items
-
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
by: Bisconti, Piercosma, et al.
Published: (2025) -
Institutional AI: A Governance Framework for Distributional AGI Safety
by: Pierucci, Federico, et al.
Published: (2026) -
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026) -
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026) -
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
by: Bisconti, Piercosma, et al.
Published: (2025)