From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bisconti, Piercosma, Galisai, Marcello, Prandi, Matteo, Pierucci, Federico, Sorokoletova, Olga, Giarrusso, Francesco, Suriani, Vincenzo, Syrnikov, Marcantonio Bracale, Nardi, Daniele |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2025)
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2025)
Institutional AI: A Governance Framework for Distributional AGI Safety
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
von: Syrnikov, Marcantonio Bracale, et al.
Veröffentlicht: (2026)
von: Syrnikov, Marcantonio Bracale, et al.
Veröffentlicht: (2026)
Agentic Microphysics: A Manifesto for Generative AI Safety
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
von: Pierucci, Federico, et al.
Veröffentlicht: (2026)
Metaphor Is Not All Attention Needs
von: Sorokoletova, Olga, et al.
Veröffentlicht: (2026)
von: Sorokoletova, Olga, et al.
Veröffentlicht: (2026)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
von: Prandi, Matteo, et al.
Veröffentlicht: (2025)
von: Prandi, Matteo, et al.
Veröffentlicht: (2025)
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
von: Galisai, Marcello, et al.
Veröffentlicht: (2026)
von: Galisai, Marcello, et al.
Veröffentlicht: (2026)
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2025)
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2025)
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2026)
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2026)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
von: Sorokoletova, Olga E., et al.
Veröffentlicht: (2026)
von: Sorokoletova, Olga E., et al.
Veröffentlicht: (2026)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
von: Giarrusso, Francesco, et al.
Veröffentlicht: (2025)
von: Giarrusso, Francesco, et al.
Veröffentlicht: (2025)
Standards for trustworthy AI in the European Union: technical rationale, structural challenges, and an implementation path
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2026)
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2026)
Shape and Style GAN-based Multispectral Data Augmentation for Crop/Weed Segmentation in Precision Farming
von: Fawakherji, Mulham, et al.
Veröffentlicht: (2024)
von: Fawakherji, Mulham, et al.
Veröffentlicht: (2024)
LLM Based Multi-Agent Generation of Semi-structured Documents from Semantic Templates in the Public Administration Domain
von: Musumeci, Emanuele, et al.
Veröffentlicht: (2024)
von: Musumeci, Emanuele, et al.
Veröffentlicht: (2024)
A Participatory Strategy for AI Ethics in Education and Rehabilitation grounded in the Capability Approach
von: Cesaroni, Valeria, et al.
Veröffentlicht: (2025)
von: Cesaroni, Valeria, et al.
Veröffentlicht: (2025)
Self-supervised Feature Extraction for Enhanced Ball Detection on Soccer Robots
von: Lin, Can, et al.
Veröffentlicht: (2025)
von: Lin, Can, et al.
Veröffentlicht: (2025)
Real-Time Multimodal Signal Processing for HRI in RoboCup: Understanding a Human Referee
von: Ansalone, Filippo, et al.
Veröffentlicht: (2024)
von: Ansalone, Filippo, et al.
Veröffentlicht: (2024)
Towards a scalable AI-driven framework for data-independent Cyber Threat Intelligence Information Extraction
von: Sorokoletova, Olga, et al.
Veröffentlicht: (2025)
von: Sorokoletova, Olga, et al.
Veröffentlicht: (2025)
Play Everywhere: A Temporal Logic based Game Environment Independent Approach for Playing Soccer with Robots
von: Suriani, Vincenzo, et al.
Veröffentlicht: (2024)
von: Suriani, Vincenzo, et al.
Veröffentlicht: (2024)
Multi Robot Coordination in Highly Dynamic Environments: Tackling Asymmetric Obstacles and Limited Communication
von: Suriani, Vincenzo, et al.
Veröffentlicht: (2025)
von: Suriani, Vincenzo, et al.
Veröffentlicht: (2025)
EMPOWER: Embodied Multi-role Open-vocabulary Planning with Online Grounding and Execution
von: Argenziano, Francesco, et al.
Veröffentlicht: (2024)
von: Argenziano, Francesco, et al.
Veröffentlicht: (2024)
Multi-Agent Planning Using Visual Language Models
von: Brienza, Michele, et al.
Veröffentlicht: (2024)
von: Brienza, Michele, et al.
Veröffentlicht: (2024)
Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
von: Drid, Abdel Hakim, et al.
Veröffentlicht: (2025)
von: Drid, Abdel Hakim, et al.
Veröffentlicht: (2025)
Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
von: Trippodo, Matteo, et al.
Veröffentlicht: (2025)
von: Trippodo, Matteo, et al.
Veröffentlicht: (2025)
Defining and Monitoring Complex Robot Activities via LLMs and Symbolic Reasoning
von: Argenziano, Francesco, et al.
Veröffentlicht: (2025)
von: Argenziano, Francesco, et al.
Veröffentlicht: (2025)
Accurate generation of stochastic dynamics based on multi-model Generative Adversarial Networks
von: Lanzoni, Daniele, et al.
Veröffentlicht: (2023)
von: Lanzoni, Daniele, et al.
Veröffentlicht: (2023)
Adversarial Camouflage
von: Borsukiewicz, Paweł, et al.
Veröffentlicht: (2026)
von: Borsukiewicz, Paweł, et al.
Veröffentlicht: (2026)
Online Price Competition under Generalized Linear Demands
von: Bracale, Daniele, et al.
Veröffentlicht: (2025)
von: Bracale, Daniele, et al.
Veröffentlicht: (2025)
Secure Diagnostics: Adversarial Robustness Meets Clinical Interpretability
von: Najafi, Mohammad Hossein, et al.
Veröffentlicht: (2025)
von: Najafi, Mohammad Hossein, et al.
Veröffentlicht: (2025)
$σ$-zero: Gradient-based Optimization of $\ell_0$-norm Adversarial Examples
von: Cinà, Antonio Emanuele, et al.
Veröffentlicht: (2024)
von: Cinà, Antonio Emanuele, et al.
Veröffentlicht: (2024)
Multi-Agent Coordination for a Partially Observable and Dynamic Robot Soccer Environment with Limited Communication
von: Affinita, Daniele, et al.
Veröffentlicht: (2024)
von: Affinita, Daniele, et al.
Veröffentlicht: (2024)
Interpretability-Guided Test-Time Adversarial Defense
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2024)
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2024)
Learning Kinetic Monte Carlo stochastic dynamics with Deep Generative Adversarial Networks
von: Lanzoni, Daniele, et al.
Veröffentlicht: (2025)
von: Lanzoni, Daniele, et al.
Veröffentlicht: (2025)
DOOM Level Generation using Generative Adversarial Networks
von: Giacomello, Edoardo, et al.
Veröffentlicht: (2018)
von: Giacomello, Edoardo, et al.
Veröffentlicht: (2018)
Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection
von: De Rosa, Vincenzo, et al.
Veröffentlicht: (2024)
von: De Rosa, Vincenzo, et al.
Veröffentlicht: (2024)
AI Agents Under EU Law
von: Nannini, Luca, et al.
Veröffentlicht: (2026)
von: Nannini, Luca, et al.
Veröffentlicht: (2026)
Human-Interpretable Adversarial Prompt Attack on Large Language Models with Situational Context
von: Das, Nilanjana, et al.
Veröffentlicht: (2024)
von: Das, Nilanjana, et al.
Veröffentlicht: (2024)
Structured Gradient-based Interpretations via Norm-Regularized Adversarial Training
von: Gong, Shizhan, et al.
Veröffentlicht: (2024)
von: Gong, Shizhan, et al.
Veröffentlicht: (2024)
Causal Interpretability for Adversarial Robustness: A Hybrid Generative Classification Approach
von: Zhao, Chunheng, et al.
Veröffentlicht: (2024)
von: Zhao, Chunheng, et al.
Veröffentlicht: (2024)
Adversarial Doodles: Interpretable and Human-drawable Attacks Provide Describable Insights
von: Nara, Ryoya, et al.
Veröffentlicht: (2023)
von: Nara, Ryoya, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2025) -
Institutional AI: A Governance Framework for Distributional AGI Safety
von: Pierucci, Federico, et al.
Veröffentlicht: (2026) -
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
von: Syrnikov, Marcantonio Bracale, et al.
Veröffentlicht: (2026) -
Agentic Microphysics: A Manifesto for Generative AI Safety
von: Pierucci, Federico, et al.
Veröffentlicht: (2026) -
Metaphor Is Not All Attention Needs
von: Sorokoletova, Olga, et al.
Veröffentlicht: (2026)