Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
Fuente:
arXiv
Saved in:
| Main Authors: | Prandi, Matteo, Suriani, Vincenzo, Pierucci, Federico, Galisai, Marcello, Nardi, Daniele, Bisconti, Piercosma |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
by: Galisai, Marcello, et al.
Published: (2026)
by: Galisai, Marcello, et al.
Published: (2026)
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Metaphor Is Not All Attention Needs
by: Sorokoletova, Olga, et al.
Published: (2026)
by: Sorokoletova, Olga, et al.
Published: (2026)
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Institutional AI: A Governance Framework for Distributional AGI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Standards for trustworthy AI in the European Union: technical rationale, structural challenges, and an implementation path
by: Bisconti, Piercosma, et al.
Published: (2026)
by: Bisconti, Piercosma, et al.
Published: (2026)
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
by: Bisconti, Piercosma, et al.
Published: (2026)
by: Bisconti, Piercosma, et al.
Published: (2026)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
by: Giarrusso, Francesco, et al.
Published: (2025)
by: Giarrusso, Francesco, et al.
Published: (2025)
LLM Based Multi-Agent Generation of Semi-structured Documents from Semantic Templates in the Public Administration Domain
by: Musumeci, Emanuele, et al.
Published: (2024)
by: Musumeci, Emanuele, et al.
Published: (2024)
A Participatory Strategy for AI Ethics in Education and Rehabilitation grounded in the Capability Approach
by: Cesaroni, Valeria, et al.
Published: (2025)
by: Cesaroni, Valeria, et al.
Published: (2025)
Play Everywhere: A Temporal Logic based Game Environment Independent Approach for Playing Soccer with Robots
by: Suriani, Vincenzo, et al.
Published: (2024)
by: Suriani, Vincenzo, et al.
Published: (2024)
Multi Robot Coordination in Highly Dynamic Environments: Tackling Asymmetric Obstacles and Limited Communication
by: Suriani, Vincenzo, et al.
Published: (2025)
by: Suriani, Vincenzo, et al.
Published: (2025)
Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
by: Drid, Abdel Hakim, et al.
Published: (2025)
by: Drid, Abdel Hakim, et al.
Published: (2025)
Can We Trust LLM Detectors?
by: Sandhan, Jivnesh, et al.
Published: (2026)
by: Sandhan, Jivnesh, et al.
Published: (2026)
Chain-of-Programming (CoP) : Empowering Large Language Models for Geospatial Code Generation
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
AI Agents Under EU Law
by: Nannini, Luca, et al.
Published: (2026)
by: Nannini, Luca, et al.
Published: (2026)
EMPOWER: Embodied Multi-role Open-vocabulary Planning with Online Grounding and Execution
by: Argenziano, Francesco, et al.
Published: (2024)
by: Argenziano, Francesco, et al.
Published: (2024)
Multi-Agent Planning Using Visual Language Models
by: Brienza, Michele, et al.
Published: (2024)
by: Brienza, Michele, et al.
Published: (2024)
LLM-REVal: Can We Trust LLM Reviewers Yet?
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
by: Deng, Hexuan, et al.
Published: (2026)
by: Deng, Hexuan, et al.
Published: (2026)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
by: Sorokoletova, Olga E., et al.
Published: (2026)
by: Sorokoletova, Olga E., et al.
Published: (2026)
Boosting Deep Reinforcement Learning with Semantic Knowledge for Robotic Manipulators
by: Güitta-López, Lucía, et al.
Published: (2026)
by: Güitta-López, Lucía, et al.
Published: (2026)
We Can't Understand AI Using our Existing Vocabulary
by: Hewitt, John, et al.
Published: (2025)
by: Hewitt, John, et al.
Published: (2025)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
Dialogue You Can Trust: Human and AI Perspectives on Generated Conversations
by: Ebubechukwu, Ike, et al.
Published: (2024)
by: Ebubechukwu, Ike, et al.
Published: (2024)
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
by: Xiong, Zixin, et al.
Published: (2026)
by: Xiong, Zixin, et al.
Published: (2026)
Shape and Style GAN-based Multispectral Data Augmentation for Crop/Weed Segmentation in Precision Farming
by: Fawakherji, Mulham, et al.
Published: (2024)
by: Fawakherji, Mulham, et al.
Published: (2024)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)
by: Jeon, Hyunjun, et al.
Published: (2026)
PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
by: Zhao, Yimin, et al.
Published: (2026)
by: Zhao, Yimin, et al.
Published: (2026)
LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench
by: Valmeekam, Karthik, et al.
Published: (2024)
by: Valmeekam, Karthik, et al.
Published: (2024)
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
by: Eriksson, Maria, et al.
Published: (2025)
by: Eriksson, Maria, et al.
Published: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
Can We Trust AI to Govern AI? Benchmarking LLM Performance on Privacy and AI Governance Exams
by: Witherspoon, Zane, et al.
Published: (2025)
by: Witherspoon, Zane, et al.
Published: (2025)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
by: Karger, Ezra, et al.
Published: (2024)
by: Karger, Ezra, et al.
Published: (2024)
Context Matters! Relaxing Goals with LLMs for Feasible 3D Scene Planning
by: Musumeci, Emanuele, et al.
Published: (2025)
by: Musumeci, Emanuele, et al.
Published: (2025)
LOST-3DSG: Lightweight Open-Vocabulary 3D Scene Graphs with Semantic Tracking in Dynamic Environments
by: Ferraina, Sara Micol, et al.
Published: (2026)
by: Ferraina, Sara Micol, et al.
Published: (2026)
Similar Items
-
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
by: Galisai, Marcello, et al.
Published: (2026) -
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
by: Bisconti, Piercosma, et al.
Published: (2025) -
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
by: Bisconti, Piercosma, et al.
Published: (2025) -
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026) -
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026)