How Catastrophic is Your LLM? Certifying Risk in Conversation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Chengxiao, Chaudhary, Isha, Hu, Qian, Ruan, Weitong, Gupta, Rahul, Singh, Gagandeep
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910011913928704
author Wang, Chengxiao
Chaudhary, Isha
Hu, Qian
Ruan, Weitong
Gupta, Rahul
Singh, Gagandeep
author_facet Wang, Chengxiao
Chaudhary, Isha
Hu, Qian
Ruan, Weitong
Gupta, Rahul
Singh, Gagandeep
contents Large Language Models (LLMs) can produce catastrophic responses in conversational settings that pose serious risks to public safety and security. Existing evaluations often fail to fully reveal these vulnerabilities because they rely on fixed attack prompt sequences, lack statistical guarantees, and do not scale to the vast space of multi-turn conversations. In this work, we propose C$^3$LLM, a novel, principled statistical Certification framework for Catastrophic risks in multi-turn Conversation for LLMs that bounds the probability of an LLM generating catastrophic responses under multi-turn conversation distributions with statistical guarantees. We model multi-turn conversations as probability distributions over query sequences, represented by a Markov process on a query graph whose edges encode semantic similarity to capture realistic conversational flow, and quantify catastrophic risks using confidence intervals. We define several inexpensive and practical distributions--random node, graph path, and adaptive with rejection. Our results demonstrate that these distributions can reveal substantial catastrophic risks in frontier models, with certified lower bounds as high as 70% for the worst model, highlighting the urgent need for improved safety training strategies in frontier LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03969
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Catastrophic is Your LLM? Certifying Risk in Conversation
Wang, Chengxiao
Chaudhary, Isha
Hu, Qian
Ruan, Weitong
Gupta, Rahul
Singh, Gagandeep
Artificial Intelligence
Cryptography and Security
Machine Learning
Large Language Models (LLMs) can produce catastrophic responses in conversational settings that pose serious risks to public safety and security. Existing evaluations often fail to fully reveal these vulnerabilities because they rely on fixed attack prompt sequences, lack statistical guarantees, and do not scale to the vast space of multi-turn conversations. In this work, we propose C$^3$LLM, a novel, principled statistical Certification framework for Catastrophic risks in multi-turn Conversation for LLMs that bounds the probability of an LLM generating catastrophic responses under multi-turn conversation distributions with statistical guarantees. We model multi-turn conversations as probability distributions over query sequences, represented by a Markov process on a query graph whose edges encode semantic similarity to capture realistic conversational flow, and quantify catastrophic risks using confidence intervals. We define several inexpensive and practical distributions--random node, graph path, and adaptive with rejection. Our results demonstrate that these distributions can reveal substantial catastrophic risks in frontier models, with certified lower bounds as high as 70% for the worst model, highlighting the urgent need for improved safety training strategies in frontier LLMs.
title How Catastrophic is Your LLM? Certifying Risk in Conversation
topic Artificial Intelligence
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2510.03969