Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rashid, Md Rafi Ur, Dasu, Vishnu Asutosh, Wang, Ye, Tan, Gang, Mehnaz, Shagufta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915619573596160
author Rashid, Md Rafi Ur
Dasu, Vishnu Asutosh
Wang, Ye
Tan, Gang
Mehnaz, Shagufta
author_facet Rashid, Md Rafi Ur
Dasu, Vishnu Asutosh
Wang, Ye
Tan, Gang
Mehnaz, Shagufta
contents Large Language Models (LLMs) exhibit impressive capabilities, but remain susceptible to a growing spectrum of safety risks, including jailbreaks, toxic content, hallucinations, and bias. Existing defenses often address only a single threat type or resort to rigid outright rejection, sacrificing user experience and failing to generalize across diverse and novel attacks. This paper introduces Adversarial Scenario Extrapolation (ASE), a novel inference-time computation framework that leverages Chain-of-Thought (CoT) reasoning to simultaneously enhance LLM robustness and seamlessness. ASE guides the LLM through a self-generative process of contemplating potential adversarial scenarios and formulating defensive strategies before generating a response to the user query. Comprehensive evaluation on four adversarial benchmarks with four latest LLMs shows that ASE achieves near-zero jailbreak attack success rates and minimal toxicity, while slashing outright rejections to <4%. ASE outperforms six state-of-the-art defenses in robustness-seamlessness trade-offs, with 92-99% accuracy on adversarial Q&A and 4-10x lower bias scores. By transforming adversarial perception into an intrinsic cognitive process, ASE sets a new paradigm for secure and natural human-AI interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17089
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
Rashid, Md Rafi Ur
Dasu, Vishnu Asutosh
Wang, Ye
Tan, Gang
Mehnaz, Shagufta
Computation and Language
Large Language Models (LLMs) exhibit impressive capabilities, but remain susceptible to a growing spectrum of safety risks, including jailbreaks, toxic content, hallucinations, and bias. Existing defenses often address only a single threat type or resort to rigid outright rejection, sacrificing user experience and failing to generalize across diverse and novel attacks. This paper introduces Adversarial Scenario Extrapolation (ASE), a novel inference-time computation framework that leverages Chain-of-Thought (CoT) reasoning to simultaneously enhance LLM robustness and seamlessness. ASE guides the LLM through a self-generative process of contemplating potential adversarial scenarios and formulating defensive strategies before generating a response to the user query. Comprehensive evaluation on four adversarial benchmarks with four latest LLMs shows that ASE achieves near-zero jailbreak attack success rates and minimal toxicity, while slashing outright rejections to <4%. ASE outperforms six state-of-the-art defenses in robustness-seamlessness trade-offs, with 92-99% accuracy on adversarial Q&A and 4-10x lower bias scores. By transforming adversarial perception into an intrinsic cognitive process, ASE sets a new paradigm for secure and natural human-AI interaction.
title Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
topic Computation and Language
url https://arxiv.org/abs/2505.17089