ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
Fuente:
arXiv
Salvato in:
| Autori principali: | Jiang, Fengqing, Xu, Zhangchen, Niu, Luyao, Lin, Bill Yuchen, Poovendran, Radha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
di: Li, Yuetai, et al.
Pubblicazione: (2024)
di: Li, Yuetai, et al.
Pubblicazione: (2024)
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
Stronger Models are NOT Stronger Teachers for Instruction Tuning
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
di: Xu, Zhangchen, et al.
Pubblicazione: (2024)
Temporal Sampling for Forgotten Reasoning in LLMs
di: Li, Yuetai, et al.
Pubblicazione: (2025)
di: Li, Yuetai, et al.
Pubblicazione: (2025)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
di: Xu, Zhangchen, et al.
Pubblicazione: (2025)
di: Xu, Zhangchen, et al.
Pubblicazione: (2025)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
di: Jiang, Fengqing, et al.
Pubblicazione: (2024)
di: Jiang, Fengqing, et al.
Pubblicazione: (2024)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
di: Feng, Yichen, et al.
Pubblicazione: (2025)
di: Feng, Yichen, et al.
Pubblicazione: (2025)
Exploring Backdoor Vulnerabilities of Chat Models
di: Hao, Yunzhuo, et al.
Pubblicazione: (2024)
di: Hao, Yunzhuo, et al.
Pubblicazione: (2024)
Double-Dip: Thwarting Label-Only Membership Inference Attacks with Transfer Learning and Randomization
di: Rajabi, Arezoo, et al.
Pubblicazione: (2024)
di: Rajabi, Arezoo, et al.
Pubblicazione: (2024)
TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
di: Shen, Qingchao, et al.
Pubblicazione: (2026)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
Small Models Struggle to Learn from Strong Reasoners
di: Li, Yuetai, et al.
Pubblicazione: (2025)
di: Li, Yuetai, et al.
Pubblicazione: (2025)
BugWhisperer: Fine-Tuning LLMs for SoC Hardware Vulnerability Detection
di: Tarek, Shams, et al.
Pubblicazione: (2025)
di: Tarek, Shams, et al.
Pubblicazione: (2025)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
di: Sahabandu, Dinuka, et al.
Pubblicazione: (2024)
di: Sahabandu, Dinuka, et al.
Pubblicazione: (2024)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
BugSweeper: Function-Level Detection of Smart Contract Vulnerabilities Using Graph Neural Networks
di: Lee, Uisang, et al.
Pubblicazione: (2025)
di: Lee, Uisang, et al.
Pubblicazione: (2025)
Exploring ChatGPT's Capabilities on Vulnerability Management
di: Liu, Peiyu, et al.
Pubblicazione: (2023)
di: Liu, Peiyu, et al.
Pubblicazione: (2023)
Clustering-Enhanced Domain Adaptation for Cross-Domain Intrusion Detection in Industrial Control Systems
di: Wang, Luyao
Pubblicazione: (2026)
di: Wang, Luyao
Pubblicazione: (2026)
An LLM Framework For Cryptography Over Chat Channels
di: Gligoroski, Danilo, et al.
Pubblicazione: (2025)
di: Gligoroski, Danilo, et al.
Pubblicazione: (2025)
Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense
di: Zhang, Jiawen, et al.
Pubblicazione: (2025)
di: Zhang, Jiawen, et al.
Pubblicazione: (2025)
EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection
di: Liu, Shigang, et al.
Pubblicazione: (2024)
di: Liu, Shigang, et al.
Pubblicazione: (2024)
Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection
di: Pushkar, Chinmay, et al.
Pubblicazione: (2025)
di: Pushkar, Chinmay, et al.
Pubblicazione: (2025)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
di: Lai, Zhenglin, et al.
Pubblicazione: (2025)
di: Lai, Zhenglin, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for Security Bug Report Prediction
di: Soltaniani, Farnaz, et al.
Pubblicazione: (2026)
di: Soltaniani, Farnaz, et al.
Pubblicazione: (2026)
SARN: Structurally-Aware Recurrent Network for Spatio-Temporal Disaggregation
di: Han, Bin, et al.
Pubblicazione: (2023)
di: Han, Bin, et al.
Pubblicazione: (2023)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
di: Peng, Benji, et al.
Pubblicazione: (2024)
di: Peng, Benji, et al.
Pubblicazione: (2024)
AbuseGPT: Abuse of Generative AI ChatBots to Create Smishing Campaigns
di: Shibli, Ashfak Md, et al.
Pubblicazione: (2024)
di: Shibli, Ashfak Md, et al.
Pubblicazione: (2024)
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
di: Yang, Xianglin, et al.
Pubblicazione: (2026)
di: Yang, Xianglin, et al.
Pubblicazione: (2026)
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
di: Li, Xinyu, et al.
Pubblicazione: (2025)
di: Li, Xinyu, et al.
Pubblicazione: (2025)
LLMpatronous: Harnessing the Power of LLMs For Vulnerability Detection
di: Yarra, Rajesh
Pubblicazione: (2025)
di: Yarra, Rajesh
Pubblicazione: (2025)
ChatGPT's Potential in Cryptography Misuse Detection: A Comparative Analysis with Static Analysis Tools
di: Firouzi, Ehsan, et al.
Pubblicazione: (2024)
di: Firouzi, Ehsan, et al.
Pubblicazione: (2024)
Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments
di: Mukherjee, Kunal, et al.
Pubblicazione: (2026)
di: Mukherjee, Kunal, et al.
Pubblicazione: (2026)
Medoid Prototype Alignment for Cross-Plant Unknown Attack Detection in Industrial Control Systems
di: Wang, Luyao
Pubblicazione: (2026)
di: Wang, Luyao
Pubblicazione: (2026)
SCoPE: Evaluating LLMs for Software Vulnerability Detection
di: Gonçalves, José, et al.
Pubblicazione: (2024)
di: Gonçalves, José, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
di: Xu, Zhangchen, et al.
Pubblicazione: (2024) -
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
di: Xu, Zhangchen, et al.
Pubblicazione: (2024) -
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
di: Li, Yuetai, et al.
Pubblicazione: (2024) -
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
di: Xu, Zhangchen, et al.
Pubblicazione: (2024) -
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)