Low-Resource Languages Jailbreak GPT-4
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yong, Zheng-Xin, Menghini, Cristina, Bach, Stephen H. |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
par: Li, Xiaochen, et autres
Publié: (2024)
par: Li, Xiaochen, et autres
Publié: (2024)
LexC-Gen: Generating Data for Extremely Low-Resource Languages with Large Language Models and Bilingual Lexicons
par: Yong, Zheng-Xin, et autres
Publié: (2024)
par: Yong, Zheng-Xin, et autres
Publié: (2024)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
par: Zheng, Xiaosen, et autres
Publié: (2024)
par: Zheng, Xiaosen, et autres
Publié: (2024)
Rethinking How to Evaluate Language Model Jailbreak
par: Cai, Hongyu, et autres
Publié: (2024)
par: Cai, Hongyu, et autres
Publié: (2024)
Jailbreaking Large Language Models with Symbolic Mathematics
par: Bethany, Emet, et autres
Publié: (2024)
par: Bethany, Emet, et autres
Publié: (2024)
EnJa: Ensemble Jailbreak on Large Language Models
par: Zhang, Jiahao, et autres
Publié: (2024)
par: Zhang, Jiahao, et autres
Publié: (2024)
Jailbreaking in the Haystack
par: Shah, Rishi Rajesh, et autres
Publié: (2025)
par: Shah, Rishi Rajesh, et autres
Publié: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
par: Chu, Junjie, et autres
Publié: (2024)
par: Chu, Junjie, et autres
Publié: (2024)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
par: Fang, Zhicheng, et autres
Publié: (2026)
par: Fang, Zhicheng, et autres
Publié: (2026)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
par: Qian, Cheng, et autres
Publié: (2024)
par: Qian, Cheng, et autres
Publié: (2024)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
par: Wei, Zeming, et autres
Publié: (2023)
par: Wei, Zeming, et autres
Publié: (2023)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
par: Yi, Sibo, et autres
Publié: (2024)
par: Yi, Sibo, et autres
Publié: (2024)
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
par: Yong, Zheng-Xin, et autres
Publié: (2025)
par: Yong, Zheng-Xin, et autres
Publié: (2025)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
par: Hua, Peichun, et autres
Publié: (2025)
par: Hua, Peichun, et autres
Publié: (2025)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
par: Boreiko, Valentyn, et autres
Publié: (2024)
par: Boreiko, Valentyn, et autres
Publié: (2024)
Jailbreaking with Universal Multi-Prompts
par: Hsu, Yu-Ling, et autres
Publié: (2025)
par: Hsu, Yu-Ling, et autres
Publié: (2025)
Jailbreaking LLMs via Calibration
par: Lu, Yuxuan, et autres
Publié: (2026)
par: Lu, Yuxuan, et autres
Publié: (2026)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
par: Lin, Shi, et autres
Publié: (2024)
par: Lin, Shi, et autres
Publié: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
par: Reddy, Aashray, et autres
Publié: (2025)
par: Reddy, Aashray, et autres
Publié: (2025)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
par: Fang, Zheng, et autres
Publié: (2026)
par: Fang, Zheng, et autres
Publié: (2026)
Using Hallucinations to Bypass GPT4's Filter
par: Lemkin, Benjamin
Publié: (2024)
par: Lemkin, Benjamin
Publié: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
par: Hu, Xiaomeng, et autres
Publié: (2024)
par: Hu, Xiaomeng, et autres
Publié: (2024)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
par: Li, Xiao, et autres
Publié: (2024)
par: Li, Xiao, et autres
Publié: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
par: Saiem, Bijoy Ahmed, et autres
Publié: (2024)
par: Saiem, Bijoy Ahmed, et autres
Publié: (2024)
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice
par: Goel, Aman, et autres
Publié: (2025)
par: Goel, Aman, et autres
Publié: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
par: Zhu, Sicheng, et autres
Publié: (2024)
par: Zhu, Sicheng, et autres
Publié: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
par: Mehrotra, Anay, et autres
Publié: (2023)
par: Mehrotra, Anay, et autres
Publié: (2023)
Universal Jailbreak Backdoors from Poisoned Human Feedback
par: Rando, Javier, et autres
Publié: (2023)
par: Rando, Javier, et autres
Publié: (2023)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
par: Mo, Yichuan, et autres
Publié: (2024)
par: Mo, Yichuan, et autres
Publié: (2024)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
par: Rando, Javier, et autres
Publié: (2024)
par: Rando, Javier, et autres
Publié: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
par: Chen, Taiye, et autres
Publié: (2025)
par: Chen, Taiye, et autres
Publié: (2025)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
par: Correia, Pedro H. Barcha, et autres
Publié: (2026)
par: Correia, Pedro H. Barcha, et autres
Publié: (2026)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
par: Hu, Xulin, et autres
Publié: (2026)
par: Hu, Xulin, et autres
Publié: (2026)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
par: Jiang, Weisen, et autres
Publié: (2025)
par: Jiang, Weisen, et autres
Publié: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
par: Liang, Buyun, et autres
Publié: (2025)
par: Liang, Buyun, et autres
Publié: (2025)
A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
par: Yao, Yang, et autres
Publié: (2025)
par: Yao, Yang, et autres
Publié: (2025)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
par: Li, Jing-Jing, et autres
Publié: (2025)
par: Li, Jing-Jing, et autres
Publié: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
par: Hasan, Adib, et autres
Publié: (2024)
par: Hasan, Adib, et autres
Publié: (2024)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
par: Sharma, Mrinank, et autres
Publié: (2025)
par: Sharma, Mrinank, et autres
Publié: (2025)
Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition
par: Nguyen, Tu, et autres
Publié: (2024)
par: Nguyen, Tu, et autres
Publié: (2024)
Documents similaires
-
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
par: Li, Xiaochen, et autres
Publié: (2024) -
LexC-Gen: Generating Data for Extremely Low-Resource Languages with Large Language Models and Bilingual Lexicons
par: Yong, Zheng-Xin, et autres
Publié: (2024) -
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
par: Zheng, Xiaosen, et autres
Publié: (2024) -
Rethinking How to Evaluate Language Model Jailbreak
par: Cai, Hongyu, et autres
Publié: (2024) -
Jailbreaking Large Language Models with Symbolic Mathematics
par: Bethany, Emet, et autres
Publié: (2024)