Jailbreaking LLMs with Arabic Transliteration and Arabizi
Fuente:
arXiv
Saved in:
| Main Authors: | Ghanim, Mansour Al, Almohaimeed, Saleh, Zheng, Mengxin, Solihin, Yan, Lou, Qian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Robustness and Accuracy of Text Watermarking Under Real-World Cross-Lingual Manipulations
by: Ghanim, Mansour Al, et al.
Published: (2025)
by: Ghanim, Mansour Al, et al.
Published: (2025)
Ar-Spider: Text-to-SQL in Arabic
by: Almohaimeed, Saleh, et al.
Published: (2024)
by: Almohaimeed, Saleh, et al.
Published: (2024)
Prompt Engineering Techniques for Context-dependent Text-to-SQL in Arabic
by: Almohaimeed, Saleh, et al.
Published: (2025)
by: Almohaimeed, Saleh, et al.
Published: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025)
by: Almohaimeed, Saad, et al.
Published: (2025)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
by: Hastuti, Rochana Prih, et al.
Published: (2025)
by: Hastuti, Rochana Prih, et al.
Published: (2025)
PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
by: Xue, Jiaqi, et al.
Published: (2025)
by: Xue, Jiaqi, et al.
Published: (2025)
AI Text Detectors and the Misclassification of Slightly Polished Arabic Text
by: Almohaimeed, Saleh, et al.
Published: (2025)
by: Almohaimeed, Saleh, et al.
Published: (2025)
Arabizi vs LLMs: Can the Genie Understand the Language of Aladdin?
by: Almaoui, Perla Al, et al.
Published: (2025)
by: Almaoui, Perla Al, et al.
Published: (2025)
Securing Transformer-based AI Execution via Unified TEEs and Crypto-protected Accelerators
by: Xue, Jiaqi, et al.
Published: (2025)
by: Xue, Jiaqi, et al.
Published: (2025)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
by: Lou, Qian, et al.
Published: (2024)
by: Lou, Qian, et al.
Published: (2024)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
CIDAR: Culturally Relevant Instruction Dataset For Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
by: Sun, Chung-En, et al.
Published: (2024)
by: Sun, Chung-En, et al.
Published: (2024)
LLMs as Compiler for Arabic Programming Language
by: Sibaee, Serry, et al.
Published: (2024)
by: Sibaee, Serry, et al.
Published: (2024)
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
by: Garcia-Gasulla, Dario, et al.
Published: (2025)
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
by: Dotsinski, Asen, et al.
Published: (2026)
by: Dotsinski, Asen, et al.
Published: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Arabic Little STT: Arabic Children Speech Recognition Dataset
by: Alkadri, Mouhand, et al.
Published: (2025)
by: Alkadri, Mouhand, et al.
Published: (2025)
Jailbreaking LLMs via Calibration
by: Lu, Yuxuan, et al.
Published: (2026)
by: Lu, Yuxuan, et al.
Published: (2026)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
by: Su, Jingtong, et al.
Published: (2024)
by: Su, Jingtong, et al.
Published: (2024)
Computational Linguistics Meets Libyan Dialect: A Study on Dialect Identification
by: Essgaer, Mansour, et al.
Published: (2025)
by: Essgaer, Mansour, et al.
Published: (2025)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Exploiting Transliterated Words for Finding Similarity in Inter-Language News Articles using Machine Learning
by: Naeem, Sameea, et al.
Published: (2022)
by: Naeem, Sameea, et al.
Published: (2022)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
Command R7B Arabic: A Small, Enterprise Focused, Multilingual, and Culturally Aware Arabic LLM
by: Alnumay, Yazeed, et al.
Published: (2025)
by: Alnumay, Yazeed, et al.
Published: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
by: Chan, Yik Siu, et al.
Published: (2025)
by: Chan, Yik Siu, et al.
Published: (2025)
Jailbreaking as a Reward Misspecification Problem
by: Xie, Zhihui, et al.
Published: (2024)
by: Xie, Zhihui, et al.
Published: (2024)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Automatic Classification of Arabic Literature into Historical Eras
by: Alhathloul, Zainab, et al.
Published: (2026)
by: Alhathloul, Zainab, et al.
Published: (2026)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
by: Mehrotra, Anay, et al.
Published: (2023)
by: Mehrotra, Anay, et al.
Published: (2023)
ArEEG_Chars: Dataset for Envisioned Speech Recognition using EEG for Arabic Characters
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
Reward-Robust RLHF in LLMs
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
Similar Items
-
Evaluating the Robustness and Accuracy of Text Watermarking Under Real-World Cross-Lingual Manipulations
by: Ghanim, Mansour Al, et al.
Published: (2025) -
Ar-Spider: Text-to-SQL in Arabic
by: Almohaimeed, Saleh, et al.
Published: (2024) -
Prompt Engineering Techniques for Context-dependent Text-to-SQL in Arabic
by: Almohaimeed, Saleh, et al.
Published: (2025) -
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
by: Almohaimeed, Saad, et al.
Published: (2025) -
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
by: Xue, Jiaqi, et al.
Published: (2024)