Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
Fuente:
arXiv
Saved in:
| Main Authors: | Rao, Abhinav, Vashistha, Sachin, Naik, Atharva, Aditya, Somak, Choudhury, Monojit |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
by: Vashistha, Sachin, et al.
Published: (2025)
by: Vashistha, Sachin, et al.
Published: (2025)
SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation
by: Pandey, Saurabh Kumar, et al.
Published: (2025)
by: Pandey, Saurabh Kumar, et al.
Published: (2025)
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025)
by: Seth, Agrima, et al.
Published: (2025)
Missing Melodies: AI Music Generation and its "Nearly" Complete Omission of the Global South
by: Mehta, Atharva, et al.
Published: (2024)
by: Mehta, Atharva, et al.
Published: (2024)
Exploring Adapter Design Tradeoffs for Low Resource Music Generation
by: Mehta, Atharva, et al.
Published: (2025)
by: Mehta, Atharva, et al.
Published: (2025)
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models
by: Mehta, Atharva, et al.
Published: (2025)
by: Mehta, Atharva, et al.
Published: (2025)
MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
by: Dutta, Aritra, et al.
Published: (2026)
by: Dutta, Aritra, et al.
Published: (2026)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
by: Agarwal, Utkarsh, et al.
Published: (2024)
by: Agarwal, Utkarsh, et al.
Published: (2024)
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test
by: Khandelwal, Aditi, et al.
Published: (2024)
by: Khandelwal, Aditi, et al.
Published: (2024)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages
by: Shahid, Farhana, et al.
Published: (2025)
by: Shahid, Farhana, et al.
Published: (2025)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
by: Hong, Pengfei, et al.
Published: (2024)
by: Hong, Pengfei, et al.
Published: (2024)
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
by: Kaur, Navreet, et al.
Published: (2023)
by: Kaur, Navreet, et al.
Published: (2023)
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
by: Du, Yanrui, et al.
Published: (2023)
by: Du, Yanrui, et al.
Published: (2023)
TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text
by: Adak, Sayantan, et al.
Published: (2024)
by: Adak, Sayantan, et al.
Published: (2024)
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
by: Das, Debrup, et al.
Published: (2024)
by: Das, Debrup, et al.
Published: (2024)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
by: Ray, Sourjyadip, et al.
Published: (2025)
by: Ray, Sourjyadip, et al.
Published: (2025)
UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
by: Choudhary, Mukund, et al.
Published: (2025)
by: Choudhary, Mukund, et al.
Published: (2025)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
DSP-MLIR: A MLIR Dialect for Digital Signal Processing
by: Kumar, Abhinav, et al.
Published: (2024)
by: Kumar, Abhinav, et al.
Published: (2024)
Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing
by: Mutsaddi, Atharva, et al.
Published: (2025)
by: Mutsaddi, Atharva, et al.
Published: (2025)
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models
by: Atwany, Hanin, et al.
Published: (2025)
by: Atwany, Hanin, et al.
Published: (2025)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
by: Luo, Man, et al.
Published: (2023)
by: Luo, Man, et al.
Published: (2023)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
by: Hawkins, John, et al.
Published: (2025)
by: Hawkins, John, et al.
Published: (2025)
NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
by: Dutta, Aritra, et al.
Published: (2025)
by: Dutta, Aritra, et al.
Published: (2025)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
by: Liu, Chris Yuhao, et al.
Published: (2024)
by: Liu, Chris Yuhao, et al.
Published: (2024)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
by: Kale, Sahil
Published: (2025)
by: Kale, Sahil
Published: (2025)
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
by: Kadali, Sri Durga Sai Sowmya, et al.
Published: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
by: Roy, Aniruddha, et al.
Published: (2025)
by: Roy, Aniruddha, et al.
Published: (2025)
Disability Across Cultures: A Human-Centered Audit of Ableism in Western and Indic LLMs
by: Phutane, Mahika, et al.
Published: (2025)
by: Phutane, Mahika, et al.
Published: (2025)
The Zeno's Paradox of `Low-Resource' Languages
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
Similar Items
-
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024) -
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
by: Vashistha, Sachin, et al.
Published: (2025) -
SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation
by: Pandey, Saurabh Kumar, et al.
Published: (2025) -
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2025) -
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025)