[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Rao, Abhinav, Choudhury, Monojit, Aditya, Somak |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
by: Rao, Abhinav, et al.
Published: (2023)
by: Rao, Abhinav, et al.
Published: (2023)
SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation
by: Pandey, Saurabh Kumar, et al.
Published: (2025)
by: Pandey, Saurabh Kumar, et al.
Published: (2025)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
by: Li, Yifan, et al.
Published: (2024)
by: Li, Yifan, et al.
Published: (2024)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
by: Singh, Joykirat, et al.
Published: (2024)
by: Singh, Joykirat, et al.
Published: (2024)
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
by: Lin, Lizhi, et al.
Published: (2024)
by: Lin, Lizhi, et al.
Published: (2024)
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
The Zeno's Paradox of `Low-Resource' Languages
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
by: Nigatu, Hellina Hailu, et al.
Published: (2024)
To Generate or Discriminate? Methodological Considerations for Measuring Cultural Alignment in LLMs
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
by: Pandey, Saurabh Kumar, et al.
Published: (2026)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
by: Dutta, Aritra, et al.
Published: (2026)
by: Dutta, Aritra, et al.
Published: (2026)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
by: Agarwal, Utkarsh, et al.
Published: (2024)
by: Agarwal, Utkarsh, et al.
Published: (2024)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test
by: Khandelwal, Aditi, et al.
Published: (2024)
by: Khandelwal, Aditi, et al.
Published: (2024)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
by: Mukherjee, Sourabrata, et al.
Published: (2025)
by: Mukherjee, Sourabrata, et al.
Published: (2025)
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
The Sword, Shield, and Achilles' Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning
by: Zhang, Xudong, et al.
Published: (2026)
by: Zhang, Xudong, et al.
Published: (2026)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
by: Hong, Pengfei, et al.
Published: (2024)
by: Hong, Pengfei, et al.
Published: (2024)
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
by: Kaur, Navreet, et al.
Published: (2023)
by: Kaur, Navreet, et al.
Published: (2023)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025)
by: Seth, Agrima, et al.
Published: (2025)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
by: Vashistha, Sachin, et al.
Published: (2025)
by: Vashistha, Sachin, et al.
Published: (2025)
TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text
by: Adak, Sayantan, et al.
Published: (2024)
by: Adak, Sayantan, et al.
Published: (2024)
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
by: Das, Debrup, et al.
Published: (2024)
by: Das, Debrup, et al.
Published: (2024)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
by: Ray, Sourjyadip, et al.
Published: (2025)
by: Ray, Sourjyadip, et al.
Published: (2025)
UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
by: Choudhary, Mukund, et al.
Published: (2025)
by: Choudhary, Mukund, et al.
Published: (2025)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models
by: Atwany, Hanin, et al.
Published: (2025)
by: Atwany, Hanin, et al.
Published: (2025)
Missing Melodies: AI Music Generation and its "Nearly" Complete Omission of the Global South
by: Mehta, Atharva, et al.
Published: (2024)
by: Mehta, Atharva, et al.
Published: (2024)
NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
by: Dutta, Aritra, et al.
Published: (2025)
by: Dutta, Aritra, et al.
Published: (2025)
Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts
by: Wang, Songping, et al.
Published: (2026)
by: Wang, Songping, et al.
Published: (2026)
Exploring Adapter Design Tradeoffs for Low Resource Music Generation
by: Mehta, Atharva, et al.
Published: (2025)
by: Mehta, Atharva, et al.
Published: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
by: Roy, Aniruddha, et al.
Published: (2025)
by: Roy, Aniruddha, et al.
Published: (2025)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models
by: Adak, Sayantan, et al.
Published: (2025)
by: Adak, Sayantan, et al.
Published: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
Weak Pareto Boundary: The Achilles' Heel of Evolutionary Multi-Objective Optimization
by: Zheng, Ruihao, et al.
Published: (2025)
by: Zheng, Ruihao, et al.
Published: (2025)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
by: Mukherjee, Sagnik, et al.
Published: (2024)
by: Mukherjee, Sagnik, et al.
Published: (2024)
Similar Items
-
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
by: Rao, Abhinav, et al.
Published: (2023) -
SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation
by: Pandey, Saurabh Kumar, et al.
Published: (2025) -
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
by: Li, Yifan, et al.
Published: (2024) -
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
by: Singh, Joykirat, et al.
Published: (2024) -
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
by: Lin, Lizhi, et al.
Published: (2024)