Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
Fuente:
arXiv
Guardado en:
| Autores principales: | Atil, Berk, Passonneau, Rebecca J., Morstatter, Fred |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Model Unlearning Objectives Vary for Distinct Language Functions
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Something Just Like TRuST : Toxicity Recognition of Span and Target
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
VerAs: Verify then Assess STEM Lab Reports
por: Atil, Berk, et al.
Publicado: (2024)
por: Atil, Berk, et al.
Publicado: (2024)
The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance
por: Salinas, Abel, et al.
Publicado: (2024)
por: Salinas, Abel, et al.
Publicado: (2024)
Knowledge Graph Analysis of Legal Understanding and Violations in LLMs
por: Jha, Abha, et al.
Publicado: (2025)
por: Jha, Abha, et al.
Publicado: (2025)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
por: Dorn, Rebecca, et al.
Publicado: (2024)
por: Dorn, Rebecca, et al.
Publicado: (2024)
Risk and Response in Large Language Models: Evaluating Key Threat Categories
por: Harandizadeh, Bahareh, et al.
Publicado: (2024)
por: Harandizadeh, Bahareh, et al.
Publicado: (2024)
Joint Training for Selective Prediction
por: Li, Zhaohui, et al.
Publicado: (2024)
por: Li, Zhaohui, et al.
Publicado: (2024)
Intention Analysis Makes LLMs A Good Jailbreak Defender
por: Zhang, Yuqi, et al.
Publicado: (2024)
por: Zhang, Yuqi, et al.
Publicado: (2024)
Defending LLMs against Jailbreaking Attacks via Backtranslation
por: Wang, Yihan, et al.
Publicado: (2024)
por: Wang, Yihan, et al.
Publicado: (2024)
"Define Your Terms" : Enhancing Efficient Offensive Speech Classification with Definition
por: Nghiem, Huy, et al.
Publicado: (2024)
por: Nghiem, Huy, et al.
Publicado: (2024)
Estimating Causal Effects of Text Interventions Leveraging LLMs
por: Guo, Siyi, et al.
Publicado: (2024)
por: Guo, Siyi, et al.
Publicado: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
por: Liu, Fan, et al.
Publicado: (2024)
por: Liu, Fan, et al.
Publicado: (2024)
Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking
por: Zhu, Junda, et al.
Publicado: (2025)
por: Zhu, Junda, et al.
Publicado: (2025)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
por: Li, Yu, et al.
Publicado: (2025)
por: Li, Yu, et al.
Publicado: (2025)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
por: Ji, Jiabao, et al.
Publicado: (2024)
por: Ji, Jiabao, et al.
Publicado: (2024)
Contextualizing Argument Quality Assessment with Relevant Knowledge
por: Deshpande, Darshan, et al.
Publicado: (2023)
por: Deshpande, Darshan, et al.
Publicado: (2023)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
por: Zhang, Zhexin, et al.
Publicado: (2023)
por: Zhang, Zhexin, et al.
Publicado: (2023)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
por: Gao, Lang, et al.
Publicado: (2024)
por: Gao, Lang, et al.
Publicado: (2024)
Defending against Jailbreak through Early Exit Generation of Large Language Models
por: Zhao, Chongwen, et al.
Publicado: (2024)
por: Zhao, Chongwen, et al.
Publicado: (2024)
Sociodemographic Bias in Language Models: A Survey and Forward Path
por: Gupta, Vipul, et al.
Publicado: (2023)
por: Gupta, Vipul, et al.
Publicado: (2023)
How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
por: Karizaki, Mahsa Sheikhi, et al.
Publicado: (2024)
por: Karizaki, Mahsa Sheikhi, et al.
Publicado: (2024)
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2025)
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2025)
Secret Keepers: The Impact of LLMs on Linguistic Markers of Personal Traits
por: Sourati, Zhivar, et al.
Publicado: (2024)
por: Sourati, Zhivar, et al.
Publicado: (2024)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
por: Zhou, Andy, et al.
Publicado: (2024)
por: Zhou, Andy, et al.
Publicado: (2024)
Non-Determinism of "Deterministic" LLM Settings
por: Atil, Berk, et al.
Publicado: (2024)
por: Atil, Berk, et al.
Publicado: (2024)
The Unequal Opportunities of Large Language Models: Revealing Demographic Bias through Job Recommendations
por: Salinas, Abel, et al.
Publicado: (2023)
por: Salinas, Abel, et al.
Publicado: (2023)
The Curious Case of Nonverbal Abstract Reasoning with Multi-Modal Large Language Models
por: Ahrabian, Kian, et al.
Publicado: (2024)
por: Ahrabian, Kian, et al.
Publicado: (2024)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
por: Gupta, Vipul, et al.
Publicado: (2023)
por: Gupta, Vipul, et al.
Publicado: (2023)
Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
por: Zhao, Yinzhi, et al.
Publicado: (2026)
por: Zhao, Yinzhi, et al.
Publicado: (2026)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
por: Zhao, Weixiang, et al.
Publicado: (2025)
por: Zhao, Weixiang, et al.
Publicado: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
por: Jiang, Weisen, et al.
Publicado: (2025)
por: Jiang, Weisen, et al.
Publicado: (2025)
Lost in Translation: Do LVLM Judges Generalize Across Languages?
por: Laskar, Md Tahmid Rahman, et al.
Publicado: (2026)
por: Laskar, Md Tahmid Rahman, et al.
Publicado: (2026)
Playing Language Game with LLMs Leads to Jailbreaking
por: Peng, Yu, et al.
Publicado: (2024)
por: Peng, Yu, et al.
Publicado: (2024)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
por: Zhao, Jiawei, et al.
Publicado: (2024)
por: Zhao, Jiawei, et al.
Publicado: (2024)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
por: Jiang, Tanqiu, et al.
Publicado: (2024)
por: Jiang, Tanqiu, et al.
Publicado: (2024)
Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English
por: Dorn, Rebecca, et al.
Publicado: (2025)
por: Dorn, Rebecca, et al.
Publicado: (2025)
Purple-teaming LLMs with Adversarial Defender Training
por: Zhou, Jingyan, et al.
Publicado: (2024)
por: Zhou, Jingyan, et al.
Publicado: (2024)
Ejemplares similares
-
Model Unlearning Objectives Vary for Distinct Language Functions
por: Atil, Berk, et al.
Publicado: (2026) -
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
por: Atil, Berk, et al.
Publicado: (2026) -
Something Just Like TRuST : Toxicity Recognition of Span and Target
por: Atil, Berk, et al.
Publicado: (2025) -
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025) -
VerAs: Verify then Assess STEM Lab Reports
por: Atil, Berk, et al.
Publicado: (2024)