How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
Fuente:
arXiv
Salvato in:
| Autori principali: | Banerjee, Somnath, Layek, Sayan, Hazra, Rima, Mukherjee, Animesh |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
Context Matters: Pushing the Boundaries of Open-Ended Answer Generation with Graph-Structured Knowledge Context
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
di: Hazra, Rima, et al.
Pubblicazione: (2024)
di: Hazra, Rima, et al.
Pubblicazione: (2024)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
di: Hazra, Rima, et al.
Pubblicazione: (2024)
di: Hazra, Rima, et al.
Pubblicazione: (2024)
Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages
di: Banerjee, Somnath, et al.
Pubblicazione: (2026)
di: Banerjee, Somnath, et al.
Pubblicazione: (2026)
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
di: Kwartler, Ted, et al.
Pubblicazione: (2024)
di: Kwartler, Ted, et al.
Pubblicazione: (2024)
Breaking Boundaries: Investigating the Effects of Model Editing on Cross-linguistic Performance
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation
di: Adak, Sayantan, et al.
Pubblicazione: (2025)
di: Adak, Sayantan, et al.
Pubblicazione: (2025)
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
di: Banerjee, Somnath, et al.
Pubblicazione: (2025)
Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
di: Banerjee, Somnath, et al.
Pubblicazione: (2026)
di: Banerjee, Somnath, et al.
Pubblicazione: (2026)
Evaluating the Ebb and Flow: An In-depth Analysis of Question-Answering Trends across Diverse Platforms
di: Hazra, Rima, et al.
Pubblicazione: (2023)
di: Hazra, Rima, et al.
Pubblicazione: (2023)
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
di: Xian, R. Patrick, et al.
Pubblicazione: (2024)
di: Xian, R. Patrick, et al.
Pubblicazione: (2024)
DistALANER: Distantly Supervised Active Learning Augmented Named Entity Recognition in the Open Source Software Ecosystem
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
di: Zhang, Rui, et al.
Pubblicazione: (2026)
di: Zhang, Rui, et al.
Pubblicazione: (2026)
How Vulnerable Are Edge LLMs?
di: Ding, Ao, et al.
Pubblicazione: (2026)
di: Ding, Ao, et al.
Pubblicazione: (2026)
Data-centric NLP Backdoor Defense from the Lens of Memorization
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
di: Wang, Zhenting, et al.
Pubblicazione: (2024)
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
di: Leong, Chak Tou, et al.
Pubblicazione: (2024)
di: Leong, Chak Tou, et al.
Pubblicazione: (2024)
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
di: Zhang, Jiankun, et al.
Pubblicazione: (2025)
di: Zhang, Jiankun, et al.
Pubblicazione: (2025)
Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy
di: Fu, Yu, et al.
Pubblicazione: (2023)
di: Fu, Yu, et al.
Pubblicazione: (2023)
A privacy preserving querying mechanism with high utility for electric vehicles
di: Atmaca, Ugur Ilker, et al.
Pubblicazione: (2022)
di: Atmaca, Ugur Ilker, et al.
Pubblicazione: (2022)
Fingerprinting LLMs via Prompt Injection
di: Hu, Yuepeng, et al.
Pubblicazione: (2025)
di: Hu, Yuepeng, et al.
Pubblicazione: (2025)
LLMs for Domain Generation Algorithm Detection
di: La O, Reynier Leyva, et al.
Pubblicazione: (2024)
di: La O, Reynier Leyva, et al.
Pubblicazione: (2024)
Mitigating Jailbreaks with Intent-Aware LLMs
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
SoK: Are Watermarks in LLMs Ready for Deployment?
di: Dang, Kieu, et al.
Pubblicazione: (2025)
di: Dang, Kieu, et al.
Pubblicazione: (2025)
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models
di: Adak, Sayantan, et al.
Pubblicazione: (2025)
di: Adak, Sayantan, et al.
Pubblicazione: (2025)
SafeMath: Inference-time Safety improves Math Accuracy
di: Basu, Sagnik, et al.
Pubblicazione: (2026)
di: Basu, Sagnik, et al.
Pubblicazione: (2026)
Fast-MIA: Efficient and Scalable Membership Inference for LLMs
di: Takahashi, Hiromu, et al.
Pubblicazione: (2025)
di: Takahashi, Hiromu, et al.
Pubblicazione: (2025)
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
di: Meeus, Matthieu, et al.
Pubblicazione: (2024)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
di: Luo, Xuan, et al.
Pubblicazione: (2025)
di: Luo, Xuan, et al.
Pubblicazione: (2025)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
di: Liu, Yepeng, et al.
Pubblicazione: (2025)
di: Liu, Yepeng, et al.
Pubblicazione: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
di: Zhang, Chiyu, et al.
Pubblicazione: (2025)
di: Zhang, Chiyu, et al.
Pubblicazione: (2025)
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
di: He, Xuanli, et al.
Pubblicazione: (2026)
di: He, Xuanli, et al.
Pubblicazione: (2026)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
di: Fu, Yu, et al.
Pubblicazione: (2026)
di: Fu, Yu, et al.
Pubblicazione: (2026)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
di: Liu, Honghao, et al.
Pubblicazione: (2026)
di: Liu, Honghao, et al.
Pubblicazione: (2026)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
di: Fu, Yu, et al.
Pubblicazione: (2024)
di: Fu, Yu, et al.
Pubblicazione: (2024)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
di: Zhao, Gejian, et al.
Pubblicazione: (2025)
di: Zhao, Gejian, et al.
Pubblicazione: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
di: Chen, Yunhao, et al.
Pubblicazione: (2025)
di: Chen, Yunhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
di: Banerjee, Somnath, et al.
Pubblicazione: (2025) -
SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
di: Banerjee, Somnath, et al.
Pubblicazione: (2024) -
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
di: Banerjee, Somnath, et al.
Pubblicazione: (2025) -
Context Matters: Pushing the Boundaries of Open-Ended Answer Generation with Graph-Structured Knowledge Context
di: Banerjee, Somnath, et al.
Pubblicazione: (2024) -
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
di: Hazra, Rima, et al.
Pubblicazione: (2024)