Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | Banerjee, Somnath, Chatterjee, Pratyush, Kumar, Shanu, Layek, Sayan, Agrawal, Parag, Hazra, Rima, Mukherjee, Animesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
by: Hazra, Rima, et al.
Published: (2024)
by: Hazra, Rima, et al.
Published: (2024)
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages
by: Banerjee, Somnath, et al.
Published: (2026)
by: Banerjee, Somnath, et al.
Published: (2026)
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models
by: Adak, Sayantan, et al.
Published: (2025)
by: Adak, Sayantan, et al.
Published: (2025)
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
by: Hazra, Rima, et al.
Published: (2024)
by: Hazra, Rima, et al.
Published: (2024)
Context Matters: Pushing the Boundaries of Open-Ended Answer Generation with Graph-Structured Knowledge Context
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Breaking Boundaries: Investigating the Effects of Model Editing on Cross-linguistic Performance
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
MemeSense: An Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation
by: Adak, Sayantan, et al.
Published: (2025)
by: Adak, Sayantan, et al.
Published: (2025)
DistALANER: Distantly Supervised Active Learning Augmented Named Entity Recognition in the Open Source Software Ecosystem
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
by: Banerjee, Somnath, et al.
Published: (2026)
by: Banerjee, Somnath, et al.
Published: (2026)
Evaluating the Ebb and Flow: An In-depth Analysis of Question-Answering Trends across Diverse Platforms
by: Hazra, Rima, et al.
Published: (2023)
by: Hazra, Rima, et al.
Published: (2023)
SafeMath: Inference-time Safety improves Math Accuracy
by: Basu, Sagnik, et al.
Published: (2026)
by: Basu, Sagnik, et al.
Published: (2026)
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
by: Hazra, Rima, et al.
Published: (2026)
by: Hazra, Rima, et al.
Published: (2026)
Redefining Developer Assistance: Through Large Language Models in Software Ecosystem
by: Banerjee, Somnath, et al.
Published: (2023)
by: Banerjee, Somnath, et al.
Published: (2023)
From Fluent to Verifiable: Claim-Level Auditability for Deep Research Agents
by: Rasheed, Razeen A, et al.
Published: (2026)
by: Rasheed, Razeen A, et al.
Published: (2026)
Duplicate Question Retrieval and Confirmation Time Prediction in Software Communities
by: Hazra, Rima, et al.
Published: (2023)
by: Hazra, Rima, et al.
Published: (2023)
Tutoring Large Language Models to be Domain-adaptive, Precise, and Safe
by: Banerjee, Somnath
Published: (2026)
by: Banerjee, Somnath
Published: (2026)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
by: Mendu, Sai Krishna, et al.
Published: (2025)
by: Mendu, Sai Krishna, et al.
Published: (2025)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
InfFeed: Influence Functions as a Feedback to Improve the Performance of Subjective Tasks
by: Banerjee, Somnath, et al.
Published: (2024)
by: Banerjee, Somnath, et al.
Published: (2024)
SCULPT: Systematic Tuning of Long Prompts
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text
by: Adak, Sayantan, et al.
Published: (2024)
by: Adak, Sayantan, et al.
Published: (2024)
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
by: Sachdeva, Rachneet, et al.
Published: (2025)
by: Sachdeva, Rachneet, et al.
Published: (2025)
Analyzing Sentiment Polarity Reduction in News Presentation through Contextual Perturbation and Large Language Models
by: Kuila, Alapan, et al.
Published: (2024)
by: Kuila, Alapan, et al.
Published: (2024)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
Benchmarking Motivational Interviewing Competence of Large Language Models
by: Jha, Aishwariya, et al.
Published: (2026)
by: Jha, Aishwariya, et al.
Published: (2026)
Integrating Large Language Models with Graph-based Reasoning for Conversational Question Answering
by: Jain, Parag, et al.
Published: (2024)
by: Jain, Parag, et al.
Published: (2024)
When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
Talent or Luck? Evaluating Attribution Bias in Large Language Models
by: Raj, Chahat, et al.
Published: (2025)
by: Raj, Chahat, et al.
Published: (2025)
Named Entity Recognition on Code-Mixed Cross-Script Social Media Content
by: Somnath Banerjee
Published: (2017)
by: Somnath Banerjee
Published: (2017)
Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework
by: Xu, Zishan, et al.
Published: (2025)
by: Xu, Zishan, et al.
Published: (2025)
On Zero-Shot Counterspeech Generation by LLMs
by: Saha, Punyajoy, et al.
Published: (2024)
by: Saha, Punyajoy, et al.
Published: (2024)
READ: Reinforcement-based Adversarial Learning for Text Classification with Limited Labeled Data
by: Sharma, Rohit, et al.
Published: (2025)
by: Sharma, Rohit, et al.
Published: (2025)
CodeMixBench: Evaluating Large Language Models on Code Generation with Code-Mixed Prompts
by: Sheokand, Manik, et al.
Published: (2025)
by: Sheokand, Manik, et al.
Published: (2025)
Exploring Safety-Utility Trade-Offs in Personalized Language Models
by: Vijjini, Anvesh Rao, et al.
Published: (2024)
by: Vijjini, Anvesh Rao, et al.
Published: (2024)
Similar Items
-
Soteria: Language-Specific Functional Parameter Steering for Multilingual Safety Alignment
by: Banerjee, Somnath, et al.
Published: (2025) -
SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language Models
by: Banerjee, Somnath, et al.
Published: (2024) -
Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language Models
by: Banerjee, Somnath, et al.
Published: (2024) -
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
by: Banerjee, Somnath, et al.
Published: (2024) -
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
by: Hazra, Rima, et al.
Published: (2024)