Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Divyanshu, Gupta, Ishita, Birur, Nitin Aravind, Baswa, Tanay, Agarwal, Sahil, Harshangi, Prashanth |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
by: Kumar, Divyanshu, et al.
Published: (2026)
by: Kumar, Divyanshu, et al.
Published: (2026)
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
VERA: Validation and Enhancement for Retrieval Augmented systems
by: Birur, Nitin Aravind, et al.
Published: (2024)
by: Birur, Nitin Aravind, et al.
Published: (2024)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024)
by: Kumar, Anurakt, et al.
Published: (2024)
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
Quantifying CBRN Risk in Frontier Models
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
by: Kumar, Divyanshu, et al.
Published: (2024)
by: Kumar, Divyanshu, et al.
Published: (2024)
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
by: Kumar, Divyanshu, et al.
Published: (2024)
by: Kumar, Divyanshu, et al.
Published: (2024)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
by: Agarwal, Utkarsh, et al.
Published: (2024)
by: Agarwal, Utkarsh, et al.
Published: (2024)
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
by: Pallagani, Vishal, et al.
Published: (2024)
by: Pallagani, Vishal, et al.
Published: (2024)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
by: Agarwal, Rajan, et al.
Published: (2025)
by: Agarwal, Rajan, et al.
Published: (2025)
Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
by: Kumar, Vanya Bannihatti, et al.
Published: (2025)
by: Kumar, Vanya Bannihatti, et al.
Published: (2025)
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
by: Vijayaraghavan, Prashanth, et al.
Published: (2025)
by: Vijayaraghavan, Prashanth, et al.
Published: (2025)
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation
by: Kumar, Tanay, et al.
Published: (2026)
by: Kumar, Tanay, et al.
Published: (2026)
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
by: Limisiewicz, Tomasz, et al.
Published: (2025)
by: Limisiewicz, Tomasz, et al.
Published: (2025)
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
by: Liu, Guangliang, et al.
Published: (2025)
by: Liu, Guangliang, et al.
Published: (2025)
Probing the Limits of Stylistic Alignment in Vision-Language Models
by: Farajidizaji, Asma, et al.
Published: (2025)
by: Farajidizaji, Asma, et al.
Published: (2025)
JEBS: A Fine-grained Biomedical Lexical Simplification Task
by: Xia, William, et al.
Published: (2025)
by: Xia, William, et al.
Published: (2025)
Do Language Models Know When They'll Refuse? Probing Introspective Awareness of Safety Boundaries
by: Gondil, Tanay
Published: (2026)
by: Gondil, Tanay
Published: (2026)
Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
by: Neplenbroek, Vera, et al.
Published: (2025)
by: Neplenbroek, Vera, et al.
Published: (2025)
Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text
by: Sommerauer, Pia, et al.
Published: (2025)
by: Sommerauer, Pia, et al.
Published: (2025)
Exploring the Limits of Fine-grained LLM-based Physics Inference via Premise Removal Interventions
by: Meadows, Jordan, et al.
Published: (2024)
by: Meadows, Jordan, et al.
Published: (2024)
Quantifying Stereotypes in Language
by: Liu, Yang
Published: (2024)
by: Liu, Yang
Published: (2024)
LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
by: Gupta, Akash, et al.
Published: (2024)
by: Gupta, Akash, et al.
Published: (2024)
Team A at SemEval-2025 Task 11: Breaking Language Barriers in Emotion Detection with Multilingual Models
by: Sahil, P Sam, et al.
Published: (2025)
by: Sahil, P Sam, et al.
Published: (2025)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
AlphaPO: Reward Shape Matters for LLM Alignment
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes
by: Jeoung, Sullam, et al.
Published: (2025)
by: Jeoung, Sullam, et al.
Published: (2025)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
LLM driven Text-to-Table Generation through Sub-Tasks Guidance and Iterative Refinement
by: C, Rajmohan, et al.
Published: (2025)
by: C, Rajmohan, et al.
Published: (2025)
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
by: Gulati, Anmol, et al.
Published: (2026)
by: Gulati, Anmol, et al.
Published: (2026)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
by: Wei, Hui, et al.
Published: (2024)
by: Wei, Hui, et al.
Published: (2024)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
by: Masoudian, Shahed, et al.
Published: (2025)
by: Masoudian, Shahed, et al.
Published: (2025)
'Since Lawyers are Males..': Examining Implicit Gender Bias in Hindi Language Generation by LLMs
by: Joshi, Ishika, et al.
Published: (2024)
by: Joshi, Ishika, et al.
Published: (2024)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
Semantic similarity estimation for domain specific data using BERT and other techniques
by: Prashanth, R.
Published: (2025)
by: Prashanth, R.
Published: (2025)
Understanding Student Sentiment on Mental Health Support in Colleges Using Large Language Models
by: Sood, Palak, et al.
Published: (2024)
by: Sood, Palak, et al.
Published: (2024)
Similar Items
-
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
by: Kumar, Divyanshu, et al.
Published: (2026) -
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
by: Kumar, Divyanshu, et al.
Published: (2025) -
VERA: Validation and Enhancement for Retrieval Augmented systems
by: Birur, Nitin Aravind, et al.
Published: (2024) -
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024) -
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)