Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
Fuente:
arXiv
Salvato in:
| Autori principali: | Kumar, Divyanshu, Gupta, Ishita, Birur, Nitin Aravind, Baswa, Tanay, Agarwal, Sahil, Harshangi, Prashanth |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
di: Kumar, Divyanshu, et al.
Pubblicazione: (2026)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2026)
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
VERA: Validation and Enhancement for Retrieval Augmented systems
di: Birur, Nitin Aravind, et al.
Pubblicazione: (2024)
di: Birur, Nitin Aravind, et al.
Pubblicazione: (2024)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
di: Kumar, Anurakt, et al.
Pubblicazione: (2024)
di: Kumar, Anurakt, et al.
Pubblicazione: (2024)
No Free Lunch with Guardrails
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
Quantifying CBRN Risk in Frontier Models
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
di: Kumar, Divyanshu, et al.
Pubblicazione: (2024)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2024)
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
di: Kumar, Divyanshu, et al.
Pubblicazione: (2024)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2024)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
di: Agarwal, Utkarsh, et al.
Pubblicazione: (2024)
di: Agarwal, Utkarsh, et al.
Pubblicazione: (2024)
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
di: Pallagani, Vishal, et al.
Pubblicazione: (2024)
di: Pallagani, Vishal, et al.
Pubblicazione: (2024)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
di: Agarwal, Rajan, et al.
Pubblicazione: (2025)
di: Agarwal, Rajan, et al.
Pubblicazione: (2025)
Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
di: Kumar, Vanya Bannihatti, et al.
Pubblicazione: (2025)
di: Kumar, Vanya Bannihatti, et al.
Pubblicazione: (2025)
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
di: Vijayaraghavan, Prashanth, et al.
Pubblicazione: (2025)
di: Vijayaraghavan, Prashanth, et al.
Pubblicazione: (2025)
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation
di: Kumar, Tanay, et al.
Pubblicazione: (2026)
di: Kumar, Tanay, et al.
Pubblicazione: (2026)
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2025)
di: Limisiewicz, Tomasz, et al.
Pubblicazione: (2025)
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
Probing the Limits of Stylistic Alignment in Vision-Language Models
di: Farajidizaji, Asma, et al.
Pubblicazione: (2025)
di: Farajidizaji, Asma, et al.
Pubblicazione: (2025)
JEBS: A Fine-grained Biomedical Lexical Simplification Task
di: Xia, William, et al.
Pubblicazione: (2025)
di: Xia, William, et al.
Pubblicazione: (2025)
Do Language Models Know When They'll Refuse? Probing Introspective Awareness of Safety Boundaries
di: Gondil, Tanay
Pubblicazione: (2026)
di: Gondil, Tanay
Pubblicazione: (2026)
Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
di: Neplenbroek, Vera, et al.
Pubblicazione: (2025)
di: Neplenbroek, Vera, et al.
Pubblicazione: (2025)
Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text
di: Sommerauer, Pia, et al.
Pubblicazione: (2025)
di: Sommerauer, Pia, et al.
Pubblicazione: (2025)
Exploring the Limits of Fine-grained LLM-based Physics Inference via Premise Removal Interventions
di: Meadows, Jordan, et al.
Pubblicazione: (2024)
di: Meadows, Jordan, et al.
Pubblicazione: (2024)
Quantifying Stereotypes in Language
di: Liu, Yang
Pubblicazione: (2024)
di: Liu, Yang
Pubblicazione: (2024)
LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
di: Gupta, Akash, et al.
Pubblicazione: (2024)
di: Gupta, Akash, et al.
Pubblicazione: (2024)
Team A at SemEval-2025 Task 11: Breaking Language Barriers in Emotion Detection with Multilingual Models
di: Sahil, P Sam, et al.
Pubblicazione: (2025)
di: Sahil, P Sam, et al.
Pubblicazione: (2025)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
di: Nejadgholi, Isar, et al.
Pubblicazione: (2024)
di: Nejadgholi, Isar, et al.
Pubblicazione: (2024)
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
di: Tomar, Aditya, et al.
Pubblicazione: (2025)
AlphaPO: Reward Shape Matters for LLM Alignment
di: Gupta, Aman, et al.
Pubblicazione: (2025)
di: Gupta, Aman, et al.
Pubblicazione: (2025)
Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes
di: Jeoung, Sullam, et al.
Pubblicazione: (2025)
di: Jeoung, Sullam, et al.
Pubblicazione: (2025)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
di: Guo, Hongyi, et al.
Pubblicazione: (2024)
di: Guo, Hongyi, et al.
Pubblicazione: (2024)
LLM driven Text-to-Table Generation through Sub-Tasks Guidance and Iterative Refinement
di: C, Rajmohan, et al.
Pubblicazione: (2025)
di: C, Rajmohan, et al.
Pubblicazione: (2025)
Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
di: Gulati, Anmol, et al.
Pubblicazione: (2026)
di: Gulati, Anmol, et al.
Pubblicazione: (2026)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
di: Wei, Hui, et al.
Pubblicazione: (2024)
di: Wei, Hui, et al.
Pubblicazione: (2024)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
di: Jain, Shomik, et al.
Pubblicazione: (2025)
di: Jain, Shomik, et al.
Pubblicazione: (2025)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
di: Masoudian, Shahed, et al.
Pubblicazione: (2025)
di: Masoudian, Shahed, et al.
Pubblicazione: (2025)
'Since Lawyers are Males..': Examining Implicit Gender Bias in Hindi Language Generation by LLMs
di: Joshi, Ishika, et al.
Pubblicazione: (2024)
di: Joshi, Ishika, et al.
Pubblicazione: (2024)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2024)
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2024)
Semantic similarity estimation for domain specific data using BERT and other techniques
di: Prashanth, R.
Pubblicazione: (2025)
di: Prashanth, R.
Pubblicazione: (2025)
Understanding Student Sentiment on Mental Health Support in Colleges Using Large Language Models
di: Sood, Palak, et al.
Pubblicazione: (2024)
di: Sood, Palak, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
di: Kumar, Divyanshu, et al.
Pubblicazione: (2026) -
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025) -
VERA: Validation and Enhancement for Retrieval Augmented systems
di: Birur, Nitin Aravind, et al.
Pubblicazione: (2024) -
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
di: Kumar, Anurakt, et al.
Pubblicazione: (2024) -
No Free Lunch with Guardrails
di: Kumar, Divyanshu, et al.
Pubblicazione: (2025)