Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
Fuente:
arXiv
Salvato in:
| Autori principali: | Iskander, Shadi, Radinsky, Kira, Belinkov, Yonatan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
di: Majumdar, Ayan, et al.
Pubblicazione: (2025)
di: Majumdar, Ayan, et al.
Pubblicazione: (2025)
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
di: Cunningham, Eoghan, et al.
Pubblicazione: (2025)
di: Cunningham, Eoghan, et al.
Pubblicazione: (2025)
Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
di: Itzhak, Itay, et al.
Pubblicazione: (2023)
di: Itzhak, Itay, et al.
Pubblicazione: (2023)
DSO: Direct Steering Optimization for Bias Mitigation
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2025)
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2025)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
di: Vargas, Francisco, et al.
Pubblicazione: (2020)
di: Vargas, Francisco, et al.
Pubblicazione: (2020)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2023)
di: Ma, Mingyu Derek, et al.
Pubblicazione: (2023)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
di: Islam, Tunazzina
Pubblicazione: (2026)
di: Islam, Tunazzina
Pubblicazione: (2026)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
di: Arzaghi', 'Mina, et al.
Pubblicazione: (2025)
di: Arzaghi', 'Mina, et al.
Pubblicazione: (2025)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
di: Xu, Xin, et al.
Pubblicazione: (2025)
di: Xu, Xin, et al.
Pubblicazione: (2025)
Fast Forwarding Low-Rank Training
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
SAEs Are Good for Steering -- If You Select the Right Features
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
Silent Tokens, Loud Effects: Padding in LLMs
di: Himelstein, Rom, et al.
Pubblicazione: (2025)
di: Himelstein, Rom, et al.
Pubblicazione: (2025)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
di: Hanna, Michael, et al.
Pubblicazione: (2024)
di: Hanna, Michael, et al.
Pubblicazione: (2024)
Fairness-Aware Graph Representation Learning with Limited Demographic Information
di: Wang, Zichong, et al.
Pubblicazione: (2025)
di: Wang, Zichong, et al.
Pubblicazione: (2025)
GG-BBQ: German Gender Bias Benchmark for Question Answering
di: Satheesh, Shalaka, et al.
Pubblicazione: (2025)
di: Satheesh, Shalaka, et al.
Pubblicazione: (2025)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
di: Abels, Axel, et al.
Pubblicazione: (2025)
di: Abels, Axel, et al.
Pubblicazione: (2025)
LIBRA: Measuring Bias of Large Language Model from a Local Context
di: Pang, Bo, et al.
Pubblicazione: (2025)
di: Pang, Bo, et al.
Pubblicazione: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
di: Katz, Shahar, et al.
Pubblicazione: (2024)
di: Katz, Shahar, et al.
Pubblicazione: (2024)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2026)
di: Itzhak, Itay, et al.
Pubblicazione: (2026)
Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study
di: Wolfe, Robert, et al.
Pubblicazione: (2024)
di: Wolfe, Robert, et al.
Pubblicazione: (2024)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
di: Contestabile, Matilde, et al.
Pubblicazione: (2025)
di: Contestabile, Matilde, et al.
Pubblicazione: (2025)
Addressing Stereotypes in Large Language Models: A Critical Examination and Mitigation
di: Kazi, Fatima
Pubblicazione: (2025)
di: Kazi, Fatima
Pubblicazione: (2025)
Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking
di: Prakash, Nikhil, et al.
Pubblicazione: (2024)
di: Prakash, Nikhil, et al.
Pubblicazione: (2024)
Leveraging Social Determinants of Health in Alzheimer's Research Using LLM-Augmented Literature Mining and Knowledge Graphs
di: Shang, Tianqi, et al.
Pubblicazione: (2024)
di: Shang, Tianqi, et al.
Pubblicazione: (2024)
Unintended Impacts of LLM Alignment on Global Representation
di: Ryan, Michael J., et al.
Pubblicazione: (2024)
di: Ryan, Michael J., et al.
Pubblicazione: (2024)
Representation Surgery: Theory and Practice of Affine Steering
di: Singh, Shashwat, et al.
Pubblicazione: (2024)
di: Singh, Shashwat, et al.
Pubblicazione: (2024)
Interpreting Latent Student Knowledge Representations in Programming Assignments
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
The Geometric Price of Discrete Logic: Context-driven Manifold Dynamics of Number Representations
di: Zhang, Long, et al.
Pubblicazione: (2026)
di: Zhang, Long, et al.
Pubblicazione: (2026)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
A Representation-Level Assessment of Bias Mitigation in Foundation Models
di: Nizhnichenkov, Svetoslav, et al.
Pubblicazione: (2026)
di: Nizhnichenkov, Svetoslav, et al.
Pubblicazione: (2026)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
Mitigating Social Desirability Bias in Random Silicon Sampling
di: Chapala, Sashank, et al.
Pubblicazione: (2025)
di: Chapala, Sashank, et al.
Pubblicazione: (2025)
Breaking Down Bias: On The Limits of Generalizable Pruning Strategies
di: Ma, Sibo, et al.
Pubblicazione: (2025)
di: Ma, Sibo, et al.
Pubblicazione: (2025)
Bias and Fairness in Large Language Models: A Survey
di: Gallegos, Isabel O., et al.
Pubblicazione: (2023)
di: Gallegos, Isabel O., et al.
Pubblicazione: (2023)
Social Determinants of Health Prediction for ICD-9 Code with Reasoning Models
di: Khan, Sharim, et al.
Pubblicazione: (2025)
di: Khan, Sharim, et al.
Pubblicazione: (2025)
Deep Learning Approaches for Detecting Adversarial Cyberbullying and Hate Speech in Social Networks
di: Azumah, Sylvia Worlali, et al.
Pubblicazione: (2024)
di: Azumah, Sylvia Worlali, et al.
Pubblicazione: (2024)
Toward Automated Detection of Biased Social Signals from the Content of Clinical Conversations
di: Chen, Feng, et al.
Pubblicazione: (2024)
di: Chen, Feng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
di: Majumdar, Ayan, et al.
Pubblicazione: (2025) -
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
di: Cunningham, Eoghan, et al.
Pubblicazione: (2025) -
Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
di: Itzhak, Itay, et al.
Pubblicazione: (2023) -
DSO: Direct Steering Optimization for Bias Mitigation
di: Paes, Lucas Monteiro, et al.
Pubblicazione: (2025) -
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
di: Vargas, Francisco, et al.
Pubblicazione: (2020)