Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Adiga, Rishabh, Nushi, Besmira, Chandrasekaran, Varun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
by: Yuksekgonul, Mert, et al.
Published: (2023)
by: Yuksekgonul, Mert, et al.
Published: (2023)
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)
by: Naik, Ranjita, et al.
Published: (2023)
Designing Informative Metrics for Few-Shot Example Selection
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025)
by: Joshi, Siddharth, et al.
Published: (2025)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
by: Zheng, Jinquan, et al.
Published: (2026)
by: Zheng, Jinquan, et al.
Published: (2026)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Parallax: Parameterized Local Linear Attention for Language Modeling
by: Zuo, Yifei, et al.
Published: (2026)
by: Zuo, Yifei, et al.
Published: (2026)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
by: Kabra, Sanchit, et al.
Published: (2025)
by: Kabra, Sanchit, et al.
Published: (2025)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
Eureka: Evaluating and Understanding Large Foundation Models
by: Balachandran, Vidhisha, et al.
Published: (2024)
by: Balachandran, Vidhisha, et al.
Published: (2024)
Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
by: Balachandran, Vidhisha, et al.
Published: (2025)
by: Balachandran, Vidhisha, et al.
Published: (2025)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
by: Hwang, Jaedong, et al.
Published: (2025)
by: Hwang, Jaedong, et al.
Published: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
by: Chuang, Yung-Sung, et al.
Published: (2024)
by: Chuang, Yung-Sung, et al.
Published: (2024)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
by: Soo, Samuel, et al.
Published: (2025)
by: Soo, Samuel, et al.
Published: (2025)
Mitigating Memorization In Language Models
by: Sakarvadia, Mansi, et al.
Published: (2024)
by: Sakarvadia, Mansi, et al.
Published: (2024)
Mitigating Extrinsic Gender Bias for Bangla Classification Tasks
by: Joy, Sajib Kumar Saha, et al.
Published: (2024)
by: Joy, Sajib Kumar Saha, et al.
Published: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
On Mitigating Code LLM Hallucinations with API Documentation
by: Jain, Nihal, et al.
Published: (2024)
by: Jain, Nihal, et al.
Published: (2024)
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
by: Zelikman, Eric, et al.
Published: (2024)
by: Zelikman, Eric, et al.
Published: (2024)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
by: Cao, Shuirong, et al.
Published: (2024)
by: Cao, Shuirong, et al.
Published: (2024)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
by: Siddique, Zara, et al.
Published: (2025)
by: Siddique, Zara, et al.
Published: (2025)
Generative Monoculture in Large Language Models
by: Wu, Fan, et al.
Published: (2024)
by: Wu, Fan, et al.
Published: (2024)
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
by: Wang, Guorun, et al.
Published: (2024)
by: Wang, Guorun, et al.
Published: (2024)
Speaking the Same Language: Leveraging LLMs in Standardizing Clinical Data for AI
by: Sett, Arindam, et al.
Published: (2024)
by: Sett, Arindam, et al.
Published: (2024)
Mitigating Degree Bias Adaptively with Hard-to-Learn Nodes in Graph Contrastive Learning
by: Hu, Jingyu, et al.
Published: (2025)
by: Hu, Jingyu, et al.
Published: (2025)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
by: Li, Miaomiao, et al.
Published: (2025)
by: Li, Miaomiao, et al.
Published: (2025)
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
by: Manchanda, Sahil, et al.
Published: (2025)
by: Manchanda, Sahil, et al.
Published: (2025)
Bias Amplification in Language Model Evolution: An Iterated Learning Perspective
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
Sociodemographic Bias in Language Models: A Survey and Forward Path
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
Soft-prompt Tuning for Large Language Models to Evaluate Bias
by: Tian, Jacob-Junqi, et al.
Published: (2023)
by: Tian, Jacob-Junqi, et al.
Published: (2023)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models
by: Maheshwary, Rishabh, et al.
Published: (2024)
by: Maheshwary, Rishabh, et al.
Published: (2024)
Similar Items
-
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024) -
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
by: Yuksekgonul, Mert, et al.
Published: (2023) -
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024) -
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
by: Bordt, Sebastian, et al.
Published: (2024) -
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)