Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Borah, Angana, Mihalcea, Rada |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Region-aware Bias Evaluation Metrics
von: Borah, Angana, et al.
Veröffentlicht: (2024)
von: Borah, Angana, et al.
Veröffentlicht: (2024)
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions
von: Borah, Angana, et al.
Veröffentlicht: (2025)
von: Borah, Angana, et al.
Veröffentlicht: (2025)
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
von: Arif, Samee, et al.
Veröffentlicht: (2026)
von: Arif, Samee, et al.
Veröffentlicht: (2026)
The Curious Case of Curiosity across Human Cultures and LLMs
von: Borah, Angana, et al.
Veröffentlicht: (2025)
von: Borah, Angana, et al.
Veröffentlicht: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
Mind the (Belief) Gap: Group Identity in the World of LLMs
von: Borah, Angana, et al.
Veröffentlicht: (2025)
von: Borah, Angana, et al.
Veröffentlicht: (2025)
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
von: Borah, Angana, et al.
Veröffentlicht: (2026)
von: Borah, Angana, et al.
Veröffentlicht: (2026)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
von: Mirza, Imran, et al.
Veröffentlicht: (2025)
von: Mirza, Imran, et al.
Veröffentlicht: (2025)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
von: Mori, Shinka, et al.
Veröffentlicht: (2024)
von: Mori, Shinka, et al.
Veröffentlicht: (2024)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
von: Backmann, Steffen, et al.
Veröffentlicht: (2025)
von: Backmann, Steffen, et al.
Veröffentlicht: (2025)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
Why AI Is WEIRD and Should Not Be This Way: Towards AI For Everyone, With Everyone, By Everyone
von: Mihalcea, Rada, et al.
Veröffentlicht: (2024)
von: Mihalcea, Rada, et al.
Veröffentlicht: (2024)
Implicit Personalization in Language Models: A Systematic Study
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
von: Ignat, Oana, et al.
Veröffentlicht: (2023)
von: Ignat, Oana, et al.
Veröffentlicht: (2023)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
von: Nwatu, Joan, et al.
Veröffentlicht: (2024)
von: Nwatu, Joan, et al.
Veröffentlicht: (2024)
Covert Bias: The Severity of Social Views' Unalignment in Language Models Towards Implicit and Explicit Opinion
von: Aldayel, Abeer, et al.
Veröffentlicht: (2024)
von: Aldayel, Abeer, et al.
Veröffentlicht: (2024)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
von: Bai, Xuechunzi, et al.
Veröffentlicht: (2024)
von: Bai, Xuechunzi, et al.
Veröffentlicht: (2024)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
von: Chand, Shireen, et al.
Veröffentlicht: (2025)
von: Chand, Shireen, et al.
Veröffentlicht: (2025)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
von: Nwatu, Joan, et al.
Veröffentlicht: (2025)
von: Nwatu, Joan, et al.
Veröffentlicht: (2025)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
Mitigating Social Desirability Bias in Random Silicon Sampling
von: Chapala, Sashank, et al.
Veröffentlicht: (2025)
von: Chapala, Sashank, et al.
Veröffentlicht: (2025)
How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
von: Bai, Longju, et al.
Veröffentlicht: (2026)
von: Bai, Longju, et al.
Veröffentlicht: (2026)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
VIGIL: An Extensible System for Real-Time Detection and Mitigation of Cognitive Bias Triggers
von: Kang, Bo, et al.
Veröffentlicht: (2026)
von: Kang, Bo, et al.
Veröffentlicht: (2026)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
von: Apsel, Molly, et al.
Veröffentlicht: (2026)
von: Apsel, Molly, et al.
Veröffentlicht: (2026)
Counterspeech for Mitigating the Influence of Media Bias: Comparing Human and LLM-Generated Responses
von: Lin, Luyang, et al.
Veröffentlicht: (2025)
von: Lin, Luyang, et al.
Veröffentlicht: (2025)
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification
von: Abzaliev, Artem, et al.
Veröffentlicht: (2024)
von: Abzaliev, Artem, et al.
Veröffentlicht: (2024)
DSO: Direct Steering Optimization for Bias Mitigation
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2025)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2025)
Rethinking Table Instruction Tuning
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
von: Deng, Naihao, et al.
Veröffentlicht: (2025)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
von: Sun, Lihao, et al.
Veröffentlicht: (2025)
von: Sun, Lihao, et al.
Veröffentlicht: (2025)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
von: Weissburg, Iain, et al.
Veröffentlicht: (2024)
Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance
von: Chen, Yupeng, et al.
Veröffentlicht: (2025)
von: Chen, Yupeng, et al.
Veröffentlicht: (2025)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
von: Piatti, Giorgio, et al.
Veröffentlicht: (2024)
von: Piatti, Giorgio, et al.
Veröffentlicht: (2024)
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias
von: Chen, Yuen, et al.
Veröffentlicht: (2022)
von: Chen, Yuen, et al.
Veröffentlicht: (2022)
Application Specific Compression of Deep Learning Models
von: Rai, Rohit Raj, et al.
Veröffentlicht: (2024)
von: Rai, Rohit Raj, et al.
Veröffentlicht: (2024)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
NLP for Social Good: A Survey and Outlook of Challenges, Opportunities, and Responsible Deployment
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2025)
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2025)
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
von: Wei, Kangda, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Region-aware Bias Evaluation Metrics
von: Borah, Angana, et al.
Veröffentlicht: (2024) -
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions
von: Borah, Angana, et al.
Veröffentlicht: (2025) -
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
von: Arif, Samee, et al.
Veröffentlicht: (2026) -
The Curious Case of Curiosity across Human Cultures and LLMs
von: Borah, Angana, et al.
Veröffentlicht: (2025) -
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)