MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
Fuente:
arXiv
Saved in:
| Main Authors: | Mirza, Imran, Huang, Cole, Vasista, Ishwara, Patil, Rohan, Akalin, Asli, O'Brien, Sean, Zhu, Kevin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
by: Borah, Angana, et al.
Published: (2024)
by: Borah, Angana, et al.
Published: (2024)
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025)
by: Hashmat, Abdullah, et al.
Published: (2025)
FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
by: Wen, Athena, et al.
Published: (2025)
by: Wen, Athena, et al.
Published: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
by: Jaipersaud, Brandon, et al.
Published: (2024)
by: Jaipersaud, Brandon, et al.
Published: (2024)
From Chat Control to Robot Control: Implications of the Chat Control Proposal for Human-Robot Interaction
by: Akalin, Neziha, et al.
Published: (2026)
by: Akalin, Neziha, et al.
Published: (2026)
AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark
by: Gupta, Abhay, et al.
Published: (2024)
by: Gupta, Abhay, et al.
Published: (2024)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
by: Liu, Joshua, et al.
Published: (2025)
by: Liu, Joshua, et al.
Published: (2025)
Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education
by: Gupta, Amogh, et al.
Published: (2026)
by: Gupta, Amogh, et al.
Published: (2026)
Unmasking Implicit Bias: Evaluating Persona-Prompted LLM Responses in Power-Disparate Social Scenarios
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
Probing Audio-Generation Capabilities of Text-Based Language Models
by: Anbazhagan, Arjun Prasaath, et al.
Published: (2025)
by: Anbazhagan, Arjun Prasaath, et al.
Published: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
by: Qu, Jinxian, et al.
Published: (2026)
by: Qu, Jinxian, et al.
Published: (2026)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
by: Bai, Xuechunzi, et al.
Published: (2024)
by: Bai, Xuechunzi, et al.
Published: (2024)
Regulating the Agency of LLM-based Agents
by: Boddy, Seán, et al.
Published: (2025)
by: Boddy, Seán, et al.
Published: (2025)
Implicit Bias in LLMs for Transgender Populations
by: Hirsch, Micaela, et al.
Published: (2026)
by: Hirsch, Micaela, et al.
Published: (2026)
Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations
by: Charnock, Jacob, et al.
Published: (2026)
by: Charnock, Jacob, et al.
Published: (2026)
BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models
by: Luo, Hanjun, et al.
Published: (2026)
by: Luo, Hanjun, et al.
Published: (2026)
Implicit Bias-Like Patterns in Reasoning Models
by: Lee, Messi H. J., et al.
Published: (2025)
by: Lee, Messi H. J., et al.
Published: (2025)
Safe-Child-LLM: A Developmental Benchmark for Evaluating LLM Safety in Child-LLM Interactions
by: Jiao, Junfeng, et al.
Published: (2025)
by: Jiao, Junfeng, et al.
Published: (2025)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
by: Apsel, Molly, et al.
Published: (2026)
by: Apsel, Molly, et al.
Published: (2026)
BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
by: Myakala, Praveen Kumar, et al.
Published: (2026)
by: Myakala, Praveen Kumar, et al.
Published: (2026)
When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
by: Wang, Zongwei, et al.
Published: (2026)
by: Wang, Zongwei, et al.
Published: (2026)
Covert Bias: The Severity of Social Views' Unalignment in Language Models Towards Implicit and Explicit Opinion
by: Aldayel, Abeer, et al.
Published: (2024)
by: Aldayel, Abeer, et al.
Published: (2024)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
by: Sun, Lihao, et al.
Published: (2025)
by: Sun, Lihao, et al.
Published: (2025)
From Bias to Balance: Detecting Facial Expression Recognition Biases in Large Multimodal Foundation Models
by: Chhua, Kaylee, et al.
Published: (2024)
by: Chhua, Kaylee, et al.
Published: (2024)
Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles
by: Haq, Irti, et al.
Published: (2026)
by: Haq, Irti, et al.
Published: (2026)
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception
by: Lin, Luyang, et al.
Published: (2024)
by: Lin, Luyang, et al.
Published: (2024)
APS: Bias-Controlled Adaptive Prototype Simulation for Population-Scale LLM Agents
by: Zheng, Quan, et al.
Published: (2026)
by: Zheng, Quan, et al.
Published: (2026)
Risk Reporting for Developers' Internal AI Model Use
by: Delaney, Oscar, et al.
Published: (2026)
by: Delaney, Oscar, et al.
Published: (2026)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025)
by: Sun, Shuzhou, et al.
Published: (2025)
Benchmarking AI Performance on End-to-End Data Science Projects
by: Hughes, Evelyn, et al.
Published: (2026)
by: Hughes, Evelyn, et al.
Published: (2026)
Demographic Benchmarking: Bridging Socio-Technical Gaps in Bias Detection
by: Clavell, Gemma Galdon, et al.
Published: (2025)
by: Clavell, Gemma Galdon, et al.
Published: (2025)
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
by: Gupta, Abhay, et al.
Published: (2025)
by: Gupta, Abhay, et al.
Published: (2025)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
by: Walsh, Cole, et al.
Published: (2026)
by: Walsh, Cole, et al.
Published: (2026)
Beyond English: Unveiling Multilingual Bias in LLM Copyright Compliance
by: Chen, Yupeng, et al.
Published: (2025)
by: Chen, Yupeng, et al.
Published: (2025)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
by: Weissburg, Iain, et al.
Published: (2024)
by: Weissburg, Iain, et al.
Published: (2024)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Social Bias in Popular Question-Answering Benchmarks
by: Kraft, Angelie, et al.
Published: (2025)
by: Kraft, Angelie, et al.
Published: (2025)
Auditing LLM Editorial Bias in News Media Exposure
by: Minici, Marco, et al.
Published: (2025)
by: Minici, Marco, et al.
Published: (2025)
Similar Items
-
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
by: Borah, Angana, et al.
Published: (2024) -
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024) -
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024) -
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025) -
FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
by: Wen, Athena, et al.
Published: (2025)