DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yin, Lake, Huang, Fan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Implicit Bias in LLMs: A Survey
par: Lin, Xinru, et autres
Publié: (2025)
par: Lin, Xinru, et autres
Publié: (2025)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
par: Mirza, Imran, et autres
Publié: (2025)
par: Mirza, Imran, et autres
Publié: (2025)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
par: Zhang, Ming, et autres
Publié: (2026)
par: Zhang, Ming, et autres
Publié: (2026)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
par: Gupta, Shashank, et autres
Publié: (2023)
par: Gupta, Shashank, et autres
Publié: (2023)
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
par: Liu, Ziyi, et autres
Publié: (2025)
par: Liu, Ziyi, et autres
Publié: (2025)
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
par: Wen, Xumeng, et autres
Publié: (2025)
par: Wen, Xumeng, et autres
Publié: (2025)
MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
par: Li, Yanlin, et autres
Publié: (2025)
par: Li, Yanlin, et autres
Publié: (2025)
Finding Words Associated with DIF: Predicting Differential Item Functioning using LLMs and Explainable AI
par: Maeda, Hotaka, et autres
Publié: (2025)
par: Maeda, Hotaka, et autres
Publié: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
par: Vedula, Bhaskara Hanuma, et autres
Publié: (2026)
par: Vedula, Bhaskara Hanuma, et autres
Publié: (2026)
BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs
par: Fan, Zhiting, et autres
Publié: (2024)
par: Fan, Zhiting, et autres
Publié: (2024)
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
par: Pelosio, Giulio, et autres
Publié: (2025)
par: Pelosio, Giulio, et autres
Publié: (2025)
eDIF: A European Deep Inference Fabric for Remote Interpretability of LLM
par: Guggenberger, Irma Heithoff. Marc, et autres
Publié: (2025)
par: Guggenberger, Irma Heithoff. Marc, et autres
Publié: (2025)
PRISM: A Unified Framework for Post-Training LLMs Without Verifiable Rewards
par: Ghimire, Mukesh, et autres
Publié: (2026)
par: Ghimire, Mukesh, et autres
Publié: (2026)
Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs
par: Yang, Chen, et autres
Publié: (2025)
par: Yang, Chen, et autres
Publié: (2025)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
par: Arbabi, Alireza, et autres
Publié: (2025)
par: Arbabi, Alireza, et autres
Publié: (2025)
Are LLMs Rational Investors? A Study on Detecting and Reducing the Financial Bias in LLMs
par: Zhou, Yuhang, et autres
Publié: (2024)
par: Zhou, Yuhang, et autres
Publié: (2024)
'Since Lawyers are Males..': Examining Implicit Gender Bias in Hindi Language Generation by LLMs
par: Joshi, Ishika, et autres
Publié: (2024)
par: Joshi, Ishika, et autres
Publié: (2024)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
par: Kumar, Divyanshu, et autres
Publié: (2024)
par: Kumar, Divyanshu, et autres
Publié: (2024)
M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency
par: Ansari, Abolfazl, et autres
Publié: (2026)
par: Ansari, Abolfazl, et autres
Publié: (2026)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
par: Hanif, Ikhlasul Akmal, et autres
Publié: (2026)
par: Hanif, Ikhlasul Akmal, et autres
Publié: (2026)
Culturally-Aware Conversations: A Framework & Benchmark for LLMs
par: Havaldar, Shreya, et autres
Publié: (2025)
par: Havaldar, Shreya, et autres
Publié: (2025)
User-Assistant Bias in LLMs
par: Pan, Xu, et autres
Publié: (2025)
par: Pan, Xu, et autres
Publié: (2025)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
par: Liu, Yanming, et autres
Publié: (2026)
par: Liu, Yanming, et autres
Publié: (2026)
A Scalable Entity-Based Framework for Auditing Bias in LLMs
par: Elbouanani, Akram, et autres
Publié: (2026)
par: Elbouanani, Akram, et autres
Publié: (2026)
The Bias is in the Details: An Assessment of Cognitive Bias in LLMs
par: Knipper, R. Alexander, et autres
Publié: (2025)
par: Knipper, R. Alexander, et autres
Publié: (2025)
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
par: Fan, Mingyuan, et autres
Publié: (2026)
par: Fan, Mingyuan, et autres
Publié: (2026)
Evaluating Gender Bias of LLMs in Making Morality Judgements
par: Bajaj, Divij, et autres
Publié: (2024)
par: Bajaj, Divij, et autres
Publié: (2024)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
par: Xu, Wenda, et autres
Publié: (2025)
par: Xu, Wenda, et autres
Publié: (2025)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
par: Liu, Shudong, et autres
Publié: (2025)
par: Liu, Shudong, et autres
Publié: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
par: Majumdar, Ayan, et autres
Publié: (2025)
par: Majumdar, Ayan, et autres
Publié: (2025)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
par: Bai, Xuechunzi, et autres
Publié: (2024)
par: Bai, Xuechunzi, et autres
Publié: (2024)
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
par: Shan, Liang, et autres
Publié: (2025)
par: Shan, Liang, et autres
Publié: (2025)
PRIME: A Process-Outcome Alignment Benchmark for Verifiable Reasoning in Mathematics and Engineering
par: Wang, Xiangfeng, et autres
Publié: (2026)
par: Wang, Xiangfeng, et autres
Publié: (2026)
MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models
par: Cai, Yuanqing, et autres
Publié: (2026)
par: Cai, Yuanqing, et autres
Publié: (2026)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
par: Huang, Jen-tse, et autres
Publié: (2025)
par: Huang, Jen-tse, et autres
Publié: (2025)
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
par: Huang, Ruiquan, et autres
Publié: (2025)
par: Huang, Ruiquan, et autres
Publié: (2025)
Lessons from Training Grounded LLMs with Verifiable Rewards
par: Sim, Shang Hong, et autres
Publié: (2025)
par: Sim, Shang Hong, et autres
Publié: (2025)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
par: Wang, Siyuan, et autres
Publié: (2024)
par: Wang, Siyuan, et autres
Publié: (2024)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
par: Zahraei, Pardis Sadat, et autres
Publié: (2025)
par: Zahraei, Pardis Sadat, et autres
Publié: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
par: Borah, Angana, et autres
Publié: (2024)
par: Borah, Angana, et autres
Publié: (2024)
Documents similaires
-
Implicit Bias in LLMs: A Survey
par: Lin, Xinru, et autres
Publié: (2025) -
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
par: Mirza, Imran, et autres
Publié: (2025) -
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
par: Zhang, Ming, et autres
Publié: (2026) -
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
par: Gupta, Shashank, et autres
Publié: (2023) -
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
par: Liu, Ziyi, et autres
Publié: (2025)