Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jhaveri, Ayush Rajesh, GX-Chen, Anthony, Sucholutsky, Ilia, Choi, Eunsol |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Interpreting and Mitigating Unwanted Uncertainty in LLMs
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
Neural Neural Scaling Laws
von: Hu, Michael Y., et al.
Veröffentlicht: (2026)
von: Hu, Michael Y., et al.
Veröffentlicht: (2026)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
von: Lake, Thom, et al.
Veröffentlicht: (2024)
von: Lake, Thom, et al.
Veröffentlicht: (2024)
Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models
von: Robinson, Sasha, et al.
Veröffentlicht: (2026)
von: Robinson, Sasha, et al.
Veröffentlicht: (2026)
Bias in Large Language Models: Origin, Evaluation, and Mitigation
von: Guo, Yufei, et al.
Veröffentlicht: (2024)
von: Guo, Yufei, et al.
Veröffentlicht: (2024)
Analyzing the Roles of Language and Vision in Learning from Limited Data
von: Chen, Allison, et al.
Veröffentlicht: (2024)
von: Chen, Allison, et al.
Veröffentlicht: (2024)
Exploring Design Choices for Building Language-Specific LLMs
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
Large Language Models Assume People are More Rational than We Really are
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2023)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
von: Bai, Xuechunzi, et al.
Veröffentlicht: (2024)
von: Bai, Xuechunzi, et al.
Veröffentlicht: (2024)
MedLM: Exploring Language Models for Medical Question Answering Systems
von: Yagnik, Niraj, et al.
Veröffentlicht: (2024)
von: Yagnik, Niraj, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
Mitigate Position Bias in Large Language Models via Scaling a Single Dimension
von: Yu, Yijiong, et al.
Veröffentlicht: (2024)
von: Yu, Yijiong, et al.
Veröffentlicht: (2024)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
von: Adiga, Rishabh, et al.
Veröffentlicht: (2024)
von: Adiga, Rishabh, et al.
Veröffentlicht: (2024)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
von: Sriram, Aniruddh, et al.
Veröffentlicht: (2024)
von: Sriram, Aniruddh, et al.
Veröffentlicht: (2024)
The Confidence Trap: Gender Bias and Predictive Certainty in LLMs
von: Sabir, Ahmed, et al.
Veröffentlicht: (2026)
von: Sabir, Ahmed, et al.
Veröffentlicht: (2026)
A Representation-Level Assessment of Bias Mitigation in Foundation Models
von: Nizhnichenkov, Svetoslav, et al.
Veröffentlicht: (2026)
von: Nizhnichenkov, Svetoslav, et al.
Veröffentlicht: (2026)
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
von: Verine, Alexandre, et al.
Veröffentlicht: (2025)
von: Verine, Alexandre, et al.
Veröffentlicht: (2025)
Mitigating Bias in RAG: Controlling the Embedder
von: Kim, Taeyoun, et al.
Veröffentlicht: (2025)
von: Kim, Taeyoun, et al.
Veröffentlicht: (2025)
Intrinsic Meets Extrinsic Fairness: Assessing the Downstream Impact of Bias Mitigation in Large Language Models
von: Arzaghi', 'Mina, et al.
Veröffentlicht: (2025)
von: Arzaghi', 'Mina, et al.
Veröffentlicht: (2025)
BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language Models
von: Chang, Yupeng, et al.
Veröffentlicht: (2024)
von: Chang, Yupeng, et al.
Veröffentlicht: (2024)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
von: Ma, Mingyu Derek, et al.
Veröffentlicht: (2023)
Mitigating Gender Bias in Contextual Word Embeddings
von: Yarrabelly, Navya, et al.
Veröffentlicht: (2024)
von: Yarrabelly, Navya, et al.
Veröffentlicht: (2024)
Combating Confirmation Bias: A Unified Pseudo-Labeling Framework for Entity Alignment
von: Ding, Qijie, et al.
Veröffentlicht: (2023)
von: Ding, Qijie, et al.
Veröffentlicht: (2023)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
von: Le, Khoi, et al.
Veröffentlicht: (2026)
von: Le, Khoi, et al.
Veröffentlicht: (2026)
Diagnosing and Mitigating System Bias in Self-Rewarding RL
von: Tan, Chuyi, et al.
Veröffentlicht: (2025)
von: Tan, Chuyi, et al.
Veröffentlicht: (2025)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
von: Kabra, Sanchit, et al.
Veröffentlicht: (2025)
von: Kabra, Sanchit, et al.
Veröffentlicht: (2025)
DSO: Direct Steering Optimization for Bias Mitigation
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2025)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2025)
Soft-prompt Tuning for Large Language Models to Evaluate Bias
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
von: Arora, Shane, et al.
Veröffentlicht: (2024)
von: Arora, Shane, et al.
Veröffentlicht: (2024)
Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
von: Xie, Qiming, et al.
Veröffentlicht: (2023)
von: Xie, Qiming, et al.
Veröffentlicht: (2023)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
von: Gueorguieva, Anna-Maria, et al.
Veröffentlicht: (2025)
von: Gueorguieva, Anna-Maria, et al.
Veröffentlicht: (2025)
Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
von: Vargas, Francisco, et al.
Veröffentlicht: (2020)
Towards Transfer Unlearning: Empirical Evidence of Cross-Domain Bias Mitigation
von: Lu, Huimin, et al.
Veröffentlicht: (2024)
von: Lu, Huimin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Interpreting and Mitigating Unwanted Uncertainty in LLMs
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025) -
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025) -
Neural Neural Scaling Laws
von: Hu, Michael Y., et al.
Veröffentlicht: (2026) -
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
von: Lake, Thom, et al.
Veröffentlicht: (2024) -
Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models
von: Robinson, Sasha, et al.
Veröffentlicht: (2026)