Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
Fuente:
arXiv
Saved in:
| Main Authors: | Itzhak, Itay, Stanovsky, Gabriel, Rosenfeld, Nir, Belinkov, Yonatan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
by: Itzhak, Itay, et al.
Published: (2025)
by: Itzhak, Itay, et al.
Published: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
by: Itzhak, Itay, et al.
Published: (2026)
by: Itzhak, Itay, et al.
Published: (2026)
Welfare as a Guiding Principle for Machine Learning -- From Compass, to Lens, to Roadmap
by: Rosenfeld, Nir, et al.
Published: (2025)
by: Rosenfeld, Nir, et al.
Published: (2025)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
by: Iskander, Shadi, et al.
Published: (2024)
by: Iskander, Shadi, et al.
Published: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Backdoor for Debias: Mitigating Model Bias with Backdoor Attack-based Artificial Bias
by: Wu, Shangxi, et al.
Published: (2023)
by: Wu, Shangxi, et al.
Published: (2023)
Mitigating Bias for Question Answering Models by Tracking Bias Influence
by: Ma, Mingyu Derek, et al.
Published: (2023)
by: Ma, Mingyu Derek, et al.
Published: (2023)
Bias and Fairness in Large Language Models: A Survey
by: Gallegos, Isabel O., et al.
Published: (2023)
by: Gallegos, Isabel O., et al.
Published: (2023)
BiasEdit: Debiasing Stereotyped Language Models via Model Editing
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Say My Name: a Model's Bias Discovery Framework
by: Ciranni, Massimiliano, et al.
Published: (2024)
by: Ciranni, Massimiliano, et al.
Published: (2024)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025)
by: Sun, Shuzhou, et al.
Published: (2025)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
by: Salinas, Alejandro, et al.
Published: (2024)
by: Salinas, Alejandro, et al.
Published: (2024)
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
by: Katz, Shahar, et al.
Published: (2024)
by: Katz, Shahar, et al.
Published: (2024)
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
by: Shvartzshnaider, Yan, et al.
Published: (2024)
by: Shvartzshnaider, Yan, et al.
Published: (2024)
Embedding Enhancement via Fine-Tuned Language Models for Learner-Item Cognitive Modeling
by: Liu, Yuanhao, et al.
Published: (2026)
by: Liu, Yuanhao, et al.
Published: (2026)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
Mitigating Gender Bias in Depression Detection via Counterfactual Inference
by: Hu, Mingxuan, et al.
Published: (2025)
by: Hu, Mingxuan, et al.
Published: (2025)
BiasGuard: Guardrailing Fairness in Machine Learning Production Systems
by: Cohen-Inger, Nurit, et al.
Published: (2025)
by: Cohen-Inger, Nurit, et al.
Published: (2025)
Unmasking Bias in AI: A Systematic Review of Bias Detection and Mitigation Strategies in Electronic Health Record-based Models
by: Chen, Feng, et al.
Published: (2023)
by: Chen, Feng, et al.
Published: (2023)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
Effective Controllable Bias Mitigation for Classification and Retrieval using Gate Adapters
by: Masoudian, Shahed, et al.
Published: (2024)
by: Masoudian, Shahed, et al.
Published: (2024)
Different Horses for Different Courses: Comparing Bias Mitigation Algorithms in ML
by: Ganesh, Prakhar, et al.
Published: (2024)
by: Ganesh, Prakhar, et al.
Published: (2024)
Demographic Bias of Expert-Level Vision-Language Foundation Models in Medical Imaging
by: Yang, Yuzhe, et al.
Published: (2024)
by: Yang, Yuzhe, et al.
Published: (2024)
Breaking Down Bias: On The Limits of Generalizable Pruning Strategies
by: Ma, Sibo, et al.
Published: (2025)
by: Ma, Sibo, et al.
Published: (2025)
Integrating Social Determinants of Health into Knowledge Graphs: Evaluating Prediction Bias and Fairness in Healthcare
by: Shang, Tianqi, et al.
Published: (2024)
by: Shang, Tianqi, et al.
Published: (2024)
Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
by: Sánchez-Cortés, Dairazalia, et al.
Published: (2024)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025)
by: Hashmat, Abdullah, et al.
Published: (2025)
Impacts of Racial Bias in Historical Training Data for News AI
by: Bhargava, Rahul, et al.
Published: (2025)
by: Bhargava, Rahul, et al.
Published: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
by: Porikli, Sedat, et al.
Published: (2025)
by: Porikli, Sedat, et al.
Published: (2025)
Addressing Selection Bias in Computerized Adaptive Testing: A User-Wise Aggregate Influence Function Approach
by: Kwon, Soonwoo, et al.
Published: (2023)
by: Kwon, Soonwoo, et al.
Published: (2023)
Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval
by: Wilson, Kyra, et al.
Published: (2024)
by: Wilson, Kyra, et al.
Published: (2024)
Soft-prompt Tuning for Large Language Models to Evaluate Bias
by: Tian, Jacob-Junqi, et al.
Published: (2023)
by: Tian, Jacob-Junqi, et al.
Published: (2023)
Anticipatory Evaluation of Language Models
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
Pro-AI Bias in Large Language Models
by: Trabelsi, Benaya, et al.
Published: (2026)
by: Trabelsi, Benaya, et al.
Published: (2026)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
Similar Items
-
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
by: Itzhak, Itay, et al.
Published: (2025) -
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
by: Itzhak, Itay, et al.
Published: (2026) -
Welfare as a Guiding Principle for Machine Learning -- From Compass, to Lens, to Roadmap
by: Rosenfeld, Nir, et al.
Published: (2025) -
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
by: Iskander, Shadi, et al.
Published: (2024) -
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)