On the Inevitability of Left-Leaning Political Bias in Aligned Language Models
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Hagendorff, Thilo |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
par: Knecht, Amelie, et autres
Publié: (2026)
par: Knecht, Amelie, et autres
Publié: (2026)
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
par: Menke, Maluna, et autres
Publié: (2025)
par: Menke, Maluna, et autres
Publié: (2025)
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
par: Jotautaitė, Monika, et autres
Publié: (2025)
par: Jotautaitė, Monika, et autres
Publié: (2025)
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
par: Vaugrante, Laurène, et autres
Publié: (2025)
par: Vaugrante, Laurène, et autres
Publié: (2025)
Deception Abilities Emerged in Large Language Models
par: Hagendorff, Thilo
Publié: (2023)
par: Hagendorff, Thilo
Publié: (2023)
"Amazing, They All Lean Left" -- Analyzing the Political Temperaments of Current LLMs
par: Neuman, W. Russell, et autres
Publié: (2025)
par: Neuman, W. Russell, et autres
Publié: (2025)
Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
par: Hagendorff, Thilo
Publié: (2024)
par: Hagendorff, Thilo
Publié: (2024)
When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models
par: Schneider, Matthias, et autres
Publié: (2024)
par: Schneider, Matthias, et autres
Publié: (2024)
Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models
par: Faulborn, Mats, et autres
Publié: (2025)
par: Faulborn, Mats, et autres
Publié: (2025)
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
par: Hagendorff, Thilo, et autres
Publié: (2025)
par: Hagendorff, Thilo, et autres
Publié: (2025)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
par: Potter, Yujin, et autres
Publié: (2024)
par: Potter, Yujin, et autres
Publié: (2024)
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
par: Meding, Kristof, et autres
Publié: (2023)
par: Meding, Kristof, et autres
Publié: (2023)
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions
par: Vaugrante, Laurène, et autres
Publié: (2024)
par: Vaugrante, Laurène, et autres
Publié: (2024)
Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment
par: Vaugrante, Laurène, et autres
Publié: (2026)
par: Vaugrante, Laurène, et autres
Publié: (2026)
Political Leaning Inference through Plurinational Scenarios
par: de Landa, Joseba Fernandez, et autres
Publié: (2024)
par: de Landa, Joseba Fernandez, et autres
Publié: (2024)
Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
par: Yang, Kai-Cheng, et autres
Publié: (2023)
par: Yang, Kai-Cheng, et autres
Publié: (2023)
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
par: Bojic, Ljubisa, et autres
Publié: (2025)
par: Bojic, Ljubisa, et autres
Publié: (2025)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
par: Sakhawat, Adib, et autres
Publié: (2026)
par: Sakhawat, Adib, et autres
Publié: (2026)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
par: Jenny, David F., et autres
Publié: (2023)
par: Jenny, David F., et autres
Publié: (2023)
Large Reasoning Models Are Autonomous Jailbreak Agents
par: Hagendorff, Thilo, et autres
Publié: (2025)
par: Hagendorff, Thilo, et autres
Publié: (2025)
Characterizing Selective Refusal Bias in Large Language Models
par: Khorramrouz, Adel, et autres
Publié: (2025)
par: Khorramrouz, Adel, et autres
Publié: (2025)
Gender Bias in Emotion Recognition by Large Language Models
par: Herbert, Maureen, et autres
Publié: (2025)
par: Herbert, Maureen, et autres
Publié: (2025)
Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
par: Dutta, Aditi, et autres
Publié: (2025)
par: Dutta, Aditi, et autres
Publié: (2025)
AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
par: Li, Kun, et autres
Publié: (2025)
par: Li, Kun, et autres
Publié: (2025)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
par: Bai, Xuechunzi, et autres
Publié: (2024)
par: Bai, Xuechunzi, et autres
Publié: (2024)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
par: Derner, Erik, et autres
Publié: (2024)
par: Derner, Erik, et autres
Publié: (2024)
Agent-Enhanced Large Language Models for Researching Political Institutions
par: Loffredo, Joseph R., et autres
Publié: (2025)
par: Loffredo, Joseph R., et autres
Publié: (2025)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
par: Fleisig, Eve, et autres
Publié: (2024)
par: Fleisig, Eve, et autres
Publié: (2024)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
par: DiGiuseppe, Matthew, et autres
Publié: (2026)
par: DiGiuseppe, Matthew, et autres
Publié: (2026)
Benchmarking Political Persuasion Risks Across Frontier Large Language Models
par: Chen, Zhongren, et autres
Publié: (2026)
par: Chen, Zhongren, et autres
Publié: (2026)
Beyond Partisan Leaning: A Comparative Analysis of Political Bias in Large Language Models
par: Peng, Tai-Quan, et autres
Publié: (2024)
par: Peng, Tai-Quan, et autres
Publié: (2024)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
par: Sun, Lihao, et autres
Publié: (2025)
par: Sun, Lihao, et autres
Publié: (2025)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
par: Yilmaz, Berk, et autres
Publié: (2025)
par: Yilmaz, Berk, et autres
Publié: (2025)
Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese
par: Lyu, Hanjia, et autres
Publié: (2025)
par: Lyu, Hanjia, et autres
Publié: (2025)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
par: Dorn, Rebecca, et autres
Publié: (2024)
par: Dorn, Rebecca, et autres
Publié: (2024)
Cross-Language Bias Examination in Large Language Models
par: Liang, Yuxuan, et autres
Publié: (2025)
par: Liang, Yuxuan, et autres
Publié: (2025)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
par: Zahraei, Pardis Sadat, et autres
Publié: (2025)
par: Zahraei, Pardis Sadat, et autres
Publié: (2025)
Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models
par: Apsel, Molly, et autres
Publié: (2026)
par: Apsel, Molly, et autres
Publié: (2026)
Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory
par: Smith-Vaniz, Nicole, et autres
Publié: (2025)
par: Smith-Vaniz, Nicole, et autres
Publié: (2025)
Theories of "Sexuality" in Natural Language Processing Bias Research
par: Hobbs, Jacob
Publié: (2025)
par: Hobbs, Jacob
Publié: (2025)
Documents similaires
-
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
par: Knecht, Amelie, et autres
Publié: (2026) -
PRIDE -- Parameter-Efficient Reduction of Identity Discrimination for Equality in LLMs
par: Menke, Maluna, et autres
Publié: (2025) -
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
par: Jotautaitė, Monika, et autres
Publié: (2025) -
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
par: Vaugrante, Laurène, et autres
Publié: (2025) -
Deception Abilities Emerged in Large Language Models
par: Hagendorff, Thilo
Publié: (2023)