Don't Change My View: Ideological Bias Auditing in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kröger, Paul, Barkett, Emilio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Status Hierarchies in Language Models
von: Barkett, Emilio
Veröffentlicht: (2026)
von: Barkett, Emilio
Veröffentlicht: (2026)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
Getting out of the Big-Muddy: Escalation of Commitment in LLMs
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
von: Barkett, Emilio, et al.
Veröffentlicht: (2025)
Language Models Don't Learn the Physical Manifestation of Language
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
von: de Carvalho, Gonçalo Hora, et al.
Veröffentlicht: (2024)
von: de Carvalho, Gonçalo Hora, et al.
Veröffentlicht: (2024)
Don't Pay Attention
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
von: Hammoud, Mohammad, et al.
Veröffentlicht: (2025)
The Compulsory Imaginary: AGI and Corporate Authority
von: Barkett, Emilio
Veröffentlicht: (2026)
von: Barkett, Emilio
Veröffentlicht: (2026)
Do Retrieval Augmented Language Models Know When They Don't Know?
von: Zhou, Youchao, et al.
Veröffentlicht: (2025)
von: Zhou, Youchao, et al.
Veröffentlicht: (2025)
Reasoning Models Reason Well, Until They Don't
von: Rameshkumar, Revanth, et al.
Veröffentlicht: (2025)
von: Rameshkumar, Revanth, et al.
Veröffentlicht: (2025)
Representation Without Control: Testing the Realization Effect in Language Models
von: Walsh, Ciarán, et al.
Veröffentlicht: (2026)
von: Walsh, Ciarán, et al.
Veröffentlicht: (2026)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
von: Yan, Shaotian, et al.
Veröffentlicht: (2025)
von: Yan, Shaotian, et al.
Veröffentlicht: (2025)
Don't Believe Everything You Read: Enhancing Summarization Interpretability through Automatic Identification of Hallucinations in Large Language Models
von: Vakharia, Priyesh, et al.
Veröffentlicht: (2023)
von: Vakharia, Priyesh, et al.
Veröffentlicht: (2023)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
von: Qiu, Peiran, et al.
Veröffentlicht: (2025)
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
von: Hariri, Mohsen, et al.
Veröffentlicht: (2025)
von: Hariri, Mohsen, et al.
Veröffentlicht: (2025)
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
von: Kim, Woojin, et al.
Veröffentlicht: (2025)
von: Kim, Woojin, et al.
Veröffentlicht: (2025)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
von: Mao, Xin, et al.
Veröffentlicht: (2024)
von: Mao, Xin, et al.
Veröffentlicht: (2024)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
von: Parmar, Mihir, et al.
Veröffentlicht: (2022)
von: Parmar, Mihir, et al.
Veröffentlicht: (2022)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
von: Bavaresco, A., et al.
Veröffentlicht: (2024)
von: Bavaresco, A., et al.
Veröffentlicht: (2024)
What's in a Name? Auditing Large Language Models for Race and Gender Bias
von: Salinas, Alejandro, et al.
Veröffentlicht: (2024)
von: Salinas, Alejandro, et al.
Veröffentlicht: (2024)
AuditWen:An Open-Source Large Language Model for Audit
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
von: Huang, Jiajia, et al.
Veröffentlicht: (2024)
Can AI Assistants Know What They Don't Know?
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
von: Chen, Xinxi, et al.
Veröffentlicht: (2024)
von: Chen, Xinxi, et al.
Veröffentlicht: (2024)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
von: Mei, Zhiting, et al.
Veröffentlicht: (2025)
Change My Frame: Reframing in the Wild in r/ChangeMyView
von: Peguero, Arturo Martínez, et al.
Veröffentlicht: (2024)
von: Peguero, Arturo Martínez, et al.
Veröffentlicht: (2024)
Reasoning Models Don't Always Say What They Think
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
von: Lacombe, Romain, et al.
Veröffentlicht: (2025)
von: Lacombe, Romain, et al.
Veröffentlicht: (2025)
sDPO: Don't Use Your Data All at Once
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
von: Chochlakis, Georgios, et al.
Veröffentlicht: (2024)
von: Chochlakis, Georgios, et al.
Veröffentlicht: (2024)
Mapping and Influencing the Political Ideology of Large Language Models using Synthetic Personas
von: Bernardelle, Pietro, et al.
Veröffentlicht: (2024)
von: Bernardelle, Pietro, et al.
Veröffentlicht: (2024)
Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load
von: Mann, Logan, et al.
Veröffentlicht: (2025)
von: Mann, Logan, et al.
Veröffentlicht: (2025)
CALM: Curiosity-Driven Auditing for Large Language Models
von: Zheng, Xiang, et al.
Veröffentlicht: (2025)
von: Zheng, Xiang, et al.
Veröffentlicht: (2025)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
Regional Bias in Large Language Models
von: Gopinadh, M P V S, et al.
Veröffentlicht: (2026)
von: Gopinadh, M P V S, et al.
Veröffentlicht: (2026)
Behavioral Bias of Vision-Language Models: A Behavioral Finance View
von: Xiao, Yuhang, et al.
Veröffentlicht: (2024)
von: Xiao, Yuhang, et al.
Veröffentlicht: (2024)
Output Scouting: Auditing Large Language Models for Catastrophic Responses
von: Bell, Andrew, et al.
Veröffentlicht: (2024)
von: Bell, Andrew, et al.
Veröffentlicht: (2024)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2026)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2026)
Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
von: Liu, Jinda, et al.
Veröffentlicht: (2025)
von: Liu, Jinda, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Status Hierarchies in Language Models
von: Barkett, Emilio
Veröffentlicht: (2026) -
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
von: Barkett, Emilio, et al.
Veröffentlicht: (2025) -
Getting out of the Big-Muddy: Escalation of Commitment in LLMs
von: Barkett, Emilio, et al.
Veröffentlicht: (2025) -
Language Models Don't Learn the Physical Manifestation of Language
von: Lee, Bruce W., et al.
Veröffentlicht: (2024) -
Large Language Models Must Be Taught to Know What They Don't Know
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)