Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Nadeem, Afrozah, Seth, Agrima, Nasim, Mehwish, Naseem, Usman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Steering Towards Fairness: Mitigating Political Bias in LLMs
por: Nadeem, Afrozah, et al.
Publicado: (2025)
por: Nadeem, Afrozah, et al.
Publicado: (2025)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
por: Nadeem, Afrozah, et al.
Publicado: (2025)
por: Nadeem, Afrozah, et al.
Publicado: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
por: Nadeem, Afrozah, et al.
Publicado: (2025)
por: Nadeem, Afrozah, et al.
Publicado: (2025)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
por: Bhandari, Pranav, et al.
Publicado: (2026)
por: Bhandari, Pranav, et al.
Publicado: (2026)
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
por: Selvaganapathy, Sanjeeevan, et al.
Publicado: (2025)
por: Selvaganapathy, Sanjeeevan, et al.
Publicado: (2025)
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires
por: Bhandari, Pranav, et al.
Publicado: (2025)
por: Bhandari, Pranav, et al.
Publicado: (2025)
Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs
por: Bhandari, Pranav, et al.
Publicado: (2025)
por: Bhandari, Pranav, et al.
Publicado: (2025)
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
por: Mustafa, Akram, et al.
Publicado: (2025)
por: Mustafa, Akram, et al.
Publicado: (2025)
Can Reasoning LLMs Enhance Clinical Document Classification?
por: Mustafa, Akram, et al.
Publicado: (2025)
por: Mustafa, Akram, et al.
Publicado: (2025)
Can LLM Agents Maintain a Persona in Discourse?
por: Bhandari, Pranav, et al.
Publicado: (2025)
por: Bhandari, Pranav, et al.
Publicado: (2025)
Competing LLM Agents in a Non-Cooperative Game of Opinion Polarisation
por: Qasmi, Amin, et al.
Publicado: (2025)
por: Qasmi, Amin, et al.
Publicado: (2025)
Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation
por: Wu, Sophie, et al.
Publicado: (2026)
por: Wu, Sophie, et al.
Publicado: (2026)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
por: Yuan, Chenchen, et al.
Publicado: (2026)
por: Yuan, Chenchen, et al.
Publicado: (2026)
Language, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
por: Pihulski, Dzmitry, et al.
Publicado: (2025)
por: Pihulski, Dzmitry, et al.
Publicado: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
por: Wachter, Jasmin, et al.
Publicado: (2025)
por: Wachter, Jasmin, et al.
Publicado: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
por: Hwang, Yerin, et al.
Publicado: (2026)
por: Hwang, Yerin, et al.
Publicado: (2026)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
por: Banayeeanzade, Amin, et al.
Publicado: (2025)
por: Banayeeanzade, Amin, et al.
Publicado: (2025)
Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
por: Yuan, Jiahao, et al.
Publicado: (2025)
por: Yuan, Jiahao, et al.
Publicado: (2025)
Reliable Reasoning Beyond Natural Language
por: Borazjanizadeh, Nasim, et al.
Publicado: (2024)
por: Borazjanizadeh, Nasim, et al.
Publicado: (2024)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
por: Pres, Itamar, et al.
Publicado: (2024)
por: Pres, Itamar, et al.
Publicado: (2024)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
por: Siddique, Zara, et al.
Publicado: (2025)
por: Siddique, Zara, et al.
Publicado: (2025)
MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
por: Wang, Pengyu, et al.
Publicado: (2025)
por: Wang, Pengyu, et al.
Publicado: (2025)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
por: DiGiuseppe, Matthew, et al.
Publicado: (2026)
por: DiGiuseppe, Matthew, et al.
Publicado: (2026)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
por: Gupta, Ashray, et al.
Publicado: (2025)
por: Gupta, Ashray, et al.
Publicado: (2025)
FIBER: A Multilingual Evaluation Resource for Factual Inference Bias
por: Munis, Evren Ayberk, et al.
Publicado: (2025)
por: Munis, Evren Ayberk, et al.
Publicado: (2025)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
por: Mahran, Mariam, et al.
Publicado: (2025)
por: Mahran, Mariam, et al.
Publicado: (2025)
Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification
por: Elbouanani, Akram, et al.
Publicado: (2025)
por: Elbouanani, Akram, et al.
Publicado: (2025)
Political Bias in LLMs: Unaligned Moral Values in Agent-centric Simulations
por: Münker, Simon
Publicado: (2024)
por: Münker, Simon
Publicado: (2024)
Ideological Bias in LLMs' Economic Causal Reasoning
por: Lee, Donggyu, et al.
Publicado: (2026)
por: Lee, Donggyu, et al.
Publicado: (2026)
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages
por: Azam, Gulfarogh, et al.
Publicado: (2025)
por: Azam, Gulfarogh, et al.
Publicado: (2025)
Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text
por: Healey, Jennifer, et al.
Publicado: (2025)
por: Healey, Jennifer, et al.
Publicado: (2025)
Cross-Lingual Activation Steering for Multilingual Language Models
por: Pokharel, Rhitabrat, et al.
Publicado: (2026)
por: Pokharel, Rhitabrat, et al.
Publicado: (2026)
Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
por: Alakeel, Yara, et al.
Publicado: (2026)
por: Alakeel, Yara, et al.
Publicado: (2026)
Don't Change My View: Ideological Bias Auditing in Large Language Models
por: Kröger, Paul, et al.
Publicado: (2025)
por: Kröger, Paul, et al.
Publicado: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
Simulating Influence Dynamics with LLM Agents
por: Nasim, Mehwish, et al.
Publicado: (2025)
por: Nasim, Mehwish, et al.
Publicado: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
por: Rystrøm, Jonathan, et al.
Publicado: (2025)
por: Rystrøm, Jonathan, et al.
Publicado: (2025)
Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up
por: Yuan, Jiahao, et al.
Publicado: (2024)
por: Yuan, Jiahao, et al.
Publicado: (2024)
Mapping and Influencing the Political Ideology of Large Language Models using Synthetic Personas
por: Bernardelle, Pietro, et al.
Publicado: (2024)
por: Bernardelle, Pietro, et al.
Publicado: (2024)
Enabling Scalable Evaluation of Bias Patterns in Medical LLMs
por: Fayyaz, Hamed, et al.
Publicado: (2024)
por: Fayyaz, Hamed, et al.
Publicado: (2024)
Ejemplares similares
-
Steering Towards Fairness: Mitigating Political Bias in LLMs
por: Nadeem, Afrozah, et al.
Publicado: (2025) -
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
por: Nadeem, Afrozah, et al.
Publicado: (2025) -
Fairness Evaluation and Inference Level Mitigation in LLMs
por: Nadeem, Afrozah, et al.
Publicado: (2025) -
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
por: Bhandari, Pranav, et al.
Publicado: (2026) -
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
por: Selvaganapathy, Sanjeeevan, et al.
Publicado: (2025)