Bias in the Mirror: Are LLMs opinions robust to their own adversarial attacks ?
Fuente:
arXiv
Saved in:
| Main Authors: | Rennard, Virgile, Xypolopoulos, Christos, Vazirgiannis, Michalis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging Discourse Structure for Extractive Meeting Summarization
by: Rennard, Virgile, et al.
Published: (2024)
by: Rennard, Virgile, et al.
Published: (2024)
CARTE: A Benchmark for Mapping Language Model Knowledge Across France
by: Carneiro, Sarah Almeida, et al.
Published: (2026)
by: Carneiro, Sarah Almeida, et al.
Published: (2026)
LLM as a Broken Telephone: Iterative Generation Distorts Information
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Effective faking of verbal deception detection with target-aligned adversarial attacks
by: Kleinberg, Bennett, et al.
Published: (2025)
by: Kleinberg, Bennett, et al.
Published: (2025)
Can adversarial attacks by large language models be attributed?
by: Cebrian, Manuel, et al.
Published: (2024)
by: Cebrian, Manuel, et al.
Published: (2024)
Markovian Generation Chains in Large Language Models
by: Geng, Mingmeng, et al.
Published: (2026)
by: Geng, Mingmeng, et al.
Published: (2026)
Difficulty Estimation and Simplification of French Text Using LLMs
by: Jamet, Henri, et al.
Published: (2024)
by: Jamet, Henri, et al.
Published: (2024)
GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
Graph Linearization Methods for Reasoning on Graphs with Large Language Models
by: Xypolopoulos, Christos, et al.
Published: (2024)
by: Xypolopoulos, Christos, et al.
Published: (2024)
Capturing Bias Diversity in LLMs
by: Gosavi, Purva Prasad, et al.
Published: (2024)
by: Gosavi, Purva Prasad, et al.
Published: (2024)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
by: Fragkathoulas, Christos, et al.
Published: (2024)
by: Fragkathoulas, Christos, et al.
Published: (2024)
Cognitive Bias in Decision-Making with LLMs
by: Echterhoff, Jessica, et al.
Published: (2024)
by: Echterhoff, Jessica, et al.
Published: (2024)
Implicit Bias in LLMs: A Survey
by: Lin, Xinru, et al.
Published: (2025)
by: Lin, Xinru, et al.
Published: (2025)
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs
by: Upadhayay, Bibek, et al.
Published: (2024)
by: Upadhayay, Bibek, et al.
Published: (2024)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
by: Pearman, Edie, et al.
Published: (2026)
by: Pearman, Edie, et al.
Published: (2026)
Enabling Scalable Evaluation of Bias Patterns in Medical LLMs
by: Fayyaz, Hamed, et al.
Published: (2024)
by: Fayyaz, Hamed, et al.
Published: (2024)
Steering Towards Fairness: Mitigating Political Bias in LLMs
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2026)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
by: Nadeem, Afrozah, et al.
Published: (2025)
by: Nadeem, Afrozah, et al.
Published: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
A Scalable Entity-Based Framework for Auditing Bias in LLMs
by: Elbouanani, Akram, et al.
Published: (2026)
by: Elbouanani, Akram, et al.
Published: (2026)
Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
by: Shang, Guokan, et al.
Published: (2025)
by: Shang, Guokan, et al.
Published: (2025)
Anti-adversarial Learning: Desensitizing Prompts for Large Language Models
by: Li, Xuan, et al.
Published: (2025)
by: Li, Xuan, et al.
Published: (2025)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
by: Arbabi, Alireza, et al.
Published: (2025)
by: Arbabi, Alireza, et al.
Published: (2025)
The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans
by: Chlapanis, Odysseas S., et al.
Published: (2026)
by: Chlapanis, Odysseas S., et al.
Published: (2026)
Political Bias in LLMs: Unaligned Moral Values in Agent-centric Simulations
by: Münker, Simon
Published: (2024)
by: Münker, Simon
Published: (2024)
Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification
by: Elbouanani, Akram, et al.
Published: (2025)
by: Elbouanani, Akram, et al.
Published: (2025)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
by: Wu, Yihao, et al.
Published: (2025)
by: Wu, Yihao, et al.
Published: (2025)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
by: Nadeem, Afrozah, et al.
Published: (2026)
by: Nadeem, Afrozah, et al.
Published: (2026)
Addressing Bias in LLMs: Strategies and Application to Fair AI-based Recruitment
by: Peña, Alejandro, et al.
Published: (2025)
by: Peña, Alejandro, et al.
Published: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Role-Play Zero-Shot Prompting with Large Language Models for Open-Domain Human-Machine Conversation
by: Njifenjou, Ahmed, et al.
Published: (2024)
by: Njifenjou, Ahmed, et al.
Published: (2024)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
by: DiGiuseppe, Matthew, et al.
Published: (2026)
by: DiGiuseppe, Matthew, et al.
Published: (2026)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025)
by: Dwivedi, Ashutosh, et al.
Published: (2025)
Similar Items
-
Leveraging Discourse Structure for Extractive Meeting Summarization
by: Rennard, Virgile, et al.
Published: (2024) -
CARTE: A Benchmark for Mapping Language Model Knowledge Across France
by: Carneiro, Sarah Almeida, et al.
Published: (2026) -
LLM as a Broken Telephone: Iterative Generation Distorts Information
by: Mohamed, Amr, et al.
Published: (2025) -
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025) -
Effective faking of verbal deception detection with target-aligned adversarial attacks
by: Kleinberg, Bennett, et al.
Published: (2025)