Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cantini, Riccardo, Cosenza, Giada, Orsino, Alessio, Talia, Domenico |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
Dynamic hashtag recommendation in social media with trend shift detection and adaptation
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
Assessing Political Bias in Large Language Models
von: Rettenberger, Luca, et al.
Veröffentlicht: (2024)
von: Rettenberger, Luca, et al.
Veröffentlicht: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
von: Lee, Isack, et al.
Veröffentlicht: (2024)
von: Lee, Isack, et al.
Veröffentlicht: (2024)
Block size estimation for data partitioning in HPC applications using machine learning techniques
von: Cantini, Riccardo, et al.
Veröffentlicht: (2022)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2022)
Detecting mental disorder on social media: a ChatGPT-augmented explainable approach
von: Belcastro, Loris, et al.
Veröffentlicht: (2024)
von: Belcastro, Loris, et al.
Veröffentlicht: (2024)
Prompt Programming for Cultural Bias and Alignment of Large Language Models
von: Eren, Maksim, et al.
Veröffentlicht: (2026)
von: Eren, Maksim, et al.
Veröffentlicht: (2026)
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
von: Kumar, Charaka Vinayak, et al.
Veröffentlicht: (2025)
von: Kumar, Charaka Vinayak, et al.
Veröffentlicht: (2025)
Is the System Message Really Important to Jailbreaks in Large Language Models?
von: Zou, Xiaotian, et al.
Veröffentlicht: (2024)
von: Zou, Xiaotian, et al.
Veröffentlicht: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
Regional Bias in Large Language Models
von: Gopinadh, M P V S, et al.
Veröffentlicht: (2026)
von: Gopinadh, M P V S, et al.
Veröffentlicht: (2026)
BiasLab: A Multilingual, Dual-Framing Framework for Robust Measurement of Output-Level Bias in Large Language Models
von: Guey, William, et al.
Veröffentlicht: (2026)
von: Guey, William, et al.
Veröffentlicht: (2026)
Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying Prompts
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
von: Govil, Priyanshul, et al.
Veröffentlicht: (2024)
von: Govil, Priyanshul, et al.
Veröffentlicht: (2024)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
von: Demchak, Nathaniel, et al.
Veröffentlicht: (2024)
GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language Models
von: Tang, Kunsheng, et al.
Veröffentlicht: (2024)
von: Tang, Kunsheng, et al.
Veröffentlicht: (2024)
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
von: Neumann, Anna, et al.
Veröffentlicht: (2025)
von: Neumann, Anna, et al.
Veröffentlicht: (2025)
DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Demonstration and Reasoning
von: Qiu, Hongye, et al.
Veröffentlicht: (2025)
von: Qiu, Hongye, et al.
Veröffentlicht: (2025)
Locating and Mitigating Gender Bias in Large Language Models
von: Cai, Yuchen, et al.
Veröffentlicht: (2024)
von: Cai, Yuchen, et al.
Veröffentlicht: (2024)
Cultural Bias and Cultural Alignment of Large Language Models
von: Tao, Yan, et al.
Veröffentlicht: (2023)
von: Tao, Yan, et al.
Veröffentlicht: (2023)
Cross-Language Bias Examination in Large Language Models
von: Liang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Liang, Yuxuan, et al.
Veröffentlicht: (2025)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
von: Park, Jean, et al.
Veröffentlicht: (2024)
von: Park, Jean, et al.
Veröffentlicht: (2024)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2023)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2023)
No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
von: Chand, Shireen, et al.
Veröffentlicht: (2025)
von: Chand, Shireen, et al.
Veröffentlicht: (2025)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
Large Language Models Are Still Misled by Simple Bias Ensembles
von: Sun, Zhouhao, et al.
Veröffentlicht: (2025)
von: Sun, Zhouhao, et al.
Veröffentlicht: (2025)
Multi-Persona Thinking for Bias Mitigation in Large Language Models
von: Chen, Yuxing, et al.
Veröffentlicht: (2026)
von: Chen, Yuxing, et al.
Veröffentlicht: (2026)
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2025)
von: Bisconti, Piercosma, et al.
Veröffentlicht: (2025)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
von: Pearman, Edie, et al.
Veröffentlicht: (2026)
von: Pearman, Edie, et al.
Veröffentlicht: (2026)
Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models
von: Ke, Shih-Wen, et al.
Veröffentlicht: (2025)
von: Ke, Shih-Wen, et al.
Veröffentlicht: (2025)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
von: Shamsi, Zafir, et al.
Veröffentlicht: (2026)
von: Shamsi, Zafir, et al.
Veröffentlicht: (2026)
The Tower of Babel Revisited: Multilingual Jailbreak Prompts on Closed-Source Large Language Models
von: Huang, Linghan, et al.
Veröffentlicht: (2025)
von: Huang, Linghan, et al.
Veröffentlicht: (2025)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Large Language Model Bias Mitigation from the Perspective of Knowledge Editing
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
von: Xu, Yuemei, et al.
Veröffentlicht: (2024)
von: Xu, Yuemei, et al.
Veröffentlicht: (2024)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
von: Pombal, José, et al.
Veröffentlicht: (2026)
von: Pombal, José, et al.
Veröffentlicht: (2026)
Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
von: Wang, Leroy Z.
Veröffentlicht: (2025)
von: Wang, Leroy Z.
Veröffentlicht: (2025)
Gender Bias in Machine Translation and The Era of Large Language Models
von: Vanmassenhove, Eva
Veröffentlicht: (2024)
von: Vanmassenhove, Eva
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025) -
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025) -
Dynamic hashtag recommendation in social media with trend shift detection and adaptation
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025) -
Assessing Political Bias in Large Language Models
von: Rettenberger, Luca, et al.
Veröffentlicht: (2024) -
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
von: Lee, Isack, et al.
Veröffentlicht: (2024)