BIPOLAR: Polarization-based granular framework for LLM bias evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Pavlíček, Martin, Filip, Tomáš, Sosík, Petr |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-tuning multilingual language models in Twitter/X sentiment analysis: a study on Eastern-European V4 languages
by: Filip, Tomáš, et al.
Published: (2024)
by: Filip, Tomáš, et al.
Published: (2024)
Two-dimensional early exit optimisation of LLM inference
by: Hůla, Jan, et al.
Published: (2026)
by: Hůla, Jan, et al.
Published: (2026)
A survey on learning models of spiking neural membrane systems and spiking neural networks
by: Paul, Prithwineel, et al.
Published: (2024)
by: Paul, Prithwineel, et al.
Published: (2024)
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024)
by: Shaikh, Ammar, et al.
Published: (2024)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
by: Kesiraju, Santosh, et al.
Published: (2023)
by: Kesiraju, Santosh, et al.
Published: (2023)
Anthropocentric bias in language model evaluation
by: Millière, Raphaël, et al.
Published: (2024)
by: Millière, Raphaël, et al.
Published: (2024)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
by: Chadimová, Milena, et al.
Published: (2024)
by: Chadimová, Milena, et al.
Published: (2024)
The Biased Samaritan: LLM biases in Perceived Kindness
by: Fagan, Jack H, et al.
Published: (2025)
by: Fagan, Jack H, et al.
Published: (2025)
NLD-LLM: A systematic framework for evaluating small language transformer models on natural language description
by: Jelodar, Hamed, et al.
Published: (2025)
by: Jelodar, Hamed, et al.
Published: (2025)
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
by: Lee, Yukyung, et al.
Published: (2024)
by: Lee, Yukyung, et al.
Published: (2024)
Source framing triggers systematic evaluation bias in Large Language Models
by: Germani, Federico, et al.
Published: (2025)
by: Germani, Federico, et al.
Published: (2025)
Assessing the quality of information extraction
by: Seitl, Filip, et al.
Published: (2024)
by: Seitl, Filip, et al.
Published: (2024)
A database to support the evaluation of gender biases in GPT-4o output
by: Mehner, Luise, et al.
Published: (2025)
by: Mehner, Luise, et al.
Published: (2025)
Large Language Models are biased to overestimate profoundness
by: Herrera-Berg, Eugenio, et al.
Published: (2023)
by: Herrera-Berg, Eugenio, et al.
Published: (2023)
Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models
by: Guey, William, et al.
Published: (2026)
by: Guey, William, et al.
Published: (2026)
RadEval: A framework for radiology text evaluation
by: Xu, Justin, et al.
Published: (2025)
by: Xu, Justin, et al.
Published: (2025)
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America
by: Ivetta, Guido, et al.
Published: (2025)
by: Ivetta, Guido, et al.
Published: (2025)
Dynamic benchmarking framework for LLM-based conversational data capture
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset
by: Shaier, Sagi, et al.
Published: (2024)
by: Shaier, Sagi, et al.
Published: (2024)
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering
by: Henriksson, Erik, et al.
Published: (2025)
by: Henriksson, Erik, et al.
Published: (2025)
An evaluation of LLMs for political bias in Western media: Israel-Hamas and Ukraine-Russia wars
by: Chandra, Rohitash, et al.
Published: (2026)
by: Chandra, Rohitash, et al.
Published: (2026)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
by: Kankowski, Florian, et al.
Published: (2025)
by: Kankowski, Florian, et al.
Published: (2025)
Emotion-cause pair extraction method based on multi-granularity information and multi-module interaction
by: Fu, Mingrui, et al.
Published: (2024)
by: Fu, Mingrui, et al.
Published: (2024)
Semantic Properties of cosine based bias scores for word embeddings
by: Schröder, Sarah, et al.
Published: (2024)
by: Schröder, Sarah, et al.
Published: (2024)
A word association network methodology for evaluating implicit biases in LLMs compared to humans
by: Abramski, Katherine, et al.
Published: (2025)
by: Abramski, Katherine, et al.
Published: (2025)
Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain
by: Hu, Xiaoyu, et al.
Published: (2026)
by: Hu, Xiaoyu, et al.
Published: (2026)
The SAME score: Improved cosine based bias score for word embeddings
by: Schröder, Sarah, et al.
Published: (2022)
by: Schröder, Sarah, et al.
Published: (2022)
HARE: an entity and relation centric evaluation framework for histopathology reports
by: Kim, Yunsoo, et al.
Published: (2025)
by: Kim, Yunsoo, et al.
Published: (2025)
Exploring LLM biases to manipulate AI search overview
by: Smirnov, Roman
Published: (2026)
by: Smirnov, Roman
Published: (2026)
LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
by: Ailem, Melissa, et al.
Published: (2024)
by: Ailem, Melissa, et al.
Published: (2024)
A comprehensive study of LLM-based argument classification: from Llama through DeepSeek to GPT-5.2
by: Pietroń, Marcin, et al.
Published: (2026)
by: Pietroń, Marcin, et al.
Published: (2026)
LLM-based feature generation from text for interpretable machine learning
by: Balek, Vojtěch, et al.
Published: (2024)
by: Balek, Vojtěch, et al.
Published: (2024)
Fine-Grained Bias Detection in LLM: Enhancing detection mechanisms for nuanced biases
by: Mohanty, Suvendu
Published: (2025)
by: Mohanty, Suvendu
Published: (2025)
A comprehensive study of LLM-based argument classification: from LLAMA through GPT-4o to Deepseek-R1
by: Pietroń, Marcin, et al.
Published: (2025)
by: Pietroń, Marcin, et al.
Published: (2025)
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations
by: Harbola, Chitranshu, et al.
Published: (2025)
by: Harbola, Chitranshu, et al.
Published: (2025)
How far can bias go? Tracing bias from pretraining data to alignment
by: Thaler, Marion, et al.
Published: (2024)
by: Thaler, Marion, et al.
Published: (2024)
Similar Items
-
Fine-tuning multilingual language models in Twitter/X sentiment analysis: a study on Eastern-European V4 languages
by: Filip, Tomáš, et al.
Published: (2024) -
Two-dimensional early exit optimisation of LLM inference
by: Hůla, Jan, et al.
Published: (2026) -
A survey on learning models of spiking neural membrane systems and spiking neural networks
by: Paul, Prithwineel, et al.
Published: (2024) -
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
by: Shaikh, Ammar, et al.
Published: (2024) -
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
by: Kesiraju, Santosh, et al.
Published: (2023)