Textual Entailment is not a Better Bias Metric than Token Probability
Fuente:
arXiv
Saved in:
| Main Authors: | Felkner, Virginia K., Lim, Allison, May, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024)
by: Felkner, Virginia K., et al.
Published: (2024)
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
by: Sami, K. M. Jubair, et al.
Published: (2026)
by: Sami, K. M. Jubair, et al.
Published: (2026)
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
by: Himmelreich, Johannes
Published: (2026)
by: Himmelreich, Johannes
Published: (2026)
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025)
by: Rivas, Pablo, et al.
Published: (2025)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)
by: Maheshwari, Aditya, et al.
Published: (2026)
Rejected Dialects: Biases Against African American Language in Reward Models
by: Mire, Joel, et al.
Published: (2025)
by: Mire, Joel, et al.
Published: (2025)
Large language models can replicate cross-cultural differences in personality
by: Niszczota, Paweł, et al.
Published: (2023)
by: Niszczota, Paweł, et al.
Published: (2023)
PromptAug: Fine-grained Conflict Classification Using Data Augmentation
by: Warke, Oliver, et al.
Published: (2025)
by: Warke, Oliver, et al.
Published: (2025)
VEAT Quantifies Implicit Associations in Text-to-Video Generator Sora and Reveals Challenges in Bias Mitigation
by: Sun, Yongxu, et al.
Published: (2026)
by: Sun, Yongxu, et al.
Published: (2026)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025)
by: Du, Yishan, et al.
Published: (2025)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
by: Mehta, Manisha, et al.
Published: (2025)
by: Mehta, Manisha, et al.
Published: (2025)
Using a cognitive architecture to consider antiBlackness in design and development of AI systems
by: Dancy, Christopher L.
Published: (2022)
by: Dancy, Christopher L.
Published: (2022)
The Company You Keep: How LLMs Respond to Dark Triad Traits
by: Lu, Zeyi, et al.
Published: (2026)
by: Lu, Zeyi, et al.
Published: (2026)
Reward Model Interpretability via Optimal and Pessimal Tokens
by: Christian, Brian, et al.
Published: (2025)
by: Christian, Brian, et al.
Published: (2025)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
by: Zahid, Anwar Hossain, et al.
Published: (2025)
by: Zahid, Anwar Hossain, et al.
Published: (2025)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
by: Hari, Vishnu, et al.
Published: (2025)
by: Hari, Vishnu, et al.
Published: (2025)
Disaster Question Answering with LoRA Efficiency and Accurate End Position
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
Collective Constitutional AI: Aligning a Language Model with Public Input
by: Huang, Saffron, et al.
Published: (2024)
by: Huang, Saffron, et al.
Published: (2024)
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
by: Loth, Alexander, et al.
Published: (2026)
by: Loth, Alexander, et al.
Published: (2026)
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026)
by: Taheri, Atieh, et al.
Published: (2026)
Assessing Crime Disclosure Patterns in a Large-Scale Cybercrime Forum
by: Hoheisel, Raphael, et al.
Published: (2026)
by: Hoheisel, Raphael, et al.
Published: (2026)
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor
by: Lehr, Steven A., et al.
Published: (2024)
by: Lehr, Steven A., et al.
Published: (2024)
Extreme Self-Preference in Language Models
by: Lehr, Steven A., et al.
Published: (2025)
by: Lehr, Steven A., et al.
Published: (2025)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
by: Guan, Xin, et al.
Published: (2024)
by: Guan, Xin, et al.
Published: (2024)
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
by: Zahraei, Pardis Sadat, et al.
Published: (2024)
Leveraging Multi-Source Textural UGC for Neighbourhood Housing Quality Assessment: A GPT-Enhanced Framework
by: Hong, Qiyuan, et al.
Published: (2025)
by: Hong, Qiyuan, et al.
Published: (2025)
Beyond the Cloud: Assessing the Benefits and Drawbacks of Local LLM Deployment for Translators
by: Sandrini, Peter
Published: (2025)
by: Sandrini, Peter
Published: (2025)
The Journal of Prompt-Engineered (Moral) Philosophy Or: Why AI-Assisted Ethics Research Requires Process Transparency
by: Loi, Michele
Published: (2025)
by: Loi, Michele
Published: (2025)
Balancing Innovation and Integrity: AI Integration in Liberal Arts College Administration
by: Read, Ian Olivo
Published: (2025)
by: Read, Ian Olivo
Published: (2025)
The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete
by: Barmettler, Joel
Published: (2026)
by: Barmettler, Joel
Published: (2026)
What are People Talking about in #BlackLivesMatter and #StopAsianHate? Exploring and Categorizing Twitter Topics Emerging in Online Social Movements through the Latent Dirichlet Allocation Model
by: Tong, Xin, et al.
Published: (2022)
by: Tong, Xin, et al.
Published: (2022)
Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems
by: Loth, Alexander, et al.
Published: (2026)
by: Loth, Alexander, et al.
Published: (2026)
Pro-AI Bias in Large Language Models
by: Trabelsi, Benaya, et al.
Published: (2026)
by: Trabelsi, Benaya, et al.
Published: (2026)
Generative UI as an Accessibility Bridge: Lessons from C2C E-Commerce
by: Ryskeldiev, Bektur
Published: (2026)
by: Ryskeldiev, Bektur
Published: (2026)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
by: Hossain, Ariyan, et al.
Published: (2025)
by: Hossain, Ariyan, et al.
Published: (2025)
Self-Anchored Attention Model for Sample-Efficient Classification of Prosocial Text Chat
by: Li, Zhuofang, et al.
Published: (2025)
by: Li, Zhuofang, et al.
Published: (2025)
Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale
by: Kocielnik, Rafal, et al.
Published: (2025)
by: Kocielnik, Rafal, et al.
Published: (2025)
The Democratic Paradox in Large Language Models' Underestimation of Press Freedom
by: Loaiza, I., et al.
Published: (2025)
by: Loaiza, I., et al.
Published: (2025)
Transforming Computer Security and Public Trust Through the Exploration of Fine-Tuning Large Language Models
by: Crumrine, Garrett, et al.
Published: (2024)
by: Crumrine, Garrett, et al.
Published: (2024)
Similar Items
-
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024) -
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
by: Sami, K. M. Jubair, et al.
Published: (2026) -
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
by: Himmelreich, Johannes
Published: (2026) -
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025) -
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)