Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Kohankhaki, Farnaz, Emerson, D. B., Tian, Jacob-Junqi, Seyyed-Kalantari, Laleh, Khattak, Faiza Khan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
Soft-prompt Tuning for Large Language Models to Evaluate Bias
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
por: Tian, Jacob-Junqi, et al.
Publicado: (2023)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
por: Hari, Vishnu, et al.
Publicado: (2025)
por: Hari, Vishnu, et al.
Publicado: (2025)
Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
por: Xu, Weijie, et al.
Publicado: (2025)
por: Xu, Weijie, et al.
Publicado: (2025)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
por: Chen, Jiaju, et al.
Publicado: (2025)
por: Chen, Jiaju, et al.
Publicado: (2025)
Do LLMs have a Gender (Entropy) Bias?
por: Prabhune, Sonal, et al.
Publicado: (2025)
por: Prabhune, Sonal, et al.
Publicado: (2025)
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
por: Han, Pengrui, et al.
Publicado: (2024)
por: Han, Pengrui, et al.
Publicado: (2024)
Innovative Tangible Interactive Games for Enhancing Artificial Intelligence Knowledge and Literacy in Elementary Education: A Pedagogical Framework
por: Sampanis, Nikolaos
Publicado: (2025)
por: Sampanis, Nikolaos
Publicado: (2025)
The GPT-4o Shock Emotional Attachment to AI Models and Its Impact on Regulatory Acceptance: A Cross-Cultural Analysis of the Immediate Transition from GPT-4o to GPT-5
por: Naito, Hiroki
Publicado: (2025)
por: Naito, Hiroki
Publicado: (2025)
Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis
por: Petrov, Nikolay B, et al.
Publicado: (2024)
por: Petrov, Nikolay B, et al.
Publicado: (2024)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
por: Ikoma, Hayato, et al.
Publicado: (2025)
por: Ikoma, Hayato, et al.
Publicado: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
por: AlquBoj, H. V., et al.
Publicado: (2025)
por: AlquBoj, H. V., et al.
Publicado: (2025)
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
por: Wibowo, Haryo Akbarianto, et al.
Publicado: (2026)
por: Wibowo, Haryo Akbarianto, et al.
Publicado: (2026)
HInter: Exposing Hidden Intersectional Bias in Large Language Models
por: Souani, Badr, et al.
Publicado: (2025)
por: Souani, Badr, et al.
Publicado: (2025)
A Survey on Natural Language Counterfactual Generation
por: Wang, Yongjie, et al.
Publicado: (2024)
por: Wang, Yongjie, et al.
Publicado: (2024)
Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
por: Souverain, Thomas
Publicado: (2025)
por: Souverain, Thomas
Publicado: (2025)
Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
por: Vasilev, Stefan, et al.
Publicado: (2025)
por: Vasilev, Stefan, et al.
Publicado: (2025)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
por: Li, Mingda, et al.
Publicado: (2024)
por: Li, Mingda, et al.
Publicado: (2024)
ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
por: Delaval, Axel, et al.
Publicado: (2025)
por: Delaval, Axel, et al.
Publicado: (2025)
Acceptable Use Policies for Foundation Models
por: Klyman, Kevin
Publicado: (2024)
por: Klyman, Kevin
Publicado: (2024)
CritiSense: Critical Digital Literacy and Resilience Against Misinformation
por: Alam, Firoj, et al.
Publicado: (2026)
por: Alam, Firoj, et al.
Publicado: (2026)
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
por: Lee, Yejin, et al.
Publicado: (2025)
por: Lee, Yejin, et al.
Publicado: (2025)
Change My Frame: Reframing in the Wild in r/ChangeMyView
por: Peguero, Arturo Martínez, et al.
Publicado: (2024)
por: Peguero, Arturo Martínez, et al.
Publicado: (2024)
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage
por: Young, Richard J., et al.
Publicado: (2026)
por: Young, Richard J., et al.
Publicado: (2026)
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
por: Brant, Thiago, et al.
Publicado: (2026)
por: Brant, Thiago, et al.
Publicado: (2026)
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
por: Yoo, Seunghyun
Publicado: (2025)
por: Yoo, Seunghyun
Publicado: (2025)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
por: Liu, Linyu, et al.
Publicado: (2024)
por: Liu, Linyu, et al.
Publicado: (2024)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
por: Anam, Rizal Khoirul
Publicado: (2025)
por: Anam, Rizal Khoirul
Publicado: (2025)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
por: Otal, Hakan T., et al.
Publicado: (2024)
por: Otal, Hakan T., et al.
Publicado: (2024)
Grandes modelos de lenguaje: de la predicción de palabras a la comprensión?
por: Gómez-Rodríguez, Carlos
Publicado: (2025)
por: Gómez-Rodríguez, Carlos
Publicado: (2025)
Value Lens: Using Large Language Models to Understand Human Values
por: Fernández, Eduardo de la Cruz, et al.
Publicado: (2025)
por: Fernández, Eduardo de la Cruz, et al.
Publicado: (2025)
The Pursuit of Empathy: Evaluating Small Language Models for PTSD Dialogue Support
por: BN, Suhas, et al.
Publicado: (2025)
por: BN, Suhas, et al.
Publicado: (2025)
Measuring Self-Rating Bias in LLM-Generated Survey Data: A Semantic Similarity Framework for Independent Scale Mapping
por: Pichardo, Eduardo Vera
Publicado: (2026)
por: Pichardo, Eduardo Vera
Publicado: (2026)
The AI Imperative: Scaling High-Quality Peer Review in Machine Learning
por: Wei, Qiyao, et al.
Publicado: (2025)
por: Wei, Qiyao, et al.
Publicado: (2025)
Adaptive Engram Memory System for Indonesian Language Model: Generative AI Based on TOBA LM for Batak and Minang Language
por: Situngkir, Hokky, et al.
Publicado: (2026)
por: Situngkir, Hokky, et al.
Publicado: (2026)
Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities
por: Levchenko, Maria
Publicado: (2025)
por: Levchenko, Maria
Publicado: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
por: Arabov, Mullosharaf K.
Publicado: (2026)
por: Arabov, Mullosharaf K.
Publicado: (2026)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
por: Yao, Huaiyuan, et al.
Publicado: (2026)
por: Yao, Huaiyuan, et al.
Publicado: (2026)
Meaning-infused grammar: Gradient Acceptability Shapes the Geometric Representations of Constructions in LLMs
por: Rakshit, Supantho, et al.
Publicado: (2025)
por: Rakshit, Supantho, et al.
Publicado: (2025)
Ejemplares similares
-
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
por: Tian, Jacob-Junqi, et al.
Publicado: (2023) -
Soft-prompt Tuning for Large Language Models to Evaluate Bias
por: Tian, Jacob-Junqi, et al.
Publicado: (2023) -
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
por: Hari, Vishnu, et al.
Publicado: (2025) -
Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
por: Xu, Weijie, et al.
Publicado: (2025) -
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
por: Chen, Jiaju, et al.
Publicado: (2025)