Toward Systematic Counterfactual Fairness Evaluation of Large Language Models: The CAFFE Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Parziale, Alessandra, Voria, Gianmario, Pontillo, Valeria, Catolino, Gemma, De Lucia, Andrea, Palomba, Fabio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCOPE: A Dataset of Stereotyped Prompts for Counterfactual Fairness Assessment of LLMs
by: Parziale, Alessandra, et al.
Published: (2026)
by: Parziale, Alessandra, et al.
Published: (2026)
Contextual Fairness-Aware Practices in ML: A Cost-Effective Empirical Evaluation
by: Parziale, Alessandra, et al.
Published: (2025)
by: Parziale, Alessandra, et al.
Published: (2025)
Bias Ahead: Sensitive Prompts as Early Warnings for Fairness in Large Language Models
by: Voria, Gianmario, et al.
Published: (2026)
by: Voria, Gianmario, et al.
Published: (2026)
Once Upon a Team: Investigating Bias in LLM-Driven Software Team Composition and Task Allocation
by: Parziale, Alessandra, et al.
Published: (2026)
by: Parziale, Alessandra, et al.
Published: (2026)
RECOVER: Toward Requirements Generation from Stakeholders' Conversations
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
From Expectation to Habit: Why Do Software Practitioners Adopt Fairness Toolkits?
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
Data Preparation for Fairness-Performance Trade-Offs: A Practitioner-Friendly Alternative?
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer Models
by: Voria, Gianmario, et al.
Published: (2026)
by: Voria, Gianmario, et al.
Published: (2026)
A Catalog of Fairness-Aware Practices in Machine Learning Engineering
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
Investigating the Role of Cultural Values in Adopting Large Language Models for Software Engineering
by: Lambiase, Stefano, et al.
Published: (2024)
by: Lambiase, Stefano, et al.
Published: (2024)
Motivations, Challenges, Best Practices, and Benefits for Bots and Conversational Agents in Software Engineering: A Multivocal Literature Review
by: Lambiase, Stefano, et al.
Published: (2024)
by: Lambiase, Stefano, et al.
Published: (2024)
How Do Communities of ML-Enabled Systems Smell? A Cross-Sectional Study on the Prevalence of Community Smells
by: Annunziata, Giusy, et al.
Published: (2025)
by: Annunziata, Giusy, et al.
Published: (2025)
Exploring Individual Factors in the Adoption of LLMs for Specific Software Engineering Purposes
by: Lambiase, Stefano, et al.
Published: (2025)
by: Lambiase, Stefano, et al.
Published: (2025)
Classification, Challenges, and Automated Approaches to Handle Non-Functional Requirements in ML-Enabled Systems: A Systematic Literature Review
by: De Martino, Vincenzo, et al.
Published: (2023)
by: De Martino, Vincenzo, et al.
Published: (2023)
Socio-Technical Well-Being of Quantum Software Communities: An Overview on Community Smells
by: Lambiase, Stefano, et al.
Published: (2026)
by: Lambiase, Stefano, et al.
Published: (2026)
GenFair: Systematic Test Generation for Fairness Fault Detection in Large Language Models
by: Srinivasan, Madhusudan, et al.
Published: (2025)
by: Srinivasan, Madhusudan, et al.
Published: (2025)
On the Robustness of Fairness Practices: A Causal Framework for Systematic Evaluation
by: Monjezi, Verya, et al.
Published: (2026)
by: Monjezi, Verya, et al.
Published: (2026)
A Methodological Framework for LLM-Based Mining of Software Repositories
by: De Martino, Vincenzo, et al.
Published: (2025)
by: De Martino, Vincenzo, et al.
Published: (2025)
LLM-Assisted Empirical Software Engineering: Systematic Literature Review and Research Agenda
by: Gomes, Victoria, et al.
Published: (2026)
by: Gomes, Victoria, et al.
Published: (2026)
REST in Pieces: RESTful Design Rule Violations in Student-Built Web Apps
by: Di Meglio, Sergio, et al.
Published: (2025)
by: Di Meglio, Sergio, et al.
Published: (2025)
Do Developers Adopt Green Architectural Tactics for ML-Enabled Systems? A Mining Software Repository Study
by: De Martino, Vincenzo, et al.
Published: (2024)
by: De Martino, Vincenzo, et al.
Published: (2024)
Green Architectural Tactics in ML-enabled Systems: An LLM-based Repository Mining Study
by: De Martino, Vincenzo, et al.
Published: (2026)
by: De Martino, Vincenzo, et al.
Published: (2026)
Advances in Artificial Intelligence forDiabetes Prediction: Insights from a Systematic Literature Review
by: Khokhar, Pir Bakhsh, et al.
Published: (2024)
by: Khokhar, Pir Bakhsh, et al.
Published: (2024)
Investigating the Performance of Small Language Models in Detecting Test Smells in Manual Test Cases
by: Lucas, Keila, et al.
Published: (2025)
by: Lucas, Keila, et al.
Published: (2025)
Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
by: Romero-Arjona, Miguel, et al.
Published: (2025)
by: Romero-Arjona, Miguel, et al.
Published: (2025)
A First Look at the Lifecycle of DL-Specific Self-Admitted Technical Debt
by: Recupito, Gilberto, et al.
Published: (2025)
by: Recupito, Gilberto, et al.
Published: (2025)
Investigating Technical Debt Types, Issues, and Solutions in Serverless Computing
by: Perera, Hasini Sumalee, et al.
Published: (2026)
by: Perera, Hasini Sumalee, et al.
Published: (2026)
A Framework for Using LLMs for Repository Mining Studies in Empirical Software Engineering
by: de Martino, Vincenzo, et al.
Published: (2024)
by: de Martino, Vincenzo, et al.
Published: (2024)
Towards Fair Machine Learning Software: Understanding and Addressing Model Bias Through Counterfactual Thinking
by: Wang, Zichong, et al.
Published: (2023)
by: Wang, Zichong, et al.
Published: (2023)
Towards Systematic Specification and Verification of Fairness Requirements: A Position Paper
by: Ramadan, Qusai, et al.
Published: (2025)
by: Ramadan, Qusai, et al.
Published: (2025)
Towards Understanding Bugs in Distributed Training and Inference Frameworks for Large Language Models
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
A Systematic Mapping on Software Fairness: Focus, Trends and Industrial Context
by: Nepomuceno, Kessia, et al.
Published: (2025)
by: Nepomuceno, Kessia, et al.
Published: (2025)
Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
by: Xiao, Yisong, et al.
Published: (2025)
by: Xiao, Yisong, et al.
Published: (2025)
Towards Automated Page Object Generation for Web Testing using Large Language Models
by: Karagöz, Betül, et al.
Published: (2026)
by: Karagöz, Betül, et al.
Published: (2026)
In Search of Metrics to Guide Developer-Based Refactoring Recommendations
by: Robredo, Mikel, et al.
Published: (2024)
by: Robredo, Mikel, et al.
Published: (2024)
STELLAR: A Search-Based Testing Framework for Large Language Model Applications
by: Sorokin, Lev, et al.
Published: (2026)
by: Sorokin, Lev, et al.
Published: (2026)
Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code
by: Della Porta, Antonio, et al.
Published: (2025)
by: Della Porta, Antonio, et al.
Published: (2025)
Testing Framework Migration with Large Language Models
by: Alves, Altino, et al.
Published: (2026)
by: Alves, Altino, et al.
Published: (2026)
Towards Transparent and Accurate Diabetes Prediction Using Machine Learning and Explainable Artificial Intelligence
by: Khokhar, Pir Bakhsh, et al.
Published: (2025)
by: Khokhar, Pir Bakhsh, et al.
Published: (2025)
Large Language Models for Unit Testing: A Systematic Literature Review
by: Zhang, Quanjun, et al.
Published: (2025)
by: Zhang, Quanjun, et al.
Published: (2025)
Similar Items
-
SCOPE: A Dataset of Stereotyped Prompts for Counterfactual Fairness Assessment of LLMs
by: Parziale, Alessandra, et al.
Published: (2026) -
Contextual Fairness-Aware Practices in ML: A Cost-Effective Empirical Evaluation
by: Parziale, Alessandra, et al.
Published: (2025) -
Bias Ahead: Sensitive Prompts as Early Warnings for Fairness in Large Language Models
by: Voria, Gianmario, et al.
Published: (2026) -
Once Upon a Team: Investigating Bias in LLM-Driven Software Team Composition and Task Allocation
by: Parziale, Alessandra, et al.
Published: (2026) -
RECOVER: Toward Requirements Generation from Stakeholders' Conversations
by: Voria, Gianmario, et al.
Published: (2024)