Why Don't Prompt-Based Fairness Metrics Correlate?
Fuente:
arXiv
Salvato in:
| Autori principali: | Zayed, Abdelrahman, Mordido, Goncalo, Baldini, Ioana, Chandar, Sarath |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Should We Attend More or Less? Modulating Attention for Fairness
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
di: Shin, Kwan Soo
Pubblicazione: (2026)
di: Shin, Kwan Soo
Pubblicazione: (2026)
Lookbehind-SAM: k steps back, 1 step forward
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
di: Khan, Imran
Pubblicazione: (2025)
di: Khan, Imran
Pubblicazione: (2025)
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Do Large Language Models Know How Much They Know?
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Fairness of ChatGPT
di: Li, Yunqi, et al.
Pubblicazione: (2023)
di: Li, Yunqi, et al.
Pubblicazione: (2023)
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback
di: Lerner, Emilia Agis, et al.
Pubblicazione: (2024)
di: Lerner, Emilia Agis, et al.
Pubblicazione: (2024)
LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
di: Bouchard, Dylan, et al.
Pubblicazione: (2025)
Bias and Fairness in Large Language Models: A Survey
di: Gallegos, Isabel O., et al.
Pubblicazione: (2023)
di: Gallegos, Isabel O., et al.
Pubblicazione: (2023)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
di: Shieh, Evan, et al.
Pubblicazione: (2024)
di: Shieh, Evan, et al.
Pubblicazione: (2024)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
Exploring Accuracy-Fairness Trade-off in Large Language Models
di: Zhang, Qingquan, et al.
Pubblicazione: (2024)
di: Zhang, Qingquan, et al.
Pubblicazione: (2024)
Reasoning Models Don't Always Say What They Think
di: Chen, Yanda, et al.
Pubblicazione: (2025)
di: Chen, Yanda, et al.
Pubblicazione: (2025)
Correlated Errors in Large Language Models
di: Kim, Elliot, et al.
Pubblicazione: (2025)
di: Kim, Elliot, et al.
Pubblicazione: (2025)
AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs
di: Ebrahimi, Sana, et al.
Pubblicazione: (2024)
di: Ebrahimi, Sana, et al.
Pubblicazione: (2024)
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models
di: Song, Kefan, et al.
Pubblicazione: (2025)
di: Song, Kefan, et al.
Pubblicazione: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
di: Zhou, Jin Peng, et al.
Pubblicazione: (2024)
di: Zhou, Jin Peng, et al.
Pubblicazione: (2024)
Inducing Group Fairness in Prompt-Based Language Model Decisions
di: Atwood, James, et al.
Pubblicazione: (2024)
di: Atwood, James, et al.
Pubblicazione: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2024)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
di: Kapoor, Sanyam, et al.
Pubblicazione: (2024)
di: Kapoor, Sanyam, et al.
Pubblicazione: (2024)
Too Big to Fool: Resisting Deception in Language Models
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models
di: Kumar, Abhishek, et al.
Pubblicazione: (2024)
di: Kumar, Abhishek, et al.
Pubblicazione: (2024)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
di: Aghajohari, Milad, et al.
Pubblicazione: (2025)
di: Aghajohari, Milad, et al.
Pubblicazione: (2025)
Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
di: Leemann, Tobias, et al.
Pubblicazione: (2024)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
di: Mayne, Harry, et al.
Pubblicazione: (2025)
di: Mayne, Harry, et al.
Pubblicazione: (2025)
Don't Shoot The Breeze: Topic Continuity Model Using Nonlinear Naive Bayes With Attention
di: Pi, Shu-Ting, et al.
Pubblicazione: (2026)
di: Pi, Shu-Ting, et al.
Pubblicazione: (2026)
Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary?
di: Jedidi, Nour, et al.
Pubblicazione: (2025)
di: Jedidi, Nour, et al.
Pubblicazione: (2025)
MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups
di: Popoola, Gideon, et al.
Pubblicazione: (2026)
di: Popoola, Gideon, et al.
Pubblicazione: (2026)
Torque-Aware Momentum
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
di: Abbes, Istabrak, et al.
Pubblicazione: (2025)
Learning Fair Invariant Representations under Covariate and Correlation Shifts Simultaneously
di: Li, Dong, et al.
Pubblicazione: (2024)
di: Li, Dong, et al.
Pubblicazione: (2024)
Uncertainty and Fairness Awareness in LLM-Based Recommendation Systems
di: Sah, Chandan Kumar, et al.
Pubblicazione: (2026)
di: Sah, Chandan Kumar, et al.
Pubblicazione: (2026)
Deep Learning Based Amharic Chatbot for FAQs in Universities
di: Hailu, Goitom Ybrah, et al.
Pubblicazione: (2024)
di: Hailu, Goitom Ybrah, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Should We Attend More or Less? Modulating Attention for Fairness
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023) -
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
di: Shin, Kwan Soo
Pubblicazione: (2026) -
Lookbehind-SAM: k steps back, 1 step forward
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023) -
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
di: Khan, Imran
Pubblicazione: (2025) -
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024)