With a Grain of SALT: Are LLMs Fair Across Social Dimensions?
Fuente:
arXiv
Guardado en:
| Autores principales: | Arif, Samee, Khan, Zohaib, Kaleem, Maaidah, Rashid, Suhaib, Raza, Agha Ali, Athar, Awais |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generalists vs. Specialists: Evaluating Large Language Models for Urdu
por: Arif, Samee, et al.
Publicado: (2024)
por: Arif, Samee, et al.
Publicado: (2024)
UQA: Corpus for Urdu Question Answering
por: Arif, Samee, et al.
Publicado: (2024)
por: Arif, Samee, et al.
Publicado: (2024)
The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation
por: Arif, Samee, et al.
Publicado: (2024)
por: Arif, Samee, et al.
Publicado: (2024)
Kahaani: A Multimodal Co-Creative Storytelling System
por: Arif, Samee, et al.
Publicado: (2024)
por: Arif, Samee, et al.
Publicado: (2024)
WER We Stand: Benchmarking Urdu ASR Models
por: Arif, Samee, et al.
Publicado: (2024)
por: Arif, Samee, et al.
Publicado: (2024)
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
por: Arif, Samee, et al.
Publicado: (2026)
por: Arif, Samee, et al.
Publicado: (2026)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
por: Arif, Samee, et al.
Publicado: (2026)
por: Arif, Samee, et al.
Publicado: (2026)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
por: Hashmat, Abdullah, et al.
Publicado: (2025)
por: Hashmat, Abdullah, et al.
Publicado: (2025)
Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention
por: Khan, Zohaib, et al.
Publicado: (2024)
por: Khan, Zohaib, et al.
Publicado: (2024)
Language Model-Driven Data Pruning Enables Efficient Active Learning
por: Azeemi, Abdul Hameed, et al.
Publicado: (2024)
por: Azeemi, Abdul Hameed, et al.
Publicado: (2024)
To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation
por: Azeemi, Abdul Hameed, et al.
Publicado: (2024)
por: Azeemi, Abdul Hameed, et al.
Publicado: (2024)
Empathy Applicability Modeling for General Health Queries
por: Randhawa, Shan, et al.
Publicado: (2026)
por: Randhawa, Shan, et al.
Publicado: (2026)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
por: Wang, Yinong Oliver, et al.
Publicado: (2025)
por: Wang, Yinong Oliver, et al.
Publicado: (2025)
Scaling Truth: The Confidence Paradox in AI Fact-Checking
por: Qazi, Ihsan A., et al.
Publicado: (2025)
por: Qazi, Ihsan A., et al.
Publicado: (2025)
FairSense-AI: Responsible AI Meets Sustainability
por: Raza, Shaina, et al.
Publicado: (2025)
por: Raza, Shaina, et al.
Publicado: (2025)
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
por: Bilal, Ahsan, et al.
Publicado: (2025)
por: Bilal, Ahsan, et al.
Publicado: (2025)
Evaluating LLMs on Generating Age-Appropriate Child-Like Conversations
por: Hassan, Syed Zohaib, et al.
Publicado: (2025)
por: Hassan, Syed Zohaib, et al.
Publicado: (2025)
Usability Study of Security Features in Programmable Logic Controllers
por: Li, Karen, et al.
Publicado: (2022)
por: Li, Karen, et al.
Publicado: (2022)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
por: Patil, Parth, et al.
Publicado: (2026)
por: Patil, Parth, et al.
Publicado: (2026)
From Press to Pixels: Evolving Urdu Text Recognition
por: Arif, Samee, et al.
Publicado: (2025)
por: Arif, Samee, et al.
Publicado: (2025)
Social Debiasing for Fair Multi-modal LLMs
por: Cheng, Harry, et al.
Publicado: (2024)
por: Cheng, Harry, et al.
Publicado: (2024)
To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMs
por: Khan, Zohaib, et al.
Publicado: (2026)
por: Khan, Zohaib, et al.
Publicado: (2026)
LLMs-Healthcare : Current Applications and Challenges of Large Language Models in various Medical Specialties
por: Mumtaz, Ummara, et al.
Publicado: (2023)
por: Mumtaz, Ummara, et al.
Publicado: (2023)
LLMs on a Budget? Say HOLA
por: Siddiqui, Zohaib Hasan, et al.
Publicado: (2025)
por: Siddiqui, Zohaib Hasan, et al.
Publicado: (2025)
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
por: Borah, Angana, et al.
Publicado: (2026)
por: Borah, Angana, et al.
Publicado: (2026)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
por: Du, Yongkang, et al.
Publicado: (2025)
por: Du, Yongkang, et al.
Publicado: (2025)
BEADs: Bias Evaluation Across Domains
por: Raza, Shaina, et al.
Publicado: (2024)
por: Raza, Shaina, et al.
Publicado: (2024)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
por: Jeong, Hyejun, et al.
Publicado: (2024)
por: Jeong, Hyejun, et al.
Publicado: (2024)
How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions
por: Kharchenko, Julia, et al.
Publicado: (2024)
por: Kharchenko, Julia, et al.
Publicado: (2024)
Revolutionizing Process Mining: A Novel Architecture for ChatGPT Integration and Enhanced User Experience through Optimized Prompt Engineering
por: Kermani, Mehrdad Agha Mohammad Ali, et al.
Publicado: (2024)
por: Kermani, Mehrdad Agha Mohammad Ali, et al.
Publicado: (2024)
Automating Thematic Analysis: How LLMs Analyse Controversial Topics
por: Khan, Awais Hameed, et al.
Publicado: (2024)
por: Khan, Awais Hameed, et al.
Publicado: (2024)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
por: Fan, Zhiting, et al.
Publicado: (2024)
por: Fan, Zhiting, et al.
Publicado: (2024)
Fine-Grained and Thematic Evaluation of LLMs in Social Deduction Game
por: Kim, Byungjun, et al.
Publicado: (2024)
por: Kim, Byungjun, et al.
Publicado: (2024)
Enhancing Naturalness in LLM-Generated Utterances through Disfluency Insertion
por: Hassan, Syed Zohaib, et al.
Publicado: (2024)
por: Hassan, Syed Zohaib, et al.
Publicado: (2024)
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups
por: Veldanda, Gautam
Publicado: (2026)
por: Veldanda, Gautam
Publicado: (2026)
The Impossibility of Fair LLMs
por: Anthis, Jacy, et al.
Publicado: (2024)
por: Anthis, Jacy, et al.
Publicado: (2024)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
por: Khalifa, Muhammad, et al.
Publicado: (2026)
por: Khalifa, Muhammad, et al.
Publicado: (2026)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
por: Chatrath, Veronica, et al.
Publicado: (2024)
por: Chatrath, Veronica, et al.
Publicado: (2024)
FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
por: Nawale, Janki Atul, et al.
Publicado: (2025)
por: Nawale, Janki Atul, et al.
Publicado: (2025)
MARBLE: A Multi-Agent Rule-Based LLM Reasoning Engine for Accident Severity Prediction
por: Qasim, Kaleem Ullah, et al.
Publicado: (2025)
por: Qasim, Kaleem Ullah, et al.
Publicado: (2025)
Ejemplares similares
-
Generalists vs. Specialists: Evaluating Large Language Models for Urdu
por: Arif, Samee, et al.
Publicado: (2024) -
UQA: Corpus for Urdu Question Answering
por: Arif, Samee, et al.
Publicado: (2024) -
The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation
por: Arif, Samee, et al.
Publicado: (2024) -
Kahaani: A Multimodal Co-Creative Storytelling System
por: Arif, Samee, et al.
Publicado: (2024) -
WER We Stand: Benchmarking Urdu ASR Models
por: Arif, Samee, et al.
Publicado: (2024)