ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
Fuente:
arXiv
Saved in:
| Main Authors: | Banyas, Peter, Sharma, Shristi, Simmons, Alistair, Vispute, Atharva |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
by: Qi, Jirui, et al.
Published: (2023)
by: Qi, Jirui, et al.
Published: (2023)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
by: Subedi, Krishna
Published: (2025)
by: Subedi, Krishna
Published: (2025)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025)
by: Schlicht, Ipek Baris, et al.
Published: (2025)
Consistency of Responses and Continuations Generated by Large Language Models on Social Media
by: Xu, Wentao, et al.
Published: (2025)
by: Xu, Wentao, et al.
Published: (2025)
Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support
by: Chiang, Sophie, et al.
Published: (2025)
by: Chiang, Sophie, et al.
Published: (2025)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
by: Zhao, Runcong, et al.
Published: (2025)
by: Zhao, Runcong, et al.
Published: (2025)
Psittacines of Innovation? Assessing the True Novelty of AI Creations
by: Mukherjee, Anirban
Published: (2024)
by: Mukherjee, Anirban
Published: (2024)
LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
by: Kirgis, Peter, et al.
Published: (2026)
by: Kirgis, Peter, et al.
Published: (2026)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
by: Liu, Hongtao, et al.
Published: (2025)
by: Liu, Hongtao, et al.
Published: (2025)
The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
by: Cohen, Ruth, et al.
Published: (2026)
by: Cohen, Ruth, et al.
Published: (2026)
AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers
by: Wuttke, Alexander, et al.
Published: (2024)
by: Wuttke, Alexander, et al.
Published: (2024)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
by: He, Li, et al.
Published: (2025)
by: He, Li, et al.
Published: (2025)
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
by: Wu, Xiaoyuan, et al.
Published: (2025)
by: Wu, Xiaoyuan, et al.
Published: (2025)
Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages
by: Biswas, Shreyan, et al.
Published: (2025)
by: Biswas, Shreyan, et al.
Published: (2025)
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
by: Shi, Quan, et al.
Published: (2025)
by: Shi, Quan, et al.
Published: (2025)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
RELIC: Investigating Large Language Model Responses using Self-Consistency
by: Cheng, Furui, et al.
Published: (2023)
by: Cheng, Furui, et al.
Published: (2023)
LVLMs and Humans Ground Differently in Referential Communication
by: Zeng, Peter, et al.
Published: (2026)
by: Zeng, Peter, et al.
Published: (2026)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
When AI Writes, Whose Voice Remains? Quantifying Cultural Marker Erasure Across World English Varieties in Large Language Models
by: Navneet, Satyam Kumar, et al.
Published: (2026)
by: Navneet, Satyam Kumar, et al.
Published: (2026)
An AI-Powered Research Assistant in the Lab: A Practical Guide for Text Analysis Through Iterative Collaboration with LLMs
by: Carmona-Díaz, Gino, et al.
Published: (2025)
by: Carmona-Díaz, Gino, et al.
Published: (2025)
On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts
by: Aremu, Toluwani, et al.
Published: (2024)
by: Aremu, Toluwani, et al.
Published: (2024)
LLMs Corrupt Your Documents When You Delegate
by: Laban, Philippe, et al.
Published: (2026)
by: Laban, Philippe, et al.
Published: (2026)
Not My Truce: Personality Differences in AI-Mediated Workplace Negotiation
by: Duddu, Veda, et al.
Published: (2026)
by: Duddu, Veda, et al.
Published: (2026)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
by: Andrusenko, Andrei, et al.
Published: (2026)
by: Andrusenko, Andrei, et al.
Published: (2026)
Sima AIunty: Caste Audit in LLM-Driven Matchmaking
by: Naik, Atharva, et al.
Published: (2026)
by: Naik, Atharva, et al.
Published: (2026)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
by: Ngueajio, Mikel K., et al.
Published: (2025)
by: Ngueajio, Mikel K., et al.
Published: (2025)
ChatBench: From Static Benchmarks to Human-AI Evaluation
by: Chang, Serina, et al.
Published: (2025)
by: Chang, Serina, et al.
Published: (2025)
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
by: He, Zeyu, et al.
Published: (2025)
by: He, Zeyu, et al.
Published: (2025)
Performance Gains of LLMs With Humans in a World of LLMs Versus Humans
by: McCullum, Lucas, et al.
Published: (2025)
by: McCullum, Lucas, et al.
Published: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
When ChatGPT is gone: Creativity reverts and homogeneity persists
by: Liu, Qinghan, et al.
Published: (2024)
by: Liu, Qinghan, et al.
Published: (2024)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
by: Xu, Haoming, et al.
Published: (2026)
by: Xu, Haoming, et al.
Published: (2026)
Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers
by: Schapiro, Samuel, et al.
Published: (2026)
by: Schapiro, Samuel, et al.
Published: (2026)
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
by: Ashkinaze, Joshua, et al.
Published: (2024)
by: Ashkinaze, Joshua, et al.
Published: (2024)
Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning
by: Xu, Yinggan, et al.
Published: (2025)
by: Xu, Yinggan, et al.
Published: (2025)
Generating Educational Materials with Different Levels of Readability using LLMs
by: Huang, Chieh-Yang, et al.
Published: (2024)
by: Huang, Chieh-Yang, et al.
Published: (2024)
Similar Items
-
Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
by: Qi, Jirui, et al.
Published: (2023) -
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
by: Subedi, Krishna
Published: (2025) -
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026) -
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025) -
Consistency of Responses and Continuations Generated by Large Language Models on Social Media
by: Xu, Wentao, et al.
Published: (2025)