Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
Fuente:
arXiv
Salvato in:
| Autori principali: | Atif, Farah, Askarbekuly, Nursultan, Darwish, Kareem, Choudhury, Monojit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
di: Saha, Sougata, et al.
Pubblicazione: (2025)
di: Saha, Sougata, et al.
Pubblicazione: (2025)
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
di: Saha, Sougata, et al.
Pubblicazione: (2025)
di: Saha, Sougata, et al.
Pubblicazione: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
Improving LLM Reliability with RAG in Religious Question-Answering: MufassirQAS
di: Alan, Ahmet Yusuf, et al.
Pubblicazione: (2024)
di: Alan, Ahmet Yusuf, et al.
Pubblicazione: (2024)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2024)
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2024)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
di: Kaur, Navreet, et al.
Pubblicazione: (2025)
di: Kaur, Navreet, et al.
Pubblicazione: (2025)
On the Credibility of Evaluating LLMs using Survey Questions
di: Libovický, Jindřich
Pubblicazione: (2026)
di: Libovický, Jindřich
Pubblicazione: (2026)
Sacred or Secular? Religious Bias in AI-Generated Financial Advice
di: Khan, Muhammad Salar, et al.
Pubblicazione: (2025)
di: Khan, Muhammad Salar, et al.
Pubblicazione: (2025)
SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection
di: Kazemi, Arefeh, et al.
Pubblicazione: (2025)
di: Kazemi, Arefeh, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
di: Kaur, Navreet, et al.
Pubblicazione: (2023)
di: Kaur, Navreet, et al.
Pubblicazione: (2023)
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
di: Najjar, Ayat A., et al.
Pubblicazione: (2025)
di: Najjar, Ayat A., et al.
Pubblicazione: (2025)
Exploring LGBTQ+ Bias in Generative AI Answers across Different Country and Religious Contexts
di: Vicsek, Lilla, et al.
Pubblicazione: (2024)
di: Vicsek, Lilla, et al.
Pubblicazione: (2024)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
di: Chen, Yuyan, et al.
Pubblicazione: (2024)
Understanding Social Support Needs in Questions: A Hybrid Approach Integrating Semi-Supervised Learning and LLM-based Data Augmentation
di: Kuang, Junwei, et al.
Pubblicazione: (2025)
di: Kuang, Junwei, et al.
Pubblicazione: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
di: Tang, Zeyu, et al.
Pubblicazione: (2026)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
di: Zugecova, Aneta, et al.
Pubblicazione: (2024)
di: Zugecova, Aneta, et al.
Pubblicazione: (2024)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
di: Liu, Yunting, et al.
Pubblicazione: (2024)
di: Liu, Yunting, et al.
Pubblicazione: (2024)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
di: Jahara, Fatima, et al.
Pubblicazione: (2025)
di: Jahara, Fatima, et al.
Pubblicazione: (2025)
Evaluating how LLM annotations represent diverse views on contentious topics
di: Brown, Megan A., et al.
Pubblicazione: (2025)
di: Brown, Megan A., et al.
Pubblicazione: (2025)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
di: Liu, Xiaoze, et al.
Pubblicazione: (2024)
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test
di: Khandelwal, Aditi, et al.
Pubblicazione: (2024)
di: Khandelwal, Aditi, et al.
Pubblicazione: (2024)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
di: Agarwal, Utkarsh, et al.
Pubblicazione: (2024)
di: Agarwal, Utkarsh, et al.
Pubblicazione: (2024)
From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes
di: Garzón, Rubén, et al.
Pubblicazione: (2026)
di: Garzón, Rubén, et al.
Pubblicazione: (2026)
Social Bias in Popular Question-Answering Benchmarks
di: Kraft, Angelie, et al.
Pubblicazione: (2025)
di: Kraft, Angelie, et al.
Pubblicazione: (2025)
Evaluating Reliability Asymmetries in Chinese Factual Search and AI Answers
di: Liu, Geng, et al.
Pubblicazione: (2025)
di: Liu, Geng, et al.
Pubblicazione: (2025)
Semantic Consistency for Assuring Reliability of Large Language Models
di: Raj, Harsh, et al.
Pubblicazione: (2023)
di: Raj, Harsh, et al.
Pubblicazione: (2023)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
di: Mittal, Avni, et al.
Pubblicazione: (2026)
di: Mittal, Avni, et al.
Pubblicazione: (2026)
Synthetic Reader Panels: Tournament-Based Ideation with LLM Personas for Autonomous Publishing
di: Zimmerman, Fred
Pubblicazione: (2026)
di: Zimmerman, Fred
Pubblicazione: (2026)
Knowledge Graph Guided Evaluation of Abstention Techniques
di: Vasisht, Kinshuk, et al.
Pubblicazione: (2024)
di: Vasisht, Kinshuk, et al.
Pubblicazione: (2024)
Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
di: Rozado, David
Pubblicazione: (2025)
di: Rozado, David
Pubblicazione: (2025)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
di: Nghiem, Huy, et al.
Pubblicazione: (2026)
di: Nghiem, Huy, et al.
Pubblicazione: (2026)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
di: Saha, Sougata, et al.
Pubblicazione: (2025)
di: Saha, Sougata, et al.
Pubblicazione: (2025)
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education
di: Singh, Shrutika, et al.
Pubblicazione: (2025)
di: Singh, Shrutika, et al.
Pubblicazione: (2025)
EduAgentQG: A Multi-Agent Workflow Framework for Personalized Question Generation
di: Jia, Rui, et al.
Pubblicazione: (2025)
di: Jia, Rui, et al.
Pubblicazione: (2025)
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
di: Steenhuis, Quinten, et al.
Pubblicazione: (2026)
di: Steenhuis, Quinten, et al.
Pubblicazione: (2026)
Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
di: Shankar, Ravi, et al.
Pubblicazione: (2025)
di: Shankar, Ravi, et al.
Pubblicazione: (2025)
What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge
di: Dutta, Arka, et al.
Pubblicazione: (2025)
di: Dutta, Arka, et al.
Pubblicazione: (2025)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
di: Zhang, Jingshen, et al.
Pubblicazione: (2024)
di: Zhang, Jingshen, et al.
Pubblicazione: (2024)
Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge
di: Usmanova, Aida, et al.
Pubblicazione: (2024)
di: Usmanova, Aida, et al.
Pubblicazione: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
di: Haider, Batool, et al.
Pubblicazione: (2025)
di: Haider, Batool, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness
di: Saha, Sougata, et al.
Pubblicazione: (2025) -
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
di: Saha, Sougata, et al.
Pubblicazione: (2025) -
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024) -
Improving LLM Reliability with RAG in Religious Question-Answering: MufassirQAS
di: Alan, Ahmet Yusuf, et al.
Pubblicazione: (2024) -
Towards Measuring and Modeling "Culture" in LLMs: A Survey
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2024)