On the Credibility of Evaluating LLMs using Survey Questions
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Libovický, Jindřich |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Gender Interacts with Political Values: A Case Study on Czech BERT Models
von: Ali, Adnan Al, et al.
Veröffentlicht: (2024)
von: Ali, Adnan Al, et al.
Veröffentlicht: (2024)
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
von: Vico, Gianluca, et al.
Veröffentlicht: (2026)
von: Vico, Gianluca, et al.
Veröffentlicht: (2026)
Lexically Grounded Subword Segmentation
von: Libovický, Jindřich, et al.
Veröffentlicht: (2024)
von: Libovický, Jindřich, et al.
Veröffentlicht: (2024)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
von: Adilazuarda, Muhammad Farid, et al.
Veröffentlicht: (2024)
von: Adilazuarda, Muhammad Farid, et al.
Veröffentlicht: (2024)
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
von: Tzachristas, Ioannis, et al.
Veröffentlicht: (2025)
von: Tzachristas, Ioannis, et al.
Veröffentlicht: (2025)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
von: Atif, Farah, et al.
Veröffentlicht: (2025)
von: Atif, Farah, et al.
Veröffentlicht: (2025)
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
von: Stephen, Abishek, et al.
Veröffentlicht: (2026)
von: Stephen, Abishek, et al.
Veröffentlicht: (2026)
CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
von: Libovický, Jindřich, et al.
Veröffentlicht: (2025)
von: Libovický, Jindřich, et al.
Veröffentlicht: (2025)
Multilingual Vision-Language Models, A Survey
von: Manea, Andrei-Alexandru, et al.
Veröffentlicht: (2025)
von: Manea, Andrei-Alexandru, et al.
Veröffentlicht: (2025)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
von: Haider, Batool, et al.
Veröffentlicht: (2025)
von: Haider, Batool, et al.
Veröffentlicht: (2025)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
von: Dawson, Fiifi, et al.
Veröffentlicht: (2024)
von: Dawson, Fiifi, et al.
Veröffentlicht: (2024)
Automated Assessment of Students' Code Comprehension using LLMs
von: Oli, Priti, et al.
Veröffentlicht: (2023)
von: Oli, Priti, et al.
Veröffentlicht: (2023)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
Unmasking and Improving Data Credibility: A Study with Datasets for Training Harmless Language Models
von: Zhu, Zhaowei, et al.
Veröffentlicht: (2023)
von: Zhu, Zhaowei, et al.
Veröffentlicht: (2023)
From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes
von: Garzón, Rubén, et al.
Veröffentlicht: (2026)
von: Garzón, Rubén, et al.
Veröffentlicht: (2026)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
von: Maurya, Sneha, et al.
Veröffentlicht: (2026)
von: Maurya, Sneha, et al.
Veröffentlicht: (2026)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
Understanding Cross-Lingual Alignment -- A Survey
von: Hämmerl, Katharina, et al.
Veröffentlicht: (2024)
von: Hämmerl, Katharina, et al.
Veröffentlicht: (2024)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
von: Shi, Yuzhen, et al.
Veröffentlicht: (2026)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
von: Li, Yuchong, et al.
Veröffentlicht: (2025)
von: Li, Yuchong, et al.
Veröffentlicht: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
von: Shin, Jisu, et al.
Veröffentlicht: (2025)
Social Bias in Popular Question-Answering Benchmarks
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
von: Kraft, Angelie, et al.
Veröffentlicht: (2025)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
von: Kaur, Navreet, et al.
Veröffentlicht: (2025)
von: Kaur, Navreet, et al.
Veröffentlicht: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
von: Mavi, John, et al.
Veröffentlicht: (2024)
von: Mavi, John, et al.
Veröffentlicht: (2024)
AIn't Nothing But a Survey? Using Large Language Models for Coding German Open-Ended Survey Responses on Survey Motivation
von: von der Heyde, Leah, et al.
Veröffentlicht: (2025)
von: von der Heyde, Leah, et al.
Veröffentlicht: (2025)
A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions
von: Ranjan, Rajesh, et al.
Veröffentlicht: (2024)
von: Ranjan, Rajesh, et al.
Veröffentlicht: (2024)
The simulation of judgment in LLMs
von: Loru, Edoardo, et al.
Veröffentlicht: (2025)
von: Loru, Edoardo, et al.
Veröffentlicht: (2025)
Measuring Teaching with LLMs
von: Hardy, Michael
Veröffentlicht: (2025)
von: Hardy, Michael
Veröffentlicht: (2025)
The Political Preferences of LLMs
von: Rozado, David
Veröffentlicht: (2024)
von: Rozado, David
Veröffentlicht: (2024)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
von: Koutcheme, Charles, et al.
Veröffentlicht: (2024)
von: Koutcheme, Charles, et al.
Veröffentlicht: (2024)
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
von: de Landa, Joseba Fernandez, et al.
Veröffentlicht: (2026)
von: de Landa, Joseba Fernandez, et al.
Veröffentlicht: (2026)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
von: Wang, Huandong, et al.
Veröffentlicht: (2025)
von: Wang, Huandong, et al.
Veröffentlicht: (2025)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
von: Steenhuis, Quinten, et al.
Veröffentlicht: (2026)
von: Steenhuis, Quinten, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Gender Interacts with Political Values: A Case Study on Czech BERT Models
von: Ali, Adnan Al, et al.
Veröffentlicht: (2024) -
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
von: Vico, Gianluca, et al.
Veröffentlicht: (2026) -
Lexically Grounded Subword Segmentation
von: Libovický, Jindřich, et al.
Veröffentlicht: (2024) -
Towards Measuring and Modeling "Culture" in LLMs: A Survey
von: Adilazuarda, Muhammad Farid, et al.
Veröffentlicht: (2024) -
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
von: Tzachristas, Ioannis, et al.
Veröffentlicht: (2025)