Gespeichert in:
| Hauptverfasser: | Akpinar, Nil-Jana, Avula, Sandeep, Lee, CJ, Dang, Brandon, Razat, Kaza, Murdock, Vanessa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.15556 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2025)
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2025)
A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text
von: Sagae, Alicia, et al.
Veröffentlicht: (2025)
von: Sagae, Alicia, et al.
Veröffentlicht: (2025)
Authenticity and exclusion: social media algorithms and the dynamics of belonging in epistemic communities
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2024)
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2024)
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2024)
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2024)
Precise Model Benchmarking with Only a Few Observations
von: Fogliato, Riccardo, et al.
Veröffentlicht: (2024)
von: Fogliato, Riccardo, et al.
Veröffentlicht: (2024)
When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries
von: Vijay, Supriti, et al.
Veröffentlicht: (2024)
von: Vijay, Supriti, et al.
Veröffentlicht: (2024)
Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
von: Prama, Tabia Tanzin, et al.
Veröffentlicht: (2025)
von: Prama, Tabia Tanzin, et al.
Veröffentlicht: (2025)
Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
von: Ashkinaze, Joshua, et al.
Veröffentlicht: (2025)
von: Ashkinaze, Joshua, et al.
Veröffentlicht: (2025)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
von: Ferrer, Robinson, et al.
Veröffentlicht: (2026)
Facts are Harder Than Opinions -- A Multilingual, Comparative Analysis of LLM-Based Fact-Checking Reliability
von: Saju, Lorraine, et al.
Veröffentlicht: (2025)
von: Saju, Lorraine, et al.
Veröffentlicht: (2025)
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
von: Alipour, Shayan, et al.
Veröffentlicht: (2024)
von: Alipour, Shayan, et al.
Veröffentlicht: (2024)
Exploring the Human-LLM Synergy in Advancing Theory-driven Qualitative Analysis
von: Meng, Han, et al.
Veröffentlicht: (2024)
von: Meng, Han, et al.
Veröffentlicht: (2024)
Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins
von: Chan, Amanda, et al.
Veröffentlicht: (2025)
von: Chan, Amanda, et al.
Veröffentlicht: (2025)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
von: Nghiem, Huy, et al.
Veröffentlicht: (2026)
von: Nghiem, Huy, et al.
Veröffentlicht: (2026)
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2023)
von: Møller, Anders Giovanni, et al.
Veröffentlicht: (2023)
Reviewing the Reviewer: Elevating Peer Review Quality through LLM-Guided Feedback
von: Purkayastha, Sukannya, et al.
Veröffentlicht: (2026)
von: Purkayastha, Sukannya, et al.
Veröffentlicht: (2026)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
von: Zhang, Bingquan, et al.
Veröffentlicht: (2025)
von: Zhang, Bingquan, et al.
Veröffentlicht: (2025)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
von: Curran, Damian, et al.
Veröffentlicht: (2025)
von: Curran, Damian, et al.
Veröffentlicht: (2025)
Culturally Adaptive Explainable LLM Assessment for Multilingual Information Disorder: A Human-in-the-Loop Approach
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
von: Jouneghani, Maziar Kianimoghadam
Veröffentlicht: (2026)
Trust, Safety, and Accuracy: Assessing LLMs for Routine Maternity Advice
von: Divya, V Sai, et al.
Veröffentlicht: (2026)
von: Divya, V Sai, et al.
Veröffentlicht: (2026)
Which Type of Students can LLMs Act? Investigating Authentic Simulation with Graph-based Human-AI Collaborative System
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency
von: Batzner, Jan, et al.
Veröffentlicht: (2025)
von: Batzner, Jan, et al.
Veröffentlicht: (2025)
"Would You Want an AI Tutor?" Understanding Stakeholder Perceptions of LLM-based Systems in the Classroom
von: Fuligni, Caterina, et al.
Veröffentlicht: (2025)
von: Fuligni, Caterina, et al.
Veröffentlicht: (2025)
Words of Warmth: Trust and Sociability Norms for over 26k English Words
von: Mohammad, Saif M.
Veröffentlicht: (2025)
von: Mohammad, Saif M.
Veröffentlicht: (2025)
Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text
von: Najjar, Ayat, et al.
Veröffentlicht: (2025)
von: Najjar, Ayat, et al.
Veröffentlicht: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
von: Kwon, Jea, et al.
Veröffentlicht: (2025)
von: Kwon, Jea, et al.
Veröffentlicht: (2025)
Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
von: Cunningham, Eoghan, et al.
Veröffentlicht: (2025)
von: Cunningham, Eoghan, et al.
Veröffentlicht: (2025)
Human Preferences for Constructive Interactions in Language Model Alignment
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
Users Mispredict Their Own Preferences for AI Writing Assistance
von: Lai, Vivian, et al.
Veröffentlicht: (2026)
von: Lai, Vivian, et al.
Veröffentlicht: (2026)
LLMs as Research Tools: A Large Scale Survey of Researchers' Usage and Perceptions
von: Liao, Zhehui, et al.
Veröffentlicht: (2024)
von: Liao, Zhehui, et al.
Veröffentlicht: (2024)
Training LLM-based Tutors to Improve Student Learning Outcomes in Dialogues
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
Generative AI in Higher Education: Seeing ChatGPT Through Universities' Policies, Resources, and Guidelines
von: Wang, Hui, et al.
Veröffentlicht: (2023)
von: Wang, Hui, et al.
Veröffentlicht: (2023)
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
von: Ivey, Jonathan, et al.
Veröffentlicht: (2024)
von: Ivey, Jonathan, et al.
Veröffentlicht: (2024)
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
von: Ehara, Yo
Veröffentlicht: (2026)
von: Ehara, Yo
Veröffentlicht: (2026)
Can Large Language Models Unlock Novel Scientific Research Ideas?
von: Kumar, Sandeep, et al.
Veröffentlicht: (2024)
von: Kumar, Sandeep, et al.
Veröffentlicht: (2024)
Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings
von: Hong, Harbin, et al.
Veröffentlicht: (2025)
von: Hong, Harbin, et al.
Veröffentlicht: (2025)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
Building Trust: Foundations of Security, Safety and Transparency in AI
von: Sidhpurwala, Huzaifa, et al.
Veröffentlicht: (2024)
von: Sidhpurwala, Huzaifa, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2025) -
A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text
von: Sagae, Alicia, et al.
Veröffentlicht: (2025) -
Authenticity and exclusion: social media algorithms and the dynamics of belonging in epistemic communities
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2024) -
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2024) -
Precise Model Benchmarking with Only a Few Observations
von: Fogliato, Riccardo, et al.
Veröffentlicht: (2024)