When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Najera, Aisha, Moon, Alvin, Srinivasan, Vedant, Veeraraghavan, Rajesh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts
by: Field, Severin
Published: (2025)
by: Field, Severin
Published: (2025)
Personalizing explanations of AI-driven hints to users' characteristics: an empirical evaluation
by: Bahel, Vedant, et al.
Published: (2024)
by: Bahel, Vedant, et al.
Published: (2024)
SeSaMe: A Framework to Simulate Self-Reported Ground Truth for Mental Health Sensing Studies
by: Choube, Akshat, et al.
Published: (2024)
by: Choube, Akshat, et al.
Published: (2024)
Comprehensive Study on Sentiment Analysis: From Rule-based to modern LLM based system
by: Gupta, Shailja, et al.
Published: (2024)
by: Gupta, Shailja, et al.
Published: (2024)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)
by: Li, Karen Jia-Hui, et al.
Published: (2025)
Not My Truce: Personality Differences in AI-Mediated Workplace Negotiation
by: Duddu, Veda, et al.
Published: (2026)
by: Duddu, Veda, et al.
Published: (2026)
Does AI Coaching Prepare us for Workplace Negotiations?
by: Duddu, Veda, et al.
Published: (2025)
by: Duddu, Veda, et al.
Published: (2025)
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
by: Kumar, Harsh, et al.
Published: (2025)
by: Kumar, Harsh, et al.
Published: (2025)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
by: Ghosh, Himel, et al.
Published: (2026)
by: Ghosh, Himel, et al.
Published: (2026)
Not someone, but something: Rethinking trust in the age of medical AI
by: Beger, Jan
Published: (2025)
by: Beger, Jan
Published: (2025)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
Unheard in the Digital Age: Rethinking AI Bias and Speech Diversity
by: Amaechi-Okorie, Onyedikachi Hope, et al.
Published: (2026)
by: Amaechi-Okorie, Onyedikachi Hope, et al.
Published: (2026)
Between Myths and Metaphors: Rethinking LLMs for SRH in Conservative Contexts
by: Humayun, Ameemah, et al.
Published: (2025)
by: Humayun, Ameemah, et al.
Published: (2025)
TUX: Measuring Human--AI Tacit Understanding
by: Li, Yueshen, et al.
Published: (2026)
by: Li, Yueshen, et al.
Published: (2026)
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
by: Konya, Andrew, et al.
Published: (2024)
by: Konya, Andrew, et al.
Published: (2024)
A Meta-Analysis of LLM Effects on Students across Qualification, Socialisation, and Subjectification
by: Huang, Jiayu, et al.
Published: (2025)
by: Huang, Jiayu, et al.
Published: (2025)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
by: Shi, Zhengliang, et al.
Published: (2025)
by: Shi, Zhengliang, et al.
Published: (2025)
When Autonomy Breaks: The Hidden Existential Risk of AI
by: Krook, Joshua
Published: (2025)
by: Krook, Joshua
Published: (2025)
Cognitive Bias Detection Using Advanced Prompt Engineering
by: Lemieux, Frederic, et al.
Published: (2025)
by: Lemieux, Frederic, et al.
Published: (2025)
Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation
by: An, Yuan
Published: (2026)
by: An, Yuan
Published: (2026)
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines
by: Pendse, Sachin R., et al.
Published: (2025)
by: Pendse, Sachin R., et al.
Published: (2025)
When combinations of humans and AI are useful: A systematic review and meta-analysis
by: Vaccaro, Michelle, et al.
Published: (2024)
by: Vaccaro, Michelle, et al.
Published: (2024)
When Visibility Outpaces Verification: Delayed Verification and Narrative Lock-in in Agentic AI Discourse
by: Shi, Hanjing, et al.
Published: (2026)
by: Shi, Hanjing, et al.
Published: (2026)
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
by: Wei, Kevin L., et al.
Published: (2025)
by: Wei, Kevin L., et al.
Published: (2025)
Evaluating Alternative Training Interventions Using Personalized Computational Models of Learning
by: MacLellan, Christopher James, et al.
Published: (2024)
by: MacLellan, Christopher James, et al.
Published: (2024)
WaLLM -- Insights from an LLM-Powered Chatbot deployment via WhatsApp
by: Eltigani, Hiba, et al.
Published: (2025)
by: Eltigani, Hiba, et al.
Published: (2025)
"What if she doesn't feel the same?" What Happens When We Ask AI for Relationship Advice
by: Manchanda, Niva, et al.
Published: (2025)
by: Manchanda, Niva, et al.
Published: (2025)
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
by: Ding, Shi, et al.
Published: (2025)
by: Ding, Shi, et al.
Published: (2025)
Modeling Public Perceptions of Science in Media
by: Pei, Jiaxin, et al.
Published: (2025)
by: Pei, Jiaxin, et al.
Published: (2025)
A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
by: Winslow, Brent, et al.
Published: (2025)
by: Winslow, Brent, et al.
Published: (2025)
Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations
by: van der Linden, Ilona, et al.
Published: (2025)
by: van der Linden, Ilona, et al.
Published: (2025)
Public Discourse Sandbox: Facilitating Human and AI Digital Communication Research
by: Radivojevic, Kristina, et al.
Published: (2025)
by: Radivojevic, Kristina, et al.
Published: (2025)
"This is not a data problem": Algorithms and Power in Public Higher Education in Canada
by: McConvey, Kelly, et al.
Published: (2024)
by: McConvey, Kelly, et al.
Published: (2024)
Public Opinion and The Rise of Digital Minds: Perceived Risk, Trust, and Regulation Support
by: Bullock, Justin B., et al.
Published: (2025)
by: Bullock, Justin B., et al.
Published: (2025)
WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis
by: Chen, Liangliang, et al.
Published: (2025)
by: Chen, Liangliang, et al.
Published: (2025)
Enhancing Large Language Models for Automated Homework Assessment in Undergraduate Circuit Analysis
by: Chen, Liangliang, et al.
Published: (2025)
by: Chen, Liangliang, et al.
Published: (2025)
Examining and Addressing Barriers to Diversity in LLM-Generated Ideas
by: Deng, Yuting, et al.
Published: (2026)
by: Deng, Yuting, et al.
Published: (2026)
Towards an LLM-powered Social Digital Twinning Platform
by: Gürcan, Önder, et al.
Published: (2025)
by: Gürcan, Önder, et al.
Published: (2025)
Knowing Your Uncertainty -- On the application of LLM in social sciences
by: Zhang, Bolun, et al.
Published: (2025)
by: Zhang, Bolun, et al.
Published: (2025)
Similar Items
-
Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts
by: Field, Severin
Published: (2025) -
Personalizing explanations of AI-driven hints to users' characteristics: an empirical evaluation
by: Bahel, Vedant, et al.
Published: (2024) -
SeSaMe: A Framework to Simulate Self-Reported Ground Truth for Mental Health Sensing Studies
by: Choube, Akshat, et al.
Published: (2024) -
Comprehensive Study on Sentiment Analysis: From Rule-based to modern LLM based system
by: Gupta, Shailja, et al.
Published: (2024) -
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)