Evaluating how LLM annotations represent diverse views on contentious topics
Fuente:
arXiv
Saved in:
| Main Authors: | Brown, Megan A., Atreja, Shubham, Hemphill, Libby, Wu, Patrick Y. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
by: Li, Lingyao, et al.
Published: (2023)
by: Li, Lingyao, et al.
Published: (2023)
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
by: Atreja, Shubham, et al.
Published: (2024)
by: Atreja, Shubham, et al.
Published: (2024)
AppealMod: Inducing Friction to Reduce Moderator Workload of Handling User Appeals
by: Atreja, Shubham, et al.
Published: (2023)
by: Atreja, Shubham, et al.
Published: (2023)
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
by: Hua, Wenyue, et al.
Published: (2023)
by: Hua, Wenyue, et al.
Published: (2023)
Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches
by: Wu, Patrick Y.
Published: (2025)
by: Wu, Patrick Y.
Published: (2025)
ALAS: Autonomous Learning Agent for Self-Updating Language Models
by: Atreja, Dhruv
Published: (2025)
by: Atreja, Dhruv
Published: (2025)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
by: Liu, Xiaoze, et al.
Published: (2024)
by: Liu, Xiaoze, et al.
Published: (2024)
Epistemic Constitutionalism Or: how to avoid coherence bias
by: Loi, Michele
Published: (2026)
by: Loi, Michele
Published: (2026)
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
by: Liu, Yunting, et al.
Published: (2024)
by: Liu, Yunting, et al.
Published: (2024)
The why, what, and how of AI-based coding in scientific research
by: Zhuang, Tonghe, et al.
Published: (2024)
by: Zhuang, Tonghe, et al.
Published: (2024)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
by: Tang, Zeyu, et al.
Published: (2026)
by: Tang, Zeyu, et al.
Published: (2026)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
by: Zugecova, Aneta, et al.
Published: (2024)
by: Zugecova, Aneta, et al.
Published: (2024)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
by: Jahara, Fatima, et al.
Published: (2025)
by: Jahara, Fatima, et al.
Published: (2025)
A closer look at how large language models trust humans: patterns and biases
by: Lerman, Valeria, et al.
Published: (2025)
by: Lerman, Valeria, et al.
Published: (2025)
From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes
by: Garzón, Rubén, et al.
Published: (2026)
by: Garzón, Rubén, et al.
Published: (2026)
How English Print Media Frames Human-Elephant Conflicts in India
by: Punith, Bonala Sai, et al.
Published: (2026)
by: Punith, Bonala Sai, et al.
Published: (2026)
Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
Leveraging Machine Learning to Detect Data Curation Activities
by: Lafia, Sara, et al.
Published: (2021)
by: Lafia, Sara, et al.
Published: (2021)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
Evaluating the Impact of Advanced LLM Techniques on AI-Lecture Tutors for a Robotics Course
by: Kahl, Sebastian, et al.
Published: (2024)
by: Kahl, Sebastian, et al.
Published: (2024)
LLM Nepotism in Organizational Governance
by: Mao, Shunqi, et al.
Published: (2026)
by: Mao, Shunqi, et al.
Published: (2026)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Mitigating Gambling-Like Risk-Taking Behaviors in Large Language Models: A Behavioral Economics Approach to AI Safety
by: Du, Y.
Published: (2025)
by: Du, Y.
Published: (2025)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
by: Ghosh, Himel, et al.
Published: (2026)
by: Ghosh, Himel, et al.
Published: (2026)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
by: Li, Chenyu, et al.
Published: (2026)
by: Li, Chenyu, et al.
Published: (2026)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
LLM_annotate: A Python package for annotating and analyzing fiction characters
by: Rosenbusch, Hannes
Published: (2025)
by: Rosenbusch, Hannes
Published: (2025)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
by: Hawkins, John, et al.
Published: (2025)
by: Hawkins, John, et al.
Published: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
by: Kwon, Jea, et al.
Published: (2025)
by: Kwon, Jea, et al.
Published: (2025)
Societal Alignment Frameworks Can Improve LLM Alignment
by: Stańczak, Karolina, et al.
Published: (2025)
by: Stańczak, Karolina, et al.
Published: (2025)
Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions
by: Hu, Yiran, et al.
Published: (2026)
by: Hu, Yiran, et al.
Published: (2026)
Clinical Note Bloat Reduction for Efficient LLM Use
by: Cahoon, Jordan L., et al.
Published: (2026)
by: Cahoon, Jordan L., et al.
Published: (2026)
Navigating LLM Ethics: Advancements, Challenges, and Future Directions
by: Jiao, Junfeng, et al.
Published: (2024)
by: Jiao, Junfeng, et al.
Published: (2024)
Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
by: Bilalpur, Maneesh, et al.
Published: (2025)
by: Bilalpur, Maneesh, et al.
Published: (2025)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
by: Kim, Jiseon, et al.
Published: (2025)
by: Kim, Jiseon, et al.
Published: (2025)
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents
by: Vatsal, Shubham, et al.
Published: (2026)
by: Vatsal, Shubham, et al.
Published: (2026)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
Similar Items
-
"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media
by: Li, Lingyao, et al.
Published: (2023) -
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
by: Atreja, Shubham, et al.
Published: (2024) -
AppealMod: Inducing Friction to Reduce Moderator Workload of Handling User Appeals
by: Atreja, Shubham, et al.
Published: (2023) -
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
by: Hua, Wenyue, et al.
Published: (2023) -
Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches
by: Wu, Patrick Y.
Published: (2025)