What Is Actually Being Annotated? Inter-Prompt Reliability as a Measurement Problem in LLM-Based Social Science Labeling
Fuente:
arXiv
Saved in:
| Main Author: | Liu, Jingyuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt Selection Matters: Enhancing Text Annotations for Social Sciences with Large Language Models
by: Abraham, Louis, et al.
Published: (2024)
by: Abraham, Louis, et al.
Published: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
by: Walsh, Cole, et al.
Published: (2026)
by: Walsh, Cole, et al.
Published: (2026)
What Work is AI Actually Doing? Uncovering the Drivers of Generative AI Adoption
by: Agarwal, Peeyush, et al.
Published: (2025)
by: Agarwal, Peeyush, et al.
Published: (2025)
LLM-Assisted Replication for Quantitative Social Science
by: Kubota, So, et al.
Published: (2026)
by: Kubota, So, et al.
Published: (2026)
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
by: Atreja, Shubham, et al.
Published: (2024)
by: Atreja, Shubham, et al.
Published: (2024)
Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
What's the Problem, Linda? The Conjunction Fallacy as a Fairness Problem
by: Colmenares, Jose Alvarez
Published: (2023)
by: Colmenares, Jose Alvarez
Published: (2023)
A More Advanced Group Polarization Measurement Approach Based on LLM-Based Agents and Graphs
by: Liu, Zixin, et al.
Published: (2024)
by: Liu, Zixin, et al.
Published: (2024)
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026)
by: Rabanser, Stephan, et al.
Published: (2026)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
by: Voudouris, Konstantinos, et al.
Published: (2026)
by: Voudouris, Konstantinos, et al.
Published: (2026)
Situated Ground Truths: Enhancing Bias-Aware AI by Situating Data Labels with SituAnnotate
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
by: Pandiani, Delfina Sol Martinez, et al.
Published: (2024)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
by: Zugecova, Aneta, et al.
Published: (2024)
by: Zugecova, Aneta, et al.
Published: (2024)
Rethinking LLM Bias Probing Using Lessons from the Social Sciences
by: Morehouse, Kirsten N., et al.
Published: (2025)
by: Morehouse, Kirsten N., et al.
Published: (2025)
Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
by: Koorndijk, Jeanice
Published: (2025)
by: Koorndijk, Jeanice
Published: (2025)
GeoAI in Social Science
by: Li, Wenwen
Published: (2023)
by: Li, Wenwen
Published: (2023)
The Landscape of AI in Science Education: What is Changing and How to Respond
by: Zhai, Xiaoming, et al.
Published: (2026)
by: Zhai, Xiaoming, et al.
Published: (2026)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
by: Wang, Bingchen, et al.
Published: (2025)
by: Wang, Bingchen, et al.
Published: (2025)
Evaluating Code Generation of LLMs in Advanced Computer Science Problems
by: Catir, Emir, et al.
Published: (2025)
by: Catir, Emir, et al.
Published: (2025)
Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem
by: Takata, Ryosuke, et al.
Published: (2025)
by: Takata, Ryosuke, et al.
Published: (2025)
Are Researchers Being Replaced by Artificial Intelligence?
by: Salatino, Angelo A., et al.
Published: (2026)
by: Salatino, Angelo A., et al.
Published: (2026)
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
by: Rupprecht, Jens, et al.
Published: (2025)
by: Rupprecht, Jens, et al.
Published: (2025)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
Homoglyph-based Adversarial Perturbation of Introductory Computer Science Theory Problems
by: Alexander, Aidan, et al.
Published: (2026)
by: Alexander, Aidan, et al.
Published: (2026)
Prompt Programming: A Platform for Dialogue-based Computational Problem Solving with Generative AI Models
by: Pădurean, Victor-Alexandru, et al.
Published: (2025)
by: Pădurean, Victor-Alexandru, et al.
Published: (2025)
neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings
by: Ramamoorthy, Arnav, et al.
Published: (2025)
by: Ramamoorthy, Arnav, et al.
Published: (2025)
Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
by: Hong, Shen Zhou, et al.
Published: (2026)
by: Hong, Shen Zhou, et al.
Published: (2026)
Arti-"fickle" Intelligence: Using LLMs as a Tool for Inference in the Political and Social Sciences
by: Argyle, Lisa P., et al.
Published: (2025)
by: Argyle, Lisa P., et al.
Published: (2025)
Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism
by: Zhong, Huixin, et al.
Published: (2025)
by: Zhong, Huixin, et al.
Published: (2025)
On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts
by: Aremu, Toluwani, et al.
Published: (2024)
by: Aremu, Toluwani, et al.
Published: (2024)
PaperRepro: Automated Computational Reproducibility Assessment for Social Science Papers
by: Zhang, Linhao, et al.
Published: (2026)
by: Zhang, Linhao, et al.
Published: (2026)
Can the Recovery Mechanism Survive AI? Skill Formation, Labor, and What Current Measurement Misses
by: Fan, Aysa Xuemo
Published: (2026)
by: Fan, Aysa Xuemo
Published: (2026)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
by: Wiedermann-Möller, Jonas, et al.
Published: (2026)
Chinese Court Simulation with LLM-Based Agent System
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts
by: Aftahee, Sabik, et al.
Published: (2025)
by: Aftahee, Sabik, et al.
Published: (2025)
Epistemic Control and the Normativity of Machine Learning-Based Science
by: Ratti, Emanuele
Published: (2026)
by: Ratti, Emanuele
Published: (2026)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
The Social Impact of Generative LLM-Based AI
by: Xie, Yu, et al.
Published: (2024)
by: Xie, Yu, et al.
Published: (2024)
Investigating Associational Biases in Inter-Model Communication of Large Generative Models
by: Dogan, Fethiye Irmak, et al.
Published: (2026)
by: Dogan, Fethiye Irmak, et al.
Published: (2026)
Similar Items
-
Prompt Selection Matters: Enhancing Text Annotations for Social Sciences with Large Language Models
by: Abraham, Louis, et al.
Published: (2024) -
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024) -
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
by: Walsh, Cole, et al.
Published: (2026) -
What Work is AI Actually Doing? Uncovering the Drivers of Generative AI Adoption
by: Agarwal, Peeyush, et al.
Published: (2025) -
LLM-Assisted Replication for Quantitative Social Science
by: Kubota, So, et al.
Published: (2026)